AI Benchmark Quality Reviewer (Remote Contractor)
Braintrust
Job description
About the role
Help review the quality and fairness of challenging tasks used to evaluate AI systems. You will inspect task instructions, model execution traces and grading behavior, then explain whether a result reflects genuine model performance or an issue with the task, grader or environment.
Key responsibilities
- Check task instructions, source materials, reference solutions and evaluation criteria for consistency and completeness.
- Review model execution traces, tool calls and deliverables to assess whether successes and failures are justified.
- Identify brittle grading checks, unsupported criteria and valid alternative solutions that may have been marked incorrect.
- Investigate discrepancies and distinguish model limitations from task, grader, tool or environment issues.
- Write concise, evidence‑backed findings and verify that revisions address the issues found.
Required profile
- At least five years of relevant technical or analytical experience.
- Ability to read Python, SQL, shell scripts, structured data and execution logs.
- Strong written English, analytical judgment and attention to detail.
- Ability to give specific, reproducible feedback and explain uncertainty clearly.
Required skills
- Python
- SQL
- Shell scripting
- Execution logs
- Structured data
What we offer
- Compensation at $23 per hour.
- Remote contractor assignment for eight weeks, 40 hours per week with eight‑hour daily overlap with Pacific Time.
Questions fréquentes
Why are you reporting this job?
Apply in 30 seconds
Enter your email to apply. An account will be created automatically.
By continuing, you accept our terms of use.
Already have an account? Login
A question about this job?
Ask it here: you will get the full job summary by e-mail, right away.
Published 1 week ago
Expires 1 month from now
13 views · 0 interested
Boost your chances
Upload your CV — we will match you with relevant openings.
Analyzing your CV...
Braintrust