26 sep
|
Braintrust
|
Chile
Help review the quality and fairness of challenging tasks used to evaluate AI systems. You will inspect task instructions, model execution traces and grading behavior, then explain whether a result reflects genuine model performance or an issue with the task, grader or environment.
What you’ll do
- Check task instructions, source materials, reference solutions and evaluation criteria for consistency and completeness.
- Review model execution traces, tool calls and deliverables to assess whether successes and failures are justified.
- Identify brittle grading checks, unsupported criteria and valid alternative solutions that may have been marked incorrect.
- Investigate discrepancies and distinguish model limitations from task, grader, tool or environment issues.
- Write concise, evidence-backed findings and verify that revisions address the issues found.
What we’re looking for
- At least five years of relevant technical or analytical experience.
- Ability to read Python, SQL, shell scripts, structured data and execution logs to understand task setup and grading behavior.
- Strong written English, analytical judgment and attention to detail.
- Ability to give specific, reproducible feedback and explain uncertainty clearly.
- Experience in AI evaluation, technical QA, data analysis or benchmark development is helpful, but not required. Familiarity with Harbor task setup is a plus.
Engagement
- Remote contractor assignment for eight weeks.
- 40 hours per week, including eight hours of daily overlap with Pacific Time.
- Shortlisted applicants may be asked to complete an interest form before final review.
Compensation
- $23 per hour.
📌 AI Benchmark Quality Reviewer (Chile)
🏢 Braintrust
📍 Chile