Open role
Research Engineer – Benchmarking
Mercor
San Francisco, California, United StatesPosted Aug 18, 2026 · 10h ago$130k – $500k
Full-time$130k – $500kMid LevelOn-siteAI / ML
About this role
Mercor is seeking a Research Engineer to own benchmarking pipelines, evaluation systems, and failure analysis workflows that directly inform how we train and improve frontier language models. You will define how we measure tool use, agentic behavior, and real-world reasoning, working at the intersection of engineering and applied AI research. This role requires strong applied research background, coding skills, and comfort with ML models and evaluation code.
What we are looking for
6- Design, implement, and maintain benchmarks and metrics for tool use, agentic behavior, and real-world reasoning
- Build and operate LLM evaluation systems end-to-end, including scoring, dashboards, and reporting
- Run systematic failure analysis on model outputs and categorize failure modes
- Create and refine rubrics, automated evaluators, and scoring frameworks
- Quantify data usability, quality, and impact on key benchmarks
- Collaborate with AI researchers and applied AI teams to align evals with training objectives
