Open role
Machine Learning Research Scientist, Evaluations
Scale AI
San Francisco, CA; Seattle, WA; New York, NYPosted Aug 26, 2026 · 9h ago$181k – $226k
Full-time$181k – $226kMid LevelHybridAI / ML
About this role
Scale AI is seeking a Machine Learning Research Scientist to join their GenAI Research Organization. This role focuses on developing benchmarks and diagnosing failure modes in large language models and multimodal systems. You will analyze model behavior, design evaluation methods, and apply post-training expertise to improve AI development.
What we are looking for
6- Develop rigorous evaluations and diagnostic methods for frontier LLMs
- Analyze model behavior to identify and diagnose failure modes
- Design and build benchmarks for text and multimodal LLM capabilities
- Apply post-training expertise (SFT, RLHF, reward modeling) to address failures
- Collaborate with researchers and engineers on evaluation best practices
- Publish research findings in top-tier AI conferences
Skills mentioned
1Machine Learning
