Open role
Member of Technical Staff - Research & Post-training
Preference Model
San Francisco, California, United StatesPosted Aug 6, 2026 · 2d ago$200k – $350k
Full-time$200k – $350kMid LevelOn-siteAI / ML
About this role
Preference Model is seeking a Research Engineer to advance self-directed learning in large language models. You will implement novel approaches, shape research directions, and train/evaluate models using proprietary RL environments. This role involves architecting and optimizing RL training infrastructure and improving research iteration cycles.
What we are looking for
6- Train and evaluate models on proprietary RL environments
- Architect and optimize RL training infrastructure
- Design, implement, and test RL training environments and methodologies
- Profile and optimize training runs for throughput
- Experience running LLM post-training pipelines (7B+ models)
- Proficiency in Python and PyTorch or JAX
