Open role
Member of Technical Staff - Research & Post-training
Preference Model
San Francisco, California, United StatesPosted Sep 11, 2026 · 5h ago$200k – $350k
Full-time$200k – $350kOn-siteAI / ML
About this role
Preference Model is seeking a Research Engineer or Scientist to advance the field of self-directed learning in large language models. This role involves implementing novel approaches, shaping research directions, and training/evaluating models on proprietary RL environments. You will architect and optimize RL training infrastructure, design training methodologies, and profile/optimize training runs for efficiency.
What we are looking for
6- Train and evaluate models on proprietary RL environments
- Architect and optimize RL training infrastructure
- Design, implement, and test RL training environments and methodologies
- Profile and optimize training runs for throughput and iteration speed
- Experience running end-to-end LLM post-training pipelines (7B+ models)
- Proficiency in Python and PyTorch or JAX
