Open role
Senior Applied Scientist, Efficient LLM Inference & Model Optimization
Nebius
Palo Alto, California, United StatesPosted Oct 7, 2026 · 7h ago$195K – $262K
Full-time$195K – $262KSenior levelRemoteSaaS
About this role
Nebius is seeking a Senior Applied Scientist to focus on efficient LLM and VLM inference and model optimization. This role involves owning research projects from hypothesis to production, publishing credible work, and collaborating with engineers to ship results. You will invent, evaluate, and productionize methods for inference optimization and build high-quality prototypes.
What we are looking for
6- Own research projects from hypothesis through experiment, ablation, prototype, and production handoff
- Prepare internal reports, technical blogs, or papers
- Partner directly with MLEs to ensure research prototypes become usable production components
- Define and execute research programs in efficient LLM and VLM inference
- Invent, evaluate, and productionize methods for quantization, distillation, speculative decoding, etc.
- Build high-quality prototypes in PyTorch, Triton, CUDA-adjacent tooling, or inference-serving frameworks
Skills mentioned
2PythonMachine Learning
