Open role
Principal Systems Software Engineer - Observability and Telemetry Platform
NVIDIA
US CA Santa ClaraPosted Aug 1, 2026 · 3h ago$272k – $431k
Full-time$272k – $431kPrincipalOn-siteConsumer Tech
About this role
NVIDIA is seeking a Principal Systems Software Engineer to design, build, and maintain large-scale production systems for their Observability and Telemetry Platform. This role focuses on ensuring high efficiency, availability, and reliability of GPU cloud services through automation, performance tuning, and proactive system management. You will be instrumental in the entire service lifecycle, from design to refinement, and will contribute to a culture of continuous improvement and innovation.
What we are looking for
6- Design, implement, and support large-scale Observability & Telemetry collection platform
- Focus on performance at scale, real-time monitoring, logging, and alerting
- Engage in the full lifecycle of services from inception to refinement
- Support services before and after launch, including system design and capacity management
- Scale systems sustainably through automation and improve reliability and velocity
- Practice sustainable incident response and blameless postmortems
Skills mentioned
4PythonDockerKubernetesLinux
