Nvidia is hiring a
Senior SRE Software Engineer, Storage and Data
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world.
Site Reliability Engineering (SRE) is an engineering discipline that involves designing, building, and maintaining large-scale production systems with high efficiency and availability.
What you will be doing:
Assist in the design, implementation, and support of large-scale storage clusters, including monitoring, logging, and alerting.
Work with AI/ML workloads to capture and correlate behavior in large clusters and workflows, which are otherwise hard to understand.
Work closely with peers on the team to improve the lifecycle of services – from inception and design, through deployment, operation, and refinement.
Support services before they go live through activities such as system design consulting, developing software and frameworks, capacity management, and launch reviews.
Maintain services once they are live by measuring and monitoring availability, latency, and overall system health, including leveraging machine learning models.
Scale systems sustainably through mechanisms like AI/ML and automation, and evolve systems by pushing for changes that improve reliability and velocity.
Practice sustainable incident response and blameless postmortems. Be part of an on-call rotation to support production systems.
What we need to see:
BS degree in Computer Science or related technical field involving coding (e.g., physics or mathematics) or equivalent experience.
At least 5+ years of practical experience.
Experience with algorithms, data structures, complexity analysis, software design, and maintaining large-scale Linux based systems.
Experience in one or more of the following: C/C++, Java, Python, Go, Perl or Ruby, AI/ML frameworks and methodologies.
Good knowledge of infrastructure configuration management tools like Ansible, Chef, Puppet, and Terraform.
Experience in using observability and tracing-related tools like InfluxDB, Prometheus, and Elastic stack.
Ways to stand out from the crowd:
Demonstrated experience in having SRE mindset, customer-first approach, and focus on customer satisfaction and passion for ensuring customer success.
Experience with Git, code review, pipelines, and CI/CD.
Interest in crafting, analyzing, and fixing large-scale distributed systems. Strong debugging skills with a systematic problem-solving approach to identify complex problems.
Thrive in collaborative environments and enjoy working with various teams.
Experience in using or running large private and public cloud systems based on Kubernetes, OpenStack, and Docker. Flexible in adapting to different working styles.
Please mention that you found the job on ARVR OK. Thanks.