NVIDIA

Senior Site Reliability Engineer, HPC and LSF

Reposted 7 Days Ago

Be an Early Applicant

4 Locations

Expert/Leader

4 Locations

Expert/Leader

As a Site Reliability Engineer, you'll improve NVIDIA's infrastructure for chip development, manage schedulers, automate processes, and troubleshoot issues in a large-scale HPC environment.

The summary above was generated by AI

NVIDIA is the leader in AI, machine learning and datacenter acceleration. NVIDIA is expanding that leadership into datacenter networking with ethernet switches, NICs and DPUs NVIDIA has continuously reinvented itself over two decades. Our invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern AI — the next era of computing. NVIDIA is a “learning machine” that constantly evolves by adapting to new opportunities that are hard to solve, that only we can tackle, and that matter to the world. This is our life’s work, to amplify human imagination and intelligence. Make the choice, join our diverse team today!

As a member of the Hardware Infrastructure Farm team, you will provide leadership in the design and implementation of ground breaking compute clusters that powers all silicon development across NVIDIA. We seek an expert to build and operate these clusters at high reliability, efficiency, and performance and drive foundational improvements and automation to improve engineer's productivity. As a Site Reliability Engineer, you are responsible for the big picture of how our systems relate to each other, we use a breadth of tools and approaches to tackle a broad spectrum of problems. Practices such as limiting time spent on reactive operational work, blameless postmortems and proactive identification of potential outages factor into iterative improvement that is key to both product quality and interesting dynamic day-to-day work. SRE's culture of diversity, intellectual curiosity, problem solving and openness is important to our success. Our organization brings together people with a wide variety of backgrounds, experiences and perspectives. We encourage them to collaborate, think big and take risks in a blame-free environment. We promote self-direction to work on meaningful projects, while we also strive to build an environment that provides the support and mentorship needed to learn and grow.

What you’ll be doing:

Manage and support workload and resource schedulers in a large-scale HPC environment.
Automate Everything: Develop automation scripts to automate deployment, configuration management, and operational monitoring.
Develop solutions for complex computing resource management requirements.
Extract and leverage grid performance metrics for troubleshooting and performance optimization.
Troubleshoot Complex Issues: Perform comprehensive troubleshooting from bare metal to application level, ensuring system reliability and efficiency.
Develop, define and document standard methodologies to share with internal teams.
Collaborate with domain experts to improve how our chip development process utilizes our infrastructure.
Directly contribute to the overall quality and improve time to market for our next generation chips.

What we need to see:

Extensive knowledge with job scheduler administration (e.g. IBM Spectrum LSF or SLURM).
Proficient in administering Centos/RHEL Linux distributions.
In depth understating of container technologies like Docker.
Proficiency in UNIX scripting languages and Python.
Excellent problem-solving skills, with the ability to analyze complex systems, identify bottlenecks, and implement scalable solutions.
Excellent communication and teamwork skills, with the ability to work effectively with diverse teams and individuals.
10+ years experience in a large, distributed Linux environment.
BS in Computer Science, similar degree or equivalent experience.

Ways to stand out from the crowd:

Experience analyzing and tuning performance for a variety of HPC or EDA workloads.
Solid understanding of cluster configuration managements tools such as Ansible.
Proficiency in Perl for maintaining legacy automation scripts.
Deep understanding of distributed system principles.

The base salary range is 184,000 USD - 287,500 USD. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions.

You will also be eligible for equity and benefits. NVIDIA accepts applications on an ongoing basis.

NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Top Skills

Ansible

Centos

Docker

Ibm Spectrum Lsf

Perl

Rhel

Slurm

Unix

Similar Jobs

NinjaOne

Senior Software Engineer C++

Yesterday

Remote

Hybrid

Austin, TX, USA

150K-220K Annually

Senior level

150K-220K Annually

Senior level

Information Technology • Productivity • Software • Infrastructure as a Service (IaaS)

As a Senior Software Engineer, you will design, develop, and maintain IT Operations products, providing mentorship while implementing scalable solutions and ensuring customer satisfaction.

Top Skills: AWSC++GrpcJavaKotlinPostgresSQLTeamcity

Capital One

Sr. Software Engineer, Full Stack (Java, AWS, Angular) - Dealer Tech

Yesterday

Hybrid

144K-181K Annually

Mid level

144K-181K Annually

Mid level

Fintech • Machine Learning • Payments • Software • Financial Services

The Sr. Software Engineer will design, develop, and support full-stack solutions, collaborating within Agile teams and utilizing various technologies, including Java and AWS.

Top Skills: AngularAWSDockerGoJavaKubernetesNode.jsNoSQLPythonRdbmsScalaSQL

The Aerospace Corporation

Docking Systems Project Engineer

Yesterday

Hybrid

Houston, TX, USA

Senior level

Aerospace • Artificial Intelligence • Cloud • Machine Learning • Software • Cybersecurity • Defense

As a Docking Systems Project Engineer, you will provide mechanical system architecture and design support for lunar rovers and docking mechanisms, manage projects, and assist in systems engineering and integration for NASA's human spaceflight missions.

Top Skills: Aerospace EngineeringMechanical Engineering

What you need to know about the Austin Tech Scene

Austin has a diverse and thriving tech ecosystem thanks to home-grown companies like Dell and major campuses for IBM, AMD and Apple. The state’s flagship university, the University of Texas at Austin, is known for its engineering school, and the city is known for its annual South by Southwest tech and media conference. Austin’s tech scene spans many verticals, but it’s particularly known for hardware, including semiconductors, as well as AI, biotechnology and cloud computing. And its food and music scene, low taxes and favorable climate has made the city a destination for tech workers from across the country.

Key Facts About Austin Tech

Number of Tech Workers: 180,500; 13.7% of overall workforce (2024 CompTIA survey)
Major Tech Employers: Dell, IBM, AMD, Apple, Alphabet
Key Industries: Artificial intelligence, hardware, cloud computing, software, healthtech
Funding Landscape: $4.5 billion in VC funding in 2024 (Pitchbook)
Notable Investors: Live Oak Ventures, Austin Ventures, Hinge Capital, Gigafund, KdT Ventures, Next Coast Ventures, Silverton Partners
Research Centers and Universities: University of Texas, Southwestern University, Texas State University, Center for Complex Quantum Systems, Oden Institute for Computational Engineering and Sciences, Texas Advanced Computing Center

Apply Save

By clicking Apply you agree to share your profile information with the hiring company.