Jobs
Benefits & Perks
•Equity
•Equity
Required Skills
Kubernetes
Docker
Jenkins
Ansible
SQL
MySQL
Prometheus
Grafana
Kibana
Splunk
System administration
Networking
Security
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world.
NVIDIA is looking for a seasoned SRE to join its complex and fast-paced Infrastructure, Planning and Processes organization where you will be working as a Senior SRE Engineer. The position will be part of a fast-paced crew that develops and maintains sophisticated NVIDIA's internal Jenkins based CI/CD product for GPUs and Tegra systems. The team works with various other business units within NVIDIA Software such as Graphics Processors, Mobile Processors, Deep Learning, Artificial Intelligence and Driverless Cars to cater to their infrastructure & systems needs. As an SRE, you’ll also be working in conjunction with various teams such as software engineering to deploy these new products and handle our infrastructure, associated processes and systems. Keen attention to detail, problem-solving abilities, and a solid knowledge base are needed.
What you’ll be doing:
-
Manage NVIDIA's on-prem infrastructure. Maintain uptime, reliability and readiness of on-prem engineering cloud spread across multiple data centers.
-
Guard service level agreements (SLAs) for critical engineering services. Implement monitoring, alerting, and incident response procedures to ensure alignment to defined performance targets. Perform root cause analysis and post-mortems of incidents for any threshold breaches.
-
Deploy, configure, and manage applications and services on Kubernetes clusters. Implement logging, monitoring, and alerting solutions (e.g., Prometheus, Grafana, ELK/EFK). Ensure high availability, fault tolerance, and disaster recovery for Kubernetes workloads.
-
Help in capacity planning, optimization and better utilization efforts.
-
Support user reported issues & issues. Monitor alerts and take necessary action. Actively participate in WAR room for critical issues
-
Drive automation of monitoring to gain more insight into applications and system health.
-
Reuse AI techniques to extract useful signals about machines and jobs from the data generated.
What we need to see:
-
Experience of maintaining cloud infrastructure and highly-available production environment.
-
Experience handling and maintaining systems installed in on-premises data centers, with strong hands-on proficiency using BMC interfaces (Redfish), KVM, and IPMI tools for hardware provisioning, remote access, and troubleshooting. Knowledge and understanding of Openstack architecture and services is a plus.
-
Proven background working with databases, including relational databases such as SQL/MySQL, as well as time-series databases like Prometheus, with experience in data querying & performance tuning.
-
Solid understanding of networking principles and protocols, including TCP/IP, DNS, DHCP, and VLANs, with the ability to diagnose connectivity issues and support complex, distributed systems.
-
Practical experience in working with data analytics and visualization tools such as Kibana, Grafana, Splunk, or similar platforms, applied to analyze logs, metrics, and system behavior for monitoring and troubleshooting purposes.
-
Strong demonstrable experience in automation tools like Jenkins and/or Temporal along with configuration tools like Ansible.
-
Proficiency with Kubernetes, Docker, and virtualization technologies, with experience deploying, managing, and operating containerized workloads and virtualized infrastructure in production environments.
-
Advanced knowledge of standard security methodologies and protocols, including system hardening, access control, vulnerability management, and secure operations across infrastructure and application layers.
-
10+ years of demonstrable experience.
-
Bachelor's degree in Computer Science, Information Technology, or related field, or equivalent experience.
Ways to stand out from the crowd:
-
Previous experience with SRE teams managing on-prem infrastructure.
-
Experience managing NVIDIA hardware like GPUs and Tegras.
-
Thrives in a multi-tasking environment with constantly evolving priorities.
-
Outstanding interpersonal skills and communication with all levels of management.
With competitive salaries and a generous benefits package, we are widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us and, due to outstanding growth, our exclusive engineering teams are rapidly growing. If you're a creative and autonomous engineer with a real passion for technology, we want to hear from you.
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 176,000 USD - 276,000 USD for Level 4, and 208,000 USD - 333,500 USD for Level 5.
You will also be eligible for equity and benefits.
Applications for this job will be accepted at least until February 21, 2026.
This posting is for an existing vacancy.
NVIDIA uses AI tools in its recruiting processes.
NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
Total Views
1
Apply Clicks
0
Mock Applicants
0
Scraps
0
Similar Jobs

Principal Mission / Systems Engineer
Collins Aerospace (RTX) · US-MD-ANNAPOLIS JUNCTION-339 ~ 306 Sentinel Dr ~ 339 BLDG

Principal AI Engineer (GenAI) - Molecular Discovery
Bristol-Myers Squibb · 4 Locations

Software Engineering Senior Analyst
Cigna · Bengaluru, India

Software Engineer, Senior
Booz Allen Hamilton · McLean, VA

Senior Python Developer
Juniper Networks · Bangalore, Karnataka, India
About NVIDIA

NVIDIA
PublicA computing platform company operating at the intersection of graphics, HPC, and AI.
10,001+
Employees
Santa Clara
Headquarters
$4.57T
Valuation
Reviews
4.1
10 reviews
Work Life Balance
3.5
Compensation
4.2
Culture
4.3
Career
4.5
Management
4.0
75%
Recommend to a Friend
Pros
Great culture and supportive environment
Smart colleagues and excellent people
Cutting-edge technology and learning opportunities
Cons
Team-dependent experience and outcomes
Work-life balance issues with long hours
Politics and influence over competence
Salary Ranges
47 data points
Junior/L3
Mid/L4
Junior/L3 · Analyst
7 reports
$170,275
total / year
Base
$130,981
Stock
-
Bonus
-
$155,480
$234,166
Interview Experience
7 interviews
Difficulty
3.1
/ 5
Experience
Positive 0%
Neutral 86%
Negative 14%
Interview Process
1
Application Review
2
Recruiter Screen
3
Online Assessment
4
Technical Interview
5
System Design Interview
6
Team Review
Common Questions
Coding/Algorithm
System Design
Technical Knowledge
Behavioral/STAR
News & Buzz
Negotiating NVIDIA's Offer
Base, stock, and sign-on negotiable. Recruiters invested in closing candidates. CEO reviews all 42K employee salaries monthly. Stock growth has made many employees millionaires.
News
·
NaNw ago
NVIDIA Company Reviews
WLB rated 3.9/5 (lowest category). 64% satisfied with WLB but 53% feel burnt out. Compensation rated 4.4-4.5/5. Experience highly team-dependent.
News
·
NaNw ago
NVIDIA Culture Discussions
Team-dependent experience; sink-or-swim culture that rewards high performers but can be overwhelming. No politics, flat structure, but demanding workload with some teams requiring evening/weekend work.
News
·
NaNw ago
NVIDIA Interview Discussions
Technical bar is high with 4-6 rounds. Process takes 4-8 weeks. Expect C++ questions, LeetCode medium, and system design. Difficulty rated 3.16/5.
News
·
NaNw ago