Senior Site Reliability Engineering, Storage (India)
NVIDIA · All India
- Experience12–13 yrs
- SalaryNot disclosed
- Work modeonsite
- Posted25 Sept 2026
About NVIDIA
NVIDIA is hiring in All India in semiconductors electronics. This role looks for around 12+ years of experience.
Skills
- NAS
- SAN
- Object Storage
- SRE
- Infrastructure as Code
- Terraform
- Ansible
- Puppet
- SaltStack
- Docker
- Kubernetes
- Python
- Go
- Shell
- CI/CD
- Version control
The role
A site reliability engineer at an accelerated computing company designs and operates highly available NAS, SAN, and Object Storage platforms, applying SRE practices and Infrastructure as Code to automate provisioning, observability, and incident response. The role also uses Python and Kubernetes to improve storage reliability and scalability.
Full job description
NVIDIA has been redefining computer graphics, PC gaming, and accelerated computing for more than 25 years. Its an outstanding legacy of innovation thats fueled by phenomenal technology and amazing people. Today, were tapping into the unlimited potential of AI to define the next era of computing. Doing whats never been done before takes vision, innovation, and the worlds best talent. As an NVIDIAN youll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world.
We are seeking a Senior Site Reliability Engineer Storage, you will own the reliability, performance, and scalability of our global NAS, SAN, and Object Storage platforms that power critical internal and external services. You will combine deep storage expertise with strong automation and SRE practices to design, build, and operate highly available storage systems at scale.
What you will be doing:
- Lead design, deployment, and operations of production NAS, SAN, and Object Storage platforms, ensuring reliability, performance, and security.
- Capture requirements from partner teams, architect storage solutions, and drive endtoend implementation for new and existing services.
- Develop, maintain, and improve automation for provisioning, configuration, monitoring, incident response, and lifecycle management of storage infrastructure.
- Participate in oncall and incident response, lead troubleshooting of complex storage and performance issues, and drive root cause analysis and preventive actions.
- Define and track SLOs/SLIs and error budgets for storage services, using observability and analytics to continuously improve reliability and efficiency.
- Build and maintain runbooks, standard operating procedures, and comprehensive documentation for storage services and automation.
- Analyze capacity and usage trends, perform forecasting, and recommend scaling or optimization strategies to support business growth.
- Collaborate closely with SRE, infrastructure, networking, and application teams in a followthesun model to deliver consistent, highquality service.
- Mentor junior engineers, share best practices, and help drive adoption of SRE principles across the team.
What we need to see:
- 12+ years of experience in Site Reliability, DevOps, or Infrastructure Engineering, with significant focus on storage systems.
- Bachelors degree in Computer Science, Computer Engineering, or a related technical field or equivalent practical experience.
- Solid handson experience with design, deployment, and operations of enterprisegrade NAS, SAN, and/or Object Storage platforms.
- Solid understanding of SRE concepts (SLOs/SLIs, error budgets, incident management, observability, postmortems).
- Proficiency with Infrastructure as Code and configuration management tools (e.g., Terraform, Ansible, Puppet, SaltStack) and source control systems.
- Experience building and operating highly available, scalable infrastructure, including automation for provisioning, monitoring, and remediation.
- Experience with container and virtualization platforms (e.g., Docker, Kubernetes, hypervisors) and modern CI/CD and version control tools.
- Strong scripting or programming skills (e.g., Python, Go, Shell) to build tools, automate workflows, and integrate systems.
- Excellent communication and collaboration skills, with the ability to work effectively across distributed and crossfunctional teams.
Ways to stand out from the crowd:
- Experience with storage for highperformance computing, AI/ML workloads, or largescale data analytics.
- Proven ability to debug complex, distributed systems and storage performance issues.
- History of driving reliability improvements through datadriven analysis and automation.
- Experience leading technical initiatives, mentoring engineers, or acting as a technical lead on critical projects. .