Senior System Software Engineer, Software Defined Networking

NVIDIA AI · Bengaluru

  • Experience5–6 yrs
  • SalaryNot disclosed
  • Work modeonsite
  • Levelsenior
  • Posted12 Sept 2026

About NVIDIA AI

NVIDIA AI is hiring in Bengaluru in semiconductors electronics. This role looks for around 5+ years of experience.

Skills

  • OVN
  • OVS
  • OpenFlow
  • C
  • Go
  • Bash
  • Python
  • Kubernetes
  • OVN-Kubernetes
  • Infrastructure as Code
  • Ansible
  • Terraform
  • ArgoCD
  • Flux
  • CI/CD
  • gRPC
  • REST
  • TLS
  • Linux
  • VM networking
  • Datacenter routing
  • Datacenter switching

The role

A software-defined networking engineer at an AI semiconductor company designs and operates multi-tenant cloud networking with OVN, OVS, and Kubernetes, building control-plane services for GPU workloads. The role develops network software in C and Go, and maintains Linux networking and CI/CD.

Full job description

We are looking for a Senior System Software Engineer, Software Defined Networking to design, build, and operate highly performant and scalable SDN solutions for NVIDIA's AI Clouds hosting GPU-accelerated workloads — including hyperscale multi-node training, inference, cloud gaming, and cloud functions. This role spans the full lifecycle of our SDN stack — from designing and developing new control and data plane software to ensuring operational excellence in production through reliability engineering, CI/CD, observability, and incident response.

What You'll Be Doing

Design and develop next-generation multi-tenant cloud SDN control and data plane software (OVS, OVN, OpenFlow) Build Infrastructure-as-a-Service virtual network orchestration and services using gRPC and REST to support tenant workload security and performance SLAs for BMaaS, VMaaS, and Kubernetes Drive upstream contributions to OVN-Kubernetes and related open-source projects; Develop software for network observability — monitoring, telemetry, intelligent metering, and performance analysis Operate and support OVS-OVN based SDN solutions in large-scale NVIDIA AI Cloud environments Own end-to-end observability for the SDN stack — build and maintain monitoring, alerting, distributed tracing, and dashboarding to ensure real-time insight into network health, performance, and tenant SLAs Design, enhance, and maintain CI/CD pipelines (GitLab) across Linux host networking, OVS, OVN, and Kubernetes CNIs Implement GitOps approaches or related experience for secure, seamless integration with cloud infrastructure; Drive reliability through incident management, resource monitoring, and performance tuning Collaborate with SRE, DevOps, and network engineering teams on production readiness and operational tooling

What We Need To See

BS/MS in Computer Science or related technical field, or a comparable blend of education and relevant experience 5+ years of proven experience in software development for large-scale distributed environments Expert-level knowledge of OVN, OVS, OpenFlow, and modern network protocols Strong programming skills in C and Go; advanced scripting in Bash and Python Deep knowledge of Kubernetes, practical experience deploying and supporting CNIs (OVN-Kubernetes) Hands-on experience with Infrastructure-as-Code and deployment tools (Ansible, Terraform, ArgoCD, Flux) Experience designing and operating complex, multi-stage CI/CD pipelines Hands-on experience developing secure, high-performance services using gRPC and REST with TLS and strong authentication Strong knowledge of datacenter routing, switching, and Linux host/VM networking

Ways To Stand Out From The Crowd

Contributions to open-source projects (especially OVS, OVN, OVN-Kubernetes, or other Kubernetes networking projects) Experience with hardware acceleration (GPU, DPU or equivalent experience) for networking Practical experience with major cloud providers (AWS, Azure, GCP) and hybrid/multi-cloud deployments SRE/DevOps top-level expertise — on-call, incident management, operations focused on service reliability targets, production ownership Experience with observability platforms and tools (Prometheus, Grafana, Jaeger, OpenTelemetry, ELK)