Infrastructure Developer
Adobe · Delhi
- Experience6–9 yrs
- SalaryNot disclosed
- Work modeonsite
- Levelsenior
- Posted2 Sept 2026
About Adobe
Adobe is hiring in Delhi in technology software. This role looks for around 6+ years of experience.
Skills
- Kubernetes
- GPU infrastructure
- Helm
- Kustomize
- GitOps
- AWS
- Terraform
- Pulumi
- Prometheus
- Grafana
- AlertManager
- OpenTelemetry
- Jaeger
- Tempo
- Loki
- Elasticsearch
- OpenSearch
- SLO/SLI design
- TCP/IP
- DNS
- TLS
- HTTP/2
- gRPC
- Cilium
- Calico
- Flannel
- Istio
- Linkerd
- Go
- Python
- Rust
- distributed systems
- Kubernetes operators
The role
An infrastructure developer at a software product company designs Kubernetes-native infrastructure for distributed GPU training, builds production platform services with Go, Python, or Rust, and automates cloud environments with Terraform. The role also applies observability and distributed systems practices to improve reliability at scale.
Full job description
We are looking for an experienced Infrastructure Developer (6-9 years) to help design, build, and scale the platform that powers our most demanding ML training workloads. This is a hands-on engineering role where you will write production-grade code, drive meaningful technical initiatives, and contribute to the reliability of an infrastructure that thousands of GPU hours depend on every day. You bring strong Kubernetes skills, solid networking fundamentals, a developer's mindset, and the ability to own projects end-to-end with limited supervision. You have operated systems at a significant scale and are ready to step up into broader technical leadership.
Responsibilities:
Build for scale: Design and improve Kubernetes-native infrastructure that runs distributed GPU training jobs reliably and efficiently. You will own significant components and drive their evolution.
Lead focused initiatives: Own meaningful projects end-to-end, write design docs, gather input from stakeholders, and deliver under realistic timelines, often collaborating with engineers across time zones.
Codify infrastructure: Define and ship cloud infrastructure through IaC (Terraform/Pulumi). Apply the same rigour, testing, and review discipline to infra changes as to application code.
Strengthen observability: Contribute to and extend deep observability stacks, metrics, distributed tracing, log aggregation, SLO/SLI frameworks that surface problems before they become incidents.
Write production code: Build automation, internal tooling, operators, and platform services in Go, Python, or Rust. This is not a YAML-only role.
Own reliability: Participate in incident response, post-mortems, and reliability reviews. Drive systemic fixes, not just workarounds. Be a strong contributor to the on-call culture.
Solve hard networking problems: Debug and resolve complex cluster networking issues, CNI, BGP, service mesh, DNS at scale, east-west traffic, and throughput tuning.
Mentor and grow: Raise the technical bar through code reviews, design feedback, and knowledge sharing with peers and more junior engineers.
The core requirements for the job include the following:
Kubernetes and GPU Infrastructure:
6-9 years in SRE, platform engineering, or infrastructure roles.
Strong working knowledge of Kubernetes internals: scheduler, kubelet, CRDs, operators, admission controllers.
Hands-on experience running GPU/accelerator training workloads in production.
Familiarity with multi-cluster management and workload placement strategies.
Helm, Kustomize, GitOps (Flux/ArgoCD) practical experience and good judgment on when to use them.
Cloud and Infrastructure as Code:
Solid hands-on AWS experience (VPC, EKS, EC2 S3 IAM; TGW a plus).
Production experience with Terraform or Pulumi, modular and tested.
CI/CD for infrastructure: drift detection, plan gating, and rollback strategies.
Working understanding of cost optimisation, reserved capacity, and spot instance management.
Observability:
Prometheus, Grafana, AlertManager production experience, not just lab setups.
Exposure to distributed tracing: OpenTelemetry, Jaeger, or Tempo.
Log aggregation: Loki, Elasticsearch/OpenSearch.
Comfort with SLO/SLI design, error budgets, and multi-tier alerting.
Networking Fundamentals:
Strong TCP/IP, DNS, TLS, HTTP/2 gRPC fundamentals.
Practical experience with CNI plugins: Cilium, Calico, or Flannel and their trade-offs.
Familiarity with service mesh (Istio/Linkerd), ingress controllers, and API gateways.
Ability to debug under load: packet captures, eBPF traces, kernel counters.
Coding and System Design:
Production-quality code in Go, Python, or Rust you ship, not just a script.
Solid grasp of distributed systems fundamentals: consistency, availability, failure modes.
Experience writing Kubernetes operators or working with controller-runtime patterns.
Engaged code reviewer, thoughtful, constructive, and consistent.
Clear technical writer: design docs, ADRs, runbooks that others can actually use.
Collaboration and Ownership:
Has delivered meaningful, cross-functional projects from design to production.
Being comfortable with ambiguity can break down a problem and make progress without a perfect spec.
Experience working async across distributed teams and time zones.
A strong communicator can explain infra trade-offs clearly to peers and partner teams.
Self-driven identifies problems, proposes solutions, and follows through to outcomes.
Bonus Points:
Azure / GCP hands-on experience.
Familiarity with ML training pipeline internals.
eBPF-based observability or networking.
Chaos engineering or game day participation.
Open-source infrastructure contributions.
Security, compliance, or audit exposure.