Site Reliability Engineer
Razorpay · Greater Bengaluru Area
- Experience10–11 yrs
- SalaryNot disclosed
- Work modeonsite
- Levelexecutive
- Posted11 Sept 2026
About Razorpay
Razorpay is hiring in Greater Bengaluru Area in financial services. This role looks for around 10+ years of experience.
Skills
- Kubernetes
- Linux
- computer networking
- Amazon Web Services
- Terraform
- relational databases
- caching
- message queues
- infrastructure as code
- observability
- capacity planning
- disaster recovery
- incident management
The role
A site reliability engineer at a payments platform builds Kubernetes infrastructure, designs chaos engineering practices, and operates cloud infrastructure for resilient production systems, using Terraform and Linux networking to automate reliability and incident response.
Full job description
About the roleYou will be the founding infrastructure SRE at Razorpay, responsible for the reliability of the platform every product team builds on: compute, orchestration, networking, data stores, CI/CD, and observability. Payment platform SREs make individual flows reliable; you make the ground they stand on reliable. Because you are one person covering a wide surface, your primary weapon is leverage: paved roads, standards, and automation that let service teams own their own reliability.
What you will doOwn the reliability of core infrastructure: Kubernetes clusters, networking and ingress, DNS, databases and caches, message queues, and the CI/CD platform.Build the paved road: golden paths for deployment, standard health checks, default dashboards and alerts, and production readiness templates that product teams adopt because they are the easiest option.Design and test failure domains: multi-AZ and multi-region strategy, failover automation, capacity planning, and disaster recovery drills that are actually run, not just documented.Own infrastructure-level rollouts and rollbacks: cluster upgrades, database migrations, and platform changes executed without customer-visible impact.Build and tune the observability platform so that every team can see its own SLIs without asking for help.Carry the pager for platform services, lead infrastructure incidents, and drive systemic fixes across teams.Establish a shared on-call model with existing platform engineers so that infrastructure coverage does not depend on one person.Drive cost-aware capacity management: headroom for peak events (festival sales, month-end settlement spikes) without permanent overprovisioning.
What we are looking for10+ years of experience, with substantial time owning production infrastructure at a company with real scale (thousands of hosts or containers, or traffic with hard peaks).Deep Kubernetes and container orchestration experience in production, including upgrades, capacity, and multi-tenant reliability.Strong Linux internals and networking depth: you can debug from a TCP dump, a flame graph, or kernel-level metrics, not just from a dashboard.Production experience with at least one major cloud (AWS preferred) and infrastructure as code (Terraform or similar).Experience operating stateful systems under load: relational databases, caches, and queues, including failover and data migration without downtime.A track record of building platform tooling or standards adopted by other teams, because a team of one scales only through software and influence.Genuine on-call and incident command experience on infrastructure that many teams depended on.The temperament for a founding role: comfortable with ambiguity, able to prioritize ruthlessly, and able to say no with data.
Nice to haveExperience running infrastructure for payments, banking, or another regulated, correctness-critical domain.Chaos engineering or failure injection experience at the infrastructure layer.Experience planning for extreme traffic events (large sale days, live sports streaming, ticket sales).