Site Reliability Engineer
Flipkart · Bengaluru
- Experience2–6 yrs
- SalaryNot disclosed
- Work modeonsite
- Levelmid
- Posted17 Sept 2026
About Flipkart
Flipkart is hiring in Bengaluru in ecommerce retail. This role looks for around 2+ years of experience.
Skills
- Kubernetes
- AWS
- GCP
- Prometheus
- Grafana
- ELK
- Loki
- Jaeger
- Tempo
- SLI
- SLO
- SLA
- error budgets
- incident management
- Terraform
- Helm
- Ansible
- Python
- Go
- Shell
- Redis
- Kafka
- Solr
- DNS
- TCP/IP
- load balancing
- TLS
The role
A site reliability engineer at a large-scale e-commerce marketplace builds reliable production systems and manages Kubernetes, cloud infrastructure, and observability. The role automates remediation and CI/CD reliability checks with Terraform and Python.
Full job description
We are looking for a Site Reliability Engineer (SRE) to build and maintain reliable, scalable, and highly available production systems. You will work across infrastructure, Kubernetes, observability, incident management, automation, and cloud platforms.
The candidate will have responsibilities across the following functions:
Observability and Monitoring:
Build and maintain monitoring and observability using Prometheus, Grafana, ELK/Loki, and Jaeger/Tempo.
Design effective alerting strategies and reduce noisy/low-value alerts.
Build service-level dashboards for real-time system visibility.
Implement SLI/SLO-based monitoring and improve observability coverage.
Set up synthetic monitoring and canary checks for critical user journeys.
Incident Management and Reliability:
Participate in on-call rotations and handle P0/P1 production incidents.
Lead incident triage, troubleshooting, and resolution.
Conduct blameless postmortems and track corrective actions.
Create and maintain runbooks for common production issues.
Monitor SLOs, SLAs, and error budgets and proactively identify reliability risks.
Cloud and Infrastructure:
Design, provision, and manage infrastructure on AWS or GCP.
Work with compute, networking, storage, IAM, and managed databases.
Manage and operate Kubernetes clusters, including autoscaling, resource management, scheduling, and upgrades.
Perform capacity planning for traffic growth and peak events.
Optimise cloud infrastructure costs through rightsizing and capacity planning.
Manage infrastructure using Terraform/Helm/Ansible and follow GitOps practices.
Automation:
Automate repetitive operational tasks and reduce engineering toil.
Build self-healing mechanisms, automated remediation, and rollback processes.
Integrate reliability checks into CI/CD pipelines.
Develop production-quality automation using Python, Go, or Shell.
Requirements:
3-6 years of SRE / DevOps / Infrastructure Engineering experience.
Strong hands-on experience with Kubernetes.
Strong experience with AWS or GCP.
Hands-on experience with Prometheus and Grafana.
Experience with ELK or Loki and distributed tracing tools such as Jaeger/Tempo.
Strong understanding of SLI, SLO, SLA, error budgets, and incident management.
Experience with Terraform, Helm, or Ansible.
Strong scripting/programming skills in Python, Go, or Shell.
Hands-on experience with distributed systems/tools such as Redis, Kafka, or Solr.
Good understanding of DNS, TCP/IP, load balancing, and TLS.
Experience handling production incidents and on-call responsibilities.
Experience working in high-traffic production environments.
Good to Have:
Experience with Istio or Linkerd.
Experience with GitOps tools and workflows.
Experience with synthetic monitoring, canary deployments, chaos engineering, and automated remediation.
Experience in large-scale e-commerce or consumer internet platforms.