Site Reliability Engineer

Groww · Bengaluru

  • Experience4–8 yrs
  • SalaryNot disclosed
  • Work modeonsite
  • Levelmid
  • Posted2 Sept 2026

About Groww

Groww is hiring in Bengaluru in financial services. This role looks for around 4+ years of experience.

Skills

  • Site Reliability Engineering
  • Python
  • Go
  • Infrastructure as Code
  • Datadog
  • Terraform
  • Ansible
  • Splunk
  • Prometheus
  • ELK Stack
  • New Relic
  • Monitoring
  • Networking
  • Load Balancing
  • Incident Management
  • DevOps
  • Kubernetes
  • AWS
  • Azure
  • Google Cloud Platform
  • DNS management
  • Microservices
  • Cloud computing
  • Chaos Engineering
  • Service Mesh

The role

A site reliability engineer at a financial services product company architects scalable infrastructure and manages Kubernetes and Terraform for reliable production systems, using Python and cloud computing to improve availability and resilience. The role also applies incident management and monitoring to guide reliability engineering.

Full job description

Responsibilities:

Architect and lead the design of scalable, reliable infrastructure solutions.

Implement strategies for high availability, scalability, and low-latency performance.

Define service-level objectives (SLOs) and indicators (SLIs) to track performance and reliability.

Drive incident management, identifying root causes and providing long-term solutions.

Mentor junior engineers and foster a collaborative, learning-focused environment.

Design advanced monitoring and alerting systems for proactive system management.

Evolve Infrastructure as Code (IaC) practices to automate infrastructure provisioning.

Collaborate on reliability roadmaps, performance benchmarks, and disaster recovery plans.

Manage Kubernetes clusters at scale, integrating service meshes like Istio or Linkerd.

Implement chaos engineering principles for system resilience.

Influence technical direction, reliability culture, and organizational strategies.

Requirements:

Mandatory Skills: SRE, Python, Go, Iac, Datadog, Terraform, Ansible, Splunk, Prometheus, ELK Stack, New Relic, Monitoring, Networking, Load Balancing, Incident Management, DevOps, Kubernetes, AWS or Azure or GCP, Bachelor's/Master's degree in Computer Science or Engineering, or equivalent experience.

6-9 years in SRE, DevOps, or system architecture roles with large-scale production systems.

Proven experience in managing complex cloud environments (AWS, GCP, Azure).

Expertise in Kubernetes, container orchestration, and microservices.

Advanced programming/scripting skills in Python, Go, or Java.

Proficiency in monitoring/logging tools (Prometheus, ELK Stack, New Relic).

Strong understanding of networking, load balancing, and DNS management.

Leadership skills with experience mentoring teams and collaborating with stakeholders.

Strong communication, critical thinking, and incident management abilities.