Site Reliability Engineer - Architect
Groww · Bengaluru
- Experience7–11 yrs
- SalaryNot disclosed
- Work modeonsite
- Levelexecutive
- Posted2 Sept 2026
About Groww
Groww is hiring in Bengaluru in financial services. This role looks for around 7+ years of experience.
Skills
- Site Reliability Engineering
- Infrastructure as Code
- Kubernetes
- AWS
- Google Cloud Platform
- Microsoft Azure
- Python
- Go
- Java
- Prometheus
- ELK Stack
- New Relic
- Istio
- Linkerd
- microservices
- networking
- load balancing
- DNS management
- chaos engineering
The role
A site reliability architect at a financial services product company designs resilient cloud infrastructure and Kubernetes platforms, applying Site Reliability Engineering, Infrastructure as Code, and chaos engineering to improve availability and performance. The role also uses Prometheus and ELK Stack for observability.
Full job description
Responsibilities:
Architect and lead the design of scalable, reliable infrastructure solutions.
Implement strategies for high availability, scalability, and low-latency performance.
Define service-level objectives (SLOs) and indicators (SLIs) to track performance and reliability.
Drive incident management, identifying root causes and providing long-term solutions.
Mentor junior engineers and foster a collaborative, learning-focused environment.
Design advanced monitoring and alerting systems for proactive system management.
Evolve Infrastructure as Code (IaC) practices to automate infrastructure provisioning.
Collaborate on reliability roadmaps, performance benchmarks, and disaster recovery plans.
Manage Kubernetes clusters at scale, integrating service meshes like Istio or Linkerd.
Implement chaos engineering principles for system resilience.
Influence technical direction, reliability culture, and organisational strategies.
Requirements:
Bachelor's/Master's degree in Computer Science, Engineering, or equivalent experience.
6-9 years in SRE, DevOps, or system architecture roles with large-scale production systems.
Proven experience in managing complex cloud environments (AWS, GCP, Azure).
Expertise in Kubernetes, container orchestration, and microservices.
Advanced programming/scripting skills in Python, Go, or Java.
Proficiency in monitoring/logging tools (Prometheus, ELK Stack, New Relic).
Strong understanding of networking, load balancing, and DNS management.
Leadership skills with experience mentoring teams and collaborating with stakeholders.
Strong communication, critical thinking, and incident management abilities.