Site Reliability Engineer III
AMERICAN EXPRESS · Bengaluru
- Experience4–11 yrs
- SalaryNot disclosed
- Work modeonsite
- Posted23 Sept 2026
About AMERICAN EXPRESS
AMERICAN EXPRESS is hiring in Bengaluru in financial services. This role looks for around 4+ years of experience.
Skills
- Splunk
- Elasticsearch
- Prometheus
- Grafana
- Kubernetes
- Docker
- microservices architecture
- logging
- monitoring
- tracing
- performance analysis
- Site Reliability Engineering
- AWS
- Microsoft Azure
- Google Cloud
- Linux/Unix
- Java
- Python
- Bash
- infrastructure as code
- chaos engineering
- disaster recovery
- business continuity
- error budgets
- service-level objectives
- service-level indicators
The role
A site reliability engineer at a financial services product company builds resilient cloud systems through Site Reliability Engineering, chaos engineering, and infrastructure as code, and applies observability practices with Kubernetes and Python.
Full job description
Site Reliability Engineer III advances efforts to enhance system resilience, scalability, and performance through feature development, automation, architectural design, chaos engineering, and disaster recovery planning, while promoting best practices for continuous improvement and reliability.
Education Qualifications:
Bachelor s degree in Computer Science, Information Technology, Engineering, and/or comparable experience; advance degree preferred
Knowledge of modern observability stack - Splunk, Elastic Search, Prometheus, Grafana
Knowledge of containerization technologies (e.g., Kubernetes, Docker) and microservices architecture
Knowledge of observability tools and methodologies, including experience with logging, monitoring, tracing, and performance analysis platforms
Knowledge of cloud-based Site Reliability Engineering (SRE) practices and experience with public cloud platforms such as AWS, Azure, or Google Cloud
Work Experience:
Experience in software development, or technology operations, with a focus on Site Reliability Engineering
Experience in Linux/Unix systems, object-oriented programming languages (e.g., Java), scripting languages (e.g., Python, Bash), and cloud platforms (e.g., AWS, Azure, GCP)
Licenses and Certifications:
Advanced certification in Site Reliability Engineering (SRE) or related is a plus
Manages the collaboration with Software Engineering teams to design, develop, and implement features that enhance system resilience, scalability, and performance, proactively identifying and resolving system bottlenecks and failure points
Develops and refines sophisticated automation tools and frameworks, including advanced infrastructure as code (IaC) practices, to streamline operational workflows, deployment processes, and infrastructure management, ensuring high system efficiency
Engages in architectural design discussions, ensuring that advanced reliability, scalability, and performance considerations are integrated into strategic decision-making processes
Designs and executes comprehensive chaos engineering experiments and advanced resiliency testing, analyzing results to implement robust improvements that enhance system robustness and recovery capabilities
Develops, optimizes, and maintains comprehensive disaster recovery plans and business continuity strategies, ensuring systems can recover quickly and effectively from complex and unexpected disruptions
Advocates for observability practices by promoting and implementing best practices such as error budgeting, service-level objectives (SLOs), and service-level indicators (SLIs), contributing to a culture of continuous improvement and reliability
Collaborates and co-creates effectively with teams in product and the business to align technology initiatives with business objectives
Disclaimer: This job posting and location has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.