Site Reliability Engineer II

American Express · Chennai

  • Experience2–5 yrs
  • SalaryNot disclosed
  • Work modehybrid
  • Levelmid
  • Posted16 Sept 2026

About American Express

American Express is hiring in Chennai in financial services. This role looks for around 2+ years of experience.

Skills

  • Splunk
  • Elasticsearch
  • Prometheus
  • Grafana
  • Kubernetes
  • Docker
  • microservices architecture
  • logging
  • monitoring
  • tracing
  • performance analysis
  • Site Reliability Engineering
  • AWS
  • Microsoft Azure
  • Google Cloud
  • Linux/Unix
  • Java
  • Python
  • Bash
  • Infrastructure as Code
  • chaos engineering
  • disaster recovery
  • business continuity
  • error budgets
  • service-level objectives
  • service-level indicators

The role

A site reliability engineer at a financial services product company designs resilient systems and implements Infrastructure as Code and chaos engineering, while applying observability and cloud platforms to improve scalability, recovery, and performance.

Full job description

Job Description

Site Reliability Engineer II collaborates with engineering teams to enhance system resilience, scalability, and performance through feature development, automation, architectural design, resiliency testing, and disaster recovery planning, while promoting best practices for continuous improvement.

Responsibilities

Collaborates with Software Engineering teams to design, develop, and implement features that enhance system resilience, scalability, and performance, while identifying and addressing potential system bottlenecks and failure points with guidance from senior colleaguesDevelops and implements automation tools and frameworks, including infrastructure as code (IaC) practices to streamline operational workflows, deployment processes, and infrastructure management, with guidance from peers and leadersCollaborates with senior engineers to contribute to the architectural design of systems, ensuring that reliability, scalability, and performance considerations are integrated into design discussions and decision-making processesCollaborates in the design and execution of chaos engineering experiments and other resiliency testing, analyzing results and implementing improvements to enhance system robustness and recovery capabilities, with guidance from peers and leadersDevelops and implements of disaster recovery plans and business continuity strategies, ensuring systems can recover quickly and effectively from unexpected disruptionsCollaborates with seniors to promote and implement best practices such as error budgeting, service-level objectives (SLOs), and service-level indicators (SLIs), contributing to a culture of continuous improvement and reliabilityCollaborates and co-creates effectively with teams in product and the business to align technology initiatives with business objectives

Qualifications

Education Qualifications:

Bachelor’s degree in Computer Science, Information Technology, Engineering, and/or comparable experience; advance degree preferredKnowledge of modern observability stack – Splunk, Elastic Search, Prometheus, GrafanaKnowledge of containerization technologies (e.g., Kubernetes, Docker) and microservices architectureKnowledge of observability tools and methodologies, including experience with logging, monitoring, tracing, and performance analysis platformsKnowledge of cloud-based Site Reliability Engineering (SRE) practices and experience with public cloud platforms such as AWS, Azure, or Google Cloud

Work Experience

Experience in software development, or technology operations, with a focus on Site Reliability EngineeringExperience in Linux/Unix systems, object-oriented programming languages (e.g., Java), scripting languages (e.g., Python, Bash), and cloud platforms (e.g., AWS, Azure, GCP)

Licenses And Certifications

Advanced certification in Site Reliability Engineering (SRE) or related is a plus

About Us

At American Express, our culture is built on a 175-year history of innovation, shared values and Leadership Behaviors, and an unwavering commitment to back our customers, communities, and colleagues. From delivering differentiated products to providing world-class customer service, we operate with a strong risk mindset, ensuring we continue to uphold our brand promise of trust, security, and service.

As part of Team Amex, you’ll experience our powerful backing with comprehensive support for your holistic well-being and many opportunities to learn new skills, develop as a leader, and grow your career. Here, your voice and ideas matter, your work makes an impact, and together, you will help us define the future of American Express.

About The Team

We back you with benefits that support your holistic well-being so you can be and deliver your best. This means caring for you and your loved ones' physical, financial, and mental health, as well as providing the flexibility you need to thrive personally and professionally:

Competitive base salariesBonus incentivesSupport for financial-well-being and retirementComprehensive medical, dental, vision, life insurance, and disability benefits (depending on location)Flexible working model with hybrid, onsite or virtual arrangements depending on role and business needGenerous paid parental leave policies (depending on your location)Free access to global on-site wellness centers staffed with nurses and doctors (depending on location)Free and confidential counseling support through our Healthy Minds programCareer development and training opportunities

American Express is an equal opportunity employer and makes employment decisions without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, veteran status, disability status, age, or any other status protected by law.

Offer of employment with American Express is conditioned upon the successful completion of a background verification check, subject to applicable laws and regulations.