Site Reliability Engineer
Visa · Bengaluru
- Experience5–9 yrs
- SalaryNot disclosed
- Work modeonsite
- Levelsenior
- Posted2 Sept 2026
About Visa
Visa is hiring in Bengaluru in financial services. This role looks for around 5+ years of experience.
Skills
- cloud infrastructure
- networking
- CI/CD
- Infrastructure as Code
- Terraform
- Docker
- Kubernetes
- distributed systems
- microservices architecture
- Linux
- TCP/IP
- DNS
- monitoring
- observability
- incident management
- root cause analysis
- Go
- Bash
- SLIs
- SLOs
- SLAs
The role
A site reliability engineer at a financial services platform builds and operates reliable cloud infrastructure and distributed systems, using Infrastructure as Code and Kubernetes to improve service performance. The role manages observability and CI/CD pipelines while driving incident response and platform reliability practices.
Full job description
As a Site Reliability Engineer, you will own and manage the end-to-end platform and infrastructure at Macromill. This includes cloud infrastructure, networking, CI/CD, observability, and platform tooling. You will work closely with engineering teams to build a reliable, scalable, secure, and cost-efficient platform while improving developer productivity and system performance.
Responsibilities:
Own and manage the end-to-end platform, including infrastructure, networking, CI/CD, observability, and operations.
Design, build, and operate scalable, reliable, and secure systems on cloud platforms.
Develop and maintain CI/CD pipelines and automation to enable faster and safer deployments.
Implement and manage Infrastructure as Code (IaC) using tools like Terraform.
Improve system reliability, availability, and performance through monitoring, alerting, and optimisation.
Build and enhance observability (logging, metrics, tracing) to reduce incidents and improve debugging.
Lead incident management, troubleshooting, and root cause analysis for production systems.
Automate infrastructure and operational processes to reduce manual work and toil.
Support service onboarding and migration to modern platform architecture.
Ensure security best practices across infrastructure and applications.
Collaborate with developers to improve developer experience and platform usability.
Drive adoption of SRE, DevOps, and engineering best practices across teams.
Define and maintain SLIs, SLOs, and SLAs to ensure service reliability.
Requirements:
4+ years of experience building and operating production systems at scale.
Strong hands-on experience with cloud platforms (GCP, AWS, or Azure).
Experience owning and managing end-to-end infrastructure and platform systems.
Proficiency in Infrastructure as Code (Terraform or similar).
Experience with containers and orchestration (Docker, Kubernetes) in production.
Solid understanding of distributed systems and microservices architecture.
Strong knowledge of Linux systems, networking (TCP/IP, DNS), and troubleshooting.
Experience with CI/CD pipelines and automation.
Experience with monitoring and observability tools (metrics, logs, tracing).
Understanding of SLIs, SLOs, and service reliability practices.
Experience in incident handling, debugging, and root cause analysis.
Ability to design systems and write technical design documents.
Programming experience in Go or a similar language, along with scripting (e. g., bash).
Basic understanding of security best practices in cloud environments.
Strong ownership mindset and ability to work across teams.