SRE 3
PhonePe · Bengaluru
- Experience7–11 yrs
- SalaryNot disclosed
- Work modeonsite
- Levelexecutive
- Posted2 Sept 2026
About PhonePe
PhonePe is hiring in Bengaluru in financial services. This role looks for around 7+ years of experience.
Skills
- Linux
- IPtables
- IPsec
- Perl
- Golang
- Python
- Hadoop
- HDFS
- HBase
- Airflow
- YARN
- Ranger
- Kafka
- Pinot
- Puppet
- Salt
- Chef
- Ansible
- Networking
- SaltStack
- Docker
- Git
- ELK stack
- Grafana
- Prometheus
- opentsdb
- OpenTelemetry
The role
A site reliability engineer at a digital payments company designs reliable big data infrastructure, operating Linux systems and Hadoop ecosystems while automating provisioning, scaling, monitoring, and incident response. The role applies Python, Ansible, and Prometheus to improve availability, performance, and resilience.
Full job description
Responsibilities:
Manage, maintain, and support incremental changes to Linux/Unix environments.
Lead on-call rotations and incident responses, conducting root cause analysis and driving postmortem processes.
Design and implement automation systems for managing big data infrastructure, including provisioning, scaling, upgrades, and patching clusters.
Troubleshoot and resolve complex production issues while identifying root causes and implementing mitigating strategies.
Design and review scalable and reliable system architectures.
Collaborate with teams to optimise overall system performance.
Enforce security standards across systems and infrastructure.
Set technical direction, drive standardisation, and operate independently.
Ensure the availability, performance, and scalability of systems and services through proactive monitoring, maintenance, and capacity planning.
Resolve, analyse, and respond to system outages and disruptions and implement measures to prevent similar incidents from recurring.
Develop tools and scripts to automate operational processes, reducing manual workload, increasing efficiency and improving system resilience.
Monitor and optimise system performance and resource usage, identify and address bottlenecks, and implement best practices for performance tuning.
Collaborate with development teams to integrate best practices for reliability, scalability, and performance into the software development lifecycle.
Stay informed of industry technology trends and innovations, and actively contribute to the organisation's technology communities.
Develop and enforce SRE best practices and principles.
Align across functional teams on priorities and deliverables.
Drive automation to enhance operational efficiency.
Requirements:
Over 7 years of experience managing and maintaining distributed big data ecosystems.
Strong expertise in Linux, including IP, Iptables, and IPsec.
Proficiency in scripting/programming with languages like Perl, Golang, or Python.
Hands-on experience with the Hadoop stack (HDFS, HBase, Airflow, YARN, Ranger, Kafka, Pinot).
Familiarity with open-source configuration management and deployment tools such as Puppet, Salt, Chef, or Ansible.
Solid understanding of networking, open-source technologies, and related tools.
Excellent communication and collaboration skills.
DevOps tools: SaltStack, Ansible, Docker, Git.
SRE Logging and monitoring tools: ELK stack, Grafana, Prometheus, opentsdb, Open Telemetry.