SRE Professional
BT Group · Bengaluru
- Experience4–6 yrs
- SalaryNot disclosed
- Work modehybrid
- Levelmid
- Posted20 Sept 2026
About BT Group
BT Group is hiring in Bengaluru in telecom. This role looks for around 4+ years of experience.
Skills
- Java
- AWS
- SQL
- Oracle
- Dynatrace
- Shell script
- Python
- AIOps
- GenAI
- incident management
- problem management
- change management
- release management
- resilience engineering
- capacity planning
The role
A site reliability engineer at a telecommunications product company ensures resilient customer-facing services across Java applications, AWS platforms, and CRM integrations. The role analyzes observability data, investigates production incidents, and improves self-healing automation and resilience engineering.
Full job description
About the role
We are looking for an experienced Site Reliability Engineer (SRE) to ensure the reliability, performance, and operational excellence of critical customer-facing applications across Siebel, Java, and AWS platforms. The role is responsible for understanding end-to-end business processes, application architecture, system integrations, and customer journeys to effectively manage service reliability and drive continuous improvement. Working with support partners, engineering teams, and business stakeholders, the role uses observability, operational data, and technical analysis to identify issues, improve application resilience, increase automation, and enhance customer and business outcomes. The role also supports the adoption of AIOps and self-healing capabilities to improve operational efficiency and service quality.
Role & responsibilities
Develop a strong understanding of customer journeys, business processes, application architecture, and integrations across CRM platform, AWS, and supporting platforms to effectively support service outcomes.
Monitor service health, performance, and availability using observability platforms and operational metrics, proactively identifying risks and improvement opportunities.
Investigate complex production issues through analysis of application logs, database queries, transaction flows, and monitoring data to support rapid resolution and root cause identification.
Provide technical oversight on investigations, validating root causes, challenging assumptions, and ensuring high-quality incident resolution.
Driving incident, problem, and change management activities, leveraging root cause analysis, preventive actions, and continuous service improvement to enhance reliability and prevent recurring issues.
Partner with engineering, business, and supplier teams to deliver application, operational, and customer experience improvements aligned to business objectives.
Develop and maintain service dashboards, operational reporting, and performance insights to support governance, decision-making, and continuous improvement.
Identify and implement opportunities for automation, self-service, and AI-driven operational capabilities to improve efficiency, resilience, and customer outcomes
Skills and Experience Required
Experience supporting or operating enterprise applications within complex production environments.
Strong understanding of application architecture, system integrations, data flows and customer-facing business processes.
Ability to analyse application, middleware and infrastructure logs to identify issues and support root cause investigations.
Knowledge of SQL/Oracle databases with experience troubleshooting production application issues.
Hands-on experience with monitoring and observability platforms such as Dynatrace, including dashboard creation, alert analysis and trend identification.
Understanding of Java-based applications, APIs, integration services and AWS-hosted platforms.
Experience working with incident, problem, change and release management processes.
Strong analytical and stakeholder management skills with the ability to challenge technical investigations, drive service improvements and influence engineering teams.
Ability to identify automation, resilience and customer experience improvement opportunities through operational insights.
Exposure to scripting or automation technologies such as Shell script, Python etc.
Experience leveraging AI/ML/GenAI capabilities to improve service reliability, operational efficiency, observability, incident response, root cause analysis, and self-healing automation.
Understanding of application and platform performance management, capacity planning, trend analysis, demand forecasting, and resilience engineering for business-critical production services.