Senior Principal SRE Engineering
Eli Lilly And Company · Hyderabad
- Experience14–19 yrs
- SalaryNot disclosed
- Work modeonsite
- Levelexecutive
- Posted17 Sept 2026
About Eli Lilly And Company
Eli Lilly And Company is hiring in Hyderabad in pharma biotech. This role looks for around 14+ years of experience.
Skills
- SLO governance
- Error-budget management
- OpenTelemetry
- Splunk
- Datadog
- New Relic
- Grafana
- Prometheus
- Terraform
- CI/CD
- Progressive delivery
- Kubernetes
- Container platforms
- AWS
- Azure
- GCP
- Incident command
- Blameless postmortems
- Root-cause analysis
- Regulated-environment compliance
- Change control
- Audit evidence
- Validated systems
- Technical leadership
- Chaos engineering
- Resilience testing
- AWS Fault Injection Service
The role
A site reliability engineer at a pharmaceutical and biotechnology company defines SLO governance and error budgets, builds observability with OpenTelemetry, and hardens infrastructure-as-code with Terraform. This person leads Kubernetes reliability, incident response, and self-healing automation while embedding regulated-environment practices and resilience engineering.
Full job description
Role summary
As the Senior Principal SRE Engineering Lead, you are the senior-most engineering authority for the reliability of the supported production estate. You own the bar for what reliable means in this organization: service-level objectives, error-budget governance, observability standards, and the engineering practices that protect both production and the teams engineering time.
This role is the judgment layer between an agents recommendation and a production change. You combine engineering rigor with operational pragmatism, and you decide which patterns surfaced from production warrant durable engineering investment. You hold the SRE engineering bar within the function - the cross-pillar reference architecture is owned by the Senior Architect, but the practice, standards, and engineering judgment of reliability at this site are yours.
You are a senior individual contributor. You do not manage people. You partner with the Reliability Leader, the Senior Principal Tech Shift Lead in the same pillar, the Senior Architect, and the senior engineering principals in the agentic automation team. Success is measured by organizational impact, sustained reliability outcomes, and the ability to scale reliability through systems and people, not heroics.
What you'll be doing
Reliability strategy and SLO governance
Define service-level objectives and indicators across the supported production estate, tiered by application risk and business impact.
Govern error-budget burn: when to slow change, when to invest in durable fixes, when to accept the budget.
Drive the adoption of reliability reporting and the disciplined operating cadence that makes SLOs real, not decorative.
Hold the SLO and error-budget conversation with product and application owners - including the conversation about what their service needs to change to meet the bar.
Engineering standards, observability, and self-healing
Establish observability and instrumentation standards as the contract every supported application must meet - and hold the line on them.
Set the bar for infrastructure-as-code, continuous-delivery hardening, and deployment safety across the supported estate.
Define the engineering work that flows from incidents and root-cause analyses into durable production change, and the architectural patterns (self-healing runbooks, graceful degradation, circuit breakers) that reduce repeat failure.
Incident learning, durable fixes, and partnership with agentic automation
Drive blameless postmortem culture; ensure root-cause analyses produce engineering work, not just narrative.
Govern which patterns from production warrant durable engineering investment, against the error-budget regime.
Partner with the production operations team in the same pillar on which patterns from the field warrant engineering attention; partner with the agentic automation engineering team on which fixes become safe agent-assisted remediations, and on the confidence thresholds, guardrails, and human-in-the-loop boundaries that make those remediations safe in production.
Lead high-severity incident response as incident commander when escalation reaches this seat, and coach the team to handle the rest.
Compliance, security, and regulated-environment readiness
Ensure reliability practices comply with Lilly standards and applicable regulatory requirements, and engineer them so that audit evidence falls out of normal operation.
Promote secure operational practices, auditability, and validated-environment-friendly engineering as part of the standards the team holds.
Act as a trusted technical leader in regulated and validated environments.
Technical leadership and talent development
Set the engineering bar through standards, expectations, and role modeling.
Mentor senior reliability engineers in the pillar; build the bench for sustained team growth and develop the next layer of principal engineers.
Influence engineering, product, and platform leaders through credibility and outcomes rather than authority.
Contribute to the evolution of enterprise-wide reliability practices in partnership with the Senior Architect and peer technical leaders.
How you will succeed
At the senior-most engineering individual-contributor level for reliability, success is defined by breadth of impact and sustained outcomes:
Be recognized as the senior reliability authority for your area.
Demonstrate measurable, sustained improvements such as: reduced major incidents, fewer recurring failures, improved time-to-recovery, and a credible error-budget regime.
Influence decisions across multiple teams and leaders through expertise and trust.
Scale reliability through systems, standards, and people, not heroics.
What you should bring
Required
14+ years of progressive engineering experience, with at least 6 years as a Site Reliability Engineer, Production Engineer, or equivalent, including a tour as the senior-most reliability engineer for a multi-application production estate - not a single product.
Production reliability experience in a regulated or audited environment (GxP, SOX, HIPAA, PCI, or equivalent), including familiarity with change-control discipline, audit evidence, and validated-system constraints.
Hands-on ownership of an SLO/SLI framework and error-budget policy across a multi-application estate, including defining SLIs, negotiating SLOs with product owners, and operationalizing burn-rate alerting and error-budget governance.
Deep, hands-on fluency across the SRE technical stack: observability and OpenTelemetry (Splunk, Datadog, New Relic, or Grafana/Prometheus); infrastructure-as-code (Terraform); CI/CD pipeline hardening and progressive delivery; Kubernetes and container platforms; and at least one major public cloud (AWS, Azure, or GCP).
Track record of leading high-severity incident response as incident commander, running blameless postmortems, and converting findings into durable engineering work - measured by reduced recurrence rather than narrative quality.
Demonstrated technical leadership at scale: mentoring senior ICs, influencing engineering and product leaders without org-chart authority, and writing the standards and engineering documents that set the bar for a discipline.
Bachelors degree or higher in Computer Science, Information Technology, or a closely related engineering field.
Preferred
Hands-on experience designing self-healing automation and running chaos engineering or resilience-testing programs (AWS Fault Injection Service,
Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.