Lead Automation Engineer

Eli Lilly And Company · Hyderabad

  • Experience10–15 yrs
  • SalaryNot disclosed
  • Work modeonsite
  • Levelexecutive
  • Posted17 Sept 2026

About Eli Lilly And Company

Eli Lilly And Company is hiring in Hyderabad in pharma biotech. This role looks for around 10+ years of experience.

Skills

  • Observability
  • Prometheus
  • Grafana
  • OpenTelemetry
  • Splunk
  • Datadog
  • Dynatrace
  • Alerting
  • Python
  • PowerShell
  • Bash
  • SRE
  • Service-level objectives
  • Error budgets
  • Blameless postmortems
  • Rollback logic
  • Exception handling
  • Canary deployment
  • ITSM
  • Incident management
  • Change management
  • Problem management
  • CAPA
  • Change control

The role

A site reliability and platform engineer at a pharmaceutical technology company builds production automation from validated remediation procedures, applying observability and SRE practices to make autonomous operations safe. The work includes Python, rollback logic, and alerting strategies for regulated enterprise IT operations.

Full job description

Job Summary

You build the automation

and the observability around it

that turns a validated fix into a production-safe, autonomous action. Sitting inside Agentic Automation Engineering, you take remediation procedures authored and validated by SRE - and automation opportunities surfaced through discovery, telemetry, and alert correlation - and codify them into automation that runs safely, is fully instrumented, respects rollback and exception handling, and earns broader autonomy over time.

This is a hands-on individual-contributor role focused on build quality, codification rigor, and instrumentation depth. You work daily with SREs runbook library and graduation criteria, with the observability stack that proves an automation is behaving as designed, and with Operations outcome validation, so automation only ever runs whats been proven safe. Success is measured by automations shipped and codified, runbooks translated into working automation, the quality of the signals you emit alongside them, zero P1/P2 incidents caused by automation youve built, and the pace at which your automation earns broader autonomy.

What you'll be doing

Automation build - from validated fix to production automation : Build production-grade automation that executes validated remediation steps end to end - detection through action - with safe rollback built in.

Runbook codification : Translate SRE-authored remediation procedures into codified, executable automation steps: safe execution order, rollback steps, and exception handling.

Observability & instrumentation - first-class, not an afterthought : Design and implement the metrics, logs, traces, and events that make every automations behavior explainable in production - inputs, decisions, actions taken, rollbacks, outcomes.

SRE pattern implementation : Apply SRE patterns in what you build: error-budget awareness, blameless-postmortem-driven fixes, and graduation-criteria-aware rollout.

Cross-team partnership & autonomy graduation : Partner with SRE to validate that an automation meets graduation criteria (accuracy over volume, zero P1/P2 caused) before it moves to a higher autonomy tier.

How you will succeed

Be recognized as a dependable builder whose automations run safely, roll back cleanly, are observable end-to-end, and rarely need rework.

Demonstrate measurable throughput: runbooks codified, automations shipped, instrumentation delivered, and time-to-production for each.

Ship automation that graduates to higher autonomy tiers on schedule, with zero P1/P2 incidents caused along the way.

Build codified runbooks and telemetry clear enough that others can extend your work - and diagnose its behavior - without you in the room.

What you should bring

Required

10+ years of hands-on automation engineering experience, building and maintaining production automation or scripted remediation in an enterprise IT operations environment.

Strong, hands-on observability skills : designing SLIs/SLOs, instrumenting code and workflows with metrics, structured logs and distributed traces, and building dashboards and alerts that drive action. Practical experience with at least one major observability stack (Prometheus / Grafana / OpenTelemetry , Splunk, Datadog, Dynatrace, or equivalent).

Alerting maturity : designing signal-based alerts, correlation rules, and noise-reduction strategies; comfort tuning alerts based on real incident data rather than intuition.

Demonstrated experience turning documented fixes or remediation procedures into reliable, repeatable automation - not just one-off scripts.

Practical scripting/programming ability (Python, PowerShell, Bash, or similar) sufficient to build, test, and maintain production-grade automation and its instrumentation.

Working knowledge of SRE concepts - SLOs, error budgets, blameless postmortems, graduation/autonomy criteria - and the ability to encode them into automation safely.

Comfort working within safe-execution guardrails: rollback logic, exception handling, and staged or canary rollout of new automations.

Familiarity with ITSM processes - incident, change, problem, CAPA - and how automation fits into that lifecycle.

Ability to work cross-functionally with SRE, Reliability, and application teams to validate that an automation is safe to promote.

Comfortable operating in a regulated, audit-ready environment (life sciences or similar), where change control and evidence matter.

Bachelor's degree or higher in Computer Science, Information Technology, or a closely related field.

Clear written and verbal communication skills, including the ability to document automation logic, telemetry, and remediation procedures for others.

Preferred

Experience with agentic or AI

Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.