Lead Automation Engineer

Eli Lilly and Company · Greater Hyderabad Area

  • Experience10–11 yrs
  • SalaryNot disclosed
  • Work modeonsite
  • Levelexecutive
  • Posted16 Sept 2026

About Eli Lilly and Company

Eli Lilly and Company is hiring in Greater Hyderabad Area in pharma biotech. This role looks for around 10+ years of experience.

Skills

  • Ansible
  • Terraform
  • Prometheus
  • Grafana
  • OpenTelemetry
  • Splunk
  • Datadog
  • Dynatrace
  • Python
  • PowerShell
  • Bash
  • SRE
  • Service-level indicators
  • Service-level objectives
  • Alert correlation
  • Rollback
  • Exception handling
  • ITSM
  • Incident management
  • Change management
  • Problem management
  • CAPA
  • Computer Science
  • Information Technology

The role

An automation engineer at a pharmaceutical technology organization builds production automation for SRE remediation, codifies runbooks, and instruments autonomous actions with observability and SRE practices. The role applies Python, Ansible, and Terraform to safe rollback, alert correlation, and staged autonomy.

Full job description

At Lilly, the work is demanding because patients are waiting. We unite caring with discovery to help make life better for people around the world, knowing that every decision, every detail, and every day matters. Headquartered in Indianapolis, Indiana, our over 50,000 employees around the globe take on complex challenges to discover and deliver life-changing medicines, strengthen how health is understood and managed, and support the communities we serve. This is hard, urgent, selfless work—but it’s work worth doing. If you’re driven by purpose and ready to bring your best to work that truly matters for patients, we invite you to join us.

About The Technology Organization

Technology at Lilly builds and operates mission-critical digital products and platforms that support the discovery, development, and delivery of medicines that make life better for people around the world. Our teams operate in highly regulated, high-availability environments, where operational excellence, reliability, and quality are non-negotiable.

Within Tech@Lilly, the Digital Core organization applies a product, platform, and reliability-first mindset, ensuring that operational capabilities scale sustainably across the enterprise.

About The Team

Tech@Lilly builds and maintains capabilities using pioneering technologies like the most prominent tech companies. What differentiates Lilly IT is that we redefine what's possible through tech to advance our purpose — creating medicines that make life better for people around the world — including data-driven drug discovery, connected clinical trials, resilient enterprise platforms, and intelligent digital operations.

The Digital Core team leads Lilly's transformation into the Digital and AI era. Lilly Capability Centre India (LCCI), Hyderabad, is Lilly's premier Global Technology Hub, harnessing data, AI, analytics, and digital solutions to revolutionize healthcare and improve patient outcomes worldwide.

Role summary

You build the automation and the observability around it that turns a validated fix into a production-safe, autonomous action. Sitting inside Agentic Automation Engineering, you take remediation procedures authored and validated by SRE — and automation opportunities surfaced through discovery, telemetry, and alert correlation — and codify them into automation that runs safely, is fully instrumented, respects rollback and exception handling, and earns broader autonomy over time.

This is a hands-on individual-contributor role focused on build quality, codification rigor, and instrumentation depth. You work daily with SRE's runbook library and graduation criteria, with the observability stack that proves an automation is behaving as designed, and with Operations' outcome validation, so automation only ever runs what's been proven safe. Success is measured by automations shipped and codified, runbooks translated into working automation, the quality of the signals you emit alongside them, zero P1/P2 incidents caused by automation you've built, and the pace at which your automation earns broader autonomy.

What you'll be doing

Automation build — from validated fix to production automationBuild production-grade automation that executes validated remediation steps end to end — detection through action — with safe rollback built in.Take automation opportunities and known-fix candidates surfaced through discovery, telemetry, and alert correlation, and turn them into shipped, tested automations.Maintain and extend the existing automation portfolio (scripted orchestration, platform tooling, and, where appropriate, RPA) as new opportunities are approved.Working proficiency with Ansible, Terraform, or equivalent IaC/configuration management tooling for building safe, repeatable remediation and provisioning automation. Runbook codificationTranslate SRE-authored remediation procedures into codified, executable automation steps: safe execution order, rollback steps, and exception handling.Partner with SRE to keep the runbook library and its automated counterparts in sync as remediation procedures evolve.Document each codified runbook clearly enough that another engineer, or an agent, can execute or extend it without relying on tribal knowledge. Observability & instrumentation — first-class, not an afterthoughtDesign and implement the metrics, logs, traces, and events that make every automation's behavior explainable in production — inputs, decisions, actions taken, rollbacks, outcomes.Define and maintain SLIs/SLOs for the automations you own, and wire them into dashboards and alerts that Operations and SRE actually use.Partner with the platform and alerting teams on signal correlation and alert tuning so automations trigger on clean, high-confidence signal rather than noise.Build the feedback loops that push automation outcome data into the knowledge base and confidence model driving what gets automated next. SRE pattern implementationApply SRE patterns in what you build: error-budget awareness, blameless-postmortem-driven fixes, and graduation-criteria-aware rollout.Implement confidence thresholds and staged autonomy — an automation earns broader scope only as it demonstrates accuracy over volume, per agreed graduation criteria.Build in the checks and evidence trail that let an automation prove it has caused zero P1/P2 incidents before it's trusted with more scope. Cross-team partnership & autonomy graduationPartner with SRE to validate that an automation meets graduation criteria (accuracy over volume, zero P1/P2 caused) before it moves to a higher autonomy tier.Work with Operations to review outcome validation and incident feedback, closing the loop between what ran and what should run differently next time.Support ticket volume and shift flexibility as operational demand requires, in line with the team's protected build-time model.

How you will succeed

Be recognized as a dependable builder whose automations run safely, roll back cleanly, are observable end-to-end, and rarely need rework.Demonstrate measurable throughput: runbooks codified, automations shipped, instrumentation delivered, and time-to-production for each.Ship automation that graduates to higher autonomy tiers on schedule, with zero P1/P2 incidents caused along the way.Build codified runbooks and telemetry clear enough that others can extend your work — and diagnose its behavior — without you in the room.

What you should bring

Required

10+ years of hands-on automation engineering experience, building and maintaining production automation or scripted remediation in an enterprise IT operations environment.Strong, hands-on observability skills: designing SLIs/SLOs, instrumenting code and workflows with metrics, structured logs and distributed traces, and building dashboards and alerts that drive action. Practical experience with at least one major observability stack (Prometheus/Grafana/OpenTelemetry, Splunk, Datadog, Dynatrace, or equivalent).Alerting maturity: designing signal-based alerts, correlation rules, and noise-reduction strategies; comfort tuning alerts based on real incident data rather than intuition.Demonstrated experience turning documented fixes or remediation procedures into reliable, repeatable automation — not just one-off scripts.Practical scripting/programming ability (Python, PowerShell, Bash, or similar) sufficient to build, test, and maintain production-grade automation and its instrumentation.Working knowledge of SRE concepts — SLOs, error budgets, blameless postmortems, graduation/autonomy criteria — and the ability to encode them into automation safely.Comfort working within safe-execution guardrails: rollback logic, exception handling, and staged or canary rollout of new automations.Familiarity with ITSM processes — incident, change, problem, CAPA — and how automation fits into that lifecycle.Ability to work cross-functionally with SRE, Reliability, and application teams to validate that an automation is safe to promote.Comfortable operating in a regulated, audit-ready environment (life sciences or similar), where change control and evidence matter.Bachelor's degree or higher in Computer Science, Information Technology, or a closely related field.Clear written and verbal communication skills, including the ability to document automation logic, telemetry, and remediation procedures for others.

Preferred

Experience with agentic or AI-assisted remediation tooling (auto-close, confidence-thresholded actions, human-in-the-loop handoffs).Hands-on experience with ServiceNow (incident, change, CMDB) or similar ITSM platforms, including Flow Designer-style automation.Exposure to automation or RPA platforms (Automation Anywhere, IBM BAW, or similar) — useful context, but not central to this role.Familiarity with CMDB and service mapping concepts (CSDM or equivalent).Exposure to analytics-driven approaches for identifying recurring incidents and tracking an automation backlog and its ROI.Experience contributing to or consuming a shared runbook library across multiple teams.Prior experience in a shift-based or 18x5/24x5 production support environment.Familiarity with CI/CD concepts for deploying, versioning, and rolling back automation and instrumentation safely.Experience operating in highly regulated industries (life sciences, financial services, healthcare).Experience mentoring junior automation engineers or codifying tribal knowledge into reusable, shared assets.Frontend experience (React, TypeScript) for internal dashboards and developer portals.

Leadership expectations

Acts with a reliability-first mindset, thinking beyond the individual fix to how it affects the wider estate — and how it will be observed once it's live.Drives accountability, clarity, and engineering rigor in every automation shipped.Builds trust through safe, well-tested, well-instrumented automation and clear documentation.Raises the team's automation coverage, telemetry quality, and codified knowledge — not just personal output.Leads through what they build and how they write, not through org-chart authority.

Additional information

Availability to work flexible work hours is/may be required. This team supports continuous operations and may require non-standard work hours, including some work on weekends and holidays. Appropriate adjustments in benefits will be provided for employees working non-standard hours where applicable.

This is an onsite role based in Hyderabad. Candidates should be open to working different shifts when required to align with global delivery partners.

Lilly is dedicated to helping individuals with disabilities to actively engage in the workforce, ensuring equal opportunities when vying for positions. If you require accommodation to submit a resume for a position at Lilly, please complete the accommodation request form for further assistance.

Lilly does not discriminate on the basis of age, race, color, religion, gender, sexual orientation, gender identity, gender expression, national origin, protected veteran status, disability, or any other legally protected status.

Lilly is dedicated to helping individuals with disabilities to actively engage in the workforce, ensuring equal opportunities when vying for positions. If you require accommodation to submit a resume for a position at Lilly, please complete the accommodation request form (https://careers.lilly.com/us/en/workplace-accommodation) for further assistance. Please note this is for individuals to request an accommodation as part of the application process and any other correspondence will not receive a response.

Lilly does not discriminate on the basis of age, race, color, religion, gender, sexual orientation, gender identity, gender expression, national origin, protected veteran status, disability or any other legally protected status.

#WeAreLilly