Sr. Product Software Engineering Engineer
The Walt Disney Company · Bengaluru
- Experience10–14 yrs
- SalaryNot disclosed
- Work modeonsite
- Levelexecutive
- Posted2 Sept 2026
About The Walt Disney Company
The Walt Disney Company is hiring in Bengaluru in media advertising. This role looks for around 10+ years of experience.
Skills
- AWS
- software development lifecycle
- incident response
- site reliability engineering
- Datadog
- Amazon CloudWatch
- Grafana
- CI/CD
- infrastructure as code
- automated testing
- Docker
- Kubernetes
- Python
- Go
- SLIs
- SLOs
- error budgets
The role
A software engineering manager at a media and entertainment company leads reliability and developer productivity platforms, directing AWS operations, incident response, and observability with Datadog and Grafana. The role sets engineering strategy, guides cross-functional teams, and drives resilient distributed services through SRE practices and infrastructure as code.
Full job description
As Senior Manager, Product Software Engineering, you will lead a large cross-functional organisation that keeps Disney Entertainment and ESPN's production platform reliable and builds the internal platforms our engineers depend on. Your team is made up of three complementary groups. A technical operations group is responsible for alert response, observability and defining alerts and thresholds, first-line eyes-on-glass support, backend capacity management and load testing, and third-layer support for technically complex customer issues. A site reliability engineering (SRE) group engineers reliability and platform resiliency into our production services and automates operational processes. A software engineering group builds and owns our internal reliability and developer productivity platforms. You will help set strategy and roadmaps across all three groups as well as ensure execution of responsibilities, own their operational and delivery outcomes, and balance day-to-day stakeholder management and people leadership. Having come up through operations and software engineering, you bring years of expertise across AWS, the software development lifecycle, incident response and management, and observability tooling such as Datadog, CloudWatch, and Grafana, but you now spend most of your time leading people: hiring, coaching, and developing engineers and leaders and partnering with product, engineering, and senior leadership as the senior point of accountability for reliability and incident outcomes.
Responsibilities:
Lead, hire, coach, schedule and retain the cross-functional team across the technical operations, SRE, and software engineering groups; own headcount planning, performance management, and career development; develop
emerging leaders and building an inclusive, blameless culture of operational and engineering excellence.
Own the technical incident response program eyes-on-glass, on-call, incident command, and blameless post-incident reviews and drive down the time to detect and resolve customer-impacting issues
Set strategy and the technical roadmap across all three groups' technical operations, SRE, and the software engineering team building our internal reliability and developer productivity platforms, balancing ongoing reliability and support work with the engineering investments that make it sustainable.
Partner with product, engineering, QA, and senior leadership to align reliability priorities, communicate operational health and risk, and represent the team in cross-functional initiatives and executive reviews
Mature the observability practice across Datadog, CloudWatch, and Grafana, directing the team to define meaningful alerts and thresholds and deliver the dashboards and metrics that speed detection and resolution
Own operational and delivery outcomes across all three groups; accountable for the reliability, availability, and performance of applications deployed to AWS and for the internal reliability, automation, and developer productivity platforms the organisation depends on.
Requirements:
10+ years in site reliability engineering, technical operations, and/or software engineering, including 5+ years managing engineering teams (hiring, performance management, and career development), with experience managing across multiple teams and managing managers or team leads.
Experience leading cross-functional or multi-team organisations spanning operations/reliability and software development, including setting strategy and owning large, complex delivery and operational outcomes across multiple teams.
Years of hands-on expertise deploying and supporting highly available, distributed applications on AWS.
Strong background across the software development lifecycle, from design and code review through build, deployment, and operations.
Proven experience owning incident response and management at scale: on-call, incident command, escalation, and blameless post-incident reviews with accountability for reliability and availability outcomes (SLIs, SLOs, and error budgets).
Experience establishing observability standards across Datadog, CloudWatch, and Grafana, including defining alerts and thresholds and driving SLO-based alerting that reduces the time to detect and resolve issues.
Familiarity with modern delivery practices, CI/CD, infrastructure as code, automated testing, and safe deployment strategies such as canary, blue/green, and feature flags.
Excellent communication and influencing skills, with the ability to align stakeholders and convey technical and operational tradeoffs at senior levels across engineering, product, and QA.
Bachelor's degree or equivalent training or work experience.
Preferred qualifications:
Experience managing operations/SRE functions alongside a software development function that supports critical systems and builds internal reliability and developer productivity platforms, organised as distinct but aligned groups.
Experience with backend capacity management, load testing, and performance tuning for high-traffic services.
Experience with containers and orchestration (Docker, Kubernetes) and supporting cloud-native distributed systems in production.
Familiarity with infrastructure-as-code tooling such as Terraform, CloudFormation, or CDK.
Building or owning internal reliability, automation, or developer productivity platforms used organisation-wide.
Programming proficiency (e. g., Python, Go, or similar) sufficient to review designs and code and mentor engineers.
Experience operating globally distributed, consumer-facing platforms with stringent availability requirements.
Familiarity with GenAI/LLM tooling and how it can accelerate reliability tooling and operational efficiency.