Staff Product Development Engineer

The Walt Disney Company · Bengaluru

  • Experience8–12 yrs
  • SalaryNot disclosed
  • Work modeonsite
  • Levelexecutive
  • Posted2 Sept 2026

About The Walt Disney Company

The Walt Disney Company is hiring in Bengaluru in media advertising. This role looks for around 8+ years of experience.

Skills

  • Python
  • microservices architectures
  • distributed systems
  • service-level objectives
  • circuit breakers
  • retry semantics
  • caching
  • graceful degradation
  • cloud environments
  • AI-powered services
  • LLM provider constraints
  • concurrency modelling
  • rate limiting
  • performance profiling
  • observability

The role

A staff product development engineer at a media and entertainment company defines reliability architecture for AI workloads and designs resilient microservices with Python, SLOs, and cloud platforms. The role leads concurrency modelling, multi-provider failover, and production incident response for high-throughput model-dependent services.

Full job description

The candidate will have responsibilities across the following functions:

Technical Leadership and Architecture:

Define reliability architecture standards for AI workloads across Ad Platforms.

Establish latency SLOs and reliability targets for AI-dependent services.

Design and implement circuit breaker, retry, fallback, and graceful degradation strategies.

Lead multi-provider failover strategies across Azure, OpenAI, Bedrock, and future model providers.

Define production readiness review standards for AI workloads.

Runtime and Scalability Engineering:

Lead concurrency modelling and throughput validation for model-dependent services.

Design scalable backend services that support high-volume AI traffic.

Partner with Infrastructure Engineering on capacity planning for AI workloads.

Drive resilience testing and load validation before beta and production releases.

Operational Standards and Governance:

Establish AI production reliability best practices.

Lead post-incident analysis for AI-related outages or degradations.

Identify systemic reliability risks and drive architectural improvements.

Partner with AI Core to ensure shared AI services meet enterprise production standards.

Engineering Excellence:

Contribute to critical reliability and runtime code paths.

Mentor engineers on distributed systems and resilience patterns.

Raise the bar for operational rigour and runtime reliability across the organisation.

Requirements:

8+ years of backend software engineering experience building distributed systems at scale.

Strong proficiency in Python (or similar backend language).

Deep experience designing resilient microservices architectures.

Strong understanding of retry semantics, circuit breakers, caching, and graceful degradation.

Experience defining and enforcing service-level objectives (SLOs).

Experience working with cloud environments such as AWS, Azure, or GCP.

Proven ability to influence architectural decisions across multiple teams.

Experience operating AI-powered or model-dependent services in production.

Familiarity with LLM provider constraints (rate limits, latency variability, concurrency caps).

Preferred Qualifications:

Experience designing multi-provider failover strategies.

Background in high-throughput, low-latency service architectures.

Experience participating in or leading incident response for distributed systems.

Scalable distributed backend systems and microservices operating at high throughput.

Experience with:

Designing and enforcing service-level objectives (SLOs) and reliability standards.

Circuit breakers, retry semantics, graceful degradation, and resilience engineering patterns.

Multi-provider integration strategies (e. g., Azure, OpenAI, Bedrock, external APIs).

Concurrency modelling, rate limiting, and performance profiling.

Observability systems including metrics, logging, tracing, and incident response workflows.

Cloud-native architectures in AWS, Azure, or GCP environments.

Production incident postmortems and systemic reliability improvement initiatives.