Autonomous Systems Engineer

eBay · Bengaluru

  • Experience5–9 yrs
  • SalaryNot disclosed
  • Work modeonsite
  • Levelsenior
  • Posted2 Sept 2026

About eBay

eBay is hiring in Bengaluru in ecommerce retail. This role looks for around 5+ years of experience.

Skills

  • Java
  • Python
  • Go
  • Shell
  • Linux
  • distributed systems architecture
  • production debugging
  • on-call operations
  • root cause analysis
  • AI-driven automation
  • agent-based automation frameworks
  • observability tools
  • CI/CD pipelines
  • platform upgrades
  • dependency lifecycle management
  • safety and validation layers

The role

An autonomous systems engineer at an e-commerce marketplace builds platform automation for distributed systems, applying AI/ML Engineer practices and Site Reliability / Platform Engineer methods to improve resilience and operability. The role designs observability tools, CI/CD Pipelines, and safety guardrails for reliable production operations.

Full job description

The role is for an Autonomous Systems Engineer (Platform) focused on improving the reliability, operability, and evolution of the internal engineering platform. It combines platform engineering, site reliability, and intelligent automation, with a strong emphasis on reducing manual work, improving observability, and enabling safe AI-driven or agent-based automation at scale.

Responsibilities:

Closed-Loop Resilience: Enable "detect, diagnose, remediate, validate" automation systems to improve system resilience and ensure customer-impacting issues are addressed proactively.

AI-Driven Automation: Build and operate platform automation and AI-powered (agent-based) workflows to minimize manual operational effort and move toward a self-managing infrastructure.

Reliability and Incident Leadership: Own the performance, availability, and reliability of the internal platform and critical services, which includes participating in on-call rotations and leading incident triage, debugging, and root cause analysis.

Safety and Guardrail Design: Design and implement validation pipelines, safety mechanisms, and guardrails specifically for automated and AI-generated changes to infrastructure and code to ensure production stability.

Requirements:

Production Operations: At least 5+ years of experience operating large-scale distributed systems or production platforms.

Programming and Scripting: Proficiency in Java, Python, Go, or Shell.

Incident Management: A proven track record in production debugging, on-call operations, and leading root cause analysis.

System Architecture: A solid understanding of distributed systems architecture and common failure modes.

Linux Expertise: Practical, working knowledge of Linux-based production environments.

AI/Automation Integration: Experience building or integrating AI-driven (agent-based) automation frameworks.

Essential Tooling and Frameworks:

Observability Tools: Hands-on experience with tools for monitoring, alerting, logging, and tracing.

CI/CD Pipelines: Expertise in building and maintaining automated release processes and deployment pipelines.

Lifecycle Management: Familiarity with managing platform upgrades and dependency lifecycles.

Safety and Governance: Skills in designing safety and validation layers for automated systems.

Nice-to-haves: SRE best practices, chaos/resilience testing, self-healing automation, governance for automated systems, and large-scale platform standardization experience.

Expectations:

Reliability and production operations expertise: This role is centered on owning the reliability, availability, and performance of the platform, so strong experience with production systems, incident response, on-call operations, debugging, and root cause analysis is essential.

Strong automation mindset: A major expectation is reducing toil through automation and building AI-powered or agent-based workflows. That makes automation-first thinking one of the most important qualities for success in the role.

Distributed systems and platform engineering knowledge: The role requires a solid understanding of distributed systems architecture, failure modes, platform upgrades, dependency management, and lifecycle operations.

Observability and operational excellence: Experience with monitoring, alerting, logging, tracing, SLIs, and SLOs is critical because the role is expected to improve visibility and guide engineering decisions using reliability data.

Cross-team collaboration and communication: Because the role partners directly with engineering teams on service design, resilience, and operability, strong communication and collaboration skills are also very important.