Machine Learning Engineer 4
Adobe · Bengaluru
- Experience7–11 yrs
- SalaryNot disclosed
- Work modeonsite
- Levelexecutive
- Posted2 Sept 2026
About Adobe
Adobe is hiring in Bengaluru in technology software. This role looks for around 7+ years of experience.
Skills
- distributed systems
- concurrency
- system design
- job scheduling
- workflow orchestration
- compute management systems
- Python
- Java
- async programming
- relational databases
- PostgreSQL
- cloud platforms
- AWS
- Kubernetes
The role
A machine learning platform engineer at a digital experiences software company designs distributed systems for AI compute infrastructure, using Python and Kubernetes to orchestrate workloads. The role builds job scheduling and resource management services, develops developer-facing platforms, and operates reliable cloud systems.
Full job description
Join a team where ambition meets brand new ideas! Adobe's Firefly Group offers an outstanding opportunity to be at the forefront of innovative technology. As leaders in digital experiences, we are determined to transform how companies interact with customers across every screen. Our mission is to empower engineers and researchers by providing world-class cloud solutions for ML workflows, using the latest technology and frameworks. If you are passionate about crafting flawless digital experiences and are eager to compete in a dynamic environment, this is the place for you.
The AI Platform team builds the compute infrastructure that enables AI workloads at scale, spanning job scheduling, resource management, and the control plane that ASML engineers at Adobe depend on daily. We sit at the intersection of distributed systems and machine learning infrastructure, building the foundational services that power model training across Adobe Firefly.
We are seeking a seasoned Senior ML Platform engineer to join our ambitious team in Noida/Bangalore. This role is ideal for a technical leader who can own and evolve complex scheduling and compute orchestration systems, work closely with ASML engineers as the primary consumers of the platform, and drive architectural decisions that meaningfully improve developer productivity and system reliability.
Responsibilities:
Design, build, and maintain core services of the AI compute control plane: job scheduling, cluster management, resource quota enforcement, and compute lifecycle management.
Lead the design and implementation of job scheduling, resource quota enforcement, and compute lifecycle management systems.
Own the control plane services that manage GPU/CPU workload orchestration from job submission through execution, monitoring, and teardown.
Design reliable, fault-tolerant worker services and supervisor patterns for long-running compute workloads.
Build and evolve the data layer that tracks job state, cluster state, and resource ownership across the platform.
Partner closely with ASML engineers to deeply understand their workflows and translate requirements into robust platform capabilities.
Develop and maintain Python SDKs and CLIs that ML engineers use to interact with the platform, prioritising developer experience and reliability.
Drive end-to-end ownership of features from API design and data modelling through deployment and production operations.
Establish observability standards (metrics, tracing, alerting) for scheduling and compute systems.
Lead incident response and root cause analysis for production issues in compute orchestration.
Mentor junior and mid-level engineers on system design, scheduling patterns, and platform engineering best practices.
Requirements:
B. Tech / M. Tech degree in Computer Science from a premier institute.
9+ years of proven experience in backend platform engineering, distributed systems, or infrastructure software.
Strong computer science fundamentals, particularly in distributed systems, concurrency, and system design.
Experience building or operating job scheduling, workflow orchestration, or compute management systems (e. g., Argo, Airflow, Ray, Slurm, or similar).
Proficiency in Python and/or Java, with strong async programming skills.
Experience designing and operating services backed by relational databases (PostgreSQL preferred) at scale.
Deep understanding of Cloud Platforms, with preference for AWS; familiarity with Azure or GCP is a plus.
Proven track record of working directly with internal engineering customers (ML engineers, researchers) to shape platform roadmap.
Strong problem-solving skills with the ability to own ambiguous, complex systems independently.
Experience with Kubernetes at the workload/scheduling layer (not just operations).
Good to Have:
Hands-on experience with ML training workflows, distributed training frameworks (PyTorch, TensorFlow), or GPU resource management.
Familiarity with gRPC/protobuf or event-driven architectures.
Experience building developer-facing internal platforms consumed by ML or research teams.
Prior work in an AI platform, MLOps, or compute infrastructure role.
Understanding of ML lifecycle experiment tracking, model versioning, and training pipelines.