Technical Lead (ML Platform & MLOps)
Myntra · Bengaluru
- Experience6–10 yrs
- SalaryNot disclosed
- Work modeonsite
- Posted28 Sept 2026
About Myntra
Myntra is hiring in Bengaluru in ecommerce retail. This role looks for around 6+ years of experience.
Skills
- Python
- Kubernetes
- Linux
- distributed systems
- cloud infrastructure
- CI/CD
- Apache Airflow
- MLflow
- Infrastructure as Code
- Terraform
- monitoring
- logging
- metrics
- alerting
- system design
- reliability
- distributed computing
The role
A technical lead at an e-commerce marketplace builds ML platforms and distributed infrastructure, applying machine learning, Kubernetes, and Infrastructure as Code to support model training and deployment. The role also uses Python and Terraform.
Full job description
Requirements:
6+ years of strong hands-on software/platform engineering experience.
Strong programming skills in Python.
Deep hands-on experience with Kubernetes, containers, and Linux.
Experience building or operating large-scale distributed systems or platform infrastructure.
Strong understanding of cloud infrastructure, including compute, storage, and networking.
Experience building production-grade CI/CD and automation platforms.
Experience with workflow orchestration such as Apache Airflow or equivalent systems.
Experience with ML lifecycle tooling such as MLflow or equivalent platforms.
Strong understanding of ML training and deployment workflows.
Experience with Infrastructure as Code, preferably Terraform.
Strong debugging and production troubleshooting skills.
Experience building systems with monitoring, logging, metrics, and alerting.
Strong fundamentals in system design, reliability, and distributed computing.
Strong Differentiators:
We would especially love to meet you if you have worked on Ray or other distributed computing frameworks.
Distributed or multi-node ML training.
GPU and multi-GPU infrastructure.
Kubernetes-based ML platforms.
ML training platforms used by multiple teams.
Model serving and inference infrastructure.
GPU scheduling and utilization optimization.
Large-scale workflow orchestration.
ML platform developer experience.
Infrastructure cost and performance optimization.
Feature platforms or feature stores.
Good to Have:
Databricks, SageMaker, Vertex AI, or similar ML platforms.
Model monitoring, data drift, and automated retraining.
NVIDIA Triton, Ray Serve, vLLM, or SGLang.
LLM training/inference and LLMOps.
Vector databases and retrieval infrastructure.
Model governance and lineage.
OpenTelemetry, Grafana, or similar observability ecosystems.