Technical Lead-Machine learning
Myntra · Bengaluru
- Experience6–7 yrs
- SalaryNot disclosed
- Work modeonsite
- Levelsenior
- Posted21 Sept 2026
About Myntra
Myntra is hiring in Bengaluru in ecommerce retail. This role looks for around 6+ years of experience.
Skills
- Python
- Kubernetes
- Linux
- Docker
- Distributed systems
- Cloud computing
- CI/CD
- Apache Airflow
- MLflow
- Infrastructure as Code
- Terraform
- Monitoring
- Logging
- Metrics
- Alerting
- System design
- Reliability engineering
- Distributed computing
The role
A machine learning platform engineer at an ecommerce marketplace builds Kubernetes infrastructure for distributed ML training, operates MLflow lifecycle tooling, and develops model serving platforms for production inference. This person also applies Python and Terraform to automate reliable cloud systems and improve ML developer productivity.
Full job description
Build the Platform Behind AI at MyntraMyntra is looking for a Technical Lead – ML Platform & MLOps to help build the next generation of infrastructure that powers machine learning across Myntra.
You will have the opportunity to design and build foundational ML platform capabilities used across multiple ML teams, influence platform architecture, solve large-scale infrastructure challenges, and establish engineering standards for how ML workloads are built and operated at Myntra.We are looking for a strong hands-on engineer who enjoys building platforms, debugging complex distributed systems, and simplifying infrastructure for hundreds of ML and engineering users.What You Will OwnBuild Myntra's ML PlatformArchitect, build and evolve scalable ML Platform and MLOps capabilities for large-scale production workloads.Create self-service infrastructure that allows Data Scientists and ML Engineers to train, experiment, deploy and operate models without managing underlying infrastructure.Design reusable platform abstractions, SDKs, APIs and tooling that dramatically improve ML developer productivity.Build systems for reproducibility, lineage, versioning, governance and lifecycle management across ML workflows.Drive architecture and technology choices for Myntra's ML infrastructure.Distributed Training & GPU InfrastructureBuild infrastructure for large-scale distributed training across CPU and GPU clusters.Design and operate multi-node and multi-GPU training environments.Work with distributed computing frameworks such as Ray or equivalent technologies.Build intelligent scheduling, resource isolation and autoscaling capabilities for ML workloads.Improve utilization of expensive GPU infrastructure through scheduling, workload optimization and capacity management.Design fault-tolerant training systems with checkpointing, retries and recovery mechanisms.ML Workflow & Training PlatformBuild a world-class training platform supporting the complete ML development lifecycle.Design scalable orchestration for training, feature engineering, validation and deployment workflows.Build reusable workflow components using technologies such as Airflow, Ray and Kubernetes.Improve scheduling, dependency management, execution isolation and reliability for thousands of ML workloads.Enable experimentation across different compute environments without exposing infrastructure complexity to users.MLOps & Model LifecycleBuild end-to-end ML lifecycle capabilities covering:Experimentation → Training → Validation → Model Registry → Deployment → Monitoring → RetrainingBuild experiment tracking and model management using MLflow or equivalent technologies.Enable reliable model versioning, approval, rollout and rollback.Build automated model validation and production-readiness workflows.Enable reproducible ML workflows across development, staging and production environments.Model Serving & AI InfrastructureBuild highly scalable infrastructure for real-time, batch and asynchronous model inference.Design model-serving platforms running on Kubernetes and GPU infrastructure.Optimize serving systems for latency, throughput, availability and cost.Explore and adopt technologies such as Ray Serve, NVIDIA Triton, vLLM, SGLang or equivalent platforms where appropriate.Enable production deployment of traditional ML, deep-learning and emerging AI/LLM workloads.Platform Reliability & ObservabilityTreat ML infrastructure as a production-grade distributed platform.Define and drive SLIs, SLOs, availability and reliability standards for ML platform services.Build deep observability across infrastructure, pipelines, training workloads and inference systems.Troubleshoot challenging production issues spanning Kubernetes, GPU workloads, distributed systems, networking, storage and ML pipelines.Drive root-cause analysis and systematically eliminate recurring operational issues.Design for high availability, fault tolerance and graceful recovery.Cloud, Kubernetes & InfrastructureDesign scalable compute, networking and storage infrastructure for ML workloads.Build and operate ML systems on Kubernetes and cloud platforms.Automate infrastructure using Terraform or equivalent Infrastructure-as-Code technologies.Build secure, isolated and reproducible runtime environments.Drive infrastructure efficiency through autoscaling, workload placement and cost optimization.ML Developer ExperienceA major part of this role is making complex ML infrastructure simple for users.You will:Build developer-facing platforms, SDKs, APIs and abstractions.Reduce the time required to move an ML experiment into production.Eliminate repetitive infrastructure work for Data Scientists and ML Engineers.Build standardized templates and paved roads for ML development.Improve debugging, discoverability and observability of ML workloads.Enable teams to focus on models and business problems rather than infrastructure.Technical LeadershipAt E3, we expect you to go beyond implementing individual components.You will:Own architecture and technical direction for major ML Platform initiatives.Lead complex system-design discussions and technical reviews.Convert ambiguous problems into scalable platform solutions.Drive engineering excellence across reliability, scalability, performance and maintainability.Mentor engineers and raise the technical bar of the team.Influence architecture across ML, Data, Platform, SRE and Infrastructure teams.Evaluate emerging technologies and make pragmatic build-vs-buy decisions.Take critical systems from concept through architecture, implementation and production adoption.What We Are Looking ForMust Have6+ years of strong hands-on software/platform engineering experience.Strong programming skills in Python.Deep hands-on experience with Kubernetes, containers and Linux.Experience building or operating large-scale distributed systems or platform infrastructure.Strong understanding of cloud infrastructure including compute, storage and networking.Experience building production-grade CI/CD and automation platforms.Experience with workflow orchestration such as Apache Airflow or equivalent systems.Experience with ML lifecycle tooling such as MLflow or equivalent platforms.Strong understanding of ML training and deployment workflows.Experience with Infrastructure as Code, preferably Terraform.Strong debugging and production troubleshooting skills.Experience building systems with monitoring, logging, metrics and alerting.Strong fundamentals in system design, reliability and distributed computing.Strong DifferentiatorsWe would especially love to meet you if you have worked on:Ray or other distributed computing frameworksDistributed or multi-node ML trainingGPU and multi-GPU infrastructureKubernetes-based ML platformsML training platforms used by multiple teamsModel serving and inference infrastructureGPU scheduling and utilization optimizationLarge-scale workflow orchestrationML platform developer experienceInfrastructure cost and performance optimizationFeature platforms or feature storesGood to HaveDatabricks, SageMaker, Vertex AI or similar ML platformsModel monitoring, data drift and automated retrainingNVIDIA Triton, Ray Serve, vLLM or SGLangLLM training/inference and LLMOpsVector databases and retrieval infrastructureModel governance and lineageOpenTelemetry, Grafana or similar observability ecosystems