Technical Lead - ML Platform & MLOps
Flipkart · Bengaluru
- Experience6–9 yrs
- SalaryNot disclosed
- Work modeonsite
- Posted29 Sept 2026
About Flipkart
Flipkart is hiring in Bengaluru in ecommerce retail. This role looks for around 6+ years of experience.
Skills
- Python
- Kubernetes
- Linux
- Distributed systems
- Cloud infrastructure
- CI/CD
- Apache Airflow
- MLflow
- Machine learning
- Infrastructure as Code
- Terraform
- Monitoring
- Logging
- Metrics
- Alerting
- System design
- Distributed computing
The role
A technical lead at a marketplace company builds ML platform infrastructure and distributed systems with Python, Kubernetes, and MLflow, while applying Terraform and Apache Airflow to production automation and model deployment. This person also works across cloud infrastructure, CI/CD, and ML training workflows.
Full job description
Requirements:
6+ years of strong hands-on software/platform engineering experience.
Strong programming skills in Python.
Deep hands-on experience with Kubernetes, containers and Linux.
Experience building or operating large-scale distributed systems or platform infrastructure.
Strong understanding of cloud infrastructure including compute, storage and networking.
Experience building production-grade CI/CD and automation platforms.
Experience with workflow orchestration such as Apache Airflow or equivalent systems.
Experience with ML lifecycle tooling such as MLflow or equivalent platforms.
Strong understanding of ML training and deployment workflows.
Experience with Infrastructure as Code, preferably Terraform.
Strong debugging and production troubleshooting skills.
Experience building systems with monitoring, logging, metrics and alerting.
Strong fundamentals in system design, reliability and distributed computing.
Strong Differentiators:
We would especially love to meet you if you have worked on:
Ray or other distributed computing frameworks
Distributed or multi-node ML training
GPU and multi-GPU infrastructure
Kubernetes-based ML platforms
ML training platforms used by multiple teams
Model serving and inference infrastructure
GPU scheduling and utilisation optimisation
Large-scale workflow orchestration
ML platform developer experience
Infrastructure cost and performance optimisation
Feature platforms or feature stores