Staff Engineer- ML Infra

ShareChat · Bengaluru

  • Experience8–9 yrs
  • SalaryNot disclosed
  • Work modehybrid
  • Levelexecutive
  • Posted16 Sept 2026

About ShareChat

ShareChat is hiring in Bengaluru in media advertising. This role looks for around 8+ years of experience.

Skills

  • distributed training
  • model serving
  • inference optimization
  • ML platform
  • MLOps
  • PyTorch
  • TensorFlow
  • TorchRec
  • TensorFlow Serving
  • ONNX Runtime
  • Triton
  • OpenVINO
  • distributed systems
  • Kubernetes
  • Flyte

The role

A machine learning infrastructure engineer at a social media platform builds distributed training, optimizes model serving, and operates Kubernetes orchestration for recommendation and advertising systems. The work also applies PyTorch and TensorFlow to production-scale AI infrastructure.

Full job description

About UsShareChat (Mohalla Tech Pvt Ltd) is India's largest homegrown social media company and the only local player to achieve profitability in the industry, with 200+ million Monthly Monetizable Users across all its platforms.Founded in 2015, ShareChat has social media brands such as the ShareChat App and Moj App and micro drama app QuickTV under its portfolio. QuickTV, the newest addition to ShareChat's family of apps, crossed the 10 million downloads mark within 3 months of launch and currently has 60Mn MAUs across the network viewing the vertical episodic content.Today, the ShareChat network maintains a whopping 1,000 Cr ARR and is India's leading social media platform servicing users across the country in 15 regional languages. This growth has led to a 28% YoY revenue growth in the July-Sept (2025-26) quarter and increased it by more than 60% in the Oct-Dec quarter.What does the team do?Serving recommendations to 200+ million users entails developing large-scale personalization and recommendation models that understand user needs and preferences in real-time, while also helping creators grow their audiences on our platforms.Underpinning all of this is a horizontal AI Infrastructure team that owns the systems every product surface depends on — training pipelines, inference infrastructure, and data pipelines spanning feed and ads. This is where you come in.

What you will doDistributed training at recommender scale. Our models are embedding-heavy — ID vocabularies in the 10⁷–10⁸ range, sharded across workers and churning constantly. That makes training a different problem from most deep learning: embedding placement and sharding strategy, eviction, multi-host data-parallel throughput, and holding cost per experiment down across ~35 models in daily training. We run on TensorFlow today and are moving to the PyTorch ecosystem.Inference optimization across hardware and runtimes. Serving is the largest line in our infrastructure cost. We run TensorFlow Serving and OpenVINO today, across CPU and GPU; testing further runtimes — ONNX Runtime, Triton, TensorRT — is on the roadmap. The work is measurement-driven: find where cost and latency actually go, then move them.Expand the GPU serving path. We run a GPU/CPU hybrid deployment today, but only for a subset of models. Extending it — and making it the obvious choice wherever it wins on cost or unlocks models CPU can't serve — is open work.Orchestration and observability for dozens of training pipelines across feed and ads, running on Flyte and Kubernetes.Work horizontally across feed and ads, and raise the ceiling around you through design review and pairing.

Who you are8+ years of professional software engineering experience — we care about the depth of what you've built more than the count. Fast progressors are very much welcome.Depth in at least one of: distributed training, model serving and inference optimization, or ML platform / MLOps — with working knowledge of the others. We're looking for T-shaped, not a tour of all three.You've owned or substantially rebuilt a piece of production ML infrastructure, not only operated one someone else built. For example: changed how a training job was parallelized, moved a model between serving runtimes and measured the result, built eval and release gating for a model fleet, or debugged a distributed training failure to root cause.Comfortable in framework internals — PyTorch, TensorFlow, TorchRec, TF Serving / ONNX Runtime / Triton / OpenVINO — not only their APIs.You reason about performance from measurement: profiles, utilization, where the time actually goes.Strong software engineering fundamentals — distributed systems, system design, building for scale and reliability.Comfortable as a senior individual contributor, owning complex problems end to end with minimal hand-holding.We welcome candidates from software engineering, ML engineering, and ML platform backgrounds. We care what you've built, not which track you sat on.

Where Will You Be?Bangalore (Hybrid)

What's In It For You?At ShareChat, our values — Ownership, Speed, User Empathy, Integrity, and First Principles — are at the core of our ways of working. We believe in hiring top talent and grooming future leaders by providing a flexible environment to aid growth and development.Our Elite Benefits Suite:Wealth & Security: ESOPs and comprehensive insurance coverage.Lifestyle: Hybrid working culture, Gym allowance, and Zomato coupons.Family Support: Monthly childcare allowance for women employees.And many more…