Machine Learning Engineer II
Nykaa · Bengaluru
- Experience5–8 yrs
- SalaryNot disclosed
- Work modeonsite
- Posted30 Sept 2026
About Nykaa
Nykaa is hiring in Bengaluru in ecommerce retail. This role looks for around 5+ years of experience.
Skills
- Python
- PyTorch
- TensorFlow
- model deployment
- deep learning
- GPU computing
- distributed systems
- APIs
- feature stores
- caching
- AWS
- Docker
- CI/CD
- MLflow
- FastAPI
- Redis
- DynamoDB
- observability
The role
An AI and machine learning engineer at an e-commerce retail company builds and operates PyTorch production models and GPU computing systems for recommendations, advertising, and search, with model deployment and deep learning as core strengths. Python and distributed systems support reliable inference, model serving, and scalable platform operations.
Full job description
Senior Applied ML Engineering | Platform Co-Ownership | Production Excellence
About the roleLead production ML engineering across Recommendations, Ads, Search, and new initiatives. Own model deployment end to end; co-own Inference Engine feature development with the platform team. Partner with Scientists, who own model behavior and business outcomes, and ensure production models meet agreed latency, reliability, quality, and cost goals.
Key responsibilitiesInference Engine development and model deploymentCo-own the Inference Engine roadmap and feature delivery with the platform team: translate Data Science needs into requirements, and contribute to design, implementation, testing, documentation, integration, reliability, and adoption.Own model deployment end to end: register artifacts and metadata, version models, validate in staging, coordinate API integration, promote to production, configure defaults, run shadow tests or go-live, and manage monitoring and rollback.Set readiness gates for dependencies, feature access, request and response contracts, compute and storage, and load-test latency, throughput, errors, and cost.Build reusable real-time and batch serving patterns for feature retrieval, model composition, and pipeline orchestration.
GPU-based inference and deep learning trainingLead GPU-backed deep learning and language-model inference; tune runtimes, batching, concurrency, precision, VRAM, and CPU-to-GPU transfers for latency, throughput, quality, and cost.Build GPU training and fine-tuning workflows with efficient data loading, mixed precision, multi-GPU execution, checkpointing, and capacity planning; support small language model adaptation and distillation.Benchmark utilization, VRAM, p95/p99 latency, throughput, and H100-class versus current GPU options; inform model, capacity, and cost decisions.
Reliability and technical leadershipOwn SLOs, dashboards, alerts, and runbooks for service health, features, models, latency, errors, and cost; lead incident response and root-cause follow-up.Improve resilience and efficiency through capacity planning, autoscaling, caching, graceful degradation, safe rollbacks, and cost reviews.Partner with Scientists, platform, and product engineering; review designs and code, mentor engineers, and maintain reusable libraries, standards, and launch guidance.Evaluate practical LLM, RAG, and AI-assisted tools for experimentation or operations, with clear quality, privacy, and cost checks.
What you will bringStrong Python and software engineering skills, with experience building APIs or distributed production services.Strong hands-on expertise training, fine-tuning, and serving deep learning models on GPUs using PyTorch or TensorFlow; optimize VRAM, precision, latency, and throughput.Experience owning production model deployments, including packaging, versioning, staged validation, monitoring, and rollback.Working knowledge of feature stores, low-latency data access, caching, data contracts, and performance profiling.Familiarity with AWS, Docker, CI/CD, MLflow, FastAPI or equivalent, Redis or DynamoDB, and observability tooling.Technical leadership across teams, clear trade-off communication, operational ownership, and focus on reliability and cost.
Useful experienceShared ML platforms, real-time feature stores, vector search, shadow deployments, model orchestration, SLM fine-tuning, or RAG systems.
How success is measuredModels are promoted safely, monitored in production, and have clear operational ownership.Inference Engine features are adopted and reduce deployment friction across Data Science teams.Serving and GPU workloads meet agreed quality, latency, reliability, and cost goals.Repeatable operating practices reduce incidents and infrastructure waste.