AI engineering Lead
Innovaccer · Delhi
- Experience6–11 yrs
- SalaryNot disclosed
- Work modeonsite
- Levelsenior
- Posted14 Sept 2026
About Innovaccer
Innovaccer is hiring in Delhi in healthcare. This role looks for around 6+ years of experience.
Skills
- Python
- Generative AI
- Natural Language Processing
- RAG
- PyTorch
- HuggingFace Transformers
- LLMOps
- Prompt Engineering
- Machine Learning
- Deep Learning
- Docker
- REST
- gRPC
- LangGraph
- LlamaIndex
- CrewAI
- Vector Stores
- PEFT
- LoRA
- QLoRA
- Databricks
- Azure ML
- Amazon SageMaker
- FastAPI
- Django
The role
A generative AI engineer at a healthcare technology company architects production agentic and RAG systems, leads applied machine learning delivery, and optimizes PyTorch and Python applications. The role designs retrieval pipelines, manages LLMOps, and guides engineers building customer-facing AI products.
Full job description
Job Summary
Your role sits exactly there. You will own a defined slice of the Applied AI roadmap - architecting the agentic and retrieval systems behind it, owning their lifecycle in production, and growing a small team of engineers who build alongside you. It is deliberately a hands-on leadership role: you keep writing and reviewing code while owning delivery, quality, and the growth of the people on your team.
A Day in the Life
You own AI systems end to end - from problem framing with product and customer teams through architecture, evaluation, deployment, and the cost and latency profile they run at.
System ownership: Architecture. Design production agentic and RAG systems that meet real customer scale and reliability requirements, not benchmark conditions.
Orchestration depth: Apply agent design patterns with judgment - memory, tool routing, multi-agent coordination, and explicit failure handling - and decide deliberately where autonomy stops and a human takes over.
Retrieval quality: Own the retrieval pipeline as a first-class system: chunking strategy, embedding model selection, vector stores, re-ranking, and relevance tuning against measured outcomes.
LLMOps lifecycle: Own model and prompt versioning, eval pipelines running in CI, observability (tracing, token and cost dashboards), and the guardrails and safety filters that ship with every release.
Model strategy: Select, fine-tune, and serve SLMs and open models - LoRA and QLoRA, quantization, inference optimization, GPU and serving trade-offs.
Engineering economics: Optimize latency, cost, and accuracy together, and make defensible build-versus-buy and model-selection calls you can explain to both engineers and executives.
Beyond the codebase
Define and execute the quarterly roadmap for your area, and translate ambiguous business problems into machine learning problems with clear solution workflows.
Work with business leaders and customers directly to understand where the workflow actually breaks, then build for that.
Partner with the data platform and applications teams so your capabilities land inside their products and workflows rather than beside them.
Set the coding and evaluation standards for your team, and lead design reviews.
Pursue published work or patents where the problem warrants it - particularly in healthcare AI.
What You Need
6+ years in data science, applied ML, or AI engineering, including 2+ years building LLM-powered products. Healthcare experience is a plus.
Deep NLP and GenAI experience. Statistical and classical machine learning is good to have on top of that.
Strong hands-on Python - building highly scalable, performant enterprise applications, plus optimization technique.
Hands-on experience with deep learning frameworks: PyTorch and/or HuggingFace transformers.
At least one shipped GenAI product with a genuinely complex architecture - multiple agents, memory, retrieval, and agent OTEL/tracing in production.
Working command of modern fine-tuning: PEFT methods, with LoRA and QLoRA preferred.
Hands-on experience with at least one ML platform - Databricks, Azure ML, or SageMaker.
Experience leading engineers, formally or as a technical lead - you have owned other people's output, not only your own.
Strong written and spoken communication, with a customer-focused instinct in both conversation and documentation.
Preferably a Master's in Computer Science, Computer Engineering, or a related field.
The engineering baseline we assume
Everything above sits on top of independent delivery. This role assumes you can already do the following without supervision:
Build production-grade RAG and LLM/SLM-powered features end to end with limited supervision.
Work fluently in at least one orchestration framework - LangGraph, LlamaIndex, CrewAI, or equivalent - to compose multi-step, tool-using flows.
Design retrieval pipelines and tune them for relevance: chunking, embeddings, vector stores, re-ranking.
Implement prompt engineering, function and tool calling, and reliable structured-output parsing.
Write and run evals - golden sets, LLM-as-judge - to measure quality and catch regressions before customers do.
Containerize and deploy services (Docker, REST/gRPC) with an eye on latency, token cost, and basic guardrails.
Document well and participate actively in code review.
Good to have
API frameworks for robust web applications - FastAPI or Django preferred.
Comfort with at least one hyperscaler cloud.
Papers or patents, especially in healthcare AI.
Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.