Principal Architect - Gen AI
Flipkart · Bengaluru
- Experience10–14 yrs
- SalaryNot disclosed
- Work modeonsite
- Levelexecutive
- Posted3 Sept 2026
About Flipkart
Flipkart is hiring in Bengaluru in ecommerce retail. This role looks for around 10+ years of experience.
Skills
- Agentic LLM systems
- Multi-agent systems
- Orchestration
- Tool-calling
- RAG
- Retrieval
- LangGraph
- CrewAI
- Vector stores
- Generative AI evaluation
- Python
- Structured output
- Prompt architecture
The role
A generative AI engineer at a large e-commerce marketplace designs agentic LLM systems for patent-legal workflows, building RAG pipelines and human-in-the-loop controls with generative AI evaluation and Python. The role also applies knowledge graphs and fine-tuning to improve grounded, attorney-grade outputs.
Full job description
Responsibilities:
Design and build the master orchestrator and its specialised sub-skills: routing an input to the right patent-legal skill, defining skill contracts, and composing partial outputs into one coherent deliverable.
Build RAG pipelines grounded in the CDL and a patent/prosecution knowledge graph, so every generated assertion traces to a located source rather than being produced free-hand.
Make grounding and hallucination control structural: back-checking generated claims and arguments against source disclosures, and enforcing attorney-grade QC.
Own the evaluation strategy for output you cannot fully trust: measurement without abundant ground truth, abstention and escalation paths, and confidence thresholds that decide when to route to a human.
Design human-in-the-loop control: an autonomous flow that pauses at the right attorney checkpoints and remains fully steerable.
Work with model selection, prompt architecture, structured output, and, where warranted, fine-tuning on the patent corpus.
Requirements:
Hands-on production experience building agentic or multi-agent LLM systems: orchestration, tool-calling, RAG, and retrieval, not only prompt engineering.
Fluency with the current toolkit: frameworks such as LangGraph, CrewAI, or equivalents; vector stores; and observability/evaluation tooling.
A real grasp of evaluation for generative systems, including how to measure quality when labels are scarce and errors are asymmetric.
Python depth, and comfort owning a system from design through production.
Bonus: knowledge graphs or GraphRAG; fine-tuning; any exposure to legal, compliance, or other high-stakes document domains.