AI Engineer (LLM Inference & AI Infrastructure)
IDFC FIRST Bank · Bengaluru
- Experience5–9 yrs
- SalaryNot disclosed
- Work modeonsite
- Levelsenior
- Posted2 Sept 2026
About IDFC FIRST Bank
IDFC FIRST Bank is hiring in Bengaluru in financial services. This role looks for around 5+ years of experience.
Skills
- Python
- LLM inference
- vLLM
- Kubernetes
- Helm
- GPU observability
- CUDA
- NCCL
- PyTorch
- SGLang
- distributed systems
- CI/CD pipelines
- Infrastructure as Code
- Tensor parallelism
- Pipeline parallelism
- KV cache management
- Continuous batching
- CUDA Streams
- Tensor Cores
- Nsight Systems
- Nsight Compute
The role
An AI engineer at a banking technology company builds production-grade LLM inference infrastructure using vLLM, Kubernetes, and GPU performance engineering. Optimizes model serving, distributed systems, and CUDA-based workloads for scalable AI platforms.
Full job description
We are looking for an experienced AI Engineer to build and optimize high-performance LLM inference infrastructure for production-scale AI systems. This role focuses on inference performance, GPU optimization, Kubernetes-native infrastructure, and large-scale model serving rather than application-level GenAI development.
The ideal candidate has hands-on experience deploying and optimizing LLM inference workloads, understands GPU architecture and performance engineering, and has worked on production AI infrastructure. Candidates whose experience is limited to RAG pipelines, chatbots, LangChain, or OpenAI API integrations without systems-level expertise will not be considered.
Responsibilities:
Design, build, and optimize production-grade LLM inference infrastructure.
Deploy, configure, and operate LLM serving frameworks such as vLLM in production environments.
Optimize inference latency, throughput, and GPU utilization across large-scale workloads.
Develop Kubernetes-native infrastructure including Operators, Helm Charts, monitoring, and GPU observability.
Analyze and improve inference performance using profiling and benchmarking tools.
Design scalable serving architectures with batching, KV cache management, tensor parallelism, and pipeline parallelism.
Work closely with platform, ML, and infrastructure teams to improve inference efficiency and production reliability.
Build tooling for benchmarking, monitoring, debugging, and performance validation.
Contribute to production-ready ML infrastructure, CI/CD pipelines, and cloud-native deployments.
Requirements:
Bachelor's degree in Computer Science, Electrical Engineering, or a related field (or equivalent experience).
5-8 years of experience building performance-critical software, ML infrastructure, distributed systems, or AI platforms.
Strong programming skills in Python (Go or Rust is a plus).
Excellent understanding of algorithms, data structures, operating systems, computer architecture, parallel programming, distributed systems, and deep learning fundamentals.
Strong understanding of modern LLM inference architecture and serving concepts.
LLM Inference: Hands-on experience with vLLM deployment and optimization, TTFT (Time to First Token), ITL (Inter-Token Latency), Prefill and Decode stages, KV Cache optimization, Speculative Decoding, Production benchmarking, and reproducible performance measurements.
LLM Serving: Experience with: Continuous batching, Chunked prefill, KV cache management, Tensor parallelism, Pipeline parallelism, and Production inference optimization.
GPU and Performance Engineering: Strong understanding of CUDA fundamentals, GPU memory hierarchy, CUDA Streams, NCCL, Tensor Cores, GPU utilization optimization, Roofline modeling, Nsight Systems, and Nsight Compute.
Infrastructure and Platform: Hands-on experience with Kubernetes, Helm, GPU observability, Production monitoring, Distributed systems, Cloud platforms (AWS/GCP/Azure), CI/CD pipelines, Infrastructure as Code.
ML Frameworks: Experience optimizing PyTorch, vLLM, SGLang, Production inference engines.
Good to Have: NVIDIA Dynamo, llm-d, Triton, TorchDynamo, MLIR, LLVM, XLA, CUTLASS, CUDA Graphs, Tensor Core optimization, Building custom inference engines, Production-scale AI infrastructure.
Preferred Skills:
Strong preference will be given to candidates from companies such as NVIDIA, Microsoft, Qualcomm, and Simplismart AI (closest fit).
Other organizations building large-scale AI infrastructure, GPU systems, or LLM inference platforms.