LLM Inference Engineer
eBay · Bengaluru
- Experience9–13 yrs
- SalaryNot disclosed
- Work modeonsite
- Levelexecutive
- Posted2 Sept 2026
About eBay
eBay is hiring in Bengaluru in ecommerce retail. This role looks for around 9+ years of experience.
Skills
- LLM inference
- Python
- Go
- Rust
- PyTorch
- vLLM
- SGLang
- TensorRT
- GPU architecture
- performance optimisation
- systems programming
- continuous batching
- KV cache management
- root-cause analysis
The role
A generative AI engineer at a global e-commerce marketplace builds and operates LLM inference services using PyTorch and vLLM, optimising GPU systems for production performance. This person applies quantisation and CUDA to improve throughput, latency, reliability, and cost across AI workloads.
Full job description
As an LLM inference engineer on our AI platform team, you'll remove the compute-scaling bottleneck for production LLMs. Your job is to make frontier-model inference fast, efficient, reliable, and observable - the last mile from GPUs to APIs that products depend on. This role sits at the intersection of HPC, GPU systems, and MLOps and requires strong intuition for how model architecture, runtimes, and hardware interact.
Responsibilities:
Own production inference: take models from handoff to production-grade serving, including release engineering, capacity planning, cost optimisation, and incident response.
Tune inference performance: reduce end-to-end latency and increase throughput across real production traffic patterns.
Optimise runtimes and servers: Scale inference across heterogeneous GPU fleets; optimise stacks such as vLLM, Triton, and related components (e. g., schedulers, KV cache, batching, and memory).
Benchmark and measure: Build benchmarking suites, metrics, and tooling to quantify latency, throughput, GPU utilisation, memory, and cost.
Reliability and observability: Improve monitoring, tracing, and alerting; participate in incident response and postmortems to harden systems.
Apply and ship new optimisations: Evaluate research and implement pragmatic inference optimisations (e. g., quantisation, paging, and kernel/runtime improvements).
Partner cross-functionally: Work with data science and product teams to translate business requirements into performance and availability SLOs.
Requirements:
Experience deploying and operating LLM inference services in production.
Strong production coding skills in Python, plus Go or Rust (systems-level implementation and debugging).
Experience with ML frameworks and runtimes: PyTorch, vLLM, SGLang (and/or TensorRT).
Knowledge of GPU architecture and performance (profiling, memory bandwidth/latency tradeoffs); CUDA/kernel programming is a strong plus.
Solid understanding of LLM inference and optimisation techniques: continuous batching, KV cache management, quantisation, speculative decoding (nice-to-have), etc.
3+ years' hands-on experience in performance optimisation and systems programming for AI/ML workloads.
Demonstrated ability to deliver measurable production improvements (e. g., 2X throughput, lower p95/p99 latency, reduced GPU cost).
Proven skill in root-cause analysis: finding bottlenecks across model, runtime, networking, and infrastructure.