GPU AI Performance Engineer
Qualcomm · Bengaluru
- Experience2–7 yrs
- SalaryNot disclosed
- Work modeonsite
- Posted23 Sept 2026
About Qualcomm
Qualcomm is hiring in Bengaluru in semiconductors electronics. This role looks for around 2+ years of experience.
Skills
- GPU architecture
- CPU architecture
- C++
- Python
- TensorFlow
- PyTorch
- Assembly
- Verilog
- SystemVerilog
- AI inference runtimes
- Neural network architectures
- GPU compiler pipelines
The role
A GPU AI performance engineer at a semiconductor technology company analyzes GPU architecture, optimizes AI inference runtimes and GPU kernels, and improves LLM performance with C++ and Python. The role also evaluates neural network workloads and compiler integration for scalable accelerator platforms.
Full job description
Company: Qualcomm India Private Limited
Job Area: Engineering Group, Engineering Group > Systems Engineering
General Summary:
We are looking for a GPU AI Performance Engineer to help design, optimize, and accelerate AI workloads on GPU-based platforms. This role focuses on improving the performance, memory efficiency, and scalability of modern AI models especially LLMs, transformer workloads, MoE models, and computer vision pipelines by working across the stack from model architecture to compiler/runtime integration and low-level kernel optimization.
Responsibilities
Analyze and evaluate GPU architecture/microarchitecture for performance optimizations of AI workloads
Work across the AI stack, from model graphs and inference runtimes down to GPU kernels and compiler IR
Collaborate with hardware, software, and ML teams to identify performance bottlenecks and propose architectural or algorithmic optimizations
Analyze AI workload characteristics and correlate their behavior across different GPU generations
Contribute to architectural trade-off studies and influence GPU roadmap decisions with data-driven insights
Preferred Skills:
Strong understanding of CPU/GPU architecture
Experience in Python, C++, and ML frameworks (e.g., TensorFlow, PyTorch)
Skills: C/C++ Programming Language, Scripting (Python/Perl), Assembly, Verilog/SystemVerilog
Familiarity with AI inference runtimes and deployment stacks for model compilation, optimization, and execution on GPU/accelerator platforms
Good understanding of common neural network layers and operations, including what they do and how they affect model behavior and performance
Nice to have:
Experience with LLM inference engines such as llama.cpp
Exposure to Triton, TTIR, TTGIR, or GPU compiler pipelines
Knowledge of quantization formats and tradeoffs
Understanding of MoE architectures
Experience with GPU driver and compiler development
Experience with OpenCL or Cuda development
Minimum Qualifications:
Bachelor's degree in Engineering, Information Systems, Computer Science, or related field and 6+ years of Systems Engineering or related work experience.
OR
Master's degree in Engineering, Information Systems, Computer Science, or related field and 5+ years of Systems Engineering or related work experience.
OR
PhD in Engineering, Information Systems, Computer Science, or related field and 4+ years of Systems Engineering or related work experience.