GPU AI Performance Engineer

Qualcomm · Bengaluru

  • Experience2–7 yrs
  • SalaryNot disclosed
  • Work modeonsite
  • Posted23 Sept 2026

About Qualcomm

Qualcomm is hiring in Bengaluru in semiconductors electronics. This role looks for around 2+ years of experience.

Skills

  • GPU architecture
  • CPU architecture
  • C++
  • Python
  • TensorFlow
  • PyTorch
  • Assembly
  • Verilog
  • SystemVerilog
  • AI inference runtimes
  • Neural network architectures
  • GPU compiler pipelines

The role

A GPU AI performance engineer at a semiconductor technology company analyzes GPU architecture, optimizes AI inference runtimes and GPU kernels, and improves LLM performance with C++ and Python. The role also evaluates neural network workloads and compiler integration for scalable accelerator platforms.

Full job description

Company: Qualcomm India Private Limited

Job Area: Engineering Group, Engineering Group > Systems Engineering

General Summary:

We are looking for a GPU AI Performance Engineer to help design, optimize, and accelerate AI workloads on GPU-based platforms. This role focuses on improving the performance, memory efficiency, and scalability of modern AI models especially LLMs, transformer workloads, MoE models, and computer vision pipelines by working across the stack from model architecture to compiler/runtime integration and low-level kernel optimization.

Responsibilities

Analyze and evaluate GPU architecture/microarchitecture for performance optimizations of AI workloads

Work across the AI stack, from model graphs and inference runtimes down to GPU kernels and compiler IR

Collaborate with hardware, software, and ML teams to identify performance bottlenecks and propose architectural or algorithmic optimizations

Analyze AI workload characteristics and correlate their behavior across different GPU generations

Contribute to architectural trade-off studies and influence GPU roadmap decisions with data-driven insights

Preferred Skills:

Strong understanding of CPU/GPU architecture

Experience in Python, C++, and ML frameworks (e.g., TensorFlow, PyTorch)

Skills: C/C++ Programming Language, Scripting (Python/Perl), Assembly, Verilog/SystemVerilog

Familiarity with AI inference runtimes and deployment stacks for model compilation, optimization, and execution on GPU/accelerator platforms

Good understanding of common neural network layers and operations, including what they do and how they affect model behavior and performance

Nice to have:

Experience with LLM inference engines such as llama.cpp

Exposure to Triton, TTIR, TTGIR, or GPU compiler pipelines

Knowledge of quantization formats and tradeoffs

Understanding of MoE architectures

Experience with GPU driver and compiler development

Experience with OpenCL or Cuda development

Minimum Qualifications:

Bachelor's degree in Engineering, Information Systems, Computer Science, or related field and 6+ years of Systems Engineering or related work experience.

OR

Master's degree in Engineering, Information Systems, Computer Science, or related field and 5+ years of Systems Engineering or related work experience.

OR

PhD in Engineering, Information Systems, Computer Science, or related field and 4+ years of Systems Engineering or related work experience.