AI System Performance Engineer (NSP/NPU, AI/ML)
Qualcomm · Greater Bengaluru Area
- Experience1–5 yrs
- SalaryNot disclosed
- Work modeonsite
- Posted24 Sept 2026
About Qualcomm
Qualcomm is hiring in Greater Bengaluru Area in semiconductors electronics. This role looks for around 1+ years of experience.
Skills
- AI/ML performance
- NPU architecture
- Deep learning
- Model optimization
- Quantization
- Pruning
- Knowledge distillation
- CPU architecture
- DDR architecture
- LPDDR4
- LPDDR5
- LPDDR6
- TensorFlow Lite
- PyTorch Mobile
- ONNX Runtime
- NPU compilers
- QNN
- Vela
- OpenVINO
- Windows Performance Toolkit
- Snapdragon Profiler
- Parallel processing
- Operating systems
- Linux
- QNX
- ARM
- x86
- Linux perf
- Intel VTune
- Dhrystone
- GeekBench
- SPECInt
- CoreMark
- lat_mem
- stream
- bw_mem
- Memory allocation
- Prefetching
- Caching
The role
A system performance engineer at a semiconductor product company profiles and optimizes AI/ML performance across NPU architecture, deep learning models, and LPDDR memory systems, using performance profiling tools and Linux. The role analyzes silicon benchmarks, SoC bottlenecks, thermal behavior, and memory access patterns for next-generation chipsets.
Full job description
About the Role
Job Area: System Performance Engineer (NSP/NPU, AI/ML)
Location: Bangalore – India
Job Overview: You will be part of System Performance team that is responsible for profiling and optimizing the System Performance on Snapdragon chipsets. This role will require a strong knowledge of AI/ML performance using NPU cores. The knowledge of CPU, DDR, NoCs will be an added advantage.
Responsibilities
Drive Performance analysis on silicon using various System and Cores (i.e. NPU, AI/ML, CPU, Memory) benchmarks like Dhrystone, GeekBench, SPECInt, CNN/GenAI ML networks etc.Use of Performance tools to analyze the load patterns across IPs and identify any performance bottlenecks in system.Analyzing Perf KPIs of SoC subsystems like NPU, CPU, Memory, and corelate performance with projectionEvaluate and characterize performance at various junction temperatures and optimize running at high ambient temperatures.Analyze and optimize the System performance parameters of SoC infrastructure like NoC, LP5 DDR, etc.Collaborate with cross-functional global teams to plan and execute performance activities on chipsets as well as make recommendations for next generation chipsets.
Qualifications
Minimum Qualifications: 1 to5 years of industry experience in the following: Experience working on any ARM/x86 based platforms, mobile/automotive/Data center operating systems and/or performance profiling tools.Experience in application or driver development in Linux\QNX and ability to create/customize make files with various compiler options is a plus.Must be quick learner and should be able to adapt to new technologies.Must have excellent communication skills.
Education Requirements: Required: Bachelor's, Computer Engineering, and/or Electrical EngineeringPreferred: Master's, Computer Engineering, and/or Electrical Engineering
Required Skills
A deep understanding of NPU, CPU and DDR architecture internals like NPU: Understanding of NPU micro-architecture, including its specialized processing elements for operations like matrix multiplications and convolutions.A strong grasp of deep learning fundamentals, including neural network architectures like Convolutional Neural Networks (CNNs), transformers, GenAI etc.Understanding the distinction between the compute-intensive training of a model (often done on GPUs) and running inference (making predictions) on NPUs.Knowledge of Model optimization and compression techniques (Quantization, Pruning, Distillation etc.)CPU caches (L1, L2, L3), instruction pipelines, branch prediction, memory hierarchy (including register, cache, and main memory) and multi-core/multi-threaded processing. You need to understand how memory access, instruction dependencies, and contention for shared resources can impact performance.DDR: Understanding of JEDEC specifications LPDDR4, LPDDR5, LPDDR6, including command and timing parameters (e.g., 𝑡𝐶𝐿, 𝑡𝑅𝐶𝐷, tRP), memory organization (rows, columns, banks), and basic view of training and initialization sequences, how a memory controller works and its specific features like command queue, port arbitration, and various control schemes. Familiarity with frameworks like TensorFlow Lite, PyTorch Mobile, and ONNX Runtime for preparing models for deployment. Experience with NPU-specific compilers, such as those from Qualcomm (QNN), Arm (Vela), Intel (OpenVINO), to optimize and orchestrate AI workloads. Evaluating hardware and software to determine the best fit for specific AI tasks and measure performance metrics like TOPS/W (trillions of operations per second perwatt). Using tools like the Windows Performance Toolkit and Qualcomm's Snapdragon Profiler to analyze and optimize NPU and system-level performance A core understanding of how to use parallel processing architectures for efficient AI computations. Expertise in how operating systems manage processes, threads, memory, MMU, and interrupt handling. This knowledge is crucial to understanding software for the kernel scheduler and system-level bottlenecks. Good understanding of Benchmarks CPU (like GeekBench, SpecInt, CoreMark etc) and DDR (like lat_mem, stream, bw_mem etc.) and how they exercise the underlying CPU/GPU/DDR architecture. Experience with a variety of performance monitoring tools like Intel VTune, Linux perf, and Utilities like top, vmstat, iostat, and netstat to monitor system resources like CPU, memory, and I/O. Experience with software tools to monitor system hotspots, command bus utilization, and identify memory traffic patterns is critical. This includes validating that the traffic generated by software is as expected. Good understanding of memory allocation policies, prefetching, and caching to minimize latency and maximize bandwidth. Understanding how an application accesses memory is vital. Skills in profiling code to analyze memory access patterns and then optimizing the code for better data locality.