Senior Engineer - Multimodal ML R&D (Audio, Vision & Language)

Qualcomm · Hyderabad

  • Experience3–8 yrs
  • SalaryNot disclosed
  • Work modeonsite
  • Posted23 Sept 2026

About Qualcomm

Qualcomm is hiring in Hyderabad in semiconductors electronics. This role looks for around 3+ years of experience.

Skills

  • Multimodal AI
  • Generative AI
  • Transformers
  • Diffusion models
  • Multimodal fusion
  • PyTorch
  • TensorFlow
  • On-device AI optimization
  • Computer vision
  • Natural language processing
  • Audio AI

The role

A multimodal AI engineer at a semiconductor technology company designs generative AI models for audio, vision, and language using Transformers, diffusion models, and PyTorch, then optimizes deployment for edge devices. The role bridges research and productization through on-device AI optimization and cross-functional engineering.

Full job description

The Advanced Technology RD team is seeking a highly motivated Multimodal AI Research Engineer to design, develop, and commercialize next-generation AI solutions across audio, video, and language domains.

In this role, you will work at the intersection of research and productization, building state-of-the-art multimodal models and enabling their efficient deployment on Snapdragon platforms powering mobile devices, XR systems, and IEOT.

You will collaborate closely with algorithm, systems, hardware, and software teams to bring cutting-edge innovations from concept and prototyping to scalable, real-world deployments.

Key Responsibilities

Design and develop multimodal AI models across audio, vision, and language

Build solutions using transformers, diffusion models, and multimodal fusion techniques

Optimize models for Snapdragon platforms (CPU, DSP, NPU)

Collaborate across hardware, systems, and software teams

Transition research into production-ready solutions and support commercialization

Preferred Qualifications

Experience in multimodal AI, generative AI, or vision-language systems

Strong understanding of Transformers, CNNs, diffusion models

Proficiency in PyTorch or TensorFlow

Experience in on-device AI optimization and deployment

Ability to work in cross-functional global teams

Education Experience

Masters or Ph.D. in Electronics, Computer Science, Electrical Engineering, or related field

2+ years of experience in CV, NLP, audio, or multimodal AI domains in academia/Industry

Keywords

Multimodal AI, Generative AI, Vision-Language Models, Audio AI, Edge AI, Vision Transformers, Transformers, diffusion models

Minimum Qualifications

Bachelor's degree in Engineering, Information Systems, Computer Science, or related field and 2+ years of Systems Engineering or related work experience.

Or Master's degree in Engineering, Information Systems, Computer Science, or related field and 1+ year of Systems Engineering or related work experience.

Or PhD in Engineering, Information Systems, Computer Science, or related field.