Amgen - Senior Data Engineer - PySpark/Azure Databricks

Amgen · Hyderabad

  • Experience3–9 yrs
  • SalaryNot disclosed
  • Work modeonsite
  • Levelmid
  • Posted5 Aug 2026

About Amgen

Amgen is hiring in Hyderabad in pharma biotech. This role looks for around 3+ years of experience.

Skills

  • Databricks
  • PySpark
  • Apache Spark
  • SparkSQL
  • AWS
  • Python
  • SQL
  • Scaled Agile
  • Data Fabric
  • Data Mesh
  • Workflow orchestration
  • Distributed computing

The role

A data engineer in a biotechnology and pharmaceutical company designs and optimizes enterprise data pipelines for research and development, using Apache Spark and Azure Databricks. They build batch and real-time ETL/ELT, metadata-driven data integration, governed data fabrics, and self-service analytics foundations across structured and unstructured sources. Apache Spark and Azure Databricks expertise, Python, SQL, AWS, data modeling, workflow orchestration, CI/CD, data governance, and distributed computing define their toolkit.

Full job description

About Amgen :

Amgen harnesses the best of biology and technology to fight the worlds toughest diseases, making peoples lives easier, fuller, and longer. We discover, develop, manufacture, and deliver innovative medicines to help millions of patients. Amgen helped establish the biotechnology industry more than 40 years ago and remains on the cutting edge of innovation, using technology and human genetic data to push beyond whats known today.

About The Role :

Lets do this. Lets change the world. We are looking for highly motivated expert Senior Data Engineer who can own the design & development of complex data pipelines, solutions and frameworks with detailed functional knowledge of R&D. The ideal candidate will be responsible to design, develop, and optimize data pipelines, data integration frameworks, and metadata-driven architectures that enable seamless data access and analytics. This role prefers deep expertise in big data processing, distributed computing, data modeling, and governance frameworks to support self-service analytics, AI-driven insights, and enterprise-wide data management.

Roles & Responsibilities :

- Design, develop, and maintain scalable ETL/ELT pipelines to support structured, semi-structured, and unstructured data processing across the Enterprise Data Engineering for Biotech or Pharma functional knowledge of R&D.

- Implement real-time and batch data processing solutions, integrating data from multiple sources into a unified, governed data fabric architecture.

- Optimize big data processing frameworks using Apache Spark, Hadoop, or similar distributed computing technologies to ensure high availability and cost efficiency.

- Work with metadata management and data lineage tracking tools to enable enterprise-wide data discovery and governance.

- Ensure data security, compliance, and role-based access control (RBAC) across data environments.

- Optimize query performance, indexing strategies, partitioning, and caching for large-scale data sets.

- Develop CI/CD pipelines for automated data pipeline deployments, version control, and monitoring.

- Implement data virtualization techniques to provide seamless access to data across multiple storage systems.

- Collaborate with cross-functional teams, including data architects, business analysts, and DevOps teams, to align data engineering strategies with enterprise goals.

- Stay up to date with emerging data technologies and best practices, ensuring continuous improvement of Enterprise Data Fabric architectures.

Must-Have Skills :

- Hands-on experience in data engineering technologies such as Databricks, PySpark, SparkSQL Apache Spark, AWS, Python, SQL, and Scaled Agile methodologies.

- Proficiency in workflow orchestration, performance tuning on big data processing.

- Strong understanding of AWS services

- Experience with Data Fabric, Data Mesh, or similar enterprise-wide data architectures.

- Ability to quickly learn, adapt and apply new technologies

- Strong problem-solving and analytical skills

- Excellent communication and teamwork skills

- Experience with Scaled Agile Framework (SAFe), Agile delivery practices, and DevOps practices.

Good-to-Have Skills :

- Good to have deep expertise in Biotech & Pharma industries

- Experience in writing APIs to make the data available to the consumers

- Experienced with SQL/NOSQL database, vector database for large language models

- Experienced with data modeling and performance tuning for both OLAP and OLTP databases

- Experienced with software engineering best-practices, including but not limited to version control (Git, Subversion, etc.), CI/CD (Jenkins, Maven etc.), automated unit testing, and Dev Ops

Education and Professional Certifications :

- Masters degree and 3 to 4 + years of relevant Computer Science, IT or related field experience

- OR

- Bachelors degree and 5 to 8 + years of relevant Computer Science, IT or related field experience

- AWS Certified Data Engineer preferred

- Databricks Certificate preferred

- Scaled Agile SAFe certification preferred

Soft Skills :

- Excellent analytical and troubleshooting skills.

- Strong verbal and written communication skills

- Ability to work effectively with global, virtual teams

- High degree of initiative and self-motivation.

- Ability to manage multiple priorities successfully.

- Team-oriented, with a focus on achieving team goals.

- Ability to learn quickly, be organized and detail oriented.

- Strong presentation and public speaking skills.

Note : For your candidature to be considered on this job, you need to apply necessarily on the company's redirected page of this job. Please make sure you apply on the redirected page as well.