Roche - Principal Data Engineer - AWS Platform
Roche Information Solutions · Pune
- Experience8–15 yrs
- SalaryNot disclosed
- Work modeonsite
- Levelexecutive
- Posted19 May 2026
About Roche Information Solutions
Roche Information Solutions is hiring in Pune in pharma biotech. This role looks for around 8+ years of experience.
Skills
- AWS
- Python
- SQL
- Apache Spark
- PySpark
- SparkSQL
- Amazon S3
- Amazon Redshift
- AWS Glue
- Amazon Athena
- Amazon EMR
- AWS Step Functions
- Amazon MWAA
- Dimensional Modeling
- Data Partitioning
- Terraform
- CI/CD
- Data Security
- Data Governance
- HIPAA
- GDPR
The role
A data platform engineer at a pharmaceutical technology company designs AWS data architectures and builds scalable data lakes for diagnostic products, using Apache Spark and Terraform to optimize processing and infrastructure. GenAI and healthcare data standards further shape the work.
Full job description
KEY RESPONSIBILITIES :
- Define and drive the future strategy for GenAI and emerging technologies to uncover hidden patterns and drive decision-making across diagnostic products.
- Design, implement, and optimize data architectures using AWS services, ensuring seamless integration and high performance.
- Drive the design and development of efficient data processing workflows using PySpark, SparkSQL, SQL, and modern formats like Iceberg and Parquet.
- Optimize the performance, scalability, and cost efficiency of data infrastructure and AWS service consumption.
- Provide technical direction and mentorship to a team of developers, enforcing code quality standards and best practices for testing and CI/CD.
- Collaborate closely with engineering and business stakeholders to translate complex requirements into impactful data-driven solutions.
- Advocate for data security, governance, and compliance, ensuring all products are compliant with HIPPA, GDPR etc.
REQUIRED EXPERIENCE, SKILLS & QUALIFICATIONS :
- Minimum 8 to 15 years of hands-on data engineering experience, including significant leadership in technical projects.
- Expert-level proficiency in AWS services including S3, Redshift, Glue, Athena, EMR, Step Functions, and MWAA.
- Mastery of Python, SQL, and data processing frameworks such as Apache Spark.
- Proven experience with Dimensional Modeling, Data Partitioning, and Infrastructure as Code using Terraform.
- Demonstrated ability to provide technical direction, guide architectural decisions, and mentor cross-functional teams.
- Excellent skills in bridging the gap between business needs and technical implementation for both technical and non-technical audiences.
DESIRED EXPERIENCE, SKILLS & QUALIFICATIONS :
- Experience in the Healthcare Laboratory domain and familiarity with regulations like HIPAA, HL7, and FHIR is a significant plus.
- Hands-on experience implementing GenAI and Machine Learning technologies within a SaaS or Cloud application environment.
- Experience in designing and implementing large-scale Data Lakes and Data Warehouses with a focus on long-term scalability.
WHY JOIN US?
- Collaborative Culture : Engage with a diverse team of talented professionals.
- Innovative Environment : Work on cutting-edge SaaS products and define the future of diagnostic data.
- Growth Opportunities : Take on challenging projects that offer continuous learning and career development.
EDUCATION :
- Bachelors or Masters degree in Engineering.