Data Platform Engineer
Yulu · Hyderabad
- Experience4–5 yrs
- SalaryNot disclosed
- Work modeonsite
- Posted23 Sept 2026
About Yulu
Yulu is hiring in Hyderabad in automotive mobility. This role looks for around 4+ years of experience.
Skills
- AWS
- Amazon S3
- AWS Glue
- Amazon Athena
- Amazon EMR
- AWS IAM
- Apache Spark
- PySpark
- Apache Iceberg
- Apache Hudi
- Delta Lake
- Apache Kafka
- Python
- SQL
- Apache Airflow
- data lakehouse architecture
- analytical data modelling
The role
A data platform engineer at a shared micro-mobility company designs and operates data lakehouse architecture, Kafka streaming, and Spark processing for reliable, cost-efficient analytics infrastructure. The role also builds AWS data services, Python tooling, and production data pipelines.
Full job description
About YuluYulu is India’s leading shared micro-mobility platform, revolutionizing urban transportation through smart, sustainable, and electric-first mobility solutions. With a rapidly growing fleet of tech-enabled electric two-wheelers and a robust battery-swapping infrastructure, Yulu makes last-mile commutes not only efficient but also planet-friendly.Our IoT-driven platform and smart electric vehicles are helping cities reduce traffic congestion and carbon emissions while empowering millions with affordable and reliable transportation.Backed by industry giants like Bajaj Auto and Magna International, Yulu operates at the intersection of mobility, technology, and sustainability. Our mission is to reduce congestion, cut emissions, and transform how India moves — one ride at a time.With millions of rides completed, thousands of EVs on the road, and a rapidly expanding footprint, we’re not just building EVs — we’re building the future of urban mobility in India.🔗 Learn more: www.yulu.bike
Role SummaryWe are looking for a Data Platform Engineer to design, build, and operate scalable and reliable data platform infrastructure at Yulu. The role will involve owning the architecture of our data lake and lakehouse, building robust data ingestion and streaming systems, and ensuring high performance, reliability, and cost efficiency across the platform.The role requires strong hands-on experience with AWS, Spark, Kafka, Python, SQL, and modern data lake technologies. You will work on large-scale data infrastructure, platform migrations, distributed data processing, production reliability, and self-service tooling that enables data engineers and analysts to work efficiently.
Key Responsibilities1. Data Platform & Lakehouse ArchitectureDesign and evolve the architecture of the data lake, including storage layout, table formats, partitioning, file sizing, and compaction.Make and document build-vs-buy and tooling decisions considering cost, operational effort, and migration risk.Own the lakehouse table format strategy across Apache Iceberg, Hudi, or Delta Lake.Manage table maintenance activities such as compaction, snapshot expiry, orphan file cleanup, and schema evolution.Plan and execute large-scale migrations across hundreds of tables with minimal downstream disruption.Design appropriate rollback mechanisms and manage concurrency, write conflicts, commit failures, locking, and idempotent writes.
2. Data Ingestion & StreamingBuild and operate data ingestion from OLTP databases such as MySQL and PostgreSQL into the data lake through batch and CDC pipelines.Own Kafka infrastructure and usage patterns, including topic design, partitioning, schema registry, consumer groups, retention, and replay.Solve streaming data challenges such as exactly-once/effectively-once delivery, out-of-order events, late-arriving data, and backfills.
3. Compute & Data ProcessingOwn Spark on EMR, including cluster configuration, autoscaling, job tuning, and cost optimisation.Identify and resolve performance issues such as data skew, shuffle, small files, spill, and memory pressure.Own the Athena and Glue layer, including workgroups, query performance, partition management, catalogue hygiene, and concurrency limits.
4. Platform Reliability & Cost OptimisationOwn the reliability of the data platform through monitoring, alerting, on-call response, and post-incident follow-through.Track platform costs by job, team, and table and identify opportunities to reduce unnecessary spend.Build cost and usage guardrails to prevent runaway queries and clusters.Build self-service tooling and internal libraries that allow data engineers and analysts to work independently.Manage infrastructure as code and CI/CD for data infrastructure.
Preferred QualificationsMust Have4+ years of experience building and operating data platforms, with hands-on ownership of production infrastructure.Strong hands-on experience with the AWS data stack, including S3, Glue, Athena, EMR, and IAM.Strong experience with Apache Spark/PySpark, including query-plan analysis and performance tuning.Production experience with at least one open table format such as Apache Iceberg, Hudi, or Delta Lake.Strong production experience with Kafka, including topic and partition design, consumer semantics, offset management, and delivery guarantees.Strong Python skills with the ability to write production-grade code.Strong SQL skills and understanding of distributed query execution, partition pruning, predicate pushdown, join strategies, and Parquet/ORC.Understanding of analytical data modelling, including dimensional modelling, slowly changing dimensions, and incremental processing.Experience with Airflow or similar workflow orchestration tools.Experience debugging production incidents in distributed systems and identifying root causes.
Good to HaveExperience with query engines such as Trino, Presto, StarRocks, ClickHouse, Druid, Pinot, or Redshift.Experience with Debezium or another CDC tool.Experience with Infrastructure as Code tools such as Terraform or CloudFormation.Experience with container orchestration platforms such as Kubernetes or ECS.Experience managing large-scale platform migrations without disrupting downstream consumers.Experience with Redis or similar technologies for caching, distributed locking, or probabilistic data structures.
What We Look ForStrong ownership of production data infrastructure.Ability to make sound platform and architectural decisions and explain the trade-offs involved.Strong focus on cost as a first-class design consideration.Ability to design for reliability, rollback, and blast-radius containment.Strong understanding of data lake/lakehouse internals and the ability to explain how different table formats handle operations such as deletes and updates.