Data Engineer - AI Labs

IDFC FIRST Bank · Bengaluru

  • Experience3–5 yrs
  • SalaryNot disclosed
  • Work modeonsite
  • Levelmid
  • Posted3 Sept 2026

About IDFC FIRST Bank

IDFC FIRST Bank is hiring in Bengaluru in financial services. This role looks for around 3+ years of experience.

Skills

  • Apache Flink
  • Spark Streaming
  • PySpark
  • Python
  • SQL
  • AWS
  • Apache Spark
  • Java
  • Scala
  • Data Structures and Algorithms
  • ETL
  • Amazon EMR
  • Amazon S3
  • AWS Lambda
  • AWS Glue
  • Neo4j

The role

A data engineer at a banking technology company designs and operates Apache Flink, Spark Streaming, and AWS pipelines for low-latency decision-making. The role builds scalable real-time data systems with PySpark and SQL, supporting analytics and machine learning teams.

Full job description

We are looking for a hands-on Data Engineer with strong experience in big data engineering and real-time streaming applications. Hands-on expertise in Apache Flink is mandatory for this role. The ideal candidate will have built and operated production-grade Flink applications, along with Spark Streaming pipelines, and solid fundamentals in distributed data systems and cloud-native data engineering on AWS. DE will work closely with data platform, analytics, and ML teams to design scalable, low-latency data pipelines that power real-time decision-making.

Responsibilities:

Design, develop, and own real-time streaming applications using Apache Flink (mandatory, core deliverable).

Tune and troubleshoot Flink and Spark jobs in production (checkpointing, state management, backpressure, resource optimisation).

Build and maintain Spark Structured Streaming pipelines for real-time data processing.

Develop and optimise batch and streaming ETL pipelines using PySpark and Python.

Write and optimise SQL for data transformation, validation, and analytics consumption.

Build and deploy data pipelines on AWS, using EMR, S3 Lambda, and Glue.

Apply strong DSA fundamentals to deliver performant, scalable, production- grade code.

Ensure data pipeline reliability through monitoring, schema validation, and data quality checks.

Deliver pipelines that meet defined SLAs for latency, throughput, and data accuracy.

Requirements:

Minimum Number of years - 3 to 5 years.

Can manage within a complex ecosystem of technology, vendors and other stakeholders.

Proficiency in Python and Java/Scala is a plus.

Experience in working on agile projects that span multiple organisations and business units.

Proficient in Spark/Pyspark, Apache Flink, Spark Streaming, Python.

Good at SQL; dashboarding experience (Superset/Tableau/Power BI); graph database Neo4j.

Good to have: Agentic AI.