Data Engineer I (R-19886)

Dun & Bradstreet · Hyderabad

  • Experience1–2 yrs
  • SalaryNot disclosed
  • Work modehybrid
  • Leveljunior
  • Posted17 Sept 2026

About Dun & Bradstreet

Dun & Bradstreet is hiring in Hyderabad in financial services. This role looks for around 1+ years of experience.

Skills

  • SQL
  • Python
  • Playwright
  • Selenium
  • ETL/ELT
  • web scraping
  • API integrations
  • Google Cloud Storage
  • Amazon S3
  • XML
  • JSON
  • PDF
  • Power BI
  • Tableau
  • Microsoft Office Suite
  • BigQuery
  • AWS
  • Google Cloud Platform
  • LangChain
  • large language models
  • prompt engineering
  • retrieval-augmented generation
  • embeddings
  • vector databases
  • AI agents
  • graph databases
  • data operations
  • data management
  • ServiceNow
  • Jira
  • NoSQL
  • SQL Server
  • R programming
  • web technologies
  • data mapping

The role

A data engineer at a business decisioning data and analytics company builds data tools and ETL/ELT pipelines for research workflows using Python, SQL, and web scraping. The role also implements AI-enabled source evaluation with LangChain and supports reliable data operations.

Full job description

Shape the Future with Dun & Bradstreet

At Dun & Bradstreet, we believe data has the power to create a better tomorrow. As a global leader in business decisioning data and analytics, we help companies worldwide grow, manage risk, and innovate. Since 1841, businesses have trusted us to turn uncertainty into opportunity. We’re a diverse, global team that values creativity, collaboration, and bold ideas. Are you ready to make an impact and help shape what’s next? Join us! Explore opportunities at dnb.com/careers.

The Data Engineer I work as part of an agile team supporting the organization's Research and Managed Services. This is a hands-on, hybrid, technical, and operational role. The team members will actively write and maintain code to build data tools, automations, and pipelines, while also performing day-to-day operational tasks that keep Research and Managed Services running reliably. The role includes identifying and evaluating internal and external data sources to feed AI-enabled and traditional research workflows. The team member will assess source relevance, quality, coverage, accessibility, reliability, compliance considerations, and applicability to defined business and technical use cases. The role drives best-in-class standards and continuous improvement and requires an adaptable professional who is willing and able to learn and adopt new technologies as they are introduced to the organization

Key Responsibilities:

Coding & Development

Write, review, test, and maintain code, including SQL and Python, to build data tools, automations, and ingestion and transformation workflows that support Research and Managed Services

Automate manual processes and develop data tools to improve efficiency, accuracy, quality, and throughput

Develop and promote coding standards and contribute to code reviews within the agile team

Build and maintain web-scraping solutions, API integrations, and reusable data-processing components

Support scalable ETL/ELT pipelines for structured and unstructured data

Source Evaluation & AI Enablement

Identify prospective sources that can feed AI solutions and Research and Managed Services workflows

Define and apply source-evaluation criteria covering relevance, authority, freshness, completeness, coverage, consistency, accessibility, legal or licensing constraints, privacy, security, and technical compatibility

Perform source profiling, sample validation, proof-of-concept testing, and comparative assessments before recommending onboarding

Document source decisions, metadata, lineage, ownership, limitations, refresh expectations, and approved use cases

Implement and support AI-enabled workflows using LangChain or equivalent orchestration frameworks, large language models, embeddings, retrieval-augmented generation, vector databases, and prompt-engineering approaches where applicable

Monitor source and AI-workflow performance and recommend remediation, replacement, or additional sources when quality or coverage falls below requirements

Operational Tasks

Perform day-to-day operational activities supporting Research and Managed Services, including monitoring, exception handling, data maintenance, and issue resolution

Perform database administration activities, including performance tuning and implementation of best practices

Implement new data-maintenance processes and provide end-to-end process ownership

Ensure data integrity by validating, reconciling, and regularly cleaning data

Investigate and resolve production incidents, pipeline failures, data-quality issues, and operational exceptions

Follow applicable data governance, security, and operational standards

Collaboration & Continuous Learning

Evaluate and implement new technology solutions, and proactively learn and adopt new tools, platforms, and methodologies introduced by the organization

Communicate with stakeholders and conduct knowledge-exchange sessions for technical and non-technical audiences

Develop and maintain data documentation, including data dictionaries, source assessments, data-flow diagrams, data mappings, runbooks, and data lineage

Collaborate with cross-functional teams across Data & Analytics, Technology, Research Services, Managed Services, Product, and Data Governance

Additional duties as assigned.

Key Skills:

Strong SQL and Python skills, with demonstrated ability to write and maintain code as a core part of daily work

Experience with Playwright, Selenium, and other web-data collection techniques

Experience developing and supporting data-ingestion, transformation, and ETL/ELT workflows

Ability to collect and interpret data from multiple sources, including web scraping and GCS/S3, and formats including delimited files, XML, JSON, and PDF

Working knowledge of data systems and databases used to maintain data pipelines

Experience with Power BI, Tableau, or other dashboard tools

Experience managing stakeholders and project plans

Proficiency in Microsoft Office Suite

Willingness and demonstrated ability to learn new technologies as they are introduced

BigQuery experience and knowledge of AWS and/or GCP

Hands-on experience implementing AI solutions using LangChain or an equivalent orchestration framework

Exposure large language models, prompt engineering, retrieval-augmented generation, embeddings, vector databases, AI agents, or graph databases

Knowledge of Data Operations methodologies, data management approaches, ServiceNow, and/or Jira

Experience with NoSQL technologies, SQL Server administration, R programming, web technologies, and data mapping from multiple sources.

All Dun & Bradstreet job postings can be found at https://jobs.lever.co/dnb. Official communication from Dun & Bradstreet will come from an email address ending in @dnb.com.

Notice to Applicants: Please be advised that this job posting page is hosted and powered by Lever, a subsidiary of Employ Inc. Your use of this page is subject to Employ's Privacy Notice and Cookie Policy, which governs the processing of visitor data on this platform.

We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please visit https://bit.ly/3LMn4CQ.