Applied Scientist
Zillow · Bengaluru
- Experience3–7 yrs
- SalaryNot disclosed
- Work modeonsite
- Levelmid
- Posted2 Sept 2026
About Zillow
Zillow is hiring in Bengaluru in real estate construction. This role looks for around 3+ years of experience.
Skills
- SQL
- Python
- Spark
- PySpark
- Databricks
- Statistics
- Data Quality
- Data Validation
- Version Control
- Code Review
- Testing
The role
An applied scientist at a housing analytics company investigates housing metrics, builds data quality checks, and improves production data pipelines using SQL, Python, and Spark. The role also applies statistical analysis and data validation to explain anomalies, assess historical changes, and deliver reliable measurement workflows.
Full job description
Zillow is looking for an Applied Scientist to join Housing Trends Metrics and Forecasting. In this role, you will own scoped, high-impact work that improves the reliability, quality, and maintainability of Zillow's published housing metrics. You will support recurring metric publication, investigate metric anomalies and possible data outages, strengthen data quality checks, and partner closely with engineers on pipeline migrations and system improvements.
This role is well suited for someone comfortable moving between analytical investigation and production data work. You should be able to use SQL, Python, and Spark to debug issues, validate upstream changes, improve metric logic, and build durable solutions for live metric systems. You should also be able to take ambiguous measurement or data quality problems, break them into manageable pieces, and clearly explain both the technical trade-offs and business impact of your recommendations.
Responsibilities:
Support recurring publication of housing metrics by monitoring outputs, validating changes, and resolving issues before they affect downstream consumers.
Investigate metric anomalies, suspected bugs, and possible data outages by tracing issues across source data, transformation logic, and publication workflows.
Design and implement stronger data quality checks, validation workflows, and monitoring patterns for production metric pipelines.
Partner with engineers to migrate pipelines, reduce fragile upstream dependencies, and improve system reliability and maintainability.
Evaluate the impact of upstream schema, logic, or source-data changes against historical baselines and communicate revisions, caveats, and trade-offs to stakeholders.
Improve metric definitions, documentation, and operational workflows so recurring processes are more reproducible, explainable, and resilient.
Requirements:
You can independently own a well-scoped problem and drive it from investigation through validated recommendation or implementation.
You know how to translate a business or measurement question into a data plan, implement the analysis, test the output, and clearly explain the result.
You are comfortable working with large operational datasets and with the engineering realities of production metrics, including schema changes, backfills, data latency, historical reproducibility, and release validation.
You care about methodological rigor, but you also know how to deliver practical improvements to a live system.
You communicate clearly with both technical and non-technical stakeholders and work effectively in close partnership with engineers.
Master's degree with 3+ years of experience, or PhD with 1+ years of experience, in statistics, economics, data science, computer science, mathematics, engineering, operations research, or a related quantitative field.
Strong SQL skills, including complex joins, window functions, aggregations, and debugging metric logic over large relational datasets.
Strong Python skills for analytical development, validation tooling, and reproducible workflows; experience with PySpark or Spark for distributed data processing.
Experience working in Databricks or a similar large-scale data platform, including warehouse-style tables, scheduled jobs, and production data workflows.
Experience investigating data quality issues, metric anomalies, or production data bugs and turning findings into durable fixes.
Experience writing maintainable code in a shared codebase using version control, code review, and testing or validation frameworks.
Good communication skills, with the ability to explain technical trade-offs, data caveats, and metric impacts to both technical and non-technical partners.
Preferred Qualifications:
Experience building or improving automated data quality checks, validation frameworks, or monitoring for production data pipelines.
Experience managing noisy operational data and complex upstream dependencies.
Experience validating backfills, historical revisions, or upstream logic changes against established baselines.
Experience designing and improving analytical datasets or metrics end-to-end, from source understanding through transformation, validation, and stakeholder communication.
Experience partnering closely with data or software engineers on pipeline migrations or system improvements.
Interest in housing market data, operational metrics, or applied measurement problems in production systems.