About the role
We're looking for a Data Engineer to build and maintain the pipelines that power analytics and data products across Chubb. You'll design ETL/ELT workflows on Databricks, turn raw source data into reliable, well-modeled datasets, and keep those pipelines fast and cost-efficient as data volumes grow.
This role suits someone who is comfortable owning a pipeline end to end — from ingestion through transformation to the tables analysts and data scientists actually query.
What you'll do
- Design, build, and maintain batch and streaming ETL/ELT pipelines using Python, SQL, and PySpark on Databricks.
- Model data across raw, cleansed, and curated layers (medallion architecture) with Delta Lake.
- Ingest data from a range of sources — relational databases, APIs, files, and event streams — including incremental and change data capture patterns.
- Tune Spark jobs and SQL queries for performance and cost: partitioning, file sizing and compaction, caching, join strategies, and shuffle reduction.
- Build data quality checks, validation rules, and monitoring so problems are caught before downstream consumers see them.
- Orchestrate and schedule workflows (Databricks Workflows, Airflow, or similar), with proper retry, alerting, and dependency handling.
- Apply software engineering practices to data work: version control, code review, testing, and CI/CD for pipeline deployments.
- Partner with analysts, data scientists, and business stakeholders to translate requirements into usable data models.
- Document pipelines, data lineage, and design decisions.
Required qualifications
- [3]+ years of experience in a data engineering or comparable role.
- Strong Python for data processing, automation, and pipeline development.
- Advanced SQL: complex joins, window functions, aggregations, and query optimization.
- Hands-on experience with Databricks and PySpark in a production environment.
- Demonstrated experience designing and operating ETL/ELT pipelines at scale.
- Practical knowledge of performance optimization — able to diagnose a slow or expensive job and explain what you changed and why.
- Solid understanding of data warehousing and modeling concepts (dimensional modeling, slowly changing dimensions, normalization trade-offs).
- Experience with Git and collaborative development workflows.
Nice to have
- Delta Lake internals: OPTIMIZE, Z-ordering, liquid clustering, time travel, VACUUM.
- Databricks features such as Unity Catalog, Delta Live Tables / Lakeflow Declarative Pipelines, Auto Loader, or Databricks SQL.
- Cloud platform experience ([AWS / Azure / GCP]) and its storage and compute services.
- Streaming experience with Structured Streaming, Kafka, or Event Hubs.
- Infrastructure as code (Terraform) and CI/CD pipelines for data workloads.
- dbt or similar transformation frameworks.
- Databricks certification (Data Engineer Associate or Professional).
- Familiarity with data governance, access control, and PII handling.