Chubb

Chubb

·Today

Data engineer - regional

Apply now

Location

onsite, San Pedro Garza García, NLE, Mexico

Commitment

Full Time

Level

Middle (2-4 years)

Required skills

PythonSQLPySparkDatabricksETL/ELTDelta LakeData ModelingData EngineeringGitCI/CDData PipelinesPerformance OptimizationData WarehousingStreamingWorkflow Orchestration

Job Description

About the role

We're looking for a Data Engineer to build and maintain the pipelines that power analytics and data products across Chubb. You'll design ETL/ELT workflows on Databricks, turn raw source data into reliable, well-modeled datasets, and keep those pipelines fast and cost-efficient as data volumes grow.

This role suits someone who is comfortable owning a pipeline end to end — from ingestion through transformation to the tables analysts and data scientists actually query.

What you'll do

  • Design, build, and maintain batch and streaming ETL/ELT pipelines using Python, SQL, and PySpark on Databricks.
  • Model data across raw, cleansed, and curated layers (medallion architecture) with Delta Lake.
  • Ingest data from a range of sources — relational databases, APIs, files, and event streams — including incremental and change data capture patterns.
  • Tune Spark jobs and SQL queries for performance and cost: partitioning, file sizing and compaction, caching, join strategies, and shuffle reduction.
  • Build data quality checks, validation rules, and monitoring so problems are caught before downstream consumers see them.
  • Orchestrate and schedule workflows (Databricks Workflows, Airflow, or similar), with proper retry, alerting, and dependency handling.
  • Apply software engineering practices to data work: version control, code review, testing, and CI/CD for pipeline deployments.
  • Partner with analysts, data scientists, and business stakeholders to translate requirements into usable data models.
  • Document pipelines, data lineage, and design decisions.

Required qualifications

  • [3]+ years of experience in a data engineering or comparable role.
  • Strong Python for data processing, automation, and pipeline development.
  • Advanced SQL: complex joins, window functions, aggregations, and query optimization.
  • Hands-on experience with Databricks and PySpark in a production environment.
  • Demonstrated experience designing and operating ETL/ELT pipelines at scale.
  • Practical knowledge of performance optimization — able to diagnose a slow or expensive job and explain what you changed and why.
  • Solid understanding of data warehousing and modeling concepts (dimensional modeling, slowly changing dimensions, normalization trade-offs).
  • Experience with Git and collaborative development workflows.

Nice to have

  • Delta Lake internals: OPTIMIZE, Z-ordering, liquid clustering, time travel, VACUUM.
  • Databricks features such as Unity Catalog, Delta Live Tables / Lakeflow Declarative Pipelines, Auto Loader, or Databricks SQL.
  • Cloud platform experience ([AWS / Azure / GCP]) and its storage and compute services.
  • Streaming experience with Structured Streaming, Kafka, or Event Hubs.
  • Infrastructure as code (Terraform) and CI/CD pipelines for data workloads.
  • dbt or similar transformation frameworks.
  • Databricks certification (Data Engineer Associate or Professional).
  • Familiarity with data governance, access control, and PII handling.

Ready to join the team?

Apply now