Thinkgrid Labs

Thinkgrid Labs

·3 days ago

Data engineer – microsoft fabric & cloud data platforms

Apply now

Location

remote, United States

Commitment

Full Time

Level

Senior (5+ years)

Required skills

Microsoft FabricSQLData ModelingPythonApache SparkETL/ELT PipelinesData EngineeringCloud Data PlatformsDatabricksData WarehouseLakehouseCDCData QualityPipeline OrchestrationGitCI/CD

Job Description

Thinkgrid Labs designs and builds custom web, mobile, cloud, data, and AI solutions. Our team brings together software engineers, architects, data engineers, and UI/UX designers to deliver reliable systems for clients around the world.

We’re expanding our data practice and building a modern data platform for a US health insurer, with Microsoft Fabric as the primary platform. You’ll design, build, and operate reliable data pipelines—from source ingestion and transformation to analytics-ready lakehouse and warehouse models. Hands-on Microsoft Fabric experience is strongly preferred, and we welcome engineers with strong Databricks or comparable cloud data platform experience who are ready to apply their skills in Fabric.

Who are you?

  • Strong in SQL and Data Modeling: You work confidently with complex relational databases and design practical data models for integration, reporting, and analytics.
  • Experienced with Cloud Data Platforms: You’ve built production solutions on Microsoft Fabric, Databricks, or comparable platforms. Experience with Fabric Mirroring, OneLake, Lakehouse/Warehouse, and Fabric Data Factory is particularly valuable.
  • A Reliable Pipeline Builder: You design ingestion and transformation pipelines that handle incremental loads, CDC, schema changes, and recovery from failures.
  • Comfortable with Python and Spark: You use code and notebooks to process data, implement reusable workflows, and validate results.
  • Focused on Data Quality: You build checks, reconciliation, and clear data contracts into your work so downstream teams can trust the data.
  • Operationally Minded: You monitor pipeline health, troubleshoot production issues, document runbooks, and balance performance with cost.
  • Security Conscious: You handle sensitive data responsibly, including PII/PHI, and apply appropriate access controls and governance practices.

What will you be doing?

  • Build Data Pipelines: Develop and maintain ingestion and ETL/ELT pipelines across relational databases, files, APIs, and other sources, primarily using Microsoft Fabric.
  • Develop Lakehouse and Warehouse Solutions: Organize data across raw, refined, and curated layers, using OneLake and Fabric Lakehouse/Warehouse to support reliable downstream consumption.
  • Transform and Model Data: Build reusable transformations and analytics-ready models that translate source data into useful, well-documented datasets.
  • Manage Incremental Data and Change: Implement CDC, replication, or watermark-based loading as appropriate. Handle deletes, late-arriving data, backfills, and schema evolution safely, using Fabric Mirroring where suitable.
  • Orchestrate Resilient Workflows: Manage dependencies, scheduling, retries, idempotency, and failure recovery using Fabric Data Factory, notebooks, and appropriate orchestration tools.
  • Ensure Quality and Observability: Implement data validation, reconciliation, logging, metrics, lineage, and alerts to maintain data accuracy and pipeline reliability.
  • Optimize Performance and Cost: Tune queries, Spark workloads, partitioning, file sizes, and compute usage to meet delivery expectations efficiently.
  • Collaborate and Document: Work with platform architects, security teams, analysts, and business stakeholders on requirements, governance, and delivery. Maintain pipeline documentation, data definitions, SLAs, and operational runbooks.

Must-have skills

  • Strong SQL and experience working with large, complex relational databases
  • Production experience with Microsoft Fabric, Databricks, or a comparable cloud data platform
  • Python and Apache Spark for data processing, transformation, and validation
  • Experience building and operating ingestion and ETL/ELT pipelines
  • Data modeling and lakehouse or data warehouse fundamentals
  • CDC/incremental loading, schema evolution, and data quality practices
  • Pipeline orchestration, monitoring, troubleshooting, and performance tuning
  • Git-based development workflows and familiarity with CI/CD for data solutions

Benefits

  • 5-day work week
  • Health Insurance
  • 100% remote setup with flexible work culture and international exposure
  • Opportunity to work on mission-critical healthcare projects impacting providers and patients globally

Ready to join the team?

Apply now

Similar Jobs: