hireVouch
hireVouch·last month

Junior data engineer

Apply now

Location

onsite, Vaughan, Canada

Commitment

Full Time

Level

Junior (<2 years)

Required skills

PaaSFrameworksInfrastructureBig Data ToolsIaaSStreaming & MessagingApplication HostingCDNSQLStorageServerlessComputeOSSDatastoresStat Tools & LanguagesOrchestration & PipelinesWorkflow ManagementData Science ToolsLibrariesApp Definition and DevelopmentHosted PlatformETLProgramming Languages

Job Description

About Us

We are dedicated to building a cleaner, more sustainable future. By applying innovative technology and data-driven insights to green environmental initiatives, we help measure, analyze, and reduce environmental impact at scale.

We are looking for a passionate, forward-thinking Junior Data Engineer to join our data team and help build the data pipelines powering our eco-focused solutions.

Position Overview

As a fresh graduate joining our team, you will work closely with senior data engineers and analysts to design, build, and maintain high-volume data pipelines. You will transform raw environmental datasets—such as energy metrics, carbon emissions data, and resource usage—into actionable insights. This role is ideal for a recent Computer Science graduate eager to apply modern data tooling (AWS, PySpark, Python) toward solving meaningful sustainability challenges.

Key Responsibilities

  • Pipeline Development: Design, build, and maintain automated batch and real-time ETL/ELT pipelines to ingest, clean, and transform large-scale environmental data.
  • Data Processing: Utilize Python and Apache Spark (PySpark) to process structured and unstructured datasets efficiently.
  • Cloud Infrastructure: Help manage and expand our cloud data infrastructure using core AWS services (e.g., S3, Glue, EMR, Redshift, Lambda).
  • Data Quality & Governance: Implement automated testing, validation, and monitoring to ensure data accuracy, reliability, and security.
  • Cross-Functional Collaboration: Partner with Data Scientists, Business Analysts, and Sustainability Specialists to deliver clean, structured data for reporting and machine learning applications.

Required Qualifications

  • Education: Bachelor’s degree in Computer Science (or a closely related core computing field, such as Computer Engineering or Software Engineering) completed within the last 0–12 months.
  • Core Programming: Strong foundation in Python and fundamental software engineering principles (OOP, data structures, algorithms, version control with Git).
  • Distributed Computing: Academic or hands-on project experience using Apache Spark / PySpark to process large datasets.
  • Cloud Fundamentals: Working knowledge or project experience with AWS core services (S3, EC2, IAM, Lambda, or managed data services).
  • Databases & SQL: Solid grasp of relational databases, SQL query writing, data modeling concepts, and basic schema design.

Nice-to-Have / Preferred Qualifications

  • Coursework, internship, or personal project focus on environmental data, sustainability, clean energy, or IoT telemetry data.
  • Exposure to workflow orchestration tools (e.g., Apache Airflow, Dagster).
  • Familiarity with containerization technologies (Docker, Kubernetes).
  • Knowledge of CI/CD practices for data infrastructure.

What We Offer

  • Mission-Driven Impact: Direct involvement in projects that combat climate change and advance sustainable practices.
  • Mentorship & Growth: A collaborative environment with dedicated mentorship from experienced senior data engineers.

Ready to join the team?

Apply now