2

24-Mag

·2 days ago

Remote | senior software engineer – llm evaluation & repository validation

Apply now

Location

remote, New York, NY, United States

Commitment

Part Time

Level

Senior (5+ years)

Required skills

OSSSaaSFrameworksCloud ComputingSoftwareContainer OrchestrationStat Tools & LanguagesContainer ManagementDevOpsProgramming LanguagesEnterprise SoftwareVersion Control

Job Description

We are sharing a specialised part-time consulting opportunity for experienced Senior Software Engineers with tech-lead-level expertise to contribute to advanced LLM evaluation, repository validation, and realistic software-engineering benchmark development. Selected professionals will work with high-quality public repositories to develop and validate verifiable software-engineering tasks, configure reproducible development environments, analyse open-source issues, assess test quality, and evaluate how advanced language models perform on realistic bug-fixing and code-modification scenarios.

Key Responsibilities

  • Repository Analysis & Issue Triage
    • Analyse issues across well-maintained public software repositories
    • Identify technically meaningful problems suitable for LLM evaluation
    • Triage issues by complexity, reproducibility, and engineering relevance
    • Navigate repository history to understand implementation context
    • Prioritise problems that require substantive software-engineering reasoning
  • Repository Setup & Environment Validation
    • Configure complex repositories for reliable local execution
    • Build reproducible development and testing environments
    • Containerise projects using Docker where appropriate
    • Resolve dependencies and environment-specific configuration issues
    • Validate that repositories can be consistently executed and tested
  • Code Modification & LLM Evaluation
    • Modify and run real-world codebases locally
    • Evaluate model performance on bug-fixing and implementation tasks
    • Assess whether generated solutions correctly address repository issues
    • Identify failure patterns across different programming languages and task types
    • Compare generated solutions against professional engineering expectations
  • Testing & Verification Quality
    • Evaluate unit-test coverage and overall test quality
    • Determine whether existing tests adequately validate intended behaviour
    • Identify missing edge cases and weak verification mechanisms
    • Develop or refine tests where stronger validation is required
    • Ensure evaluation tasks have objective and reproducible outcomes
  • Research Collaboration & Technical Leadership
    • Collaborate with research teams on repository and task selection
    • Identify software-engineering problems that remain challenging for LLMs
    • Contribute insight into benchmark difficulty and dataset coverage
    • Support expansion across programming languages and task complexity
    • Provide technical leadership or guidance to junior engineers where required

Ideal Profile

  • Senior or tech-lead-level software-engineering experience
  • Strong expertise in at least one of Python, JavaScript, Java, Go, Rust, C, C++, C#, or Ruby
  • Comfortable working across unfamiliar and complex codebases
  • Strong proficiency with Git and repository-based development workflows
  • Practical Docker and environment-setup experience
  • Ability to run, modify, debug, and test production-quality software locally
  • Strong understanding of software testing and unit-test design
  • Ability to evaluate code quality, test coverage, and implementation correctness
  • Experience working with high-quality public repositories
  • Familiarity with widely used repositories with 500+ stars
  • Experience contributing to or evaluating open-source software is advantageous
  • Strong debugging and issue-triage capabilities
  • Ability to identify technically challenging software-engineering tasks
  • Clear written communication and technical reasoning skills
  • Comfortable collaborating remotely with research and engineering teams
  • Previous LLM research or evaluation experience is advantageous but not required
  • Experience with developer tools or software-automation agents is also beneficial

Engagement Details

  • Part-time independent contractor engagement
  • Fully remote
  • Commitment: approximately 20 hours per week
  • Some overlap with PST working hours is required
  • Initial contract duration is approximately 3 months
  • Expected project start is the following week
  • No medical or paid-leave benefits are included under the contractor arrangement
  • Work will involve public repository analysis, issue triage, environment setup, testing, and LLM evaluation
  • Assignments may span Python, JavaScript, Java, Go, Rust, C/C++, C#, Ruby, and related technologies
  • Project work may include development-environment automation and repository validation
  • Contributors may have opportunities to lead or support junior engineers
  • Work must be completed without using confidential, proprietary, unreleased, employer-restricted, client-restricted, or otherwise protected code, datasets, architecture materials, or technical information belonging to any employer, client, institution, or other third party

About the Platform

This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams. By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: https://www.24-mag.com/privacy-policy

Ready to join the team?

Apply now

Similar Jobs: