2

24-Mag

·2 days ago

Remote | senior software engineer – llm evaluation (us/canada/weu based)

Apply now

Location

remote, United States

Commitment

Part Time

Level

Senior (5+ years)

Required skills

LibrariesJavaScript UI LibrariesStat Tools & LanguagesProgramming LanguagesOSS

Job Description

We are sharing a specialised part-time consulting opportunity for experienced software engineers to contribute to advanced large language model evaluation, coding benchmark development, and AI-assisted software-engineering research. Selected professionals will curate and evaluate code, develop verification mechanisms, assess AI-generated software across multiple programming languages, and help research teams understand how advanced models perform throughout realistic software-development workflows.

Key Responsibilities

  • Code Curation & Solution Development
    • Curate high-quality code examples for model training and benchmarking
    • Develop precise solutions to software-engineering tasks
    • Correct and improve code across multiple programming languages
    • Work with Python, JavaScript, ReactJS, C/C++, Java, Rust, and Go
    • Maintain strong standards for correctness and maintainability
  • AI-Generated Code Evaluation
    • Evaluate AI-generated code for technical correctness
    • Assess solutions for efficiency, scalability, and reliability
    • Identify implementation weaknesses and recurring error patterns
    • Review code quality against professional engineering standards
    • Provide structured rationales supporting evaluation decisions
  • Verification & Automated Assessment
    • Build agents that assess code quality
    • Design mechanisms for automatically verifying software solutions
    • Identify recurring model-generated coding errors
    • Develop reliable checks for engineering tasks
    • Support reproducible evaluation across repeated assignments
  • Software Engineering Lifecycle Evaluation
    • Evaluate model capabilities across the software-development lifecycle
    • Assess reasoning around prototyping and architecture design
    • Review API design and production implementation decisions
    • Evaluate launch, experimentation, monitoring, and maintenance scenarios
    • Identify areas where models struggle with real-world engineering workflows
  • Research & Benchmark Collaboration
    • Collaborate with research and cross-functional technical teams
    • Contribute to datasets used for training and benchmarking
    • Help define engineering evaluation strategies
    • Compare model performance against professional engineering expectations
    • Support iterative improvements to coding-focused evaluation systems

Ideal Profile

  • 3+ years of professional software-engineering experience
  • Strong full-stack development capabilities
  • Experience building scalable, production-grade software
  • Strong understanding of software architecture and system design
  • Deep knowledge of development, debugging, and code-quality assessment
  • Experience reviewing and improving complex software implementations
  • Proficiency in one or more of Python, JavaScript, Java, C++, Rust, or related languages
  • ReactJS, C, or Go experience may also be relevant to project assignments
  • Strong understanding of API design and production implementation
  • Familiarity with software monitoring and operational maintenance
  • Ability to reason across the complete software-engineering lifecycle
  • Strong analytical and problem-solving capabilities
  • Excellent written and verbal communication skills
  • Ability to provide clear, structured evaluation rationales
  • Comfortable collaborating remotely with research and technical teams

Engagement Details

  • Part-time independent contractor engagement
  • Fully remote
  • Candidates must be based in the United States, Canada, or eligible Western European (WEU) countries
  • Source examples of WEU locations include Austria, Belgium, France, and Germany
  • Minimum commitment: 10 hours per week
  • Flexible workload of up to 40 hours per week
  • Initial project duration is approximately 1 month
  • Extension may be available depending on performance and project fit
  • No medical or paid-leave benefits are included under the contractor arrangement
  • Application process takes approximately 15–30 minutes
  • Completion of an AI video interview is required
  • Compensation is not specified in the source materials
  • Work must be completed without using confidential, proprietary, unreleased, employer-restricted, client-restricted, or otherwise protected code, datasets, architecture materials, or technical information belonging to any employer, client, institution, or other third party

About the Platform

This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams. By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: https://www.24-mag.com/privacy-policy

Ready to join the team?

Apply now

Similar Jobs: