Search
Search places

374,525 Jobs

No logo available
Sparrow Company
locationIndependence, MO, USA
PublishedPublished: 8/19/2026
No logo available
EDAG Mexico
locationRochester, MI, USA
PublishedPublished: 8/19/2026
No logo available
MY House
locationWasilla, AK 99654, USA
PublishedPublished: 8/19/2026
No logo available
Allied Universal
locationNew York, NY, USA
PublishedPublished: 8/19/2026
No logo available
Allied Universal
locationNew York, NY, USA
PublishedPublished: 8/19/2026
No logo available
Allied Universal
locationSt Paul, MN, USA
PublishedPublished: 8/19/2026
No logo available
Essential Access Health
locationLos Angeles, CA, USA
PublishedPublished: 8/19/2026
No logo available
B & G Industries Inc
locationWatertown, SD 57201, USA
PublishedPublished: 8/19/2026
No logo available
Adara Communities
locationGarland, TX, USA
PublishedPublished: 8/19/2026
No logo available
Mondo
locationNew York, NY, USA
PublishedPublished: 8/19/2026

Research Engineer - AI Evaluation & Benchmark Development

PublishedPublished: 6/14/2022
Technology

Job Description

Hiring: Research Engineer – AI Evaluation & Benchmark Development (USA | Remote)

\n

We are seeking exceptional Research Engineers, Applied Scientists, AI Researchers, and Machine Learning Researchers to join a cutting-edge AI evaluation initiative focused on developing next-generation research engineering benchmarks for frontier AI systems.

\n


\n

This opportunity is ideal for candidates with a strong research background who enjoy solving complex technical problems, designing experiments, and contributing to AI evaluation and benchmarking.

\n


\n

Target Domains

\n

We are specifically looking for candidates with expertise in one or more of the following domains:

\n

Computer Science & Software Engineering

\n

    \n
  • Python Programming
  • \n

  • Software Development
  • \n

  • Git & Development Infrastructure
  • \n

  • AI Coding Assistants / Agentic Workflows
  • \n

  • Software Quality & Debugging
  • \n

\n


\n

Machine Learning & Artificial Intelligence

\n

    \n
  • Machine Learning
  • \n

  • Deep Learning
  • \n

  • Large Language Models (LLMs)
  • \n

  • Reinforcement Learning
  • \n

  • ML Experimentation
  • \n

  • Model Training & Evaluation
  • \n

  • AI Evaluation
  • \n

\n


\n

Data Science & Quantitative Analysis

\n

    \n
  • Data Science
  • \n

  • Statistical Analysis
  • \n

  • Data Analytics
  • \n

  • Experimental Data Analysis
  • \n

  • Jupyter Notebook / Google Colab
  • \n

\n


\n

STEM Research & Experimental Methodology

\n

    \n
  • Research Engineering
  • \n

  • Scientific Computing
  • \n

  • Experimental Design
  • \n

  • Hypothesis Testing
  • \n

  • Computational Sciences
  • \n

  • Computational Mathematics
  • \n

  • Computational Physics
  • \n

  • Computational Biology
  • \n

  • Computational Social Sciences
  • \n

\n


\n

Preferred Experience

\n

    \n
  • AI Safety
  • \n

  • LLM Red Teaming
  • \n

  • Benchmark Development
  • \n

  • Test Engineering
  • \n

  • Quality Assurance
  • \n

  • AI System Evaluation
  • \n

\n


\n

Key Responsibilities

\n

    \n
  • Design and develop complex research engineering tasks that simulate real-world AI and machine learning challenges.
  • \n

  • Create problem statements and benchmark scenarios for evaluating advanced AI systems.
  • \n

  • Design, execute, and analyze machine learning experiments.
  • \n

  • Evaluate AI model performance and document experimental findings.
  • \n

  • Collaborate with research teams to improve benchmark quality and evaluation methodologies.
  • \n

\n


\n

Minimum Qualifications

\n

    \n
  • Master's or PhD in Computer Science, Artificial Intelligence, Machine Learning, Statistics, Mathematics, Physics, Data Science, Computational Sciences, or another STEM discipline.
  • \n

  • Minimum 1–2 years of experience in research or research engineering.
  • \n

  • Strong Python programming skills.
  • \n

  • Hands-on experience with Machine Learning, LLMs, experimentation, and data analysis.
  • \n

  • Proficiency with Git, IDEs, Jupyter Notebook, or Google Colab.
  • \n

  • Strong analytical thinking, scientific methodology, and problem-solving abilities.
  • \n

  • Excellent written and verbal communication skills in English.
  • \n

  • Candidates from top-tier universities are highly preferred
  • \n

\n


\n

Preferred Qualifications

\n

    \n
  • Experience working alongside Research Scientists or AI Research teams.
  • \n

  • Publications in AI, Machine Learning, Data Science, or related research areas.
  • \n

  • Experience with AI benchmarking, LLM evaluation, or model testing.
  • \n

  • Familiarity with agentic AI systems and research workflows.
  • \n

\n


\n

Work Details

\n

Location: Remote (United States)

\n

Availability: 35 hours per week

\n

Schedule: Monday to Friday

\n


\n

Working Hours: 7 hours per day between 6:00 AM and 6:00 PM PST

\n

If you have a passion for research, experimentation, and advancing AI through rigorous evaluation and benchmarking, we would be pleased to hear from you.

\n


\n

To apply, please send your updated resume to sajid.ahmed@truelancer.com or connect with me on LinkedIn for additional information.

\n


\n

#Hiring #ResearchEngineer #ResearchScientist #AppliedScientist #ArtificialIntelligence #MachineLearning #LLM #DataScience #Python #AIEvaluation #Benchmarking #ResearchJobs #RemoteJobs #USAJobs #Masters #PhD