PitchMeAI
hackajob

Human Baseliner for Open-Ended ML Research Tasks (Train AI Models Part Time!)

hackajob · United States

  • Hybrid
  • Contract
  • $80,000 / year
  • United States
Tailored resumekeyword-matched to this role.
Hiring managerwe find who's hiring.
Intro emaildrafted to reach them directly.

Job highlights

  • Evaluate AI agents on ML research tasks.
  • Act as a human benchmark for AI performance.
  • Work independently with preferred tools.
  • Minimum 20 hours per week commitment.
  • Requires 3+ years ML experience.

About the role

Machine Learning Researcher Baseliner at hackajob (with Mercor)

hackajob is collaborating with Mercor to connect them with exceptional professionals for this role.

Overview

We are hiring experienced machine learning engineers and researchers to serve as human baseliners for evaluations of open-ended machine learning research tasks. These evaluations measure how well AI agents perform on realistic AI R&D problems. To interpret agent performance, we also need strong human reference points: skilled practitioners attempting the same tasks under the same time and compute constraints. As a baseliner, you will complete self-contained ML research tasks in a sandboxed environment, working independently with your preferred tools and workflow. Your performance will be used as a benchmark against which frontier-model agents are evaluated.

What You’ll Do

  • Attempt open-ended machine learning research tasks under a fixed time and compute budget (work trial)
  • Work independently in a sandboxed Linux environment with internet access
  • Use your preferred tooling, including IDEs and AI coding assistants such as Cursor, Claude Code, and ChatGPT
  • Record your full working session via screen recording
  • Complete a short pre-task and post-task questionnaire
  • Submit your final work product, screen recording, and completed questionnaires: Post this you will be hired for a longer commitment

Commitment

  • Minimum 20 hours per week if selected
  • More availability is strongly preferred

Requirements

Candidates must meet all of the following:

  • 3+ years of machine learning experience (Time spent in a PhD program counts toward this requirement; undergraduate and master’s experience does not count)
  • Attended a top-100 university or worked at FAANG or a comparable company
  • Experience with at least one major ML framework such as PyTorch, JAX, or TensorFlow
  • Deep, hands-on expertise in at least one of the focus areas below:
    • Pretraining under tight data and compute budgets
    • PPO, reward shaping, custom `gym` / `gymnasium` environments, and throughput tuning
    • Full fine-tuning, LoRA, QLoRA, DPO, RLHF, RLAIF, and distillation
    • Large-scale corpus filtering, deduplication, subsampling, and benchmark contamination avoidance
    • Architecture design under strict parameter-count or size constraints
    • Modifying pretrained architectures, including attention patterns, pooling heads, or training objectives
    • Contrastive training for embedding or retrieval models
    • Generative vision or video modeling
    • Multilingual or low-resource language experience
    • Image or video data pipelines at scale
    • Experience balancing competing model objectives such as safety and capability
    • Prior work as an ML evaluator, red-teamer, or baseliner

Required Domain Expertise

Candidates must have strong practical experience in at least one of the following:

  • Pretraining: training transformer language models from scratch
  • Reinforcement learning: training agents in custom or existing environments
  • Post-training: fine-tuning and aligning LLMs
  • Dataset curation: building and cleaning large text corpora for LLM training
  • Model architecture: designing and modifying neural network architectures

Logistics (work trial requirements)

  • One baseline attempt per contractor per task
  • Each task may only be attempted once by a given contractor
  • All work is confidential and covered by NDA
  • Compute and environment are provided; no personal GPU is required

Key skills/competency

  • Machine Learning
  • AI Research
  • Model Evaluation
  • Baselining
  • PyTorch
  • TensorFlow
  • JAX
  • Reinforcement Learning
  • LLM Fine-tuning
  • Model Architecture

Skills & topics

  • Machine Learning Engineer
  • AI Researcher
  • Baseliner
  • PyTorch
  • TensorFlow
  • JAX
  • Reinforcement Learning
  • LLM Fine-tuning
  • Model Architecture
  • ML Evaluation

How to get hired

  • Tailor your resume: Highlight 3+ years of ML experience, top university/company background, and specific framework expertise (PyTorch, JAX, TensorFlow).
  • Showcase domain expertise: Emphasize practical experience in pretraining, reinforcement learning, post-training, dataset curation, or model architecture.
  • Prepare for the trial: Understand that tasks are timed, compute-limited, and require screen recording. Practice using your preferred ML tools and AI coding assistants.
  • Highlight relevant experience: Mention any prior work as an ML evaluator, red-teamer, or baseliner to demonstrate your suitability.
  • Be ready to commit: Ensure you can meet the minimum 20-hour per week availability, with more preferred for longer commitments.

Technical preparation

Review PyTorch, JAX, or TensorFlow documentation.,Practice ML task execution under time limits.,Familiarize yourself with common ML frameworks.,Prepare for independent work in Linux.

Behavioral questions

Describe a complex ML problem you solved.,How do you work independently on tasks?,How do you manage time and compute budgets?,Explain your experience with ML evaluation.

Frequently asked questions

What is a Human Baseliner for Open-Ended ML Research Tasks at hackajob?
A Human Baseliner acts as a human benchmark to evaluate the performance of AI agents on realistic machine learning research problems. You will attempt these tasks independently using your preferred tools within a sandboxed environment. Your results set the standard against which AI models are measured.
What are the primary responsibilities of a Human Baseliner?
Your main responsibilities include attempting open-ended ML research tasks within a fixed time and compute budget, working independently in a Linux environment, using your preferred tools and AI coding assistants, recording your sessions, and submitting your work product along with questionnaires.
What qualifications are essential for this Machine Learning Researcher Baseliner role?
Essential qualifications include 3+ years of machine learning experience (PhD time counts), attendance at a top-100 university or experience at FAANG/comparable company, proficiency in PyTorch, JAX, or TensorFlow, and deep expertise in at least one of the listed ML focus areas or domain expertise.
Can PhD and Master's program experience count towards the 3+ years of ML experience requirement?
Time spent in a PhD program counts towards the 3+ years of machine learning experience requirement. However, undergraduate and master's program experience does not count towards this specific requirement.
What is the expected time commitment for a Human Baseliner?
The minimum commitment for a selected Human Baseliner is 20 hours per week. However, more availability is strongly preferred, as this role can lead to longer commitments.
Will I need my own GPU to perform these ML research tasks?
No, you will not need your own GPU. The compute and environment are provided for the tasks. You will work within a sandboxed Linux environment with internet access.
How is my performance as a Human Baseliner evaluated?
Your performance will be used as a benchmark against which frontier-model AI agents are evaluated. You will attempt the same tasks under the same time and compute constraints as the AI agents.
What kind of ML focus areas or domain expertise is required for this role?
You need deep, hands-on expertise in at least one of the focus areas such as pretraining, reinforcement learning, fine-tuning, dataset curation, or model architecture design. Specific skills like PPO, RLHF, LoRA, and corpus filtering are also highly valued.
Is the work confidential and are there any NDAs involved for the Human Baseliner role?
Yes, all work performed as a Human Baseliner is confidential and covered by a Non-Disclosure Agreement (NDA).
How can I apply for this Machine Learning Researcher Baseliner position at hackajob?
To apply, you'll typically need to submit your resume through the hackajob platform, highlighting your relevant ML experience and qualifications. Ensure your application details align with the requirements, especially regarding experience, education, and specific ML expertise. You may also be asked to complete an initial assessment or trial task.