PitchMeAI
Alignerr

Senior Software Engineer — AI Evaluation & Benchmarks

Alignerr · United States

  • Hybrid
  • Contract
  • $150,000 / year
  • United States
Tailored resumekeyword-matched to this role.
Hiring managerwe find who's hiring.
Intro emaildrafted to reach them directly.

Job highlights

  • Design and build AI model evaluation benchmarks.
  • Develop and maintain scalable data pipelines.
  • Analyze AI-generated code for quality.
  • Work with large codebases and Python.
  • Shape the future of AI development.

About the role

About The Role

What if the code you write could determine how smart the next generation of AI truly is? We're looking for experienced Software Engineers to design and build the coding benchmarks and data pipelines used to evaluate frontier AI models — the systems that decide whether an AI can actually reason, debug, and write production-quality software.

This is high-impact, technically demanding work at the intersection of software engineering and AI research. You'll work with large codebases, multiple programming languages, and scalable infrastructure to create evaluation systems that push the boundaries of what AI can do.

This is a fully remote contract role. If you thrive in fast-paced engineering environments and want your work to directly shape the trajectory of AI — this is the role.

Organization: Alignerr
Type: Hourly Contract
Location: Remote
Contract Length: 3 Months
Commitment: Full-time availability preferred

What You'll Do

  • Design and implement coding benchmarks used to evaluate frontier AI models across real-world programming tasks
  • Build and maintain scalable data pipelines for AI evaluation workflows
  • Analyze AI-generated code for correctness, reliability, and edge-case failures
  • Create structured evaluation scenarios that rigorously test reasoning, debugging, and code quality
  • Work with large code repositories and multi-language environments
  • Collaborate on systems that improve how AI models understand and generate software
  • Provide detailed technical feedback on model performance and failure patterns
  • Contribute to the design of evaluation frameworks that set industry standards

Who You Are

  • 4+ years of professional software engineering experience — this is non-negotiable
  • Experience working at a high-growth tech company or top-tier software organization
  • Expert proficiency in Python — you write clean, performant, well-tested Python code
  • Hands-on experience with code repositories and working in large, complex codebases
  • Proven experience designing and implementing LLM coding benchmarks and data pipelines
  • Track record of working in high-performance engineering environments with large-scale products or platforms
  • Strong command of version control systems (Git) and modern development workflows
  • Bilingual or native English speaker with strong written communication skills
  • Self-directed, technically rigorous, and comfortable operating with autonomy

What Makes a Perfect Match

  • Senior or Lead-level engineering profiles with a history of technical ownership
  • Bachelor's or Master's degree in Computer Science, Machine Learning, or a related field — or equivalent professional experience
  • Proficiency in one or more additional languages: JavaScript, Go, C++, or other relevant languages
  • Experience with CI/CD pipelines and writing robust unit tests (pytest, Mocha, JUnit)
  • Background in security engineering or significant open-source contributions
  • Familiarity with AI/ML evaluation methodologies or model benchmarking

Why Join Us

  • Work on cutting-edge AI evaluation projects alongside world-class research teams
  • Fully remote — work from anywhere with a reliable internet connection
  • Your benchmarks directly influence how the most advanced AI systems in the world are measured and improved
  • Freelance autonomy with meaningful, high-stakes engineering work
  • Collaborate with a global community of elite engineers and researchers
  • Potential for contract extension and ongoing engagement as new evaluation challenges emerge

Key skills/competency

  • Senior Software Engineer AI Evaluation
  • Python
  • AI Benchmarks
  • Data Pipelines
  • LLM Evaluation
  • Software Engineering
  • Git
  • CI/CD
  • AI Research
  • Remote Work

Skills & topics

  • Software Engineer
  • AI Evaluation
  • Python
  • Data Pipelines
  • LLM Benchmarks
  • Remote Work
  • Git
  • CI/CD
  • AI Research
  • Full-time Contract

How to get hired

  • Tailor your resume: Highlight 4+ years of Python and AI evaluation experience.
  • Showcase technical skills: Emphasize Git, CI/CD, and large codebase experience.
  • Demonstrate autonomy: Provide examples of self-directed work and technical rigor.
  • Prepare for interviews: Be ready to discuss AI benchmarking and LLM evaluation.

Technical preparation

Master Python for clean, performant code.,Practice with Git and large codebases.,Build sample AI evaluation benchmarks.,Familiarize with CI/CD and testing.

Behavioral questions

Describe a complex codebase you navigated.,How do you ensure code quality autonomously?,Share an example of rigorous technical feedback.,How do you handle ambiguity in projects?

Frequently asked questions

What is the primary focus of the Senior Software Engineer AI Evaluation role at Alignerr?
The Senior Software Engineer AI Evaluation role at Alignerr focuses on designing and building coding benchmarks and data pipelines to evaluate frontier AI models, specifically assessing their ability to reason, debug, and write production-quality software.
Is this a remote position with Alignerr?
Yes, this is a fully remote contract role. Alignerr allows you to work from anywhere with a reliable internet connection.
What are the key technical skills required for the Senior Software Engineer AI Evaluation role?
Expert proficiency in Python is required, along with hands-on experience in large codebases, version control systems (Git), and modern development workflows. Experience with CI/CD pipelines and unit testing is also highly valued.
What kind of experience is considered non-negotiable for this role?
A minimum of 4 years of professional software engineering experience is non-negotiable for this position at Alignerr.
Are there any preferred qualifications for the Senior Software Engineer AI Evaluation position?
Preferred qualifications include senior or lead-level engineering experience, a degree in Computer Science or a related field, proficiency in additional languages like JavaScript or Go, and familiarity with AI/ML evaluation methodologies.
What is the contract length and commitment for this role?
This is a 3-month hourly contract role, with full-time availability being preferred by Alignerr.
How does this role contribute to the field of AI?
Your work will directly influence how advanced AI systems are measured and improved, contributing to the trajectory of AI development by setting industry standards for evaluation frameworks.
What opportunities for future engagement exist at Alignerr for this role?
There is potential for contract extension and ongoing engagement as new AI evaluation challenges emerge, offering continued opportunities for collaboration with Alignerr.