
Senior Software Engineer — AI Evaluation & Benchmarks
Alignerr · United States
- Hybrid
- Contract
- $150,000 / year
- United States
Job highlights
- Design and implement AI model evaluation benchmarks.
- Build scalable data pipelines for AI workflows.
- Analyze AI-generated code for quality.
- Work with large codebases and languages.
- Shape the future of AI development.
About the role
About The Role
What if the code you write could determine how smart the next generation of AI truly is? We're looking for experienced Software Engineers to design and build the coding benchmarks and data pipelines used to evaluate frontier AI models — the systems that decide whether an AI can actually reason, debug, and write production-quality software.
This is high-impact, technically demanding work at the intersection of software engineering and AI research. You'll work with large codebases, multiple programming languages, and scalable infrastructure to create evaluation systems that push the boundaries of what AI can do.
This is a fully remote contract role. If you thrive in fast-paced engineering environments and want your work to directly shape the trajectory of AI — this is the role.
Organization Details
- Alignerr
- Type: Hourly Contract
- Location: Remote
- Contract Length: 3 Months
- Commitment: Full-time availability preferred
What You'll Do
- Design and implement coding benchmarks used to evaluate frontier AI models across real-world programming tasks
- Build and maintain scalable data pipelines for AI evaluation workflows
- Analyze AI-generated code for correctness, reliability, and edge-case failures
- Create structured evaluation scenarios that rigorously test reasoning, debugging, and code quality
- Work with large code repositories and multi-language environments
- Collaborate on systems that improve how AI models understand and generate software
- Provide detailed technical feedback on model performance and failure patterns
- Contribute to the design of evaluation frameworks that set industry standards
Who You Are
- 4+ years of professional software engineering experience — this is non-negotiable
- Experience working at a high-growth tech company or top-tier software organization
- Expert proficiency in Python — you write clean, performant, well-tested Python code
- Hands-on experience with code repositories and working in large, complex codebases
- Proven experience designing and implementing LLM coding benchmarks and data pipelines
- Track record of working in high-performance engineering environments with large-scale products or platforms
- Strong command of version control systems (Git) and modern development workflows
- Bilingual or native English speaker with strong written communication skills
- Self-directed, technically rigorous, and comfortable operating with autonomy
What Makes a Perfect Match
- Senior or Lead-level engineering profiles with a history of technical ownership
- Bachelor's or Master's degree in Computer Science, Machine Learning, or a related field — or equivalent professional experience
- Proficiency in one or more additional languages: JavaScript, Go, C++, or other relevant languages
- Experience with CI/CD pipelines and writing robust unit tests (pytest, Mocha, JUnit)
- Background in security engineering or significant open-source contributions
- Familiarity with AI/ML evaluation methodologies or model benchmarking
Why Join Us
- Work on cutting-edge AI evaluation projects alongside world-class research teams
- Fully remote — work from anywhere with a reliable internet connection
- Your benchmarks directly influence how the most advanced AI systems in the world are measured and improved
- Freelance autonomy with meaningful, high-stakes engineering work
- Collaborate with a global community of elite engineers and researchers
- Potential for contract extension and ongoing engagement as new evaluation challenges emerge
Key skills/competency
- Software Engineering
- AI Evaluation
- Coding Benchmarks
- Data Pipelines
- Python
- Large Codebases
- Version Control (Git)
- AI Research
- Scalable Infrastructure
- LLM
Skills & topics
- Senior Software Engineer
- AI Evaluation
- Benchmarks
- Data Pipelines
- Python
- LLM
- Remote
- Contract
- AI Research
- Software Development
How to get hired
- Tailor your resume: Highlight your 4+ years of software engineering, Python expertise, and experience with benchmarks.
- Showcase your experience: Emphasize work with large codebases, Git, and high-growth tech environments.
- Demonstrate AI/ML knowledge: Mention any experience with LLM benchmarks or evaluation methodologies.
- Prepare for technical interviews: Be ready to discuss Python, coding challenges, and system design.
- Highlight autonomy: Show you can operate independently and provide technical feedback.
Technical preparation
Behavioral questions
Frequently asked questions
- What specific AI models will I be evaluating as a Senior Software Engineer AI Evaluation at Alignerr?
- As a Senior Software Engineer AI Evaluation at Alignerr, you will be working with frontier AI models, focusing on those designed to reason, debug, and write production-quality software. While specific model names aren't listed, the role centers on evaluating the capabilities of advanced AI systems in coding-related tasks.
- Is the 3-month contract length for the Senior Software Engineer AI Evaluation role at Alignerr negotiable?
- The job description states a 3-month contract length for the Senior Software Engineer AI Evaluation role. While not explicitly stated as negotiable, it also mentions potential for contract extension and ongoing engagement, suggesting flexibility based on performance and project needs.
- What level of English proficiency is required for the Senior Software Engineer AI Evaluation position at Alignerr?
- The role requires bilingual or native English speaker proficiency with strong written communication skills. This is crucial for providing detailed technical feedback and collaborating effectively within a global team.
- Can I apply for the Senior Software Engineer AI Evaluation role at Alignerr if I have a degree in a field other than Computer Science or Machine Learning?
- Yes, the qualifications state that a Bachelor's or Master's degree in Computer Science, Machine Learning, or a related field is preferred, but equivalent professional experience will also be considered. Your extensive professional software engineering experience is a key requirement.
- What kind of technical feedback will I provide as a Senior Software Engineer AI Evaluation at Alignerr?
- As a Senior Software Engineer AI Evaluation, you will provide detailed technical feedback on AI model performance and failure patterns. This includes analyzing AI-generated code for correctness, reliability, and edge-case failures, and contributing to the design of evaluation frameworks.
- Does Alignerr provide opportunities for growth or extension beyond the initial 3-month contract for the Senior Software Engineer AI Evaluation role?
- Yes, the job description explicitly mentions 'Potential for contract extension and ongoing engagement as new evaluation challenges emerge.' This indicates that strong performance in the Senior Software Engineer AI Evaluation role could lead to continued opportunities.