Software Engineer, AI Code Evaluation and Benchmarking
$50 per hourSaidGig
Evaluate and benchmark the coding abilities of frontier AI models by reviewing AI-generated solutions, validating them against real-world software engineering tasks, and shaping high-quality evaluation datasets and benchmarks. This role suits experienced engineers who enjoy code review, debugging, and applying software engineering judgment to improve model correctness and reliability.
Key Responsibilities- Review AI-generated code for correctness, efficiency, maintainability, and compliance with task requirements.
- Analyze software engineering tasks and validate whether proposed solutions meet expected outcomes.
- Debug code, reproduce issues, and verify fixes across multiple programming environments.
- Evaluate model-generated explanations and reasoning for technical accuracy and soundness.
- Create, refine, and maintain evaluation datasets, benchmarks, and grading rubrics for coding tasks.
- Identify edge cases, failure modes, and areas where models struggle with software engineering problems.
- Document findings clearly and provide structured feedback to improve evaluation consistency and quality.
- Collaborate with project teams to establish and maintain quality standards and evaluation methodologies.
- Bachelor''s or Master’s degree in Computer Science, Software Engineering, or a related technical field.
- Minimum 3 years of professional software engineering experience.
- Strong proficiency in one or more of the following languages: Python, Java, C/C++, Go, Swift, Objective-C, PHP, or SQL.
- Solid understanding of data structures, algorithms, software design principles, and debugging methodologies.
- Experience performing code reviews and evaluating code quality in production or large-scale codebases.
- Ability to analyze complex technical problems and assess solution correctness with minimal supervision.
- Familiarity with version control systems such as Git and with modern software development workflows.
- Strong written communication skills and attention to detail.
- Experience with AI or ML data annotation, NLP, prompt engineering, model evaluation, or LLM-related projects is a plus.
- Experience evaluating AI-generated code, creating benchmarks, or assessing software quality is highly preferred.
- Engagement type: Contractor assignment, no medical or paid leave provided.
- Minimum commitment: at least 4 hours per day and at least 20 hours per week, with a required daily overlap of 4 hours aligned to PST.
- Contract length: 1 month, expected start date is next week.
- Work location: Remote, United States only.
- Perks: fully remote work and the opportunity to contribute to cutting-edge AI coding evaluation projects.
- Applicants must be located in the United States. This role is open to US-based candidates only.
Candidates will complete an online automated coding assessment covering Python and a Docker-based test identified as RHLF as part of the selection process.
- ...Role Overview Evaluate and benchmark the coding abilities of frontier AI models by reviewing AI-generated solutions, validating them against real-world software engineering tasks, and shaping high-quality evaluation datasets and benchmarks. This role suits experienced...SuggestedContract workFor contractorsRemote work
- ...Surge was founded by engineers and researchers who dreamed... ...building the next generation AI. We're building a... ...to conducting rigorous evaluations that go beyond benchmarks. We've run a profitable... .... The Role As a Software Engineer, Coding Evaluation & Training Data...SuggestedFull time
$80 - $100 per hour
...employment verification for this role. What You'll Be Doing Design and build the coding benchmarks and evaluation pipelines used to test frontier AI models on real software engineering work: Design coding benchmarks that evaluate frontier models on real-world...SuggestedRemote jobFull timeContract workFor contractors- Turing is seeking a Software Engineering evaluator to create cutting-edge datasets for training and benchmarking large language models. You will curate code examples and provide precise solutions mainly... ...languages, while evaluating AI-generated code for efficiency and...SuggestedFor contractors10 hours per weekFlexible hours
- ...Responsibilities Review and refine AI-generated prompts, responses, and code Validate algorithms and software concepts for technical... ...or language Support benchmarking efforts to evaluate and compare model... ...experience in software engineering, technical research, or...SuggestedPart timeRemote work
$80 - $100 per hour
...client who is a leading AI platform that enables organizations... ...human feedback, AI evaluation, and model alignment.... ...designing programming benchmarks, evaluating AI-generated code, and helping improve the... ...for experienced software engineers who enjoy solving complex...Hourly payWeekly payContract workRemote work10 hours per week- ...5, we started Handshake AI and built the fastest-growing... ...researchers to create evaluations, publish benchmarks, and push the boundary... ...Work together with engineers, scientists, operators,... ...the Role As a Senior Software Engineer on our Coding Pod, you'll lead the design...Full timeWork at officeFlexible hours
- ...Handshake is the career network for the AI economy. 20 million knowledge... ...About the Role As a Software Engineer on our Coding Pod, you will build the data infrastructure... ...large-scale, high-quality benchmark datasets that evaluate how models perform on real-world,...Full timeFreelanceInternshipWork at officeRemote workFlexible hours
$60 - $100 per hour
...technical talent with leading AI research labs.... ...our investors include Benchmark , General Catalyst ,... ...Dorsey . Position: Software Engineering, Data Science, and Systems... ...Responsibilities Evaluate LLM-generated responses to coding and software engineering...Full timeContract workSummer workRemote work$152k - $241.5k
...now looking for a Senior Software Engineer for Agent Simulation & Evaluation! Today, NVIDIA is... ...unlimited potential of AI to define the next era of... ...driven by reasoning models, coding agents, parallel subagents... ...critical LPU workflows such as benchmarking, workload...Full time- ...new era, we seek AI-native thinkers across... ...The Cortex Code team is building... ...ll own the full AI engineering lifecycle: design,... ...coding harnesses, evaluation pipelines. Build... ...orchestration, or software with substantial state... ...—not only one-off benchmarks. ~ Proficiency...Full time
$50 - $150 per hour
...technical talent with leading AI research labs.... ...Francisco, our investors include Benchmark , General Catalyst , Peter... ...Dorsey . Position: Software Engineering Expert Type: Contract... ...Write and evaluate code across diverse programming...Full timeContract workSummer workRemote work- ...the role We're looking for a Software Engineer to build and run the platform behind our AI model evaluations. You'll work closely with our research team to prepare benchmark datasets, build the pipelines... ...You write robust, maintainable code and are comfortable diving deep...Full time
- ...Senior Software Engineer, AI Code Modernization Location: Hybrid – Exton, PA/Philadelphia... ...conversion/test/correction process. Evaluate and improve AI systems for accuracy,... ...frameworks, simulation systems, benchmarking, or validation tooling. Experience...Worldwide
- ...development of high-quality datasets and evaluation pipelines that improve and benchmark large language models for code generation and software engineering tasks. You will curate and author reference code, evaluate and refine AI-generated solutions across multiple programming...Full timeFor contractorsRemote work10 hours per weekFlexible hours
$35 - $120 per hour
...technical talent with leading AI research labs.... ..., our investors include Benchmark , General Catalyst ,... ...Dorsey . Position: Code-Data Eval Author — Software Engineer Type: Contract... ...structured review. Evaluate the accuracy and depth of...Full timeContract workSummer workRemote work$65 - $105 per hour
...Apply deep engineering judgment to help frontier AI models reason more accurately... ...about real-world software development and engineered... ...engineering work, evaluate model performance,... ...standards and benchmarks. Key Responsibilities... ..., subtly incorrect code, unaddressed edge...Hourly payFull timeFreelanceInternshipLive inRelocationRelocation package- ...Senior Software Engineer — AI Coding Evaluator is a remote engineering review track for evaluating production code, debugging traces, and developer-facing AI outputs against real-world correctness standards. Reviewers reproduce failures, write the unit test the model should...Remote jobHourly payFor contractors10 hours per week
$320k
...interpretable, and steerable AI systems. We want AI to... ...committed researchers, engineers, policy experts, and... ...looking for a Software Engineer to help us build... ...and maintain new agentic coding tools for developers. The... ...experimenting with new tools, evaluating emerging techniques,...Full timeWork experience placementWork at officeVisa sponsorshipFlexible hours$204k - $259k
...state-of-the-art Generative AI to create a training ground for... ...Waymo Driver. The Simulator Evaluation team faces the ultimate data... ...We are looking for a Senior Software Engineer to build the metrics and systems... ...: you write clean, testable code that is built to last. Data...Full timeRemote work$175k - $215k
...state-of-the-art Generative AI to create a training ground for... ...Waymo Driver. The Simulator Evaluation team faces the ultimate data... ...real"? We are looking for a Software Engineer to build the metrics and pipelines... ...and AI, writing the code that processes petabytes of simulation...Full timeRemote work$139k - $204k
...CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers,... ...role We’re looking for a Senior Engineer for CoreWeave’s Benchmarking & Performance team. You will have an... ...cross-team designs and elevate coding/testing standards. Help ensure reproducible...Permanent employmentFull timeTemporary workCasual workWork at officeRemote workFlexible hours$170k - $216k
...the Waymo Driver. Our software allows the Waymo Driver... ...sensors, enabling software engineers like you to develop... ...critical automation and evaluation frameworks that establish... ...in industrial AI applications involving... ...with robust and efficient code We Prefer MS or...Full timeRemote work- ...organize human intelligence to power the AI economy. We partner with leading AI... ...context that can't be captured in code alone. Today, more than 30,000 experts... ...About the Role As a Senior Software Engineer (AI Data & Evaluation) at Mercor, you will be at the core of...Full timeWork at officeRelocation package
- ...Deepgram is looking for a Senior Software Engineer - Model Evaluation & AI Systems to join the team responsible... ...fail criteria grounded in Research benchmarks, and build the monitoring that... ...manually. Help raise the bar through code reviews, technical design...Full time
- AI Systems Modernization DeveloperBentley Systems... ...dedicated expert team in code modernization. This... ...support and guide other software developers in the company... ...: Manual evaluation of the quality of the conversion... ...solutions for architecture, engineering, and construction -...Worldwide
- ...About the Organization The Evaluation team builds and evolves... ...clear feedback for engineering and leadership, and help... ...autonomous driving software performance at subsystem... ...thoughtful system design, code reviews, testing,... ...Experience leveraging AI-assisted development and...Full timeLocal areaWork from home
$152k - $241.5k
...tapping into the unlimited potential of AI to define the next era of computing.... ...impact on the world.We are seeking a Software Engineer - Scientific Evaluation to own a shared platform for classical testing, scientific benchmarking, and agentic evaluation. The portfolio...Full time$224k - $356.5k
...autonomous driving, and evaluation is how we know the drive... ...our organization develops AI drivers!We are looking for a senior engineer to own the engine that... ...end by this person, in code and in the room, not from... ...experience).12+ years building software, with significant time...Full timeRemote work- ...incident at a time. Ambient.ai is the category... ...infrastructure for inference, evaluation, and continuous model... ...of infrastructure engineering, production ML systems... ...harnesses and benchmarking systems to measure model... ...in Python, with solid software engineering fundamentals...Full timeWork at officeLocal areaFlexible hours3 days per week
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Software Engineer, AI Code Evaluation and Benchmarking. Be the first to apply!
- software engineer internship United States
- software development engineer aws United States
- software developer internship no experience United States
- real time software engineer United States
- financial software developer United States
- oracle software engineer United States
- part time software developer United States
- software engineer co-op United States
- graduate software developer United States
- software engineer travel United States




