Evaluations Engineer
Vals AI, Inc.
About the Role We are looking for strong engineers to join our team and own the leaderboards that appear on Vals AI. You will be responsible for testing and benchmarking new models as they are released on tasks in law, tax, coding, finance, and more. You will analyze error modes of models, evaluate their strengths and weaknesses, and work with our communications team to release results. Our results are used by startups, enterprises, and research labs alike. We work with all the major foundation model labs, some of the largest financial institutions, and hospital systems in the world. Our work has been featured by the Wall Street Journal, Washington Post, and Bloomberg. We are building the standard for evaluating the ability of LLMs to perform real-world tasks. You will contribute directly to the leaderboards that make this possible. What You’ll Do Evaluate new LLM model releases across the Vals AI suite of benchmarks Work directly with both open-source and closed-source foundation model labs in evaluating model performance Use tools like Docent to analyze common failure modes and patterns in model performance Work directly with our social media team to post interesting findings and results Add new models and maintain integrations in our model library Help improve and maintain the infrastructure we use to run benchmarks (agentic and non-agentic). Collaborate closely with our research team on the creation of new benchmarks This role follows the rhythm of model releases. Expect intense sprints in the days following a major launch, and calmer stretches in between releases. Requirements Familiarity with the LLMs: You should already be familiar with the space - the current leading models, relative performance across them, how to use large language models in practice. Strong engineering fundamentals : You can build and ship quickly with high quality. You should have a track record of building things of significant scope (at jobs, side projects, open source, etc.) Python expertise : Significant experience in Python, especially in a professional setting. Team collaboration : Experience working in development sprints, Git workflows, and pull request reviews. Location : We are an in-person team based in San Francisco. We will support your relocation or transportation as needed. Nice-to-Haves Previous experience with benchmarking large language models, or creating benchmarks Previous experience working at a startup or starting your own company Technical writing experience and ability Machine learning research experience What We Offer Highly competitive salary and meaningful ownership. Excellence is well rewarded. Relocation and transportation support Health/dental insurance coverage Lunch and dinner provided, free snacks/coffee/drinks 401K plan Unlimited PTO About Us Founding team : The core methodology behind this platform comes from NLP evaluation research we had done at Stanford. We raised a $5M seed from some of the top institutional and angel investors in the valley. Our team has prior work experience at NVIDIA, Meta, Microsoft, Palantir and HRT. Collectively, we have over 300 citations in our published work. Our early team include Stanford PhDs, ex-Jane Street quants, and the first designer at Snorkel. Tech stack : We use Python for most things at Vals. Our platform is built on Django, with a React frontend. All of the infra is on AWS using CDK for IaC. What We\'re Looking For Learning velocity: The role encompasses a wide variety of tasks. Rather than expecting you to be an expert on Day 1, we are looking for someone who can learn new skills and technologies extremely quickly. Ownership : Working in a small, talent-dense team, we expect everyone to show initiative to build where it\'s needed, not where it\'s asked. We strive for autonomy over consensus. This is especially true for this role. Intensity : The LLM landscape is constantly changing. Foundation model labs are continuously pushing the frontier. The unicorn companies that will emerge from this technology shift are being built now. Those that win will have an incredibly high speed of execution. Solution-oriented mindset : We\'re looking for people who see opportunities to craft solutions at each juncture, not those who pass hard problems to others or admit defeat. Further Reading: Hugging Face blog on evaluation Anthropic’s blog on challenges in evaluation New York Times article on issues in benchmarking Stanford HAI report showing hallucinations in legal tech tools #J-18808-Ljbffr Vals AI, Inc.
- YO AI Labs is seeking experienced Civil Engineers for remote, contractor-type work focused on AI evaluation. You will craft realistic civil engineering tasks, prepare technical materials, and assess code compliance across design domains. The role emphasizes constraint reasoning...SuggestedRemote jobFor contractors
$315k
We are looking for Research Engineers to build “gold standard” evaluations for catastrophic risks, in order to understand what AI Safety Level (ASL) to assign to models. Research leads on this team collaborate with engineers in one of our focus areas: CBRN, Cyber, Autonomy...SuggestedCurrently hiringWork at officeImmediate startHome officeVisa sponsorshipRelocation package- Synthires in the United States seeks experienced Mechanical Engineering professionals to contribute to AI evaluation and training projects. You will shape enterprise-grade scenarios, evaluation frameworks, and reference solutions for advanced AI systems. The work focuses...SuggestedRemote jobFlexible hours
- Obsidian is seeking a candid evaluator to assess the quality, correctness, and reproducibility of software-engineering benchmark tasks used to train frontier AI labs. You will review repository-level tasks, reference patches, test harnesses, and grading integrity, and provide...Suggested
- Artificial Analysis is seeking a Member of Technical Staff to lead the design and execution of frontier language model evaluations. You will build datasets, evaluation harnesses, and scoring systems that resist contamination at scale, and you will publish findings that...Suggested
- ...increasingly integrated into various aspects of our daily lives. Arize AI is the leading AI observability and evaluation platform , empowering AI engineers to build and deploy high-performing, reliable models. As the AI landscape shifts from traditional ML to...Full timeLocal areaShift workAfternoon shift
$150k
Tzafon is seeking a skilled engineer to enhance their machine intelligence systems in San Francisco. As part of the team, you'll be responsible for building evaluation infrastructure, designing data pipelines, and implementing fine-tuning processes. Ideal candidates have...- Ambral is hiring to build a replayable enterprise history engine and deploy it in real workflows. You will work closely with the CTO to design production systems, process real enterprise data, and improve agent performance through reinforcement learning and context engineering...
- Dynamo AI is seeking a candidate to lead LLM evaluation and benchmarking in San Francisco, California. You will generate high-quality data and develop innovative methods for assessing the safety and helpfulness of LLMs. The role requires domain knowledge in evaluation...
- Vals AI in San Francisco is seeking engineers to own leaderboards that evaluate LLMs across tasks including law, tax, coding, finance, and more. You will test and benchmark new models as released, analyze error modes, and work with our communications team to publish results...
$140k - $185k
...About the Role Join Vals AI to own and operate leaderboards that evaluate LLMs: test new model releases against benchmarks, analyze... ...knowledge of leading models and their relative strengths. Strong engineering fundamentals and a track record of building and shipping...Full timeRelocation package- AIUC in San Francisco is hiring an engineer to own end-to-end evaluation work for enterprise agents. You will sit with customers, understand how their agent works, and integrate it into our evaluation system. You will work across delivery, core engineering, and sales to...
$129.5k - $145k
...looking for an experienced and driven GRC Controls Automation Engineer who is looking to put their demonstrated awareness of regulatory... ...executionImplement improvements by assessing the current environment, evaluate trends, and anticipating future enhancements/requirementsAssess...Live inWork at officeShift work3 days per week$206.4k - $379.1k
...Firefly’s Generative AI Services team is seeking a Principal Service Engineer to serve as the technical lead for our GenAI Services domain.... ...architecture, design, implementation, and best practices.Evaluate and incorporate emerging MLOps technologies to improve engineering...Full timeTemporary workLocal areaWorldwide- ...the office. We are looking for a Hardware Signal Quality Test Engineer to join our San Francisco team. As Oura’s first HSQT engineer in... ...hardware products, executing defined test activities and helping evaluate signal quality across relevant sensor domains.Contribute to...Contract workWork experience placementWork at officeLocal areaRemote workFlexible hours
- MaxIT Consulting - Max Corporate Group in San Francisco is seeking an Agent Evaluation Infrastructure Engineer to build the environments, evaluation systems, and supporting infrastructure used to train and assess long-horizon enterprise AI agents. You will work on the...
$120k - $150k
...seeking a dynamic individual for the role of Commissioning Project Engineer to provide reliable, timely and efficient support to our... ...site/project specific Systems Manuals.Ability to develop/review/evaluate training programs for installed equipment and systems.The candidate...Full timeFor contractorsSeasonal workLocal areaRemote work$117.8k - $176.8k
....Stantec is currently seeking an Instrumentation and Controls Engineer for any one of our California state office locations - San Francisco... ...listed areas.Design activities including:- Field review and evaluation of existing equipment, control systems (hardware and software)...Full timeContract workTemporary workPart timeFor contractorsCasual workWork at officeLocal areaFlexible hours- ...Field Engineer Our mission is to automate coding. The first step in our journey is to build the best tool for professional programmers... ...technical face of Anysphere in the field, helping customers evaluate Cursor, guiding them through proofs of concept, and ensuring they...Full time
$211.7k - $302.4k
...it's a competitive advantage.We're looking for an Automation Engineering Managerto lead the team responsible for building and operating... ...summarization), prototype them, and codify them into team standards.Evaluate and selectively adopt emerging testing tools and AI-for-...Full timeTemporary workWork at officeLocal areaImmediate startFlexible hours2 days per week$155k - $269k
...Be part of a multidisciplinary team of Research Scientists and Engineers building the content backbone of a best-in-class multi-sensor... ...diversity, realism, and coverage demands of training and closed-loop evaluation. Qualifications: ~ Shipping Production Software. You...Full timeWork at officeWork from homeFlexible hours$36.06 - $40.87 per hour
...Technical Support Field Engineer - San Francisco, CA Dentsply Sirona is the world’s largest manufacturer of professional dental products... ...Treatment Center customers and authorized dealer technicians. Evaluates and analyzes hardware and software issues and use technology...Hourly payWork experience placementWork at officeRemote workWorldwideFlexible hoursNight shift$168.75k - $270k
...every single day.Your ImpactAs a PrincipalSecurity Operations Engineer, you will play a key role in building secure, reliable, and developer... ...environment.We collect personal information from applicants to evaluate candidates for employment. You may request access, deletion, or...Work experience placementWork at office- Artificial Intelligence Underwriting Company in San Francisco seeks an engineer to own end-to-end evaluation work for enterprise AI agents. You will read API docs, integrate with customer systems, and run evaluations in our platform to prove agent safety and reliability...
- ...in-person five days a week in our San Francisco, NYC, or London offices. About the Role As a Senior Software Engineer (AI Data & Evaluation) at Mercor, you will be at the core of building the data infrastructure and evaluation systems that power the next generation...Full timeWork at officeRelocation package
- Obsidian is seeking a senior engineer/researcher to design realistic evaluation tasks in materials science or engineering. You will author prompts, assemble data rooms, and create objective grading criteria to test AI models' expert performance. You’ll collaborate with...
- Walden Robotics is seeking an engineer to build tooling for the autonomy team to measure whether our robot policies are actually getting... ...senior engineers. You will own end-to-end development of evaluation pipelines, dashboards, and checks, enabling researchers to quantify...
- ...the Role We are looking for a Principal RF Systems & Hardware Engineer to lead the definition and execution of our communication... ...ensuring link budgets and coverage goals are met. Lead the evaluation and selection of critical active RF components, including Power...Work at officeRemote workShift work
$113k - $170k
Kennedy Jenks is seeking a licensed Civil Engineer with structural focus in Northern California to lead design and consulting services... ...coordinated project delivery.Analyze infrastructure systems: Design and evaluate treatment plant buildings, tanks, clarifiers, pump stations,...Work at officeWork from home2 days per week$60k - $95k
...escalated help desk tickets. Diagnose and resolve technical hardware and software issues. Gather information to determine the issue by evaluating and analyzing the symptoms. Partners with other departments to provide technical assistance. Knowledge in handling Wi-Fi and LAN...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Evaluations Engineer. Be the first to apply!




