Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Evaluations Engineer

Vals AI, Inc.

About the Role We are looking for strong engineers to join our team and own the leaderboards that appear on Vals AI. You will be responsible for testing and benchmarking new models as they are released on tasks in law, tax, coding, finance, and more. You will analyze error modes of models, evaluate their strengths and weaknesses, and work with our communications team to release results. Our results are used by startups, enterprises, and research labs alike. We work with all the major foundation model labs, some of the largest financial institutions, and hospital systems in the world. Our work has been featured by the Wall Street Journal, Washington Post, and Bloomberg. We are building the standard for evaluating the ability of LLMs to perform real-world tasks. You will contribute directly to the leaderboards that make this possible. What You’ll Do Evaluate new LLM model releases across the Vals AI suite of benchmarks Work directly with both open-source and closed-source foundation model labs in evaluating model performance Use tools like Docent to analyze common failure modes and patterns in model performance Work directly with our social media team to post interesting findings and results Add new models and maintain integrations in our model library Help improve and maintain the infrastructure we use to run benchmarks (agentic and non-agentic). Collaborate closely with our research team on the creation of new benchmarks This role follows the rhythm of model releases. Expect intense sprints in the days following a major launch, and calmer stretches in between releases. Requirements Familiarity with the LLMs: You should already be familiar with the space - the current leading models, relative performance across them, how to use large language models in practice. Strong engineering fundamentals : You can build and ship quickly with high quality. You should have a track record of building things of significant scope (at jobs, side projects, open source, etc.) Python expertise : Significant experience in Python, especially in a professional setting. Team collaboration : Experience working in development sprints, Git workflows, and pull request reviews. Location : We are an in-person team based in San Francisco. We will support your relocation or transportation as needed. Nice-to-Haves Previous experience with benchmarking large language models, or creating benchmarks Previous experience working at a startup or starting your own company Technical writing experience and ability Machine learning research experience What We Offer Highly competitive salary and meaningful ownership. Excellence is well rewarded. Relocation and transportation support Health/dental insurance coverage Lunch and dinner provided, free snacks/coffee/drinks 401K plan Unlimited PTO About Us Founding team : The core methodology behind this platform comes from NLP evaluation research we had done at Stanford. We raised a $5M seed from some of the top institutional and angel investors in the valley. Our team has prior work experience at NVIDIA, Meta, Microsoft, Palantir and HRT. Collectively, we have over 300 citations in our published work. Our early team include Stanford PhDs, ex-Jane Street quants, and the first designer at Snorkel. Tech stack : We use Python for most things at Vals. Our platform is built on Django, with a React frontend. All of the infra is on AWS using CDK for IaC. What We\'re Looking For Learning velocity: The role encompasses a wide variety of tasks. Rather than expecting you to be an expert on Day 1, we are looking for someone who can learn new skills and technologies extremely quickly. Ownership : Working in a small, talent-dense team, we expect everyone to show initiative to build where it\'s needed, not where it\'s asked. We strive for autonomy over consensus. This is especially true for this role. Intensity : The LLM landscape is constantly changing. Foundation model labs are continuously pushing the frontier. The unicorn companies that will emerge from this technology shift are being built now. Those that win will have an incredibly high speed of execution. Solution-oriented mindset : We\'re looking for people who see opportunities to craft solutions at each juncture, not those who pass hard problems to others or admit defeat. Further Reading: Hugging Face blog on evaluation Anthropic’s blog on challenges in evaluation New York Times article on issues in benchmarking Stanford HAI report showing hallucinations in legal tech tools #J-18808-Ljbffr Vals AI, Inc.

Vacancy posted 5 days ago
Similar jobs that could be interesting for youBased on the Evaluations Engineer in San Francisco, CA vacancy
  • YO AI Labs is seeking experienced Civil Engineers for remote, contractor-type work focused on AI evaluation. You will craft realistic civil engineering tasks, prepare technical materials, and assess code compliance across design domains. The role emphasizes constraint reasoning... 
    Suggested
    Remote job
    For contractors

    YO AI Labs

    San Francisco, CA
    3 days ago
  • $315k

    We are looking for Research Engineers to build “gold standard” evaluations for catastrophic risks, in order to understand what AI Safety Level (ASL) to assign to models. Research leads on this team collaborate with engineers in one of our focus areas: CBRN, Cyber, Autonomy... 
    Suggested
    Currently hiring
    Work at office
    Immediate start
    Home office
    Visa sponsorship
    Relocation package

    Anthropic

    San Francisco, CA
    1 day ago
  • Synthires in the United States seeks experienced Mechanical Engineering professionals to contribute to AI evaluation and training projects. You will shape enterprise-grade scenarios, evaluation frameworks, and reference solutions for advanced AI systems. The work focuses... 
    Suggested
    Remote job
    Flexible hours

    Synthires

    San Francisco, CA
    4 days ago
  • Obsidian is seeking a candid evaluator to assess the quality, correctness, and reproducibility of software-engineering benchmark tasks used to train frontier AI labs. You will review repository-level tasks, reference patches, test harnesses, and grading integrity, and provide... 
    Suggested

    Obsidian

    San Francisco, CA
    1 day ago
  • Artificial Analysis is seeking a Member of Technical Staff to lead the design and execution of frontier language model evaluations. You will build datasets, evaluation harnesses, and scoring systems that resist contamination at scale, and you will publish findings that... 
    Suggested

    Artificial Analysis, Inc.

    San Francisco, CA
    3 days ago
  •  ...increasingly integrated into various aspects of our daily lives. Arize AI is the leading AI observability and evaluation platform , empowering AI engineers to build and deploy high-performing, reliable models. As the AI landscape shifts from traditional ML to... 
    Full time
    Local area
    Shift work
    Afternoon shift

    Arizeai

    San Francisco, CA
    13 days ago
  • $150k

    Tzafon is seeking a skilled engineer to enhance their machine intelligence systems in San Francisco. As part of the team, you'll be responsible for building evaluation infrastructure, designing data pipelines, and implementing fine-tuning processes. Ideal candidates have... 

    Tzafon

    San Francisco, CA
    4 days ago
  • Ambral is hiring to build a replayable enterprise history engine and deploy it in real workflows. You will work closely with the CTO to design production systems, process real enterprise data, and improve agent performance through reinforcement learning and context engineering... 

    Ambral (YC S25)

    San Francisco, CA
    3 days ago
  • Dynamo AI is seeking a candidate to lead LLM evaluation and benchmarking in San Francisco, California. You will generate high-quality data and develop innovative methods for assessing the safety and helpfulness of LLMs. The role requires domain knowledge in evaluation... 

    Capitolis

    San Francisco, CA
    2 days ago
  • Vals AI in San Francisco is seeking engineers to own leaderboards that evaluate LLMs across tasks including law, tax, coding, finance, and more. You will test and benchmark new models as released, analyze error modes, and work with our communications team to publish results... 

    Vals AI

    San Francisco, CA
    1 day ago
  • $140k - $185k

     ...About the Role Join Vals AI to own and operate leaderboards that evaluate LLMs: test new model releases against benchmarks, analyze...  ...knowledge of leading models and their relative strengths. Strong engineering fundamentals and a track record of building and shipping... 
    Full time
    Relocation package

    Vibehackers

    San Francisco, CA
    1 day ago
  • AIUC in San Francisco is hiring an engineer to own end-to-end evaluation work for enterprise agents. You will sit with customers, understand how their agent works, and integrate it into our evaluation system. You will work across delivery, core engineering, and sales to... 

    Socket

    San Francisco, CA
    1 day ago
  • $129.5k - $145k

     ...looking for an experienced and driven GRC Controls Automation Engineer who is looking to put their demonstrated awareness of regulatory...  ...executionImplement improvements by assessing the current environment, evaluate trends, and anticipating future enhancements/requirementsAssess... 
    Live in
    Work at office
    Shift work
    3 days per week

    Box

    San Francisco, CA
    19 hours ago
  • $206.4k - $379.1k

     ...Firefly’s Generative AI Services team is seeking a Principal Service Engineer to serve as the technical lead for our GenAI Services domain....  ...architecture, design, implementation, and best practices.Evaluate and incorporate emerging MLOps technologies to improve engineering... 
    Full time
    Temporary work
    Local area
    Worldwide

    Adobe Systems

    San Francisco, CA
    2 days ago
  •  ...the office. We are looking for a Hardware Signal Quality Test Engineer to join our San Francisco team. As Oura’s first HSQT engineer in...  ...hardware products, executing defined test activities and helping evaluate signal quality across relevant sensor domains.Contribute to... 
    Contract work
    Work experience placement
    Work at office
    Local area
    Remote work
    Flexible hours

    Oura

    San Francisco, CA
    3 days ago
  • MaxIT Consulting - Max Corporate Group in San Francisco is seeking an Agent Evaluation Infrastructure Engineer to build the environments, evaluation systems, and supporting infrastructure used to train and assess long-horizon enterprise AI agents. You will work on the... 

    MaxIT Consulting - Max Corporate Group

    San Francisco, CA
    4 days ago
  • $120k - $150k

     ...seeking a dynamic individual for the role of Commissioning Project Engineer to provide reliable, timely and efficient support to our...  ...site/project specific Systems Manuals.Ability to develop/review/evaluate training programs for installed equipment and systems.The candidate... 
    Full time
    For contractors
    Seasonal work
    Local area
    Remote work

    Jones Lang LaSalle

    San Francisco, CA
    2 days ago
  • $117.8k - $176.8k

     ....Stantec is currently seeking an Instrumentation and Controls Engineer for any one of our California state office locations - San Francisco...  ...listed areas.Design activities including:- Field review and evaluation of existing equipment, control systems (hardware and software)... 
    Full time
    Contract work
    Temporary work
    Part time
    For contractors
    Casual work
    Work at office
    Local area
    Flexible hours

    Stantec

    San Francisco, CA
    3 days ago
  •  ...Field Engineer Our mission is to automate coding. The first step in our journey is to build the best tool for professional programmers...  ...technical face of Anysphere in the field, helping customers evaluate Cursor, guiding them through proofs of concept, and ensuring they... 
    Full time

    Anysphere

    San Francisco, CA
    5 days ago
  • $211.7k - $302.4k

     ...it's a competitive advantage.We're looking for an Automation Engineering Managerto lead the team responsible for building and operating...  ...summarization), prototype them, and codify them into team standards.Evaluate and selectively adopt emerging testing tools and AI-for-... 
    Full time
    Temporary work
    Work at office
    Local area
    Immediate start
    Flexible hours
    2 days per week

    Tubi TV

    San Francisco, CA
    4 days ago
  • $155k - $269k

     ...Be part of a multidisciplinary team of Research Scientists and Engineers building the content backbone of a best-in-class multi-sensor...  ...diversity, realism, and coverage demands of training and closed-loop evaluation. Qualifications: ~ Shipping Production Software. You... 
    Full time
    Work at office
    Work from home
    Flexible hours

    Waabi

    San Francisco, CA
    12 days ago
  • $36.06 - $40.87 per hour

     ...Technical Support Field Engineer - San Francisco, CA Dentsply Sirona is the world’s largest manufacturer of professional dental products...  ...Treatment Center customers and authorized dealer technicians. Evaluates and analyzes hardware and software issues and use technology... 
    Hourly pay
    Work experience placement
    Work at office
    Remote work
    Worldwide
    Flexible hours
    Night shift

    Wellspect HealthCare

    San Francisco, CA
    3 days ago
  • $168.75k - $270k

     ...every single day.Your ImpactAs a PrincipalSecurity Operations Engineer, you will play a key role in building secure, reliable, and developer...  ...environment.We collect personal information from applicants to evaluate candidates for employment. You may request access, deletion, or... 
    Work experience placement
    Work at office

    Axon

    San Francisco, CA
    2 days ago
  • Artificial Intelligence Underwriting Company in San Francisco seeks an engineer to own end-to-end evaluation work for enterprise AI agents. You will read API docs, integrate with customer systems, and run evaluations in our platform to prove agent safety and reliability... 

    Artificial Intelligence Underwriting Company

    San Francisco, CA
    4 days ago
  •  ...in-person five days a week in our San Francisco, NYC, or London offices. About the Role As a Senior Software Engineer (AI Data & Evaluation) at Mercor, you will be at the core of building the data infrastructure and evaluation systems that power the next generation... 
    Full time
    Work at office
    Relocation package

    Mercor

    San Francisco, CA
    19 hours ago
  • Obsidian is seeking a senior engineer/researcher to design realistic evaluation tasks in materials science or engineering. You will author prompts, assemble data rooms, and create objective grading criteria to test AI models' expert performance. You’ll collaborate with... 

    Obsidian

    San Francisco, CA
    1 day ago
  • Walden Robotics is seeking an engineer to build tooling for the autonomy team to measure whether our robot policies are actually getting...  ...senior engineers. You will own end-to-end development of evaluation pipelines, dashboards, and checks, enabling researchers to quantify... 

    Walden Robotics

    San Francisco, CA
    1 day ago
  •  ...the Role We are looking for a Principal RF Systems & Hardware Engineer to lead the definition and execution of our communication...  ...ensuring link budgets and coverage goals are met. Lead the evaluation and selection of critical active RF components, including Power... 
    Work at office
    Remote work
    Shift work

    AdAstra

    San Francisco, CA
    23 days ago
  • $113k - $170k

    Kennedy Jenks is seeking a licensed Civil Engineer with structural focus in Northern California to lead design and consulting services...  ...coordinated project delivery.Analyze infrastructure systems: Design and evaluate treatment plant buildings, tanks, clarifiers, pump stations,... 
    Work at office
    Work from home
    2 days per week

    Kennedy Jenks

    San Francisco, CA
    14 hours ago
  • $60k - $95k

     ...escalated help desk tickets. Diagnose and resolve technical hardware and software issues. Gather information to determine the issue by evaluating and analyzing the symptoms. Partners with other departments to provide technical assistance. Knowledge in handling Wi-Fi and LAN... 

    Tata Consultancy Services

    San Francisco, CA
    5 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Evaluations Engineer. Be the first to apply!