Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Member of Technical Staff — Research Engineering, Evaluation

Causal Labs

Our mission is general causal intelligence; AI that is capable of (1) predicting the future and (2) identifying the actions to alter it. To achieve this breakthrough, we are building a Large Physics foundation Model (LPM) because physical systems, unlike text or images, are governed by verifiable cause and effect. We believe that scaling on physics will enable an understanding of causality required to predict and control physical systems, starting with weather. Our founding team has built and deployed AI against the physical world in robotics, drug discovery, and particle physics at institutions like DeepMind, Waymo, Cruise, Insitro, Nabla Bio, and CERN. We look for research engineers who are excited to tackle unsolved problems. Progress is only as trustworthy as its measurement. As our model, data, and reasoning efforts multiply, every team needs to know, precisely and comparably, whether a change made the model better. Your mission is to build the central evaluation framework that the entire research organization runs on: the pipelines, the metrics, and the tools that turn results into shared understanding. Responsibilities Design and build a central, reusable evaluation framework that every model and every team runs through Implement evaluation pipelines, benchmark suites, and baselines that make model quality measurable and comparable across efforts Build the visualization and dashboard tools that turn raw results into shared, actionable understanding for the whole team Establish sound statistical methodology for evaluation, so teams can distinguish real improvements from noise Partner with research and domain teams to translate what "good" means in each domain into standardized, automated metrics What we\'re looking for We value a relentless approach to problem-solving, rapid execution, and the ability to quickly learn in unfamiliar domains. Strong software engineering skills and experience building data or evaluation pipelines at scale Experience turning research or model outputs into metrics, benchmarks, and visualizations that teams rely on Solid grasp of probability and statistics, with the judgment to design evaluations that measure what they claim to Full-stack range: comfortable building both backend pipelines and the frontend tools people read results in Owns deliverables end-to-end, from collecting requirements to autonomously driving execution #J-18808-Ljbffr Causal Labs

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Member of Technical Staff — Research Engineering, Evaluation in San Francisco, CA vacancy
  •  ...preference and judgment. That lets us evaluate models on what people actually care about...  ...actually want. We’re a small, deeply technical team with people from Harvard,...  ...About the Role We’re looking for an ML Research Engineer to help us build better ways to evaluate... 
    Suggested

    Arcada Labs Incorporated

    San Francisco, CA
    6 hours ago
  • $180k - $290k

     ...staying true to what makes us different: research excellence, open science, and building...  ...cleanly separated into “research” or “engineering.” A promising architecture only matters...  ...for this role. The common thread is deep technical ownership: you should be able to make... 
    Suggested
    Remote work
    Worldwide
    2 days per week

    Black Forest Labs

    San Francisco, CA
    2 days ago
  •  ...Github, Gmail, Notion, Salesforce, etc. We are a small team of engineers wrangling problems from context to search, that help us...  ...top of our agentic harness and app sandboxes Qualifications research you can independently execute against the research goals you... 
    Suggested

    Composio

    San Francisco, CA
    6 hours ago
  • About Polymath Polymathisan applied research lab focused on advancing long‑...  ...team. About the role We’re hiring a Member of Technical Staff - Engineering to build the infrastructure and systems...  ...makes it possible to train and evaluate autonomous agents in complex, realistic... 
    Suggested

    Polymath

    San Francisco, CA
    4 days ago
  •  ...architecting what comes next. We believe great talent powers great technology. The Liquid team is a community of world-class engineers, researchers, and builders creating the next generation of AI. Whether you're helping shape model architectures, scaling our dev... 
    Suggested

    Liquid AI

    San Francisco, CA
    1 day ago
  • $125k - $200k

    Member of Technical Staff - Product Engineer Patronus AI is a frontier lab developing simulation research and infrastructure to accelerate progress toward human-aligned AGI. We are on...  ...and most influential research in AI evaluation like FinanceBench , Lynx , SimpleSafetyTests... 
    Work at office

    Patronus AI, Inc.

    San Francisco, CA
    2 days ago
  • $150k - $250k

     ...Product Hunt), Charlie Songhurst (Board Member, Meta), and Michael Jones (Former...  ...United Nations, UChicago, and Oxford engineers and researchers. Our omnichannel agents are supporting...  ...custom performance and quality evaluations for our agents Minimum Qualifications... 
    Full time
    Internship
    Worldwide

    Krew

    San Francisco, CA
    more than 2 months ago
  •  ...states. Our team of AI researchers and company builders...  ...model capabilities. As a member of the Data Team, your...  ...role is ideal for engineers who love building distributed...  ...solving the unique technical challenges of...  ...used in training and evaluating large language models... 
    Relocation package

    Reflection

    San Francisco, CA
    4 days ago
  • $200k - $400k

     ...agents based on real humans. Our research pioneered the field of AI-based simulation...  ...Rauch. About the Role As a Member of Technical Staff, Model Evaluations at Simile, you will build the...  ...evaluation, model training, and research engineering. Others may bring exceptional... 
    Flexible hours

    Simile

    San Francisco, CA
    6 hours ago
  •  ...Location Type On-site Department Engineering Our Mission Reflection’s...  ...states. Our team of AI researchers and company builders come from...  ...but from better data. As a member of the Data Team, your mission...  ...Run experiments to evaluate crawling strategies, extraction... 
    Full time
    Relocation package

    B Capital

    San Francisco, CA
    4 days ago
  • Member of Technical Staff - Applied AI Engineer Valthos | Posted Mar 3 Full-time Negotiable Advanced (5-10 yrs)...  ...and executing on the Valthos-wide research and development roadmap Develop,...  ...Develop, build, and apply evaluation frameworks to rigorously assess model... 
    Full time
    Work at office

    Valthos

    San Francisco, CA
    1 day ago
  •  ...Perplexity is seeking experienced ML engineers to design, build, and...  ...user. Build the data and evaluation foundations that let these...  ...with usage. Help shape the technical direction of ranking,...  ...constantly learn outside of it. Research is essential to them and never... 
    Full time

    Perplexity®️

    San Francisco, CA
    16 hours ago
  •  ...company. We support labs, engineers and enterprises to understand...  ...Language model evaluation is the sharpest question in...  ...industry uses. We’re hiring Members of Technical Staff to build the next generation...  ...working directly with their research teams; our commercial team... 

    Artificial Analysis

    San Francisco, CA
    1 day ago
  •  ...states. Our team of AI researchers and company builders...  ...from better data. As a member of the Data Team, your...  ...used to train and evaluate our models meets a high...  ...campaigns. We’re looking for engineers who combine strong...  ...articulate complex technical concepts across teams.... 
    Relocation package

    Reflection

    San Francisco, CA
    3 days ago
  •  ...high-impact former founders, ML engineers, roboticists, and data leads...  ...intelligence. THE ROLE As a Member of Technical Staff, you will work directly with...  ...fluent : Deep understanding of evaluations, RL, and how LLMs and Agents work. Research-oriented — Able to read... 

    Socket.dev

    San Francisco, CA
    3 days ago
  • $300 per month

     ...team run by scientists and engineers from leading institutions across...  ..., tool integrations, and evaluation frameworks. Develop reusable...  ...field insights and internal research into product direction –...  ...Data Science, or a related technical field. Proficiency in Python... 
    Full time
    Work at office

    Edison Scientific Inc.

    San Francisco, CA
    16 hours ago
  • Our Mission Reflection is a research lab making intelligence open and accessible for everyone to use, customize, and build...  ...to all. Role Overview Reflection AI is looking for a Member of Technical Staff - IT Engineer. In this role, you’ll be expected to manage a broad... 
    Work at office
    Visa sponsorship

    Reflection

    San Francisco, CA
    3 days ago
  • $240k - $295k

     ...mechanism. The real product is a scalable risk engine, our Stand World Model . We stay when...  ...engine designed to manage agents and evaluations at scale. This includes coordinating...  ...Experience partnering with an applied‑science or research team and shipping their work into a... 
    Full time
    Contract work
    Temporary work
    H1b
    Work at office
    Visa sponsorship
    Work visa
    Flexible hours

    Stand Insurance

    San Francisco, CA
    3 days ago
  •  ...to notice whats broken Comfortable writing technical design docs and operational runbooks Bachelors degree in CS, Engineering, or related field (or equivalent practical experience...  ...impact Visa sponsorship As a founding member, youll help define the technical foundation... 
    Visa sponsorship
    Flexible hours

    NeoSigma Telecom Solutions Private Limited

    San Francisco, CA
    4 days ago
  • $150k - $250k

     ...from C to Safe Rust. Our team has published award-winning AI research and is backed by top-tier investors including Eric Schmidt,...  ...make agents work, from data processing, training, inference, evaluation, to other components. You work closely with research and product... 
    Work at office
    Flexible hours

    Asari AI

    San Francisco, CA
    8 days ago
  •  ...engagement, working directly with the engineer who leads our embedded...  ...and turn them into trained, evaluated, production-ready model...  ...: You can explain something technical you built from first principles...  ...autonomous vehicle, or defense research labs. What Success Looks... 
    Full time
    Internship
    Shift work

    Liquid Ai Inc

    San Francisco, CA
    16 hours ago
  • $150k - $275k

     ...lease origination. We're building a better system to evaluate consumers consistently and fairly. We prevent...  .... Role overview Two Dots is looking for Software Engineers of varying experience levels (new grads to staff), to help build out our platform to serve more customers... 

    Two Dots

    San Francisco, CA
    2 days ago
  • $200k - $280k

    Member of the Technical Staff - Product Engineer Why Join Stand: At Stand, you’ll help build a new class of global property protection. We use advanced physics and AI to model catastrophic risk at the asset level, then automate underwriting and mitigation before loss occurs... 
    Full time
    Temporary work
    H1b
    Visa sponsorship
    Work visa
    Flexible hours

    Stand

    San Francisco, CA
    6 hours ago
  • $200k - $300k

    Member of Technical Staff - Agent Engineer About Phylo Phylo is an applied research lab building agentic intelligence to accelerate discovery for every biomedical scientist...  ...production experience in AI agents to build and evaluate systems that make our agents capable,... 
    Work at office

    Phylo

    South San Francisco, CA
    4 days ago
  •  ...model providers like OpenAI to rigorously evaluate the capabilities of state-of-the-art...  ...evaluation platform and next-generation data engine. You'll train human preference models,...  ...you, and work directly with the researchers at frontier labs to creatively scale new... 
    Full time
    Relocation
    Visa sponsorship

    Design Arena

    San Francisco, CA
    3 days ago
  • $100k - $300k

     ...cutting edge, we blend frontier research with real-world execution....  ...for talented, ambitious AI/ML Engineers who are excited to build in...  ...in place frameworks to train, evaluate, and stress-test ML systems,...  ...support and uplevel future team members Mentor and grow future... 

    Cogent

    San Francisco, CA
    3 days ago
  • $285.55k

     ...METR We are a nonprofit research organization that...  ...We're Looking For The Evaluation Execution team at METR...  ...execution and software engineering skills. Research Execution...  ...and make sound technical decisions. You lead large...  ...Requirements Our technical team members are in our office in... 
    H1b
    Work at office
    Work from home
    Home office
    Relocation package
    3 days per week

    METR

    Berkeley, CA
    2 days ago
  •  ...agents, enterprises, and even nation states. Our team of AI researchers and company builders come from DeepMind, OpenAI, Google Brain...  ...advance our understanding of model capabilities Build and refine evaluation systems and processes that create tight feedback loops... 
    Relocation package

    Reflection

    San Francisco, CA
    3 days ago
  • $240k - $295k

     ...current delivery mechanism. The real product is a scalable risk engine, our Stand World Model. We stay when traditional insurers exit....  ...should be prepared to discuss how they use AI, how they evaluate its output, and how it informs their work. Strong communication... 
    Hourly pay
    Permanent employment
    Full time
    Temporary work
    H1b
    Work at office
    Local area
    Visa sponsorship
    Work visa
    Flexible hours

    Stand Insurance

    San Francisco, CA
    1 day ago
  • Member of Technical Staff, Transparency (Backend Engineer) At Anchorage Digital, we are building the world’s most advanced digital asset platform for institutions to participate in crypto. Anchorage Digital is a crypto platform that enables institutions through custody,... 

    Crypto Pro Network

    San Francisco, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Member of Technical Staff — Research Engineering, Evaluation. Be the first to apply!