Evaluations Engineering - Member of Technical Staff
$200k - $400kSimile
About the Company Pilots don't train with real passengers. Actors don't rehearse with real audiences. Yet, the most consequential decisions in society are often pushed straight to production. Simile is changing that. We have built the first AI simulation of society, populated by generative agents based on real humans. Our research pioneered the field of AI-based simulation, proving it is possible to model human behavior with high accuracy. Today, we are developing a Foundation Model to predict human behavior in any situation, at any scale. We are backed by $100M in funding led by Index Ventures, with participation from Hanabi, A*, Bain Capital Ventures, and AI visionaries including Andrej Karpathy, Fei-Fei Li, Adam D'Angelo, and Guillermo Rauch. About the Role As a Member of Technical Staff in Evaluations Engineering, you will build the systems that enable Simile to evaluate whether our simulations of human behavior are accurate, trustworthy, and improving over time. You will work across data and evaluation infrastructure, evaluation execution workflows, backend services, automation, and internal tooling. Your initial focus will include streamlining how evaluations are run across models; strengthening evaluation versioning, data models, and access controls; and automating customer validations, survey operations, and human data workflows. Evaluation at Simile presents unusual engineering challenges. Our models predict distributions of human behavior, and the ground truth used to evaluate them can be noisy and heterogeneous. You will partner closely with Evals, Modeling, Product Engineering, and Data Operations to turn complex methods and inputs into systems that are reproducible, scalable, and useful for model development and business decisions. In this role, you will: Build evaluation execution infrastructure: Develop the services, pipelines, and orchestration needed to run evaluations efficiently across datasets, model versions, populations, and use cases. Strengthen evaluation data systems: Design relational schemas, versioning, provenance, permissions, and quality controls that make evaluation results reproducible and trustworthy. Automate validation and data collection: Partner with Evals and Data Operations to streamline customer validations, survey deployment, response ingestion, and the integration of new ground truth. Build human data workflows: Create labeling and review tools that enable external experts and operators to contribute high-quality judgments to evaluation campaigns. Develop evaluation tooling: Build interfaces that help teams manage evals, compare models, investigate results, and identify regressions. Requirements Must Haves Strong Engineering Fundamentals: Several years of experience building and maintaining production-quality software, with sound judgment in system design, testing, debugging, and maintainability. Data and Systems Experience: Experience building backend services, data pipelines, automation workflows, and relational data models. End-to-End Execution: Ability to work across data, backend, and interface layers and take ambiguous projects from technical design through deployment and adoption. Evaluation Judgment: Strong intuition for what makes evaluation infrastructure reliable, including versioning, provenance, reproducibility, holdout integrity, noisy ground truth, and meaningful model comparisons. ML and LLM Fluency: Familiarity with modern model-development and evaluation workflows sufficient to partner effectively with modeling and evaluation researchers. Product and User Judgment: Ability to build clear, efficient tools for researchers, engineers, data operators, and other expert users. Ownership and Communication: A track record of independently driving important technical work and collaborating effectively across engineering, research, and operations. Nice to Haves We do not expect one person to have all of these. We are hiring a team with complementary strengths. Model-Evaluation Infrastructure: Experience building LLM or ML evaluation systems, benchmark platforms, regression suites, experiment-tracking tools, or model-quality dashboards. Research and Internal Tools: Experience developing technical surfaces for ML engineers, researchers, data scientists, or operations teams. Human Data Systems: Experience with labeling platforms, expert-review workflows, LLM-as-judge systems, grader calibration, or other human-in-the-loop evaluation methods. Data-Collection Automation: Experience automating surveys, experiments, customer-data ingestion, or other human data collection workflows. Statistical Fluency: Comfort reasoning about sampling error, uncertainty, calibration, confidence intervals, and distributional metrics. Sensitive Data and Access Controls: Experience designing permissions, auditability, and data-governance systems for human or customer data. Agentic Engineering: Experience using modern AI coding tools to accelerate development while independently testing and validating their output. You might be a great fit if you have worked in LLM evals, applied ML research, data science, research engineering, human data, market research, UXR, polling, behavioral science, computational social science, or behavioral economics. You might also be a recent graduate or self-directed builder with unusually strong taste in evaluation, statistics, and AI tools. You do not need to match every bullet. If you do not perfectly see yourself in this JD but believe you would be exceptional at building the measurement layer for behavioral simulation, we would love to hear from you. Compensation & Benefits At Simile, we provide competitive compensation packages that include base salary, equity, and comprehensive benefits. Salary Range: $200,000 – $400,000 USD Note: Final offers are based on experience, specialized skills, interview performance, and relevant training. Equity: Grants are available for eligible roles, subject to board approval. Health & Wellness: Comprehensive medical, dental, and vision coverage. Time Off: Flexible time off policies to support work-life balance. Our Process We prioritize thoughtful conversations and clear examples of past work. Our hiring journey is designed to help both sides align on fit, working style, and expectations. Reapplication Policy: To ensure a fair and thorough evaluation for all applicants, Simile observes a 90-day waiting period before reconsidering candidates for the same role. Commitment to Diversity & Inclusion Equal Opportunity: Simile is an equal opportunity workplace. We welcome applicants of all backgrounds and identities, valuing an environment where everyone can contribute authentically. Accommodations: If you require support or reasonable accommodations during the application process due to a disability, please let us know. We are happy to assist. #J-18808-Ljbffr Simile
- ...Employment Type Full time Location Type On-site Department Engineering Our Mission Reflection’s mission is to build open superintelligence... ...our understanding of model capabilities Build and refine evaluation systems and processes that create tight feedback loops...SuggestedFull timeRelocation package
- ...As a Member of Technical Staff (MTS), you'll build production-grade systems that power continuous... ...optimization loops for AI agents—from evaluation pipelines and data/trace... ...This role is a blend of MLE + backend engineering with a strong customer empathy component...Suggested
$227.5k - $401k
...Adyen, everything we do is engineered for ambition. We create an... ...who tackle unique technical challenges at scale and solve... ...financial technology sector. As a Member of Technical Staff , you will operate with a... ...(DABStep) , which evaluates AI agents on real-world data...SuggestedWork at officeImmediate startRelocationFlexible hours$160k - $220k
...backed startup building the simulation, evaluation, and observability layer for voice AI... ...with real customers. Founded by engineers who previously built core infrastructure... ...AI infrastructure The Role As a Member of Technical Staff, you\'ll architect and build the systems...SuggestedFull time- ..., Cruise, Insitro, Nabla Bio, and CERN. We look for research engineers who are excited to tackle unsolved problems. Progress is only... ...change made the model better. Your mission is to build the central evaluation framework that the entire research organization runs on: the...Suggested
- ...Adyen, everything we do is engineered for ambition. We create an... ...who tackle unique technical challenges at scale and solve... ...financial technology sector. As a Member of Technical Staff , you will operate with a... ...(DABStep), which evaluates AI agents on real-world data...Flexible hours
- ...and candidates. Title of the Role: Member of Technical Staff, iOS Company Stage of Funding: Series... ...labs and Fortune 500 companies capture, evaluate, and operationalize human knowledge... ...world-class native iOS platform. This engineer will help build the mobile...Work at officeRelocation package
$100k - $150k
Founding Member of Technical Staff (Security) Location: San Francisco • Singapore • Hyderabad • London Engineering • Hybrid • Full-time We're looking for a founding security researcher... ...in our blog. Create benchmarks to evaluate agent performance on real-world scenarios...Full timeFor contractorsWork at office$3,000 - $5,000 per month
Member of Technical Staff Intern San Francisco, Singapore, Hyderabad, London Engineering Hybrid Full-time We're looking for interns to join us to build the future of security... ...and are therefore framework-agnostic when evaluating past experience. You have deep experience...Full timeFor contractorsInternshipWork at office$200k - $350k
Member of Technical Staff, LLM Evaluation Infra Bay Area Research In office Full-time Inception creates the world’s fastest, most efficient AI models.... ...best-in-class quality. We are the AI researchers and engineers behind such breakthrough AI technologies as diffusion...Full timeWork at officeImmediate startFlexible hours- Title of the Role: Member of Technical Staff, Android Company Stage of Funding: Series C AI Infrastructure... ...and Fortune 500 companies capture, evaluate, and operationalize human knowledge... ...while establishing the architecture, engineering standards, and best practices that...Work at officeWorldwideRelocation package
$160k - $220k
Bluejay — Member of Technical Staff Type: Full-time | On-site | San Francisco, CA Compensation: $1... ...modality for AI, and Bluejay builds the evaluation infrastructure to make voice agents... ...creating hard real-time product and engineering problems. Founded: 2025 | Team size:...Full time$200k - $400k
...D'Angelo, and Guillermo Rauch. About the Role As a Member of Technical Staff, Model Evaluations at Simile, you will build the measurement systems that... ...deep in LLM evaluation, model training, and research engineering. Others may bring exceptional strength in statistics,...Flexible hours$225k - $300k
Bluejay - Senior Member of Technical Staff Type: Full-time | On-site | San Francisco, CA Compensation... ...the simulation, observability, and evaluation infrastructure to make voice the... ...month, which creates hard, real-time engineering problems. Founded: 2025 | Team size...Full timeWork at office- ...About Us Preference Model is building automated ML research engineering. Existing frontier models are brittle when applied to real... ...and implementing RL environments, conducting experiments and evaluations, delivering your work into production training runs, and collaborating...Full timeVisa sponsorshipRelocation package
$240k - $280k
...recruiter to learn more. Base pay range $240,000.00/yr - $280,000.00/yr Direct message the job poster from Cabana Senior Engineer / Member of Technical Staff @ AI Healthcare Startup $240,000 - $280,000 You know how breakthrough gene therapies for cancer, Alzheimer's, and...Full timeRemote workWorldwideRelocation- ...Pixeltable Inc. Member of Technical Staff San Francisco, CA·Full time Apply for Member of Technical Staff As a founding member of the engineering team, you will impact the design and direction of Pixeltable at a formative stage, contributing to some of our most foundational...Full timePart timeWork at officeWork from homeFlexible hours2 days per week
- ...related opportunities. Career Launch is hiring candidates for Member of Technical Staff roles and similar opportunities with fast-moving teams.... ...across cross-functional teams. Work with product, engineering, operations, growth, customer, or business stakeholders depending...Remote work
$320k
...machine — post-training, agentic architectures, the data and evaluation layers beneath them — and we hire people who can build any part... ...the instrument needs next. We do not separate research from engineering, and we do not carve the system into territories. You run...Full timeVisa sponsorshipRelocation package- ...out our founding team. About the role We’re hiring a Member of Technical Staff - Engineering to build the infrastructure and systems that power our... ...technical foundation that makes it possible to train and evaluate autonomous agents in complex, realistic environments. You...
$256k - $276k
...vision at Postman. The Opportunity As a Member of Technical Staff and AI Agent Development Lead, you... ...closely with research, product, and engineering teams to build safe, interpretable, and... ...of collaboration and innovation. Evaluate new tools, frameworks, and methodologies...Work at officeFlexible hours3 days per week$125k - $200k
Member of Technical Staff - Product Engineer Patronus AI is a frontier lab developing simulation research and infrastructure to accelerate progress toward... ...of the earliest and most influential research in AI evaluation like FinanceBench , Lynx , SimpleSafetyTests ,...Work at office- ...influence model capabilities. As a member of the Data Team, your... ...scale. This role is ideal for engineers who love building distributed... ..., and solving the unique technical challenges of collecting data... ...data is used in training and evaluating large language models Experience...Relocation package
$150k - $300k
...agreed to acquire our first firm, and are rapidly growing our engineering team to help us build and deploy our agents at scale. We’re a... ...outcomes with our design partners — your work will hit production Evaluate and integrate emerging AI frameworks, tools, and best...Work at office- ...6-month product roadmap, so we are expanding our engineering team. We're looking for someone highly technical (our current team includes 3 IOI medalists) who wants... ...balance and trust. Room to Grow: As an early member of the team, you’ll have the opportunity to take on...Flexible hours
$150k - $300k
...on a distributed system with performance engineering at its core. The role will draw on the... ...fast, robust, and reliable at scale. Core Technical Responsibilities Infrastructure... ...in open development and encourage team members to contribute to the broader AI community...Work at officeRemote workVisa sponsorshipRelocation packageFlexible hours- About Us Preference Model is building automated ML research engineering. Existing frontier models are brittle when applied to real-world... ...and shape research directions. What You Will Do: Train and evaluate models on our proprietary RL environments to validate data quality...Visa sponsorshipRelocation package
$285.55k
...What We're Looking For The Evaluation Execution team at METR focuses... ...execution and software engineering skills. Research Execution... ...scalable systems and make sound technical decisions. You lead large projects... ...Our technical team members are in our office in Berkeley...H1bWork at officeWork from homeHome officeRelocation package3 days per week- Perplexity is seeking an intrepid, polymathic Member of Technical Staff to take on one of the AI industry’s most unique engineering roles. You’ll work directly with Perplexity’s senior leadership to spearhead a broad portfolio of strategic technology initiatives across...
$200k
Join to apply for the Member of Technical Staff role at Listen Labs . TL;DR: We are seeing strong market demand and an aggressive 6‑month product roadmap, so we are expanding our engineering team. We're looking for someone highly technical (our current team includes 3...Flexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Evaluations Engineering - Member of Technical Staff. Be the first to apply!
- assistant engineering manager San Francisco, CA
- assistant civil engineer San Francisco, CA
- assistant mechanical engineer San Francisco, CA
- assistant engineer San Francisco, CA
- staff engineer San Francisco, CA
- staff data engineer San Francisco, CA
- software engineer staff San Francisco, CA
- assistant electrical engineer San Francisco, CA
- staff design engineer San Francisco, CA
- senior staff engineer San Francisco, CA


