Member of Technical Staff (Evals & Post-Training)
Ambral
What we do Ambral Labs helps enterprises own the intelligence behind their most important workflows. Every company has years of historical evidence showing how work gets done: the context people had, the decisions they made, the actions they took, and the outcomes that followed. Today, most of that history is inert. It isn’t structured in a way that companies can use to evaluate models and improve agent behavior. Ambral turns this history into replayable environments and eval sets grounded in real workflows and observed outcomes. We use those environments to improve model performance through reinforcement learning and other post-training techniques, alongside context engineering, harness design, and agent engineering. The result is better, more cost-efficient AI for each enterprise’s specific work, powered by open-weight models that the company owns and controls. This allows each company to retain ownership of its core intelligence instead of outsourcing it to a model provider. We graduated from YC S2025, raised millions in funding, and are already deployed within multi-billion dollar enterprises. Now we're growing the founding team What you’ll do At the center of Ambral Labs is a replayable environment engine for the enterpirse. The system reconstructs a company’s world as it existed at a particular moment in the past then exposes that state through the same tools an agent would use in production. This lets us place new policies and agent configurations inside real historical environments, observe how they reason and act, and grade their performance against real outcomes. You'll work across research, infrastructure, and production systems including: Building an environment factory that converts recorded enterprise data and task definitions into runnable environments Designing graders that turn ambiguous business objectives into verifiable rewards Developing methods for mining useful tasks, trajectories, and evaluation cases from historical workflows Creating eval sets that are representative, reproducible, and resistant to overfitting Finding the right combinations of models, tools, context, and policies to maximize performance while reducing inference cost Training and evaluating agents that operate over long horizons, incomplete information, and large tool spaces Building replay and observability systems that make agent behavior explainable and measurable Scaling from individual environments to thousands of concurrent training and evaluation runs These problems are wide open. You’ll have significant ownership over both the research direction and the production systems that make it real. You’ll work directly with the CTO, deploy into real enterprise workflows, and see your research tested against consequential problems and observable outcomes. Who you are You have 1-7 years of experience building production software or machine-learning systems (we're hiring at multiple levels for this role). Bonus points for working on reinforcement-learning environments, LLM post-training, evaluation infrastructure, agent harnesses, or closely related systems You understand how environment design, reward design, context, tooling, and policy behavior interact You’re comfortable turning fuzzy business objectives into tasks and signals that can be evaluated reliably You can diagnose whether a model’s limitations come from the model itself, its context, its tools, its harness, or its training You can move between research questions and production implementation without treating them as separate jobs You write strong software and can build systems that process large, messy datasets at scale You care about reproducibility, observability, and understanding why a model behaves the way it does You’re looking to do the best work of your life and build something you’ll be proud of for decades We care much more about what you’ve built and how you think than credentials or conventional career paths. Benefits Significant equity and ownership Equinox membership Free meals, coffee, and snacks Health insurance Unlimited PTO #J-18808-Ljbffr Ambral
- ...and field repair. About the role As a Member of Technical Staff focused on Agentic Reasoning & Core... ...domain-specific evaluations - Build evals, graders, traces, and diagnostics that... ...-use agents, multi-agent systems, RL/post-training, synthetic data generation, harness engineering...TrainingLocal areaRelocation package
- About the Role As a Member of Technical Staff - Evals at Entendre, you will play a key role in ensuring the quality and reliability of our AI-powered features. Your primary responsibilities will include: Designing and maintaining evaluation frameworks to measure the accuracy...Suggested
- ...As a Member of Technical Staff at Quadrillion Labs, you'll build the systems that power Qualia, our research agent. This is a mix of exciting... ...knowledge. Scale environments for agent experiments and training. Help us build the sandboxed compute where agents run...TrainingWork at officeLocal area
- ...are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are... ...York City, Montreal, Seoul, Germany and Paris. Join us! Member of Technical Staff, Search Why this role? We are looking for talented individuals...TrainingFull timeWork at officeLocal areaRemote workHome office
- ...and improve ML components across data, training, evaluation, and inference. Fine-tune and... ...engineering and product excellence. All members are expected to be hands-on and to... ...interviews. Applications are evaluated by our technical team members. Interviews will be...Training
- ...and acted upon by intelligent systems. Role Summary: As a Member of Technical Staff, you’ll help bridge the gap between cutting‑edge research... ...deployments Own the full stack of ML workflows: data ingestion, training, evaluation, deployment, and monitoring Improve platform...Training
- As a Member of Technical Staff (Research) at Quadrillion Labs, you'll help train our next-generation model for research creativity. This is genuinely hard: very few people have post-trained models at the scale of several hundred billion parameters, and we're doing it with...TrainingWork at officeLocal area
- .... Our work sits at the intersection of post-training, agent environments, data quality, and research... .... You should be interested in where evals lie, where rewards get hacked, and how... ...have spent years shipping high‑impact technical products. #J-18808-Ljbffr Chakra Data WarehouseTraining
- ...reliability Wall Street demands is one of the hardest applied AI problems out there. Our technical staff works across the stack: agent architecture, evals, model post-training, and new product surfaces. You'll ship to the largest financial firms in the world. Example...TrainingFull timeShift work
- ...and processes that create tight feedback loops between data, evals, and model behavior Develop generalizable evaluation frameworks... ..., alignment, and usefulness Collaborate closely with pre‑training, post‑training, and applied teams to translate insights into model improvements...TrainingRelocation package
- ...financial investors, distils our deep technical research and knowledge into key... ...We are looking for a highly motivated member of technical staff to join our engineering team to work... ...chip projects across both frontier LLM training & inference models Implement modern...TrainingFull timeWork at officeRemote workWorldwide
$250k - $350k
Member of Technical Staff - Quantitative Research New York City (Remote possible for exceptional candidates) About Uncharted/Udio Udio builds... ...loops and apply your findings to our pretraining, post-training and inference systems as applicable. Drive product & research...TrainingWork experience placementRemote workFlexible hours- ...performance of popular frontier models MI300X vs H100 vs H200 Training : Extensive Benchmarking & Deep Dive into CUDA & ROCm... ...Position Overview We are seeking a highly motivated & skilled Member of Technical Staff to join our growing engineering team. Member of Technical...TrainingFull timeWork at officeRemote workWorldwide
- ...towards financial investors, distils our deep technical research and knowledge into key insights on technology... ..., and government agencies. Position Overview Member of Technical Staff will play a crucial role in developing training & inference benchmarks & system modelling. You...TrainingFull timeWork at officeRemote workWorldwide
- Member of Technical Staff - Computational Biologist Valthos | Posted Mar 3 Full-time Negotiable Advanced (5-10 yrs) Computational Biologist Valthos Inc. Valthos... ...large-scale biological datasets appropriate for training frontier biological models, and build workflows...TrainingFull timeWork at office
- ...Stripe, DoorDash, and Ramp. About the Role Members of Technical Staff (MTS) are the senior engineers who... ...the same primitive. Observability and evals. The harness that tells us whether the... ...around a model someone else trained, and have an informed opinion on where...
- ...including Y Combinator. You will advance the core architecture and training of Output's foundation model, the system that learns biological... ..., physics, mathematics, or a related field with 2+ years of post-doctoral or industry research experience, or a Bachelor's or Master...Training
$150k - $350k
...generation for practical chemistry. You will continue developing and training models that incorporate knowledge of chemical synthesis routes... ...chemistry, cheminformatics, or a related field with 2+ years of post-doctoral or industry research experience, or a Master's degree...Training- ...advances come not from new architectures, but from better data. As a member of the Data Team, your mission is to build and operate the... ...‑scale data sources into reliable, well‑structured corpora for training frontier models. You will own the machinery that acquires,...TrainingRelocation package
- ...We're building a platform that covers the whole life of an LLM: training it, deploying it, and observing it in production. We already run... ...infrastructure underneath. We’re looking for research depth in post-training to sit alongside our systems and product work. You will...TrainingWork at office
- ...performance debugging at this scale presents genuinely hard systems problems. More broadly, you will work closely with Reflection's training teams to co-design fault tolerance, node health checks, and remediation strategies. What You'll Do Cluster Management: Build and...TrainingWork at officeVisa sponsorship
- ...platform for specialized intelligence, enabling companies to build, train, and serve AI models tailored to their own data, workflows, and... ...'ll architect scalable, resilient backend infrastructure, lead technical design discussions, mentor engineers, and establish best...Training
- ...platform for specialized intelligence, enabling companies to build, train, and serve AI models tailored to their own data, workflows, and... ...the world's most innovative AI products.This is a highly technical role requiring deep expertise in distributed systems, cloud-native...Training
$220k
...platform are paid over $4 million per day to train frontier AI models. Mercor's APEX... ...systems Comfort with core ML/LLM concepts - post training, evaluation metrics, reward modeling... ..., fast-moving environments with sound technical judgment Proven ability to coach...TrainingWork at officeRelocation package- ...ship, and maintain the software yourself and through the other technical staff you coordinate on-site and at HQ. Write production code... ...so the two evolve in lockstep, including the cultural work: training, change management, work instructions, quiet one-on-ones with...TrainingWork at office
- About Us: Fireworks is the platform for specialized intelligence, enabling companies to build, train, and serve AI models tailored to their own data, workflows, and products. Founded by the team behind PyTorch and backed by AMD, Atreides, Benchmark Capital, Index Ventures...TrainingShift work
- ...Job Description Job Description The role As a Member of Technical Staff focused on Applied AI, you'll own our AI stack end to end. One framing... .... Evaluations that work for financial services. Most evals are built for tasks with a single right answer. Financial...Work at office1 day per week
- ...R&D, integration testing, production assembly, and field repair. About the role As a Machine Learning Engineer focused on Data & Training Infrastructure, you will join Arena's Platform team to build the data generation and training substrate behind the first electromagnetic...TrainingLocal areaRelocation package
$140k - $270k
...resonates, mention it in your application. About The Role Members of Technical Staff at Anterior own problems end-to-end — from system design through... ...to frontend — to ship complete solutions Writing evals to identify weaknesses and drive improvements in our LLM-powered...ApprenticeshipFlexible hours- Member of Technical Staff - Applied AI Engineer Valthos | Posted Mar 3 Full-time Negotiable Advanced (5-10 yrs) Valthos Inc. Valthos is an applied biological intelligence... ...modeling approaches—including adapting and post-training biological frontier models—for tasks in...TrainingFull timeWork at office
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Member of Technical Staff (Evals & Post-Training). Be the first to apply!
- mri tech aide New York, NY
- salesforce technical analyst New York, NY
- lead technical specialist New York, NY
- service desk assistant New York, NY
- end user support technician New York, NY
- operations support technician New York, NY
- help desk technical support New York, NY
- technical assistant New York, NY
- support analyst New York, NY
- technical associate New York, NY


