AI Evaluation Engineer - Reinforcement Learning & Agents
MaxIT Consulting - Max Corporate Group
Job Description
Job Description
San Francisco, California | Primarily On-site
We are seeking an AI Evaluation Engineer – Reinforcement Learning & Agents to build the environments, evaluation systems, and supporting infrastructure used to train and assess long-horizon enterprise AI agents.
The OpportunityYou will work on the engineering and research problems behind realistic agent environments, post-training systems, and reliable evaluation of complex multi-step workflows.
Key Responsibilities- Design evaluation environments for long-horizon enterprise agent workflows.
- Define tasks, state, tools, graders, and reward signals used to evaluate and improve agents.
- Build high-fidelity representations of complex enterprise software environments.
- Develop infrastructure for rollouts, orchestration, trajectory inspection, and grader pipelines.
- Measure both correctness and efficiency across multi-step agent behavior.
- Investigate evaluation failures, reward-quality issues, and agent behavior.
- Build production-quality systems rather than notebook-only research prototypes.
- Hands-on experience with AI environments, evaluations, reinforcement learning infrastructure, or related agent-training systems.
- Strong software engineering fundamentals.
- Demonstrated ability to build and ship technical infrastructure.
- Understanding of evaluation methodology, reward design, graders, and agent trajectories.
- Ability to work across languages and technology stacks based on system requirements.
A PhD is not required. Strong engineering and shipped environment or evaluation systems are more important than academic credentials or publication history.
SeniorityThe opportunity is open to exceptional new graduates, early-career engineers, and experienced senior candidates. Selection is based primarily on engineering strength and relevant technical work.
Work ArrangementThe role is anchored in San Francisco with a strong preference for in-person collaboration. Limited flexibility may be considered case by case for exceptional candidates.
$250.8k - $286.2k
AI Engineer 5 (AI Foundations: LLM Customization, Finetuning, Reinforcement Learning) At Capital One, we are creating responsible and reliable... ...language model inference, agents and multi-agent workflows,... ...search, guardrails, model evaluation, experimentation, governance...SuggestedFull timePart timeLocal area- ...software development with AI-powered formal... ...developed groundbreaking agents that provide mathematical... ...Join our team as an AI Engineer and help us push the boundaries... ...and models Evaluate reasoning approaches, including... ...to accelerate machine learning development Derive...SuggestedContract work
- ...Description Senior Software Engineer Job Type: Contractor (~15... ...Software Engineers to support an AI training project by creating reinforcement learning environments that evaluate AI models on complex... ...reference solutions. Evaluate AI agents' ability to reason through...SuggestedRemote jobFor contractors
- ...role sits at the intersection of product engineering and AI infrastructure, building the tools and systems that power a reinforcement learning data platform used by frontier AI labs... ...services, enabling partners to create, evaluate, and iterate on RL training data. Your...SuggestedFull timeWork at officeRemote workVisa sponsorshipRelocation package
$220k - $405k
...Perplexity is seeking energetic engineers to join our highly driven Agents engineering team. The... ..., full-stack, and AI/ML engineers who... ...reliability, code quality, AI evaluation, testing, and... ...models Post-training and reinforcement learning (particularly for multimodal...SuggestedFull timeFlexible hours$275k - $305k
...role, level, and location. Learn more about ourTotal Rewards philosophy.AI is a fundamental part of... ...Gusto builds, deploys, evaluates, and scales AI/ML... ...spanning Machine Learning Engineering, ML Platform, Risk Data... ...retrieval, evaluation, agents, and observability should...Full timeWork at officeLocal areaRemote work2 days per week3 days per week- ...AI Research Engineer (Robot Learning) San Francisco AI & Software In office Full-time mimic is an early stage deep tech robotics & AI... ...specification, data collection, model training and model evaluation in the real world. As an early employee, you will...Full timeWork at officeImmediate start
- ...Origin is building Physical AI for the built world - starting... ...allows our modular robots to learn, adapt and work in unstructured... ...Role Our system runs a Multi Agent Action Expert architecture: classical... ...detection. Design offline evaluation metrics that predict real-...
- ...training data and evals for frontier AI agents, as well as a marketplace to sell... ...re looking for a Full-Stack Software Engineer, Reinforcement Learning to build the product surfaces, backend... ...task authoring, data collection, evaluation, QA/QC, and RL training infrastructure...Full timeWork at officeRemote workRelocationVisa sponsorship
- ...at the intersection of AI, biology, chemistry, and large-scale engineering. Our goal is to translate... ...carefully designed learning systems that can scale... ...Design, train, and evaluate large-scale models, including... ...Have Experience with reinforcement learning, fine-tuning,...Full timeRemote workFlexible hours
$250k - $325k
...frontier of applied AI At Artisan, we're... .... You'll work across agent behavior, complex multi... ...tool use, context, and evaluation, taking promising... ...closely with our existing engineers and leadership. You... ..., skill acquisition, reinforcement learning or post-training,...Full timeWork at officeVisa sponsorship- ...intelligent buying in the age of AI. With 200M+ combined annual... ...for each other’s successes, learn from our mistakes, and... ...of Responsibilities: G2 evaluates AI agents on real tasks and publishes... ...product, data science, and engineering. Detailed Responsibilities...Internship
$179.4k - $245.6k
Machine Learning Engineer 4 (Python, AWS, SQL, GenAI) (Enterprise Platforms... ...and pioneering in the AI and technology space? Do you... ...supervised, and unsupervised, reinforcement learning, etc.) model types... ...regularization), and how to evaluate model accuracy and diagnose...Full timePart timeH1bLocal area$264.8k - $331k
Scale AI is the data foundation for AI, helping... ....About the General Agents TeamThe General Agents... ...Senior/Staff Machine Learning Engineer (MLE) on the General... ...and system design to evaluation, deployment, and iteration... ...fine-tuning (SFT), reinforcement learning with...Full time$250.8k - $286.2k
Sr. Manager, AI Engineer (Gen AI Platform Services: Agentic AI, Guardrails, Evaluation) At Capital One, we are creating responsible and reliable AI systems, changing... ...has been an industry leader in using machine learning to create real-time, personalized customer experiences...Full timePart timeLocal area$150k - $250k
About Distyl AI Distyl is an applied AI technology company partnering... ..., we build AI systems using Evaluation-Driven Development—an approach... ...in production.AI Evaluation Engineers focus on designing and... ...results inform prompt design, agent logic, model selection, and release...Work at office3 days per week$242.25k - $285k
...connected platform and AI-native ecosystem for... ...’ll join Carta’s ML Engineering team, embedded in Carta... ...around autonomous AI agents, specialized legal... ...from post-training and evaluation through model serving... ...optimization, reinforcement learning, and related methods,...Full timeContract work- ...a highly technical Forward Deployed AI Engineer – Agents & Research Systems to build and deploy... ...pipelines, post-training workflows, and evaluation infrastructure required for production... ...to work across technologies and learn new stacks as required by the problem....
$55k - $151.47k
...data and analytics engineering focus on... ...science and machine learning engineering at PwC... ...recommendations. Uphold and reinforce professional and... ...and implement AI solutions that enhance... ...experience with agent-based systems... ...team members. We evaluate these factors thoughtfully...H1b- ...sits within Machine Learning Platform and... ...powered products, agents, automation, and personalization... ...for Generative AI at DoorDash,... ...and inference engines, fine-tuning and training... ...such as reinforcement learning (RLHF/RLVR... ...recruiting team by evaluating job related qualifications...Hourly payWork at officeLocal areaRemote workFlexible hours
$123.98k - $201.45k
...connectivity, autonomy, and AI, we deliver solutions... ...that continuously learn, improve, and deliver... ...are seeking exceptional engineers who thrive at the... ...planning, and action. Evaluate emerging technologies... ...background in deep learning, reinforcement learning, imitation...Full timePart timeRelocationRelocation packageFlexible hours- ...Senior AI Engineer New York City or San Francisco |... ...how AI systems reason, learn, retrieve evidence, and... ...retrieval algorithms, and evaluation infrastructure that... ...in which specialized agents and models work together... ...tuning, distillation, reinforcement learning, preference...Full time
$73.5k - $212.28k
...in data and analytics engineering focus on leveraging advanced... ...science and machine learning engineering at PwC... ...appropriate.Uphold and reinforce professional and... ...teams to incorporate AI into various applications... ...with team members. We evaluate these factors thoughtfully...Full timeH1b$73.5k - $212.28k
...in data and analytics engineering focus on leveraging advanced... ...science and machine learning engineering at PwC... ...appropriate. Uphold and reinforce professional and... ...development of innovative AI solutions that drive... ...with team members. We evaluate these factors...H1b- ...building the leading AI-native platform... ...Senior Applied AI Engineer to design and build... ...you may build an agent that collects clinical... ...should be evaluated, where deterministic... ...engineering, machine learning, or Applied AI... ...tuning, post-training, reinforcement learning,...Full timeWork experience placementRelocation
$250k
...partnering with a high-growth AI company building the next generation... ...looking for a strong software engineer who has hands-on experience building and shipping AI agents into production. What You'll... ...infrastructure. Debug and evaluate agent behaviour across real-...$120k - $220k
...Liberate builds AI agents to automate manual tasks for the $2.7T insurance industry. We started... ...The Role We are seeking an AI Agent Engineer to bridge the gap between our cutting‑... ...Nice to have Familiarity with AI, machine learning, and data analytics technologies and how...Work at officeFlexible hours2 days per week$190.9k - $334.1k
...DescriptionIt all started when engineer Fred Luddy wrote code that... ...work. Today, ServiceNow is the AI control tower for business reinvention... ...ways. Develop novel AI Agent frameworks, architectural patterns... ...reusable templates and share learnings broadly across the team....Work at officeImmediate startRemote workFlexible hoursDay shift$225k - $320k
...unprecedented speed and accuracy. Our AI-enabled platform turns... .... As the Applied AI Engineering Lead, you will own how artificial... ...intelligence is applied, evaluated, and operationalized across... ...fine tuning strategies and reinforcement learning to improve decision...Local area$187.9k - $252k
...global organization of engineers, product developers,... ...applied science and machine learning engineering team that... ...approaches, evaluation methodology, and how data... ...experience (RecSys, ML, AI/LLM) who can help bridge... ...scaleExperience with reinforcement learning or related sequential...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Evaluation Engineer - Reinforcement Learning & Agents. Be the first to apply!
- ai engineer San Francisco, CA
- ai engineer remote San Francisco, CA
- ai prompt engineer San Francisco, CA
- ai developer San Francisco, CA
- senior ai engineer San Francisco, CA
- machine learning ai engineer San Francisco, CA
- ai ml engineer San Francisco, CA
- tsa agent San Francisco, CA
- operations agent San Francisco, CA
- special agent San Francisco, CA



