Research Engineer -- AI Alignment & Evaluation
W3 Sourcing
Job Description
Job Description
Research Engineer — AI Alignment & Evaluation
AI Safety / Research Engineering | San Francisco, CA | Hybrid / In-Person
About the CompanyWe are representing a high-growth AI research organization working at the intersection of frontier model evaluation, AI safety, and security.
The team develops sophisticated evaluation environments designed to surface undesirable or misaligned model behavior and help leading AI organizations better understand how advanced systems behave under complex, long-horizon conditions.
This is a technically rigorous environment for engineers who are interested in AI alignment, agent behavior, model evaluation, and building systems that help make increasingly capable AI more reliable and controllable.
The RoleThis is an opportunity to join a small, highly technical team as a Research Engineer with significant end-to-end ownership.
You will independently design and build evaluation environments that test frontier AI systems for subtle forms of undesirable behavior. You will own the full lifecycle of each environment, from initial concept and failure-mode identification through implementation, grader development, testing, measurement, and refinement.
A significant part of the role involves working directly with advanced LLM agents: prompting them to perform technical tasks, reviewing their output, identifying subtle errors, and making judgment calls where current models still fall short.
The role is ideal for a strong software engineer or technical researcher who enjoys ambiguous problems, learns new domains quickly, and is deeply interested in AI alignment and security.
What You'll Do- Design and build complex evaluation environments for frontier AI models.
- Own evaluation projects end to end, including ideation, implementation, testing, grading, measurement, and iteration.
- Investigate potential model failure modes and identify ways advanced agents may exploit or circumvent intended constraints.
- Develop and improve software infrastructure used to isolate, reproduce, and evaluate model behavior.
- Work extensively with LLM-based agents to accelerate implementation and research workflows.
- Review agent-generated work critically and identify subtle technical or conceptual errors.
- Build long-horizon tasks that operate near the edge of current model capabilities.
- Apply strong qualitative judgment when evaluating behavior that cannot be captured through simple automated metrics.
- Rapidly learn unfamiliar technical domains as required by individual evaluation environments.
- Share findings, lessons, and technical context with a highly collaborative research and engineering team.
- 1+ years of experience in software engineering, machine learning engineering, technical research, or a closely related field.
- Strong traditional software engineering fundamentals.
- Proficiency with Python .
- Strong interest in AI alignment, AI safety, or AI security.
- Ability to reason carefully about complex systems and ambiguous failure modes.
- Strong conceptual judgment and the ability to think through how an autonomous agent may interpret or exploit a task.
- Ability to learn new technical domains quickly.
- Experience using LLMs or AI agents effectively as part of technical workflows.
- Strong ability to assess whether agent-generated work is correct, including when errors are subtle.
- Comfortable taking full ownership of technically demanding projects with limited oversight.
- High standards for quality, execution, and accountability.
- Experience building evaluation frameworks, benchmarks, simulation environments, or agent-based systems.
- Exposure to frontier language models or autonomous agent workflows.
- Background in AI safety, alignment research, adversarial testing, or security.
- Experience designing tasks that require multi-step or long-horizon reasoning.
- Research experience involving model behavior, reward hacking, robustness, or control mechanisms.
- Own technically challenging research environments from concept through final evaluation.
- Work directly with state-of-the-art AI systems and agentic workflows.
- Tackle problems at the frontier of AI safety, model behavior, and alignment.
- Join a small technical team where individual work has meaningful visibility and impact.
- Operate with substantial autonomy while receiving frequent technical feedback.
- Build expertise across a wide range of domains rather than working within a narrow product surface.
- Contribute to work focused on understanding and mitigating undesirable AI behavior rather than simply increasing model capabilities.
- Full-time position.
- San Francisco-based role with regular in-office collaboration expected.
- Flexibility around hybrid working arrangements.
- Open to candidates willing to relocate.
- Visa transfers and new visa sponsorship may be available.
- Work is highly ownership-driven, with emphasis on the quality of what you ship.
Confidential details removed: salary, client name, founder names, exact address, company links, investor names, funding details, exact team size, founding year, and highly identifiable wording.
$136.44k - $265.11k
...are rebuilding biotech for the AI era.When a breakthrough is... ...here.You’ll build the datasets, evaluations, and systems that help close... ...the intersection of software engineering, biology, and frontier AI:... ...engineers, scientists, and external research partners.Desire to work in a...SuggestedWork at officeLocal areaMonday to FridayShift work$174k - $252k
Research new alignment methods, studying alignment failures, and applying AGI-scalable alignment techniques... ...techniques to understand what AI systems are thinking.Work with product... ...Computer Science, a related Software Engineering field, or equivalent practical experience...Suggested$110.7k - $379.2k
Position Summary Research Engineer — Post-Training & Small Language Models (SLMs), Healthcare AI Three hundred fifty million Americans rely on a healthcare system... ...-training team, you will design, train, evaluate, and align the models that reason about healthcare —...SuggestedLocal areaVisa sponsorship$250k - $350k
...is to develop reliable AI systems for the world’... ..., combining rigorous evaluation with full-stack deployment... ...with applied ML research, design, and evaluation... ...Machine Learning Research Engineer, you will operate... ...training methods, LLM alignment, or applied MLExperience...SuggestedFull time$164.6k - $313.3k
...OpportunityAdobe's Sound Design AI group (SODA) is looking for a driven Data/ML engineer to push the boundaries of audio... ...small, collaborative and efficient research team looking for highly... ...finetune models on pipeline outputs, evaluate their behavior, and use those findings...SuggestedFull timeTemporary workLocal areaWorldwide- Factory is seeking innovative Research Engineers to design and integrate advanced AI and ML capabilities that revolutionize productivity and accelerate innovation... ..., focusing on retrieval systems, code generation evaluation, agentic user experience development, and agent...Work at office
$150k - $250k
About Distyl AI Distyl is an applied AI technology company partnering... ...social organizations.We research and deploy technologies that... ...ForAt Distyl, Research Engineers build the bridge between frontier... ...understand system behavior, build evaluation frameworks, and collaborate...Work at office3 days per week$197.3k - $313.7k
...SalesforceSalesforce is the #1 AI CRM, where humans with agents... ...software and platform engineers to embed in our AI team to bridge... ...directly enable world-class research and products used by millions... ...services, like APIs, UIs, agentic evaluators, and more.Roll out scalable...Full time- ...iterate on, and innovate on the AI brains behind The Path’s AI Therapist. Combine research, data science, and engineering to create models, orchestration, and evaluation systems that make therapy... ...ideas. # Improve Safety, Alignment, and Clinical Guardrails Work...
$150k - $250k
About Distyl AI Distyl is an applied AI technology company partnering... ...social organizations.We research and deploy technologies that... ...ForAt Distyl, Research Engineers build the bridge between frontier... ...understand system behavior, build evaluation frameworks, and collaborate...Work at office3 days per week$175k - $250k
...Research Engineer About Scorecard We’re a small, nimble team backed by top-tier investors... ...engineers to help shape the future of AI development. This role is part of the... ...resulting in unrealistic users that corrupt evaluation results. To fix this, we’re building...Work at office- ...security lab. We protect the world as AI systems grow more capable and more... ...wrong hands. We build high-fidelity research infrastructure that evaluates and strengthens frontier models on critical... .... About the role As a Research Engineer, you will be responsible for post-...
- ...OpenAI is seeking an exceptional health AI researcher to build frontier capabilities that translate into real-world impact... ...contribute to pretraining, RL/post-training, and evaluation, working with researchers, engineers, clinicians, and product teams to deliver scalable...
- ...the Data Foundation & AI team within Plaid’s Data... ...production serving, evaluation, and monitoring. As part... ...Machine Learning Engineer, you will lead the technical... ...that translate research into production impact... ...Ability to drive technical alignment across teams: setting...Full timeWork experience placementLocal areaImmediate start
$122k - $215k
...Description Job Description Waabi, founded by AI visionary Raquel Urtasun, is the leader... ...way. To learn more visit: As a Research Engineer, you will be at the forefront of... ...apply. You will... - Prototype, evaluate, and iterate on solutions, using real-world...Full timeWork at officeWork from homeFlexible hours- ...Francisco, California. The Role: As a Research Engineer - Brain Computer Interface Models ,... ...-scale EEG model training runs and evaluation Architecture and training methodology... ...all enjoy what we do and love discussing AI Benefits and Perks: Comprehensive...Work at officeRelocation package
- ...for the demands of advanced AI. AI for Chips connects these... ...to the work of semiconductor engineering.Our goal is to help engineers... ...design cycles. This work brings research, model training, and hardware... ...learning, tool use, and evaluation.You’ll own experiments from the...
- Get AI-powered advice on this job and more exclusive features... ...Agentic AI that empowers software engineers by automating production... ...workflows end‑to‑end, balancing research and engineering to create... ...unstructured data for training and evaluation Design and execute...Full timeWork at officeVisa sponsorshipFlexible hours
- ...Description We are Genmo, a research lab dedicated to building open... ...us in shaping the future of AI and pushing the boundaries of... ...seeking an exceptional Software Engineer to join our research team in... ...Employer. Candidates are evaluated without regard to age, race,...Work at office
- ...Francisco, California. The Role: As a Research Engineer - Audio & Speech Models , you will be... ...dataset collection, processing, and evaluation Architecture and training methodology... ...enjoy what we do and love discussing AI Benefits and Perks: Comprehensive...Work at officeRelocation package
$140k - $160k
...OVERVIEW The Experimental Engineering team leads the engineering R&... ...seeking a highly capable Senior Research Engineer to join our team and... ...innovative applications of AI, machine learning, and... ...architectural decisions on retrieval, evaluation, and orchestration...Full timeLocal areaRemote workRelocation package- ...multi-agent systems, harnesses, models, and evaluations for code validation. Study and... ...approaches into production systems. Track research in LLMs, information retrieval, and developer... ...assistance. Opportunity to work on AI-powered code review used by thousands of...Full timeWork at officeRelocation package
- ...world reasoning. Build and operate end-to-end LLM evaluation systems, including runs, scoring, dashboards, and... ...augmentation, and curation. Collaborate with AI researchers, applied AI teams, and data producers to align evaluations with training objectives. Own...Full timeWork at officeRelocation package
- ...reliable, interpretable, and steerable AI systems. We want AI to be safe and... ...a quickly growing group of committed researchers, engineers, policy experts, and business leaders... ...of the Anthropic Institute . We design evaluations of AI R&D capabilities, build the internal...Full time
- ...Job Description Job Description We are looking for a hybrid Systems Engineer and AI Researcher to lead the development of our agent evaluation framework and post-training data pipelines. You will design sandboxed execution environments, high-throughput reinforcement...Work at office
$180k - $240k
Our client, a venture-backed AI Startup, is hiring a talented ML/AI Research Engineer to join their team in San Francisco... ...will lead the design, training, evaluation and optimization of agent-native... ...observability, drift detection and alignment strategies across production...Full time$200k - $330k
...Machine Learning Research Engineer Emeryville, California, United States; Hybrid (2-3 days... ...-site) Profluent is the frontier AI lab for biology. Profluent builds powerful... ...for automated model fine-tuning, alignment and evaluation Design and implement modular, easy...$200k - $350k
...re partnering with a frontier AI startup We’re working with an... ...building at the intersection of research, product, and creativity . The... ...a Machine Learning Research Engineer , you’ll own end-to-end research... ...experiments that help models evaluate style and subjective quality....- ...Francisco, California. The Role: As a Research Engineer - Language Model Pre-Training , you'll... ...Dataset collection, processing, and evaluation Architecture and methodology... ...all enjoy what we do and love discussing AI Benefits and Perks: Comprehensive...Work at officeRelocation package
- ...skilled Cyber Security Web Application Research Engineer to join our Technology Cybersecurity department... ...a strong knowledge and experience with AI enhanced scanners and Tools.In this role... ...considerationsEnsure AI usage aligns with security, compliance, privacy, and...Full timeWork experience placementWork at officeFree visa2 days per week3 days per week
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Research Engineer -- AI Alignment & Evaluation. Be the first to apply!
- research software engineer San Francisco, CA
- research engineer San Francisco, CA
- deep learning research engineer San Francisco, CA
- research programmer San Francisco, CA
- research statistician San Francisco, CA
- climate research San Francisco, CA
- research and development San Francisco, CA
- outcomes research San Francisco, CA
- vice president research San Francisco, CA
- biochemistry research San Francisco, CA




