Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Research Engineer — Reinforcement Learning

$180k - $290k

Firecrawl

Research Engineer — Reinforcement Learning You'll bring reinforcement learning to Firecrawl's core product — building the training infrastructure, reward pipelines, and fine-tuning systems that make our models meaningfully better at extracting, understanding, and structuring web data. This isn't theoretical RL research. You'll build your own training infra, run fast experiments, ship models to production, and bridge the gap between classical RL approaches and modern LLM agent systems. If you care as much about training throughput as you do about reward design, this is the role. Salary Range: $180,000–$290,000/year (Range shown is for U.S.-based employees. Compensation outside the U.S. is adjusted fairly based on your country's cost of living. You can explore how we calculate this here: Equity Range: Up to 0.15% Location: San Francisco, CA or Remote (Americas, UTC-3 to UTC-10) Job Type: Full-Time Experience: 3+ years in applied RL, ML engineering, or model training — with production systems Visa: US Citizenship/Visa required for SF; N/A for Remote About Firecrawl Firecrawl is the easiest way to extract data from the web. Developers use us to reliably convert URLs into LLM-ready markdown or structured data with a single API call. In just a year, we've hit 8 figures in ARR and 100k+ GitHub stars by building the fastest way for developers to get LLM-ready data. We're a small, fast-moving, technical team building essential infrastructure superintelligence will use to gather data on the web. We ship fast and deep. What You'll Do Build training infrastructure and reward pipelines from scratch. Design and operate the systems that train and evaluate Firecrawl's models. You'll own the full loop — data collection, reward modeling, training runs, evaluation, and deployment. You build the infra yourself because you're the one who needs it to work. Fine-tune models to achieve state-of-the-art results. Take foundation models and make them dramatically better at web data extraction, content understanding, and structured output generation. You know how to get from "decent fine-tune" to "best-in-class" and you have the patience and rigor to close that gap. Bridge LLM agents and classical RL. The most interesting problems at Firecrawl sit at the intersection of modern LLM-based agents and classical RL techniques. You'll design reward signals for agent behaviors, apply RL methods to improve multi-step agent workflows, and figure out where traditional RL approaches outperform prompting — and vice versa. Run fast experiments and iterate. You design experiments that test meaningful hypotheses, run them quickly, and make decisions based on results. You don't spend weeks on experiment infrastructure before getting a single result. Speed of iteration is a core part of how you work. Communicate clearly to non-RL people. RL can be opaque. You translate your work into language that engineers, product people, and leadership can understand and act on. You know how to explain why a reward function matters without requiring everyone to read the paper. Collaborate closely with the team. Work directly with the Search/IR-focused Research Engineer and the engineering team to connect RL improvements with search, ranking, and the broader product roadmap. What We're Looking For Builds their own training infra and reward pipelines. You don't wait for an ML platform team to set things up. You build the training loops, reward models, data pipelines, and evaluation frameworks yourself — because you understand that infra choices directly affect the quality of results. You've operated GPU clusters, managed training runs, and debugged convergence issues in production. Can fine-tune models to SOTA. You've taken models from baseline to best-in-class on tasks that matter. You understand the full fine-tuning lifecycle — data curation, training dynamics, hyperparameter sensitivity, evaluation methodology — and you have the taste to know when a model is actually good versus when the eval is flattering. Bridges LLM agents and classical RL. You're fluent in both worlds. You understand PPO, RLHF, reward modeling, and policy optimization — and you understand how modern LLM agents work, where they fail, and how RL techniques make them better. You see connections between these domains that most people miss. Production-minded. You care about whether your models work in production, not just on benchmarks. You've deployed models that serve real traffic and made hard tradeoffs between model quality, latency, and cost. Research that doesn't ship isn't research that matters here. Runs fast experiments and communicates clearly. You'd rather run three rough experiments this week than one polished one next month. When you have results, anyone on the team can understand what they mean — no decoder ring required. Backgrounds that tend to do well: RL engineers at AI labs or applied ML teams who've shipped models to production. Researchers who've done RLHF or reward modeling for LLM systems. ML engineers who've built training infrastructure at startups and cared as much about the pipeline as the model. People who've worked at the intersection of RL and language models — whether in academic labs with a production bent or at companies building agent systems. What We're NOT Looking For Pure theorists. If your best RL work lives in a paper and you've never trained a model on real data at real scale, this isn't the role. We need someone who builds and ships. Researchers who need a platform team. If you expect training infrastructure, data pipelines, and evaluation frameworks to be set up before you can be productive, you'll be frustrated here. You build the tools you need. People who only know one paradigm. Deep in classical RL but never worked with LLMs? LLM fine-tuner who's never touched RL? You'll be missing half the picture. This role requires fluency in both. Slow iterators. If your standard experiment cycle is measured in weeks, not days, you'll struggle with the pace. We need someone who can run a meaningful experiment, interpret results, and decide next steps within a day or two. Black-box communicators. If your typical update is a wall of metrics only another RL researcher can parse, this isn't the right fit. We need someone who can explain what's working, what's not, and why it matters — to people without RL PhDs. A Note On Pace We operate at an absurd level of urgency because the window for what we're building won't stay open forever. If that excites you, keep reading. If it doesn't, no hard feelings — but this role probably isn't for you. Benefits & Perks Available to all employees Salary that makes sense — $180,000–$290,000/year, based on impact, not tenure Own a piece — Up to 0.15% equity in what you're helping build Generous PTO — 15 days mandatory, anything after 24 days, just ask (holidays excluded); take the time you need to recharge Parental leave — 12 weeks fully paid, for moms and dads Wellness stipend — $100/month for the gym, therapy, massages, or whatever keeps you human Learning & Development — Expense up to $1,000/year toward anything that helps you grow professionally Team offsites — A change of scenery, minus the trust falls Sabbatical — 3 paid months off after 4 years, do something fun and new Available to US-based full-time employees Full coverage, no red tape — Medical, dental, and vision (100% for employees, 50% for spouse/kids) — no weird loopholes, just care that works Life & Disability insurance — Employer-paid short-term disability, long-term disability, and life insurance — coverage for life's curveballs Supplemental options — Optional accident, critical illness, hospital indemnity, and voluntary life insurance for extra peace of mind Doctegrity telehealth — Talk to a doctor from your couch 401(k) plan — Retirement might be a ways off, but future-you will thank you Pre-tax benefits — Access to FSAs and commuter benefits (US-only) to help your wallet out a bit Pet insurance — Because fur babies are family too Available to SF-based employees SF HQ perks — Snacks, drinks, team lunches, intense ping pong, and peak startup energy E-Bike transportation — A loaner electric bike to get you around the city, on us Interview Process Application Review — Send us your work and a quick note on why this excites you. Show us what you've trained — models, reward systems, training pipelines. Published work is great; shipped production models are better. Intro Chat (~20 min) - A quick conversation to get to know each other before we go deep. We'll talk about what you've been working on, what drew you to Firecrawl, and what you're looking for in your next role. Time for your questions too. Technical Deep Dive (~60 min) — Go deep on RL and model training work you've done: training infrastructure decisions, reward design, fine-tuning approaches, production deployment. We'll explore a live problem — how you'd apply RL to improve an LLM agent workflow at Firecrawl. We're looking for depth across classical RL and modern LLM techniques, production instincts, and fast reasoning. Founder Chat (~30 min) — Culture, pace, ownership, and how you like to work. Time for your questions too. Paid Work Trial (1–2 weeks) — Tackle a real RL/fine-tuning problem with production implications. We evaluate on technical depth, experiment velocity, and how clearly you communicate results. Decision — We move fast after the trial. If you want to bring RL to one of the most interesting applied problems in AI — making agents smarter at understanding and extracting web data at scale — this is your shot. Apply now. #J-18808-Ljbffr Firecrawl

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Research Engineer — Reinforcement Learning in San Francisco, CA vacancy
  •  ...and society. About the RL Teams The Reinforcement Learning teams are critical to advancing our AI...  ...reinforcement learning Pioneering fundamental RL research for large language models Building...  ...The Role We are hiring a Research Engineer for the Code RL team. You will design... 
    Suggested
    Work at office
    Visa sponsorship

    Anthropic

    San Francisco, CA
    1 day ago
  • $350k

     ...whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. About the RL Teams Our Reinforcement Learning teams lead Anthropic's reinforcement learning research and... 
    Suggested
    Work at office
    Visa sponsorship
    Flexible hours

    job-boards.greenhouse.io- JobBoard

    San Francisco, CA
    1 day ago
  • $300k

     ...can truly reason and act? This team are creating complex reinforcement learning environments — simulations where advanced agents learn to plan...  ...” looks like for intelligent behaviour. It’s open-ended, research-driven work where the task definition, data, and reward structure... 
    Suggested

    Techire Ai

    San Francisco, CA
    1 day ago
  • About the Role You'll bring reinforcement learning to Firecrawl's core product — building the training...  ...web data. This isn't theoretical RL research. You'll build your own training infra,...  ...your work into language that engineers, product people, and leadership can understand... 
    Suggested

    Gravity Engineering Services Pvt Ltd.

    San Francisco, CA
    1 day ago
  • $300k - $405k

     ...whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to...  .... About Horizons The Horizons team leads Anthropic's reinforcement learning (RL) research and development, playing a critical role in... 
    Suggested
    Work at office
    Visa sponsorship
    Flexible hours

    Anthropic

    San Francisco, CA
    1 day ago
  •  ...whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. About the RL Teams Our Reinforcement Learning teams play a critical role in advancing our AI systems.... 

    Anthropic

    San Francisco, CA
    1 day ago
  • $300k - $405k

     ...hiring for the Cybersecurity RL team within Horizons. As a Research Engineer, you'll help to safely advance the capabilities of our models...  ...experience in cybersecurity research. Have experience with machine learning. Have strong software engineering skills. Can balance... 
    Full time
    Work at office
    Visa sponsorship
    Flexible hours

    Anthropic

    San Francisco, CA
    1 day ago
  • About the Role As a Research Engineer in our Reasoning team, you'll play a crucial role in shaping our technological direction, focusing on our test-time compute scaling research ideas. This role is ideal for individuals who enjoy working with synthetic data and teaching... 
    Worldwide

    Gravity Engineering Services Pvt Ltd.

    San Francisco, CA
    1 day ago
  • $192.6k - $344.85k

    ## AI Research Manager/Scientist, Reinforcement LearningApplylocations: San Francisco, CA, USA: AMER - United...  ...AI Scientist Manager Reinforcement Learning** at Autodesk Research, you will be...  ...Architecture, Civil or Mechanical Engineering, Construction, Manufacturing,... 
    Remote work

    Autodesk, Inc.

    San Francisco, CA
    1 day ago
  • $264.8k - $331k

     ...About the Role As a Senior/Staff Machine Learning Engineer (MLE) on the General Agents team, you'...  ...ambiguous problem spaces, balancing research‑driven approaches with pragmatic...  ...methods like supervised fine‑tuning (SFT), reinforcement learning with verifiable rewards (RLVR... 
    Full time

    Scale AI, Inc.

    San Francisco, CA
    1 day ago
  • $200k - $350k

     ...stage AI company that’s redefining how models learn to understand subjective quality , from...  ...—building at the intersection of research, product, and creativity . The Role As a Machine Learning Research Engineer , you’ll own end-to-end research cycles—designing... 

    Coders Connect

    San Francisco, CA
    2 days ago
  • $340k

     ...whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to...  ...About Horizons The Horizons team leads Anthropic's reinforcement learning research and development, playing a critical role in advancing... 
    Work at office
    Visa sponsorship
    Flexible hours

    Anthropic

    San Francisco, CA
    more than 2 months ago
  • $176k - $255k

     ...accelerate progress in GenAI research. We are looking for Research Scientists and Research Engineers with expertise in LLM post-training...  ...in Computer Science, Machine Learning, AI, or a related field....  ...of deep learning, reinforcement learning, and large-scale model... 
    Full time
    Shift work

    Scale AI

    San Francisco, CA
    more than 2 months ago
  • $264.8k - $331k

     ...around the world. The Enterprise ML Research Lab works on the front lines of this AI...  ...clients. As an ML Sys Research Engineer, you'll work on building out the algorithms...  ...vision coverage, retirement benefits, a learning and development stipend, and generous PTO... 
    Full time

    Scale AI

    San Francisco, CA
    7 hours ago
  • $208k - $260k

     ...This position will be a key contributor in conducting applied research in Robotics and developing ML pipelines for training and...  ...robotics, computer vision, embodied AI, sim-to-real, imitation learning, reinforcement learning, and vision language actions models  ~ PhD or... 
    Full time
    Shift work

    Scale AI

    San Francisco, CA
    more than 2 months ago
  • $200k - $350k

     ...training), second-time technical founders, engineers that made 100+ games for Voodoo,...  ...games & 3D environments. Our current research spans: Distributed multi-agent orchestration...  ...population, and behavior trees Reinforcement learning pipelines for adaptive, open-ended... 
    Visa sponsorship
    Relocation package

    ROAM

    San Francisco, CA
    1 day ago
  •  ...Research Engineer On Physical Ai Team Hedra is a pioneering generative modeling company — first models to market — now building...  ...and refine training methodologies, including fine-tuning, reinforcement learning, and large-scale multimodal learning Design and generate... 
    Work at office

    HEDRA INC

    San Francisco, CA
    5 days ago
  • $300k

     ...provider in San Francisco seeks an experienced professional to design and build large-scale reinforcement learning environments. The role involves collaborating with researchers to tackle complex problems, refining learning processes, and scaling infrastructure. Ideal candidates... 

    Jobleads-US

    San Francisco, CA
    4 days ago
  • $350k

     ...a quickly growing group of committed researchers, engineers, policy experts, and business leaders...  ...and run elegant and thorough machine learning experiments to help us understand...  ...our interventions. * Run multi-agent reinforcement learning experiments to test out techniques... 
    Work at office
    Visa sponsorship
    Flexible hours

    Anthropic

    San Francisco, CA
    5 days ago
  • $350k

     ...whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to...  ...domains or real-world use cases Have experience with reinforcement learning, reward design, or training data curation for LLMs Are... 
    For contractors
    Work at office
    Visa sponsorship
    Flexible hours

    job-boards.greenhouse.io- JobBoard

    San Francisco, CA
    1 day ago
  •  ...building Agentic AI that empowers software engineers by automating production engineering...  ...workflows end‑to‑end, balancing research and engineering to create production‑ready...  ...with novel techniques, including reinforcement learning, retrieval‑augmented generation, and autonomous... 
    Full time
    Work at office
    Visa sponsorship
    Flexible hours

    Resolve AI

    San Francisco, CA
    2 days ago
  • About Liquid Labs Research has been core to Liquid AI from the beginning...  ...research and practical engineering. We work hand-in-hand with...  ...Have experience in machine learning research or production-grade...  ...Science and Impact Liquid Labs reinforces our commitment to... 
    Immediate start
    Work from home

    Liquid AI

    San Francisco, CA
    1 day ago
  • $350k

     ...whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to...  ..., academia, or other projects Have experience with reinforcement learning, reward design, or training data curation for large language... 
    Visa sponsorship

    Anthropic

    San Francisco, CA
    1 day ago
  • $280k

     ...and run elegant and thorough machine learning experiments to help us understand and...  ...yourself as both a scientist and an engineer. As a Research Engineer on Alignment Science, you'll...  ...our interventions. Run multi‑agent reinforcement learning experiments to test out techniques... 
    Work at office
    Home office
    Relocation package

    Anthropic

    San Francisco, CA
    1 day ago
  • As an Applied Research Engineer , you'll design and build advanced systems to collect, analyze, and optimize human-in-the-loop data...  ...edge AI models. Your work will focus on techniques such as Reinforcement Learning from Human Feedback (RLHF), Direct Preference... 
    Work from home

    Sterling Inspired Staffing.

    San Francisco, CA
    4 days ago
  •  ...Meta and Google DeepMind. About the Role As an RL Engineer at BioStack, you will help build the reinforcement learning infrastructure for healthcare AI. BioStack is...  ..., diagnostic decision-making, and biomedical research tasks. We’re looking for someone with strong reinforcement... 
    Contract work

    BioStack

    San Francisco, CA
    1 day ago
  • Research Engineer, Post-Training (All Industry Levels) Join to apply for the Research Engineer, Post-Training (All Industry Levels...  ...infrastructure Strong understanding of modern machine learning techniques (reinforcement learning, transformers, etc) Track-record of... 

    Character.AI

    San Francisco, CA
    2 days ago
  • $200k - $275k

     ...Writer. By uniting applied AI research, flexible infrastructure,...  ...and help build the platform engineers turn to to ship AI products....  ...strong experience in machine learning and solid foundations in maths...  ..., whether that be through reinforcement learning, supervised finetuning... 
    Flexible hours

    Baseten

    San Francisco, CA
    1 day ago
  • Research Engineer - Interpretability Systems An AI research lab working at the frontier of interpretability, alignment, and reinforcement learning is hiring Research Engineers focused on understanding what’s happening inside large language models This role is for engineers... 

    Acceler8 Talent

    San Francisco, CA
    1 day ago
  • Fundamental AI Research Institute Come join one of the only research institutions globally with resources to compete with top...  ...lab is a playground for state-of-the-art research in LLMs, Reinforcement Learning and Agentic AI. Hiring for those experienced in LLM Post-Training... 

    Storm3

    San Francisco, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Research Engineer — Reinforcement Learning. Be the first to apply!