Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Reinforcement Learning Engineer

$176.4k - $242.55k
Full-time

Bugcrowd

Founded in 2012, Bugcrowd is the preemptive security platform that unifies exposure discovery and assessment, offensive testing, and intelligence shaped by AI and human insight to help organizations avoid, discover, and validate real-world risk. Bugcrowd helps security teams move faster by identifying the exposures that matter most so they can act first and stay ahead of attackers. By combining the power of humans and AI, teams can preempt attack paths and prevent breaches. Based in San Francisco and New Hampshire, Bugcrowd is supported by General Catalyst, Rally Ventures, Costanoa Ventures, and others. Visit

Job Summary

The Bugcrowd RL and Reasoning Team focuses on pushing the boundaries of autonomous cybersecurity by building authentic, verifiable reinforcement learning environments for world-leading foundational AI companies. As a Reinforcement Learning Engineer specializing in Reinforcement Learning from Verifiable Rewards (RLVR), you will design and scale automated verification pipelines that transform real-world software vulnerabilities into deterministic reward functions. In this role, you will bridge the gap between low-level security analysis and modern LLM reasoning models, engineering environments where AI agents learn to discover, exploit, and remediate software vulnerabilities with mathematical certainty. Instead of relying on subjective human feedback, your work directly powers the rigorous, verifiable reward signals that teach next-generation frontier AI models how to master complex cybersecurity domain logic. You will work at the intersection of fuzzing, dynamic program analysis, system exploitation, and scalable ML infrastructure to shape the safety and offensive/defensive capabilities of future artificial intelligence.

Essential Duties and Responsibilities

  • Design, build, and deploy high-throughput RLVR (Reinforcement Learning from Verifiable Rewards) environments that evaluate LLM action sequences against deterministic execution outcomes.
  • Develop automated test harnesses, sandboxes, and verification engines that convert complex vulnerability research (e.g., memory corruption, web security, logic bugs) into binary pass/fail reward signals.
  • Integrate Bugcrowd’s Mayhem automated analysis platform and real-world vulnerability feeds into continuous, scalable RL environment generation pipelines.
  • Architect safe, isolated, and highly reproducible execution environments (using Docker, BuildKit, or Nix) capable of running thousands of simultaneous agent-driven exploitation and patching trajectories.
  • Collaborate directly with researchers at frontier AI labs including Anthropic, OpenAI, and Cohere to define standard benchmark formats, observation spaces, and verifiable evaluation metrics for cybersecurity tasks.
  • Implement precise telemetry, ground-truth verification algorithms, and trajectory logging to analyze agent reasoning paths and prevent reward hacking or false positives.
  • Build low-level instrumentation and debugging tools to monitor memory states, process executions, and network behaviors during agent interaction cycles.
  • Optimize infrastructure performance and environment reset latency to support massive-scale parallel sampling and distributed RL training workflows.
  • Benchmark and evaluate frontier AI model performance across diverse offensive and defensive security challenges, such as automated fuzzing, exploit payload generation, and patch validation.

Education, Experience, Knowledge, Skills, and Abilities

  • Understanding of RL training workflows used by modern LLM systems, specifically execution-based feedback or Reinforcement Learning from Verifiable Rewards (RLVR).
  • Proficiency developing applications in Python and low-level systems programming in C, with Rust experience being a strong plus.
  • Solid understanding of software vulnerabilities, binary exploitation, fuzzing methodologies, or program analysis.
  • Experience with DevOps pipelines (e.g., GitHub Actions), reproducible builds (Docker, BuildKit, Nix), and comfort working with Linux systems and low-level debugging.
  • Experience working with or building benchmark environments (e.g., CTFs, SWE-bench, security challenges, or execution sandboxes).

Preferred Experience

  • Experience designing custom reward functions, ground-truth verifiers, or automated grading engines for AI safety and reasoning models.
  • Background in low-level program analysis tools, sanitizers (e.g., ASan/MSan), compiler instrumentation, or automated exploit generation tools.
  • Proven track record of participating in or developing competitive cybersecurity benchmarks, CTFs, or open-source AI evaluation frameworks.

Working Conditions and Physical Requirements

The ideal candidate must be able to complete all physical requirements of the job with or without reasonable accommodation.

Sitting and / or standing - Must be able to remain in a stationary position 50% of the time

Carrying and / or lifting - Must be able to carry / move laptop as needed throughout the work day.

Environment - remote, work-from-home 100% of the time.

Pay Range Disclosure

At Bugcrowd, we strive for fairness, equality and to create an environment that allows our people to perform at their very best. Our compensation philosophy is to foster a collaborative community that rewards, attracts and retains the best possible talent. The provided salary details are based on US national averages and we retain the flexibility to tailor to the needs of the business.

The national estimate for the current base range for the position of $176,400 - $242,550.

This position may also be eligible to participate in a discretionary bonus program or commission plan, subject to the rules governing the program, whereby an award, if any, depends on various factors, including, without limitation, individual and organizational performance.

Culture

  • At Bugcrowd, we understand that diversity in the workplace is vital to a company’s success and growth. We strive to make sure that people are included and have a sense of being part of making Bugcrowd not only a great product but a great place to work.
  • We regularly hear from both customers and researchers that Bugcrowd feels like a family, and we strive to maintain that internally as well.
  • Our team consists of a broad range of people: musicians, adventure sports junkies, nature lovers, parents, cereal enthusiasts, night owls, cyclists, artists—you get the point.

At Bugcrowd, we are solving security threats and vulnerabilities that are relevant to everyone, therefore we believe solving these problems takes all kinds of backgrounds. We value the perspectives and experiences people from underrepresented backgrounds bring.

Disclaimer

This position has access to highly confidential, sensitive information relating to the technologies of Bugcrowd. It is essential that the applicant possess the requisite integrity to maintain the information in the strictest confidence.

The company is authorized to obtain background checks for employment purposes under state and federal law. Background checks will be conducted for positions that involve access to confidential or proprietary information (including trade secrets).

Background checks may include Social Security verification, prior employment verification, personal and professional references, educational verification, and criminal history. Applicants with conviction histories will not be excluded from consideration to the extent required bylaw.

Any personal data you submit in connection with your application will be processed in compliance with Bugcrowd's Privacy Policy, which you may review here: .

Equal Employment Opportunity:

Bugcrowd is EOE, Disability/Age Employer.

Individuals seeking employment at Bugcrowd are considered without regards to race, color, religion, national origin, age, sex, marital status, ancestry, physical or mental disability, veteran status, gender identity, or sexual orientation.

Bugcrowd is committed to the full inclusion of all qualified individuals. In keeping with our commitment, Bugcrowd will take the steps to assure that people with disabilities are provided reasonable accommodations. Accordingly, if reasonable accommodation is required to fully participate in the job application or interview process, to perform the essential functions of the position, and/or to receive all other benefits and privileges of employment, please contact HR at ADA at bugcrowd.com.

Apply at:

Vacancy posted 12 days ago
Similar jobs that could be interesting for youBased on the Reinforcement Learning Engineer in Remote vacancy
  • $96k - $120k

     ...Reinforcement Learning Engineer – Remote Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States. This is a fantastic opportunity to join an established and... 
    Suggested
    Full time
    H1b
    Local area
    Immediate start
    Remote work
    Visa sponsorship

    Bright Vision Technologies

    Remote
    1 day ago
  • Codertal is hiring a Reinforcement Learning Engineer for a remote opportunity within the European Union on a B2B contract . If you’re passionate about designing intelligent agents that learn through interaction, optimize rewards, and shape autonomous decision‑making systems... 
    Suggested
    Daily paid
    Contract work
    Remote work
    Flexible hours

    Codertal

    Union, NJ
    3 days ago
  • $100k - $150k

     ...Reinforcement Learning Engineer - Remote Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States. This is a fantastic opportunity to join an established and... 
    Suggested
    Full time
    H1b
    Local area
    Immediate start
    Remote work
    Visa sponsorship

    Bright Vision Technologies

    Remote
    6 days ago
  • $168k - $247k

     ...Planning & Decision-Making group is investing heavily in deep reinforcement learning to move beyond classical planning, learning policies that...  ...in minutes.About the RoleAs a Senior/Staff Deep RL Engineer, you will design, train, and deploy deep reinforcement learning... 
    Suggested
    Hourly pay
    Work at office
    Local area
    Remote work
    Flexible hours

    Doordash

    San Francisco, CA
    3 days ago
  • $183k - $247.6k

    Our Machine Learning Acceleration (MLA) team develops the Inferentia and Trainium SOCs that...  ...seeking experienced Design Verification Engineers to build the next generation of our...  ...conferences. Amazon’s culture of inclusion is reinforced within our 14 Leadership Principles,... 
    Suggested
    Local area
    Work from home
    Flexible hours

    Amazon

    Cupertino, CA
    10 hours ago
  •  ...spin-out, and our unique combination of learning-based automation and augmented remote...  ...Explore the usage of adaptive and online reinforcement learning in deployed systems Provide...  ...with excavation and motion planning engineers Build tools for analysing and evaluating... 
    Remote work

    Gravis Robotics

    Austin, TX
    26 days ago
  • $99k - $225k

    Reinforcement Learning AI EngineerThe Opportunity:Are you an innovative and experienced artificial intelligence (AI) developer specializing...  ...expertise in AI, data science, and machine learning (ML) engineering to train, test, deploy, and maintain models that learn from... 
    Full time
    Contract work
    Part time
    Work at office
    Local area
    Remote work

    Booz Allen Hamilton

    Colorado Springs, CO
    3 days ago
  • $197.3k - $225.1k

    AI Engineer 4 (AI Foundations: LLM Customization, Finetuning, Reinforcement Learning) At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine learning to create real... 
    Full time
    Part time
    Local area

    Capital One

    McLean, VA
    2 days ago
  •  ...take initiative, streamline complex workflows, and continuously learn and adapt.Moveworks is trusted by over 5.5 million employees at...  ...ServiceNow’s leading workflow automation with Moveworks’ Reasoning Engine and natural language capabilities, we deliver the AI platform... 
    Work at office
    Immediate start
    Remote work
    Flexible hours

    Moveworks

    Mountain View, CA
    1 day ago
  •  ...About the Role You'll be a Full-Stack Software Engineer building the product surfaces, backend systems, and internal tools that...  ...and external partners to create, evaluate, and iterate on reinforcement learning training data. You don't need to be an ML researcher, but you... 
    Full time
    Work at office
    Remote work
    Visa sponsorship
    Relocation package

    hud

    Singapore
    16 days ago
  •  ...Helix team is responsible for developing the core AI systems that power humanoid autonomy. We are looking for a Helix AI Engineer, Reinforcement Learning to develop learning systems that enable robots to acquire skills through interaction, feedback, and experience.... 
    Full time
    Work at office

    Figure

    California
    10 days ago
  •  ...take initiative, streamline complex workflows, and continuously learn and adapt.Moveworks is trusted by over 5.5 million employees at...  ...ServiceNow’s leading workflow automation with Moveworks’ Reasoning Engine and natural language capabilities, we deliver the AI platform... 
    Permanent employment
    Work at office
    Remote work
    Flexible hours

    Moveworks

    Mountain View, CA
    1 day ago
  •  ...Earth. We are seeking a talented  Senior Autonomy Machine Learning Engineer . As a Senior Autonomy Machine Learning Engineer, you will...  ...foundation models ~ Robot learning, imitation learning, or reinforcement learning ~ Natural-language planning, tool use, or... 
    Remote work
    Monday to Friday

    Lunar Outpost

    Arvada, CO
    15 days ago
  • $184k - $287.5k

     ...intelligence to autonomous cars.We are looking for a motivated Deep Learning engineer to bring advanced communication technologies into AI stacks...  ...more of these areas: Training, Distributed inference, MoE, Reinforcement Learning, kernel authoring (on CUDA, Triton, cuTe, etc).... 
    Full time
    Remote work

    Nvidia

    Durham, NC
    4 days ago
  •  ...customers can expect, and how we deliver. Learn more about our extraordinary teams...  ...pipeline. Job Summary:The Quality Control Engineer (QCE) will report to the Lead Quality Control...  ..., reduce nonconforming conditions, reinforce quality expectations, and support right-... 
    Full time
    Work experience placement
    Work at office
    Local area
    Remote work
    Relocation

    Bechtel

    New Albany, OH
    10 hours ago
  • $397.46k

     ...of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences-...  ...Monetization, Payments, and Avatar.We’re looking for a Distinguished Engineer/Technical Director to lead the strategy and technical direction... 
    Full time
    Work experience placement
    H1b
    Work at office
    Local area
    Visa sponsorship
    Monday to Friday

    Roblox

    San Mateo, CA
    4 days ago
  •  ...inference. Turn advances in information retrieval and machine learning into production systems through rigorous evaluation, experimentation...  ...multiple teams, simplify fragmented systems, mentor senior engineers, and align technical and product leaders around long-term... 
    Work at office
    Local area

    Atlassian

    New York, NY
    2 days ago
  • $158k - $294k

     ...connection, education, and inclusion. About the RoleWe are seeking a hands-on Principal MLOps Engineer & Functional Lead to architect, build, and scale our next-generation Machine Learning and Generative AI infrastructure. In this role, you will design a unified, automated,... 
    Remote work

    Definitive Healthcare

    Framingham, MA
    3 days ago
  • $250k - $350k

     ...started — this team will define what dependable, production-grade agentic AI looks like.About the RoleAs a Staff Machine Learning Research Engineer, you will operate across the full breadth of AIS’s technical needs — wherever the hardest ML problem in agentic AI happens... 
    Full time

    Scale AI

    San Francisco, CA
    4 days ago
  • $89k - $130k

     ...yourself through meaningful work projects and learning opportunities. We strive to provide our...  .... This role will provide Pre-sales engineering, design and estimation during project...  ...for projects across North America. Reinforces methodology and skills used to collect,... 
    Full time
    Contract work
    Work experience placement
    For subcontractor
    Local area
    Remote work
    Worldwide

    Johnson Controls

    Austin, TX
    2 days ago
  •  ...you’re not excited to experiment, adapt, think on your feet, and learn constantly, or if you’re seeking something highly prescriptive...  ...is seeking a highly skilled and versatile Machine Learning Engineer to join our Research team. As a Member of the Research Staff, you... 
    Full time

    Deepgram

    Remote
    27 days ago
  • $130k - $180k

     ...tremendous career growth potential. Job Title Autonomous Learning Engineer Location: 100% Remote (U.S.) Position Type: Full-...  ...with 10+ years of experience in Artificial Intelligence, Reinforcement Learning (RL), and Deep Learning to design, train, and deploy... 
    Full time
    H1b
    Local area
    Immediate start
    Remote work
    Visa sponsorship

    Bright Vision Technologies

    Remote
    2 days ago
  •  ...OpportunitiesWhy This JobCMC provides an excellent opportunity to learn the steel, construction reinforcement and ground stabilization industries and to grow in...  ..., etc.Your EducationBachelor's degree in Electrical Engineering or a closely related disciplineWe are CMC, a Fortune... 
    Work at office
    Local area
    Immediate start
    Remote work

    Commercial Metals Company

    Morrow, GA
    1 day ago
  •  ...customers can expect, and how we deliver. Learn more about our extraordinary teams...  ...this role, you will provide project field engineering support with work planning and packaging...  ...and documentation for concrete placement, reinforcing steel, formwork, and embedded items.Technical... 
    Full time
    For contractors
    Work experience placement
    For subcontractor
    Work at office
    Local area
    Remote work
    Relocation

    Bechtel

    Brownsville, TX
    2 hours ago
  • $40 - $45 per hour

     ...environment where everyone can thrive. If you're committed to learning and advancing your career while being able to work in your local...  ...Safety Help create a safe and productive jobsite by reinforcing safety expectations, supporting employee training, addressing... 
    For contractors
    Work at office
    Local area

    Willmar Electric

    Elgin, OK
    23 days ago
  •  ...Controls to lead a small team of automation engineering professionals. This key position...  ...the Company. Embody, demonstrate, and reinforce the ANDRITZ values. What We Have...  ...and certifications, including a LinkedIn Learning license. Clear career paths for... 
    Full time
    Immediate start
    Remote work

    ANDRITZ

    Morrilton, AR
    1 day ago
  •  ...growing and responsible for building machine learning models and systems to identify and...  ...scientists and machine learning engineers who can take initiative, design and develop...  ...preferred.- Experiences in active learning, reinforcement learning, and LLM are a plus.-... 

    TikTok

    Seattle, WA
    1 day ago
  •  ...and IT control environment by designing, engineering, implementing, assessing, and...  ...GRC and technical platforms.Clarify and reinforce control ownership and accountability across...  ...can focus on realizing your ambitions. Learn how GM supports a rewarding career that... 
    Full time
    Local area
    Work from home
    Relocation package

    General Motors

    Center Line, MI
    4 days ago
  • $147.6k - $274k

     ...worldwide.The OpportunityWithin AI for Drug Discovery, the Software Engineering team builds and operates software platforms that put advanced...  ...across drug discovery. We are seeking a very talented Machine Learning Engineer to help build our agentic platform for molecule... 
    Full time
    Local area
    Worldwide
    Relocation package
    3 days per week

    Genentech

    New York, NY
    1 day ago
  • $144k - $192k

    Mission Summary:We are looking for a Machine Learning Systems Engineer to join our ML Acceleration team. In this role, you will be responsible for the core systems that enable our researchers to train frontier models at scale, focusing obsessively on speed, cost, reliability... 
    Work at office
    Remote work

    Motional

    Boston, MA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Reinforcement Learning Engineer. Be the first to apply!