SWE (RL Environments) "Reinforcement Learning"
AI Talent Now
San Francisco, United States | Posted on 06/15/2026 AI Talent Now, LLC is a powerhouse in direct‑hire talent acquisition, connecting exceptional talent with industry‑leading organizations across IT, Engineering, Financial Services & Fintech, and Manufacturing & Robotics. Headquartered in Atlanta, Georgia, we serve clients and candidates nationwide. Job Description AI Talent Now Job #ZR 72 About us This dynamic company is helping push the frontier of LLMs and AI Agents through novel datasets and experimentation. We build the most complex infrastructure that powers frontier data creation for agentic and hard‑reasoning workflows. Working with all five leading AI labs, we are becoming the go‑to partner for data infrastructure for YC companies. Our sharp hockey‑stick growth and talent density come from a founding team with backgrounds in top IB and quant firms. We build the training data and evaluation infrastructure that frontier AI labs use to improve their models. Working with the world's leading labs, we design high‑signal datasets and run rigorous evaluations that go beyond static benchmarks. At a small, early post‑Series A team, individual contributors directly influence how the next generation of models learn and improve. As an RL Environment Engineer you will design datasets that directly influence how frontier models learn and work hands‑on with research teams at top AI labs. Responsibilities Design and develop datasets that shape frontier model learning. Collaborate with research teams at leading AI labs to build and refine reinforcement learning environments. Implement high‑quality benchmarks and evaluate model performance beyond static tests. Qualifications Recent graduates from top schools focused on excellence and depth; extensive track record not required. First‑author publications at top venues such as NeurIPS or ICML are highly desirable. Open to profiles from data companies with benchmarking experience. Ideal candidates have created iconic benchmarks and possess experience in supervised fine‑tuning (SFT) or reinforcement learning (RL). Experience 1–6 years of experience as a software engineer. Explicit experience building reinforcement learning environments. Signals of excellence in software engineering (e.g., work at an RL company, top VC‑backed startup, strong side projects, quant or hedge‑fund background, or founding/early‑stage startup engineer). CS degree from a top‑30 school (US/CA/EUR only). Strong full‑stack skills with depth in Python, Typescript, and other backend languages. Developed quantitative frameworks for measuring dataset quality/diversity. Bias for action and execution; willing to tackle difficult and tedious work. Compensation and Logistics Base salary: $150k–$250k, with significant bonus potential based on performance. Bonuses are uncapped and can substantially increase total compensation. Visa sponsorships are possible (experience in H1B and other types). Work Arrangement Flexible working hours; most team members in the office from noon to midnight. Work environment encourages results over strict hours. EEO Statement We are an equal opportunity employer and welcome applications from all qualified individuals regardless of race, color, religion, sex, gender identity, sexual orientation, national origin, disability, or veteran status. #J-18808-Ljbffr AI Talent Now
- ...Now in San Francisco is seeking a skilled RL Environment Engineer to design impactful datasets that influence the learning of frontier AI models. You will collaborate... ...research teams at leading AI labs to refine reinforcement learning environments and create comprehensive...Suggested
$200k
...direct impact on how the next generation of models learn and improve. The Role As a SWE (Environments), you will design the datasets and evaluation rubrics... ...~ Major plus if they've worked for/interned for any RL environment companies in the past or any AI safety or...SuggestedFull time- Traverse is a research data lab building reinforcement learning environments for frontier AI labs. We focus on non-deterministic, taste-dependent work and train models for real-world tasks. You will design reward functions, evaluation pipelines, and the tooling needed to...Suggested
- ...RippleMatch Inc. is seeking an innovative and motivated individual to design and refine reinforcement learning tasks in San Francisco. This role requires a strong command of Python and the ability to work independently with coding agents. Responsibilities include the...Suggested
- Scale AI, Inc. in San Francisco seeks an AI Product Manager to own the Agent & Reinforcement Learning Environments data vertical, focusing on Computer Using Agent data. You will drive the roadmap, define data as a product, and shape researcher-facing tools for labs to train...Suggested
- Preference Model is hiring new graduate Machine Learning Engineers to design and build reinforcement learning environments. You will blend research and engineering roles in a dynamic startup environment that values diverse perspectives. The ideal candidate should possess...
$204k - $259k
...Research Scientist, RL for Autonomous Planning & World Modeling Waymo is an autonomous... ...Foundations team is to develop machine learning solutions addressing open problems in... ...we are currently focusing on include reinforcement learning, learning from demonstration, generative...Full timeTemporary workRemote work$218.4k - $273k
...team The Agent Capabilities & Environments (ACE) team, part of Scale’s Research... ...on agent environments and RL reward signals, benchmarking... ...art agents, such as browser and SWE agents. The ideal candidate... ...or GCP) and developing machine learning models in a cloud environment....Full time$1,750 per week
...Peru, Mexico, and more! Backed by substantial funding and a passionate, collaborative team, we offer a rewarding work environment where you'll learn and make a significant impact, no matter where you are in your career. Environmental Test Associate Engineer...Permanent employmentFull timePart timeInternshipWork at office- ...abilities, we solve our clients’ complex challenges in water, environment, energy, transportation and buildings. Our teams partner with public... ...00 firm that had revenue of $16.1 billion in fiscal year 2025. Learn more at aecom.com. What makes AECOM a great place to workYou...Permanent employmentWork at officeLocal areaImmediate startRemote workWorldwideFlexible hours
- ...abilities, we solve our clients’ complex challenges in water, environment, energy, transportation and buildings. Our teams partner with public... ...00 firm that had revenue of $16.1 billion in fiscal year 2025. Learn more at aecom.com. What makes AECOM a great place to workYou...Permanent employmentWork experience placementWork at officeLocal areaWorldwideFlexible hours
- David Joseph & Company is seeking a founding Member of Technical Staff for platform engineering in San Francisco. You will own the RL environment infrastructure, build scalable inference pipelines, and help stand up the engineering org and its culture from scratch. We...
- ...get made. We build the data, environments, graders, training methods, and... ...in this role include GDPval, SWE-bench Verified, MLE-bench, PaperBench... ..., you might Create ambitious RL environments to push our... ...technical fundamentals in machine learning, software engineering, systems...
- ...equipment. sort recyclable materials. maintain clean and safe work environment. follow safety and environmental regulations. respond to... ...supply of waste containers. assist with billing and documentation. learn company waste services and pricing. work closely with finance...
- .... ~ Proficiency with Microsoft Office Suite and ability to learn new project management and field data systems quickly. ~ Experience... ...are pursued in responsible ways where both people and the environment thrive. As an organization, we believe that independence and...Full timePart timeSecond jobWork at officeLocal areaFlexible hoursShift workNight shift
$98k - $126k
...the firm. Proficiency with Microsoft Office Suite and ability to learn new project management and field data systems quickly.... ...opportunity employer, and we are committed to creating a diverse environment. We recruit, employ, train, compensate and promote regardless of...Part timeWork at officeShift workNight shift- ...providing constructive feedback on project tasks.About AECOM’s Environment Business Line Join AECOM to be part of an expert global team who... ..., internal technical practice network through which you can learn from and brainstorm with the best in the world. AECOM is an industry...Local areaWorldwideFlexible hours
- ...three days a week to contribute, connect and excel in our vibrant environment.Working with an energetic and high performing team, this... ...delivery.Close out projects with complete documentation, lessons learned, and handoff as requiredWhat you will bring to the team:...For subcontractorWork at officeLocal area3 days per week
$28 per hour
...our company’s commitment to maintaining a safe and healthy work environment Must be eligible to work in the United States without future... ...assistance discount plans, discounted movie passes & more! To learn more about our business, culture, and the exciting work that we...Work at officeRelocationMonday to FridayFlexible hoursRotating shift- Epsilon Labs, Inc. is seeking a Research Scientist with deep expertise in post-training and reinforcement learning to advance multimodal models for clinical radiology use. You will own stages after pretraining, including supervised fine-tuning, reward modeling, and inference...
- The Voleon Group seeks a Reinforcement Learning Researcher to join its ML research group in Berkeley, California.... ...relevant field and a strong publication record in RL are required. Join a team of experts in a dynamic environment with opportunities for relocation. #J-18808-...Relocation
- ...customer pipeline already in place. Role Overview We're seeking a Research Scientist with deep expertise in post‑training and reinforcement learning to join our ML Research team . You'll be at the forefront of developing and deploying state‑of‑the‑art multimodal models...
- ...focus. About the Role LMArena is seeking a variety of Machine Learning Scientist to help advance how we evaluate and understand AI... ...learning architectures (e.g., Transformers, diffusion models, reinforcement learning with human feedback) Proficiency in Python and ML research...Permanent employmentWork at office
- ...manage multiple matters. Team-oriented mindset and eagerness to learn from leading practitioners. Education Juris Doctor (JD) from an... ...aptitude. Benefits Competitive salary. Opportunities for professional development. Collaborative work environment. #J-18808-Ljbffr...Work at officeLocal area
- ...AI Research Scientist (Robot Learning) San Francisco AI & Software In office Full... ...and ROS2 experience Experience with RL fine-tuning of generative models Experience... ...towards creating a diverse and inclusive environment we encourage everyone to apply and we are...Full timeWork at officeImmediate start
$230k - $400k
...models at scale. Building Infrastructure We work on many infrastructure projects including: Complex multimodal reinforcement learning environments. High-performance RPC servers for processing image inputs. Sandboxing infrastructure for securely collecting...Work experience placementWork at officeHome officeVisa sponsorshipRelocation packageFlexible hours$117.2k - $176.7k
...dedicated to improving developer productivity in cloud development environments. Our mission is to build a robust Cloud development platform... ...developer productivity and experience.Proven ability to learn new technologies and adapt to evolving project needs.Experience...Full timeRemote work$84k - $113k
...COMPANY OVERVIEW EKI Environment & Water, Inc. (EKI) is an employee-owned, full service, engineering and environmental sciences consulting... ...as professionally Must have a great attitude and eager to learn Must have a current valid driver's license PHYSICAL...Full timeFor subcontractorLocal areaNight shift- ...research and engineering, allowing you to take ownership from ideation to production. Locally and remotely, you will focus on reinforcement learning, evaluation, and building robust training infrastructure. Ideal candidates possess strong machine learning fundamentals and...Remote work
- ...project outcomes. The ideal candidate will thrive in a dynamic environment, balancing staff management responsibilities with hands on project... ...00 firm that had revenue of $16.1 billion in fiscal year 2025. Learn more at aecom.com. What makes AECOM a great place to workYou...Permanent employmentLocal areaWorldwideFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to SWE (RL Environments) "Reinforcement Learning". Be the first to apply!
- quality environment health safety manager San Francisco, CA
- senior environment artist San Francisco, CA
- project manager environment San Francisco, CA
- environment health safety San Francisco, CA
- environment San Francisco, CA
- environment artist San Francisco, CA
- environmental work from home
- quality environment health safety manager
- 3d game environment artist
- junior environment artist




