Research Intern, Agent RL Training
$35 - $50 per hourNewsBreak
Research Intern, Agent RL Training
We are looking for a Research Intern to join our Agent RL Training team. You will be paired with a full-time employee as your mentor, working together to explore, from zero to one, how to apply large language models to NewsBreak's core business, including content understanding, recommendation, agentic web browsing, and autonomous multi-step task completion.
This is a hands-on research role. You are expected to independently drive experiments, propose novel ideas, and iterate quickly. We value self-starters with deep intellectual curiosity and the drive to push boundaries in LLM post-training and agent capabilities.
Location: Onsite in Mountain View, CA office
What You'll Work On
- Collaborate with your full-time mentor to identify high-impact research directions for applying LLMs to NewsBreak's products
- Independently run end-to-end SFT experiments on LLM-based agents, and assist with RL-related exploration such as reward design and training iteration
- Curate and build high-quality training datasets: instruction-following, preference pairs, agent trajectories, and synthetic data
- Contribute to public publications; we encourage and support top-venue submissions during your internship
What We're Looking For
Requirements
- Highly motivated and committed: willing to put in extra hours when needed to push projects across the finish line
- Genuine passion for research: you read papers for fun, tinker with models on weekends, and care deeply about advancing the field
- Independently capable of end-to-end model SFT: with basic understanding of RL-based post-training methods (RLHF, DPO, PPO, GRPO, etc.)
- Excellent taste in model behavior: able to reason about what "good" looks like across user-facing domains and articulate why
- Strong Python and PyTorch skills
Preferred Qualifications
- Publication at a top-tier venue (NeurIPS, ICML, ICLR, ACL, EMNLP, or equivalent)
- Experience with multi-node distributed training (FSDP, DeepSpeed, Megatron-LM)
- Proficiency in writing custom GPU kernels with Triton or CUDA
- Experience building synthetic data pipelines for agent training
- Familiarity with open-source RL frameworks: TRL, OpenRLHF, veRL/vLLM
Hourly Pay: $35- $50
The US base salary range for this full-time position is listed below. Pay may vary based on a number of factors including job-related skills, level, experience, geographic location and relevant education or training. At NewsBreak, we design our overall rewards package to attract top talents. Depending on the position, the role may also be eligible for discretionary bonus and options. Your recruiter can share more details during the hiring process.
Annual Base Pay Range
$35 - $50 USD
CPRA Privacy Notice for California Candidates
$193.93k - $291.15k
...member of the Prediction and Smart Agents team, you will focus on... ...enable effective closed-loop training in simulation.If you are passionate... ...problems, leading impactful research, and seeing your work deployed... ...training via Reinforcement Learning (RL).Mitigate accumulated...TrainingImmediate startFlexible hours$300k - $365k
...company’s codebase. About the role The Research Scientist will be responsible for designing and implementing large-scale RL experiments for improving code agents. The ideal candidate will have extensive experience in post-training state-of-the art large language models...TrainingFull timeTemporary workFlexible hours- ...Foundation Models We are a dedicated research lab for building,... ...cutting‑edge foundation model training, alongside world‑class researchers... ...Position Summary As a member of the Agents team, you will tackle research... ...knowledge of literature on RL, LLM reasoning, and tool use...TrainingVisa sponsorship
- ...Architect Architect is an AI research and product lab for chip design... ...What You’ll Do As a Research Intern at Architect, you will spend 3... ...Learning experiments (GRPO/PPO/DPO), training data mixes and reward signal... ...also encouraged to apply. RL Knowledge: Strong academic...InternshipTraining
- CoreWeave, a leading AI infrastructure company, is seeking an accomplished ML researcher/engineer for its OpenPipe team. You will develop and validate advanced LLM training methods, including RL and supervised fine-tuning, and drive production-ready solutions with a...Training
$244.14k - $413.16k
...We are looking for exceptional Research Engineers / Scientists to... ...design learning systems that allow agents to plan over long horizons,... ...feedback (RLHF / RLAIF).Agent training pipelines built on top of our... ...learning.Experience implementing RL algorithms such as PPO, Actor-...TrainingFull time- Silimate in Mountain View, CA is hiring software interns with strong AI research experience to push AI-native tools and agent-driven workflows for chip design. This is an... ...methods across models, algorithms, and training frontiers for AI-native chip design. #J-18808...InternshipTrainingFull timePart timeWork at officeImmediate start
- ...We are looking for multiple passionate Research Interns to join the Research Group at Applied Intuition... ...Conduct research on reinforcement learning (RL) related topics including large-scale closed-loop RL and VLA post-training with applications to robotics with emphasis...InternshipTrainingFor contractorsFor subcontractorCasual workWork at officeImmediate startRemote workDay shift
- ...safely and ethically—for AI labs, pharma companies, and researchers. This isn’t a chatbot, or an AI agent replacing clinicians or automating paperwork. It’s... ...through hands‑on research and implementation. Design, train, fine‑tune, and evaluate models with a focus on real‑...InternshipTrainingNight shift
$200k - $300k
A leading financial technology firm in California is looking for AI Research Scientists and Engineers to advance their AI capabilities. You will focus on post-training SOTA LLMs and work with multiple specialized teams. Applicants should have experience in large-scale LLMs...Training- ...a new era of biomedicine, with our LBM training leading to ground-breaking advancements... ...cutting-edge AI and computational biology research. Your primary tasks will include improving... ..., or Cell) journals and conferences Intern experience in industry (e.g., OpenAI, FAIR...InternshipTrainingWork at office
$30 - $94 per hour
NVIDIA Corporation is seeking a Deep Learning PhD Research Intern focused on Reinforcement Learning for LLMs. This role involves developing and prototyping novel RL algorithms to enhance LLM behavior, alongside exploration of reasoning and alignment methods. An ideal candidate...InternshipHourly pay- ...readiness. We are building an LLM‑powered agent to help MLEs collect experiment context,... ...development work. We’re looking for an intern to help build the data foundation for this... ...with machine learning workflows, model training, evaluation metrics, or MLOps concepts....InternshipTraining
- We’re hiring software interns with significant experience with AI research (across models, algorithms, and training methodology) to help push the frontier of our AI-native tools and agent-driven workflows for chip design. This is an in-person role at our Mountain View office...InternshipTrainingFull timePart timeWork at officeImmediate start
- CoreWeave, The Essential Cloud for AI, is hiring for an applied research role in Sunnyvale, CA. You will generate and validate research... ...tackle large-scale AI tasks using extensive GPU resources. You have trained LLMs, evaluated sequence- vs token-level strategies, and driven...
- CoreWeave is seeking an applied research engineer for the OpenPipe team to advance continuous learning in production for autonomous agents. You will generate and evaluate research ideas, validate them on real customer tasks, and leverage extensive GPU resources to accelerate...
$30 - $94 per hour
Applied Deep Learning PhD Research Intern, Reinforcement Learning for LLMs - Fall 2026 page is loaded##... ...design, implement, and evaluate new RL-based methods for improving LLM behavior... ...models, and working with large-scale training pipelines**Ways to stand out from the crowd...InternshipTrainingHourly pay$192.2k - $260k
...development tools. Join the Amazon Kiro LLM-Training team and help create groundbreaking... ...forefront of innovation, where cutting-edge research meets real-world application:- Push the... ...systemsAbout the teamThe AWS Developer Agents and Experiences (DAE) team is...TrainingWork at officeLocal areaWorldwideFlexible hours- ...AI Research Intern We're looking for an AI Research Intern to join our AI team and explore... ...AI, LLMs, computer vision, speech, and agent systems. You'll work across OpusClip, AgentOpus... ...and generation LLM post-training (e.g. SFT, RLHF, DPO) Applied AI...InternshipTrainingLocal areaRemote workWorldwideFlexible hours3 days per week
$120k - $250k
...leading organizations. About the Role As a Research Engineer at Orbifold AI, you will be at... ...scale, pushing the boundaries of training, RAG, and reinforcement learning for enterprise... ...weeks ago Software Engineer: Fullstack Intern Opportunities for University Students,...InternshipTrainingFull timeFlexible hours$142.8k - $193.2k
...re seeking brilliant minds to join us as interns and contribute to the development of... ...you immerse yourself in groundbreaking research, exploring novel machine learning models... ...featurize diverse data sources for model training - Conduct research into the latest advancements...InternshipTrainingFull timeWorldwideRelocation$180k
...skills are important. All engineers and researchers are expected to have strong... ...infrastructure team at xAI builds an end to end RL training framework to enable pretrain scale RL.... ...-scale reinforcement learning and multi-agent reinforcement learning Location ~...TrainingFull timeWork at officeWork from homeRelocation$193.93k - $352.29k
...platform that lets AI agents operate autonomously inside... ...of every engineer and researcher at Nuro by 100x. Not a... ...loop runs against the training pipelines behind the driving... ...own models where an internal workload justifies it,... ...optimization, and RL — and you can reason about...TrainingImmediate startFlexible hours- Centific Global Solutions, Inc. is seeking a Senior Staff Research Scientist in Agentic AI & Reinforcement Learning to... ...mentoring, and design responsibilities in building governed RL environments and LLM post-training pipelines. The ideal candidate will excel at both...Training
- ...advancement into a permanent position Training & development Bonus based on performance... ...DESCRIPTION: Yama Ebrat - State Farm Agent is seeking an organized and efficient... ...agency in gaining and keeping customers. As a Intern - State Farm Agent Team Member with our...InternshipTrainingPermanent employmentFlexible hours
$174k - $252k
...generalization, and uncertainty quantification. Research experience with first-author... ...publishing your findings, with ideas inspired by internal projects as well as from collaborations... ..., experience, and relevant education or training. US: $174000 - $252000 (USD) + 15% bonus...InternshipTrainingShift work- ...asset resolution west of the Mississippi. Our people are expertly trained in every aspect of the business, from assignments to remarketing... .... We are seeking a highly motivated and skilled Recovery Agent to join our team. The successful candidate will be responsible for...TrainingHourly payFull timeNight shiftWeekend work
$195.2k - $361.2k
...servesRole SummaryThe model is only half an agent; the harness is the other half. You build... ...experiences and or schoolwork/classes/research.Benefits at IntelOur total rewards package... ..., experience, and relevant education or training. Your recruiter can share more about the...InternshipTrainingFull timeLocal areaImmediate startShift work- ...About the Institute of Foundation Models We are a dedicated research lab for building, understanding, using, and risk-managing foundation... ...to work on the core of cutting-edge foundation model training, alongside world-class researchers, data scientists, and engineers...InternshipTrainingTemporary work
- ...Secret Service. During the course of their careers, special agents carry out assignments in both investigations and protection and... ...while you occupy the position. Complete 13 weeks of intensive training at the Federal Law Enforcement Training Center(FLETC) in Glynco...TrainingOverseas
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Research Intern, Agent RL Training. Be the first to apply!
- work from home chat agent Mountain View, CA
- operations agent Mountain View, CA
- state farm agent Mountain View, CA
- tsa agent Mountain View, CA
- cruise agent Mountain View, CA
- import export agent Mountain View, CA
- commissioning agent Mountain View, CA
- agent Mountain View, CA
- airport agent Mountain View, CA
- remote chat agent Mountain View, CA




