Member of Technical Staff, Post-training
$180kHark
About the Role We are looking for a Member of Technical Staff, Post-Training to lead the development of post-training strategies that define how our models acquire coding, computer use, and agentic capabilities at scale. This role sits at the frontier of a rapidly emerging discipline — one where reinforcement learning, simulation, and large-scale model training converge to produce agents that can reason, plan, and act over long horizons. There is no established playbook here. We’re looking for researchers and engineers who can bring rigor and creativity from adjacent fields — RL, robotics, game-playing systems, compiler tooling, formal verification, or program synthesis — and apply them to the next generation of coding and agentic AI. Responsibilities Design and implement post-training strategies, primarily RL-based, to develop strong coding agents capable of multi-step reasoning, tool use, and long-horizon task completion. Build and scale simulation and scaffolding environments for agentic RL: code execution sandboxes, computer use environments, tool-calling harnesses, and verifiable reward signals. Develop reward modeling pipelines — including outcome-based, execution-based, and process-based reward signals — and iterate on them based on training dynamics. Scale synthetic data generation and trajectory distillation pipelines that feed RL training and improve sample efficiency. Design and run rigorous ablations to understand how algorithm choice, data mixture, reward shaping, and scale interact in the agentic setting. Build evaluation frameworks grounded in real agent tasks — code correctness, execution success, multi-step tool use — to measure progress and guide iteration. Collaborate with mid-training, infrastructure, and product teams to translate research insights into durable improvements on the model. Requirements Strong background in machine learning, with hands-on experience training or fine-tuning large models — LLMs, multimodal, or equivalent systems. Deep understanding of reinforcement learning: policy optimization, reward design, exploration, and the interplay between environment design and agent behavior. Experience building or working within simulation or execution environments (e.g., code interpreters, sandboxed execution, game environments, robotics simulators). Proven ability to design and execute rigorous experiments, with strong intuition for diagnosing training failures and scaling bottlenecks. Proficiency in Python and PyTorch; comfort working across research and systems code. Ability to work in a fast-moving, research-forward environment where the right approach is often unknown at the outset. We expect strong candidates to come from a range of backgrounds — RL research, robotics, competitive programming systems, compilers, formal methods, or large-scale ML — rather than post-training specifically. The field is new enough that directly relevant experience is rare; what matters is depth, rigor, and transferability. Bonus Qualifications Experience with RL algorithms applied to language or code: RLHF, DPO, GRPO, PPO, or similar paradigms in the LLM setting. Familiarity with coding agent benchmarks and evaluation environments (e.g., SWE-bench, HumanEval, LiveCodeBench, competitive programming judges). Background in reward modeling — outcome-based, process-based, or learned reward signals. Experience with trajectory-based training, imitation learning, or data distillation from stronger models or human demonstrations. Prior work on computer use, GUI agents, or tool-using LLMs (e.g., OSWorld, WebArena-style tasks). Experience training or scaling models at 10B+ parameters, with attention to efficiency, stability, and GPU utilization. Contributions to open-source ML projects or publications at top venues (NeurIPS, ICML, ICLR, EMNLP, COLM, etc.). Compensation The US base salary range for this full-time position is between $180,000 - $450,000 annually. The pay offered for this position may vary based on several individual factors, including job-related knowledge, skills, and experience. The total compensation package may also include additional components and benefits depending on the specific role. This information will be shared if an employment offer is extended. #J-18808-Ljbffr Hark
$180k
...for a new era of intelligent systems. About the Role We are looking for a Member of Technical Staff - Mid-Training to lead the development of training strategies that bridge pre‑training and post‑training, shaping how models acquire reasoning, planning, and tool‑use capabilities...TrainingFull time- Member of Technical Staff, Computer Vision DeepReach is building the next-generation data infrastructure for robotics. We help bridge the gap... ...across robot deployment, teleoperation, data generation, model training, and evaluation, with a strong bias toward hands‑on...Training
$180k
Member of Technical Staff, Multimodal Vision San Jose About Hark Hark is an artificial intelligence company building advanced, personalized intelligence... ...working across the full stack—from data and modeling to training, serving, and product integration. You will contribute to...TrainingFull time$141.57k - $208.19k
...,188.75 Salary Job Shift: Days TITLE: MEMBER TECHNICAL STAFF (Design Integration) SUMMARY Under the... ...enforcement, documentation, and peer training; and preparing and maintaining documentation... ...listed in the job description or posting. TDK/Headway Technologies, Inc. provides...TrainingFull timeFor contractorsWork at officeMonday to FridayFlexible hoursShift work- ...also working with massive, ever-growing datasets and models in training. Your focus will be ensuring our models deliver exceptional... ...development and inference infrastructure. Have significant autonomy in technical decisions. Use the latest-generation GPUs. Who you are 8+...Training
$180k
...large‑scale pretraining systems and foundation models. This includes working across the full stack—from data curation and large‑scale training infrastructure to model architecture and optimization. You will play a key role in advancing the core capabilities of our models...TrainingFull time$180k
...multimodal world models. This includes working across the full stack—from data and modeling to training, serving, and product integration. You will contribute to both pre‑training and post‑training efforts while collaborating closely with product teams to push the boundaries...TrainingFull time- ...looking for We’re looking for a deeply technical and creative researcher who thrives on invention... ...architectures, learning objectives, and training paradigms that move beyond today’s... ...generative models. Who you are A staff-level or senior researcher with deep expertise...Training
$180k
...speech and audio capabilities within multimodal foundation models. You will work across the full stack—from data and modeling to training, evaluation, and real‑time serving—pushing the boundaries of speech intelligence and human‑computer interaction. Responsibilities...TrainingFull time$230k
...Cerebras to deliver industry-leading training and inference speeds and empowers machine... ...Inc. has multiple openings for Sr. Member of Technical Staff. Title: Sr. Member of Technical Staff... ...preprocessing, inference execution, and post-processing for real-time inference...TrainingRemote work- Member of Technical Staff (Software Engineer) Sunnyvale, CA Cerebras Systems builds the world’s largest... ...delivers industry‑leading training and inference speeds, empowering machine... ...preprocessing, inference execution, and post‑processing for real‑time inference tasks...TrainingFull timePart timeInternship
- ...minimize risk in the cloud. You will mentor junior engineers, new-grads, and interns to help them grow as engineers and become productive members of the team. You will primarily write code in Java (using spring boot framework) and work with data pipelines using Kafka/SQL or...Immediate start
- ...implementation. You will also mentor junior engineers, new-grads, and interns to help them grow as engineers and become productive members of the team. You will primarily write code in Go and work with data pipeline using SQL or other types of interfaces. We leverage Kubernetes...Immediate start
$217.6k - $290.5k
...innovators, and dreamers — and help us connect people and build communities to create economic opportunity for all. Senior Technical Staff Member, Applied Research (T27) San Jose, CA Position Overview eBay is looking for a Senior Member of Technical Staff, Applied...Local areaImmediate startRemote workVisa sponsorshipShift work- Member of Technical Staff, Lead Researcher San Francisco, CA; Sunnyvale, CA About the Role DoorDash is... ...can match High compute budgets for training and inference, sized to support frontier... ...large-model pre-training and post-training, RL training runs, and large-...TrainingLocal area
- ...approximation Design efficient architectures and attention mechanisms suited to real-time inference on edge and robot hardware Develop training strategies that produce better accuracy-efficiency tradeoffs from the start Profile and benchmark models across hardware targets...Training
- ...and manufacturing scale‑up to make generalist robotics a reality. We're looking for a Research Engineer to build and maintain the training platform that powers our model development — experiment orchestration, job management, observability, and the tooling that lets...Training
- ...scoring at web scale Collaborate closely with pre-training and post-training teams to ensure data quality and... ...that are diagnostic, reproducible, and actionable Staff-level candidates are expected to define technical direction and drive research strategy independently...Training
- ...About the Role We're looking for a senior or staff-level Research Engineer or ML Systems... ...collection through dataset generation, model training, inference, and real-robot evaluation.... ..., data ingestion and compilation, post-training, checkpoint generation, inference...Training
$27.5 - $32.5 per hour
...Technical Support RepresentativeSan Jose, California, United States$ 27.50 - 32.50 (US Dollar... ...effectively with clients and team members to understand and resolve technical problems... ...of FAQs, support documentation, and training materials for users.Product Knowledge:Maintain...TrainingRemote workMonday to FridayFlexible hours- Overview Onwards Together! Illumio is the leader in ransomware and breach containment, redefining how organizations contain cyberattacks and enable operational resilience. Powered by the Illumio AI Security Graph, our breach containment platform identifies and contains ...Immediate start
- Onwards Together! Illumio is the leader in ransomware and breach containment, redefining how organizations contain cyberattacks and enable operational resilience. Powered by the Illumio AI Security Graph, our breach containment platform identifies and contains threats across...Immediate start
- ...will be applying AI at the scale and quality consumers expect to make real impact. Come be a part of what’s next. We're seeking a technical leader to help shape the strategy, development & delivery of GenAI tools and agentic systems across Netflix Games. This is a pivotal...Hourly payFull timeImmediate startFlexible hours
- We are looking for a Member of Technical Staff with strong Python skills and a passion for building scalable platforms for AI and ML workloads. As MTS, you'll influence strategic decisions, partner closely with the founding team, and play a critical role in shaping Activeloop...
- ...forefront of AI—backed by world-class institutional investors and strategic partners. We are looking for an exceptional Member of Technical Staff to help design, build, and scale core components of our next-generation AI compute platform. Key Responsibilities Core Engineering...
- ...seeking experienced Senior ASIC Designers. This role demands proven technical expertise in advanced ASIC design flows, and leadership in... ..., verification, physical design, firmware, DFT, and post silicon domains to ensure successful system level functionality...Visa sponsorshipRelocation package
- Member of Technical Staff - Applied AI Research San Francisco, CA; Sunnyvale, CA DoorDash’s mission is to empower local economies. With AI, we believe... ...to support - rather than replace - human decision‑making; trained personnel make final decisions with meaningful human review...Hourly payWork at officeLocal areaImmediate startRemote workFlexible hours
$120k - $200k
...full-cycle data engineering, Abaka AI provides the foundation for building high-performance AI systems. About the Role As a Member of Technical Staff, Platform, you'll build full-stack product features Expers Talent platform for our end-to-end, from a rough idea in a...Flexible hours$17.2 - $17.7 per hour
...McDonald's Crew Team Member McDonald's and its independent franchisees care about their... ...education programs and world-class training, we provide opportunities that inspire confidence... ...a Crew Member at McDonald's. This job posting contains some information about what it...TrainingHourly payFull timePart timeNight shiftWeekend work$20 - $21.5 per hour
...This job posting is for a position in a restaurant owned and operated by an independent franchisee... ...education programs and world-class training, we provide opportunities that inspire... ...working at one of our restaurants. A Crew Team Member at McDonald’s is more than just a...TrainingHourly payFull timePart timeLocal areaNight shiftWeekend work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Member of Technical Staff, Post-training. Be the first to apply!
- operations support technician San Jose, CA
- product support technician San Jose, CA
- senior technical analyst San Jose, CA
- systems support technician San Jose, CA
- user support analyst San Jose, CA
- technical support specialist San Jose, CA
- help desk assistant San Jose, CA
- mri tech aide San Jose, CA
- desktop support analyst San Jose, CA
- senior IT support technician San Jose, CA

