Senior Machine Learning Engineer
$188.5k - $282.7kRubrik
About the Team & Role:We're building SAGE, Rubrik's Semantic AI Governance Engine, which is the first system designed to monitor, govern, and remediate autonomous AI agents in real time. SAGE powers Rubrik Agent Cloud: enterprises define governance policies in natural language, and SAGE's custom small language models act as judges on every agent action. These models are fast enough to sit in the live request path and accurate enough that customers trust them with allow/block decisions on production traffic.At its core, SAGE is "LLM-as-judge" applied to AI governance, utilizing the same technique most teams use for offline evaluation but productionized for real-time enforcement at enterprise scale. Our first-generation SLM Policy Guard already outperforms the larger frontier models we've benchmarked against on accuracy while running approximately 5x faster on the same workload. We're hiring to push that lead even further.As an Applied ML Engineer on the SAGE team, you'll work end-to-end across the model lifecycle: curating data, training small models, serving them at production latency, and closing the feedback loop with real customer signals. The models you build don't just enforce policies in the live request path; they will also drive Agent Rewind, Rubrik's capability to instantly and precisely undo destructive autonomous-agent actions and restore the affected data to a trusted state. We're a collaborative, applied team that ships models to enterprise customers within weeks, and we're passionate about proving that small, specialized models can outperform frontier LLMs at the problems that matter most for AI safety and governance.Nature of the Specialized DutiesTraining, Fine-Tuning, and Distilling Production Small Language Models and Classifiers (25% of time)Owning the full training lifecycle for the SLMs and classifiers in SAGE's real-time enforcement path, including base-model selection, supervised fine-tuning, preference optimization (DPO/RLAIF), and distillation from frontier teacher models.Training anomaly and action-severity models that catch novel agent-side attack patterns at real-time decision latency, such as supply-chain compromises or emergent destructive behaviors not covered by any explicit policy. Severity scores route the highest-impact events to Agent Rewind for precise remediation.Designing adversarial training pipelines like purpose-built adversarial agents and automated red-teams whose outputs feed directly into the next training run, turning every discovered weakness into a permanent model improvement.Pushing the pareto frontier of accuracy, latency, and cost for governance-specific tasks through deliberate post-training choices (LoRA, quantization-aware training, distillation recipes, GRPO, etc.) and validating the wins on production traffic patterns.Engineering High-Performance Model Serving and Inference Infrastructure (25% of time)Designing multi-stage inference pipelines that handle both real-time enforcement (inline prompt, response, and tool-call blocking) and high-throughput batch workloads (offline scoring, back-testing, corpus mining) while processing billions of tokens daily across Global 2000 customer agent fleets.Optimizing live deployments through shared GPU pools, KV-cache-aware routing, continuous batching, FP8/INT8 quantization, and speculative decoding to minimize inference cost while holding sub-second P99 SLOs.Building serving-layer infrastructure that lets SAGE block agent prompts, responses, and tool calls in real time without becoming a latency bottleneck. This includes model gateway design, request routing, and graceful degradation.Owning canary, shadow, and A/B traffic patterns so new model variants are validated against live customer traffic before they take enforcement decisions.Building Synthetic Data Pipelines and Online + Offline Evaluation Frameworks (20% of time)Designing automated data curation pipelines that mine live customer environments (with privacy and tenancy guarantees) for high-value per tenant training examples, such as long-tail violations, near-miss policy edges, or novel agent behaviors, and routing them back into the training loop for each customer.Building automated policy back-testing by replaying historical agent traffic against new model and policy versions to catch regressions and recommend policy improvements before customer-visible deployment.Building online evaluation systems for live model decisions, including shadow scoring, drift detection, calibration monitoring, and policy-coverage gap analysis, ensuring quality regressions surface in minutes rather than weeks.Generating synthetic data using frontier teachers (adversarial prompts, policy-edge cases, multi-turn interactions) with evaluation that confirms synthetic data improves downstream quality, not just dataset size.Insights Mining, Failure Diagnosis, and Adaptive Model Improvement (15% of time)Building memory and context harnesses that fuse data sensitivity, identity, and historical agent behavior into real-time enforcement decisions to ensure SAGE reasons from each customer's specific context.Mining agent insights across millions of sessions to surface security gaps, which are then turned into new policy proposals, refinements to existing policies, and signals about upstream issues across the agent ecosystem (Google ADK, Azure AI Foundry, Vertex AI, and others).Building feedback loops that turn production decisions, customer-flagged false positives, and missed violations into one-click natural-language policy refinements to drive false-positive rates down without sacrificing recall.Diagnosing model failures end-to-end and distinguishing data, training-recipe, architecture, and serving-layer root causes so fixes land in the right layer the first time.Cross-Functional Collaboration and Translating Customer Reality into Modeling Problems (15% of time)Providing technical leadership on a pillar of the SAGE model stack (training infrastructure, eval methodology, serving architecture, or insights pipeline), mentoring engineers ramping into ML, and shaping the team's technical roadmap.Partnering with Product Management, customer-facing teams, and security analysts to translate customer agent-governance requirements into well-scoped modeling problems, and pushing back when ML is the wrong tool.Communicating model behavior, tradeoffs, and limitations clearly to non-ML stakeholders, such as product managers and enterprise security leaders, so model decisions are made with full context.Collaborating with Agent Cloud platform, security engineering, and AI research teams to integrate new SLMs into the real-time enforcement path with the right latency, observability, rollback, and tenancy guarantees.Minimum Requirements for the PositionEducation: A Bachelor's degree (or higher) in Computer Science, Machine Learning, Computer Engineering, Statistics, or a closely related technical field is required. Designing production SLM training and serving systems requires a deep theoretical understanding of modern deep learning, optimization, and systems performance.Specialized Technical Knowledge:2+ years of professional ML experience with demonstrable end-to-end production ownership; you have taken models from training to serving real customer traffic and stayed accountable for them through post-launch iteration.Proficiency in Python and PyTorch (or equivalent) for production-grade training and evaluation.Hands-on experience training, fine-tuning, or distilling language models or classifiers in a production setting, including SFT and at least one preference-optimization technique (DPO, RLAIF, or RLHF).Production experience with serving frameworks (vLLM, SGLang, TensorRT-LLM, or equivalent), including optimization involving continuous batching, KV-cache strategy, and inference-time quantization.Experience designing closed-loop ML systems, including the eval, telemetry, data-curation, and synthetic-data infrastructure that turns production signals back into training data and the next model release. You have built (not just used) at least one such loop.Comfort operating at production scale, including debugging models that handle high QPS in safety-critical request paths where errors have customer-visible consequences.Preferred Qualifications:Deep background in AI safety and red-teaming, including hands-on experience with adversarial ML, prompt injection defense strategies, and automated evaluation suites for enterprise-grade LLM safety.Expertise in model evaluation methodology, specifically building "LLM-as-judge" pipelines, calibration monitoring, and adversarial benchmarks that surface the subtle failure modes static metrics often overlook.Experience with context-fusion and retrieval systems that synthesize disparate signals - such as data sensitivity, user identity, and behavioral history - into high-fidelity model decisions.Production experience with low-latency inference for streaming or safety-critical request paths where model throughput and P99 SLOs are paramount.Mastery of label-efficient training and data mining, utilizing weak supervision, active learning, and embedding-based retrieval to surface the production examples that drive the most significant quality improvements.Hands-on knowledge distillation experience, successfully transferring capabilities from frontier teacher models to specialized, small-scale student models for production serving.Familiarity with the agentic ecosystem, including tool-use frameworks, model gateway architectures (MCP, LiteLLM, or equivalent), and autonomous agent patterns.Active open-source contributions to mainstream ML training, serving, or evaluation libraries.The minimum and maximum base salaries for this role are posted below; additionally, the role is eligible for bonus potential, equity and benefits. The range displayed reflects the minimum and maximum target for new hire salaries for the role based on U.S. location. Within the range, the salary offered will be determined by work location and additional factors, including job-related skills, experience, and relevant education or training.US Pay Range$188,500—$282,700 USDJoin Us in Securing and Accelerating the World's AI TransformationRubrik (RBRK), the Security and AI Operations Company, leads at the intersection of data protection, cyber resilience, and enterprise AI acceleration. Rubrik Security Cloud delivers complete cyber resilience by securing, monitoring, and recovering data, identities, and workloads across clouds. Rubrik Agent Cloud accelerates trusted AI agent deployments at scale by monitoring and auditing agentic actions, enforcing real-time guardrails, fine-tuning for accuracy and undoing agentic mistakes. Linkedin | X (formerly Twitter) | Instagram | Rubrik.comInclusion @ RubrikAt Rubrik, we are dedicated to fostering a culture where people from all backgrounds are valued, feel they belong, and believe they can succeed. Our commitment to inclusion is at the heart of our mission to secure the world’s data.Our goal is to hire and promote the best talent, regardless of background. We continually review our hiring practices to ensure fairness and strive to create an environment where every employee has equal access to opportunities for growth and excellence. We believe in empowering everyone to bring their authentic selves to work and achieve their fullest potential.Our inclusion strategy focuses on three core areas of our business and culture:Our Company: We are committed to building a merit-based organization that offers equal access to growth and success for all employees globally. Your potential is limitless here.Our Culture: We strive to create an inclusive atmosphere where individuals from all backgrounds feel a strong sense of belonging, can thrive, and do their best work. Your contributions help us innovate and break boundaries.Our Communities: We are dedicated to expanding our engagement with the communities we operate in, creating opportunities for underrepresented talent and driving greater innovation for our clients. Your impact extends beyond Rubrik, contributing to safer and stronger communities.Equal Opportunity Employer/Veterans/DisabledRubrik is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, or protected veteran status and will not be discriminated against on the basis of disability.Rubrik provides equal employment opportunities (EEO) to all employees and applicants for employment without regard to race, color, religion, sex, national origin, age, disability or genetics. In addition to federal law requirements, Rubrik complies with applicable state and local laws governing nondiscrimination in employment in every location in which the company has facilities. This policy applies to all terms and conditions of employment, including recruiting, hiring, placement, promotion, termination, layoff, recall, transfer, leaves of absence, compensation and training. Federal law requires employers to provide reasonable accommodation to qualified individuals with disabilities. Please contact us at View email address on click.appcast.io if you require a reasonable accommodation to apply for a job or to perform your job. Examples of reasonable accommodation include making a change to the application process or work procedures, providing documents in an alternate format, using a sign language interpreter, or using specialized equipment.EEO IS THE LAWNOTIFICATION OF EMPLOYEE RIGHTS UNDER FEDERAL LABOR LAWS
$195k - $230k
...fast, and make a difference, we’d love to hear from you! For more information, visit About the RoleWe are looking for a Senior Machine Learning Engineer to help evolve our large-scale recommendation systems and apply AI / LLM technologies to real-world production...SeniorFull timeLocal areaWork from home$230k - $265k
...meetings and conversations? Join our core AI team responsible for ML and work alongside industry-veteran scientists and engineers. As a Senior Machine Learning Engineer, you’ll bring your strong software engineering mindset to machine learning in order to scale and optimize...SeniorPermanent employment$174k - $299k
...Role Overview We are looking for experienced and innovative ML engineers with an entrepreneurial mindset to be part of the technical... ...years of experience in AI/ML and familiarity with the latest deep learning techniques Preferred Qualifications 10+ years of relevant...SeniorTemporary workFlexible hours$262k - $361k
...experts in energy, AI, software, engineering, and product to build tools... ...matters at global scale. Learn more about our team and our mission... ...of Tapestry’s multi-year machine learning strategy, bridging cutting... ...organization by mentoring senior and staff-level engineers,...SeniorFull timeRemote workFlexible hours- ...Aibreakingwire is seeking a Senior Software Engineer for our AI Platform in Menlo Park, CA. You will develop robust and scalable software systems to enhance AI model deployment, collaborating closely with research scientists to integrate cutting-edge algorithms. The ideal...Senior
$281k - $356k
...current solutions with future innovations. You'll build active learning and ML-aided labeling workflows to tackle rare, "longtail"... ...technology designs. You'll partner closely with Product and Engineering teams, especially those focused on data and automation. Your responsibilities...SeniorFull timeTemporary workRemote work- ...graph + metadata lake Experience: 6+ years industry overall experience with 3+ in ML Infra or MLE Expertise: back end software engineering strength with recent industry exp making ML systems more reliable/scalable (with opportunities to help improve model quality in...SeniorFull timeImmediate start
- ...We have a Senior Machine Learning Engineer role. Below are the key details: Main skill: Time Series, ML models training production experience, Vision Models - 2D/3D Project duration: 12 months Location: 3 days per week onsite in Mountain View, CA Recruitment...SeniorImmediate start3 days per week
- ...About the Role We are seeking a Senior Data / AI / ML Software Engineer with 7+ years of experience building data-intensive systems. This role is ideal for someone who enjoys designing and improving core platform components at the intersection of software engineering,...SeniorFull timeContract workInternship
$213k - $263k
...Senior Machine Learning Engineer, Computer Vision/VLM Waymo is an autonomous driving technology company with the mission to be the world's most trusted driver. Since its start as the Google Self-Driving Car Project in 2009, Waymo has focused on building the Waymo Driver...SeniorFull timeRemote work$164.2k - $205.2k
...products continue to evolve, we are seeking multiple GenAI Engineers from junior levels to more senior levels to drive the next phase of development. In 20... ...performance. What We’re Looking For 2-8 years of machine learning engineering experience in high-velocity, high-growth...SeniorWork at officeLocal area- ...streamline complex workflows, and continuously learn and adapt.Moveworks is trusted by over 5.... ...automation with Moveworks’ Reasoning Engine and natural language capabilities, we... ...DescriptionThe RoleWe are looking for a Machine Learning Engineer to help build cutting edge...SeniorWork at officeRemote workFlexible hours
$210.3k - $273.4k
...and planning in dynamic environments - is transforming how games and simulations come to life. We're looking for a Senior Machine Learning Engineer to lead the development of these foundational AI systems within the Unity engine, empowering creators to build smarter...SeniorWork at officeWorldwide$115k - $230k
...Senior ML Engineer At GEICO, we offer a rewarding career where your ambitions are met with endless possibilities. Every day we honor our iconic brand by offering quality coverage to millions of customers and being there when they need us most. We thrive on relentless...SeniorHourly payWork experience placementLocal area$175k - $230k
...Who we are Atoms is building the machines that power the next era of progress. Over... ...them into real environments, operate them, learn from them, and improve them until they work at scale. We are roboticists, engineers, operators, and builders. We believe the next...SeniorFull timeTemporary workWork at officeFlexible hours$214k - $289.5k
...Overview Come join Intuit as a Senior Staff Machine Learning Engineer (MLE). Senior Staff MLEs deliver end-to-end AI solutions that span multiple domains and products, influencing the strategic direction of machine learning and AI across the company. You will identify...Senior$183.7k - $248.6k
...The opportunity Unity is looking for a Senior Machine Learning Infrastructure Engineer to join our Vector Ads team, where we build the real-time systems that power Unity's global advertising platform. This is a high-scale, low-latency environment — processing billions...SeniorWork at officeRemote workWorldwideRelocation package$150k - $300k
..., Great Rewards and Great Careers. GEICO is seeking a Senior Staff AI engineer to join our AI org. This person will play key senior technical... ...lives. Great Careers: We offer a career where you can learn, grow, and thrive through personalized development programs...SeniorHourly payWork experience placementLocal areaFlexible hours$193.93k - $291.15k
...growing and we are looking for a Software Engineer to join our Sensor Data and Calibration... ...for an engineer with robotics and machine learning expertise to develop synthetic sensor simulation... ...and requirementsRole is scoped as a Senior/Staff IC with the flexibility to grow...SeniorImmediate startFlexible hours$184k - $287.5k
Intelligent machines powered by Artificial Intelligence computers that can learn, reason, and interact with people are no longer science fiction. Today, a self-driving... ...AV. We are seeking the best Machine Learning Engineers with a background in computer vision, LiDAR &...SeniorFull timeWorldwideNight shift$175.8k - $312.2k
SummaryWe are looking for a Machine Learning Engineer who will be converting abstract, high-level goals into concrete, measurable requirements. They will be proposing, implementing, evaluating, and shipping different AI/ML technologies and resulting data to achieve a given...SeniorRelocation$174.72k - $295.68k
...is dedicated to reshaping the future of transportation through cutting-edge R&D in AI, machine learning, and smart connectivity.We are looking for a full-time Machine Learning Engineer - AI Foundation, with deep knowledge and strong enthusiasm towards establishing a...SeniorFull time- ...delivered for millions of patients worldwide.We’re a team of engineers, clinicians, and innovators united by one purpose: to make... ...and implementing user-facing software and computer vision / machine learning algorithmsIterating with user feedback and delivering production...SeniorLocal areaWorldwideFlexible hours
$242k - $290k
...Software /Full-time /HybridAs a Perception Engineer, you will be instrumental in designing... ...ML frameworksExperience deploying learned models into productionExcellent collaboration... ...Sitting at the intersection of robotics, machine learning, and design, Zoox aims to...SeniorFull timeTemporary workRelocation package$145k - $200k
Santa Clara, CAUS Research and Development - Perception /Full-time /HybridWe are seeking a highly skilled Machine Learning Engineer with deep expertise in developing Bird’s Eye View (BEV) fusion models using multimodal sensor inputs, particularly LiDAR. You will play a...SeniorFull time- ...delivered for millions of patients worldwide.We’re a team of engineers, clinicians, and innovators united by one purpose: to make... ...structure of the lung, to take a biopsy at a target location. As a machine learning engineer on the Ion project, you will join a small team of...SeniorLocal areaWorldwideFlexible hours
$184k - $287.5k
We are seeking a Senior Machine Learning Engineer to join our end‑to‑end autonomous driving team! You will help build, train, and deploy large‑scale E2E driving models that leverage VLM/VLA architectures, and build a data flywheel that continuously improves our systems...SeniorFull time$189.72k - $332.01k
...in our recruiting process here.With more than 500 million users around the world and 300 billion ideas saved, Pinterest Machine Learning engineers build personalized experiences to help Pinners create a life they love. With just over 4,000 global employees, our teams...SeniorLocal areaRelocation package$224k - $356.5k
We are looking for outstanding Machine Learning Engineers to join our Physical AI teams. As the pioneers of the GPU—the visual cortex of modern computing—we are building the foundation for the next wave of AI that interacts with the physical world.This role is at the forefront...SeniorFull time$153.75k - $225k
...that values ownership, collaboration, and high standards. Our engineers, product leaders, and go-to-market teams work closely... ...command execution.Qualifications:Knowledge and passion in machine learning algorithms, GenAI, LLMs, and Agentic AIUnderstanding of agent...SeniorWork experience placementWork at office3 days per week
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Machine Learning Engineer. Be the first to apply!
- machine learning engineer Palo Alto, CA
- computer vision machine learning engineer Palo Alto, CA
- senior manager tax Palo Alto, CA
- senior devops Palo Alto, CA
- senior recruiter Palo Alto, CA
- senior director digital marketing Palo Alto, CA
- senior international accountant Palo Alto, CA
- senior vmware engineer Palo Alto, CA
- sr marketing manager Palo Alto, CA
- sr technical product manager Palo Alto, CA


