Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Platform Engineer, Training and Inference

$240k - $260k

Saviynt

Job Description

Job Description

AI Platform Engineer – Training & Inference


Saviynt's AI-powered identity platform manages and governs human and non-human access to all of an organization's applications, data, and business processes. Customers trust Saviynt to safeguard their digital assets, drive operational efficiency, and reduce compliance costs. Built for the AI age, Saviynt is today helping organizations safely accelerate their deployment and usage of AI. Saviynt is recognized as the leader in identity security, with solutions that protect and empower the world's leading brands, Fortune 500 companies and government institutions. For more information, please visit

The AI Platform team is building the compute layer that trains, evaluates, and serves every AI model at Saviynt. We need an ML Platform Engineer to own distributed training on Ray + H100s, the multi-engine LLM inference mesh (vLLM, SGLang, NVIDIA Triton), and the full model promotion lifecycle — from shadow mode through canary rollout to GA.

The AI Platform team's mission is to build a secure, scalable, product-agnostic AI foundation that enables Saviynt's identity products to deliver measurable AI-powered outcomes. Training & Inference is the engine — it turns data into deployed models that make Saviynt's products smarter.

 

What You Will Be Doing


• Own the Ray ecosystem end-to-end: manage KubeRay on GKE, tune Ray Core Task/Actor scheduling, operate the Plasma distributed object store, and configure Ray Data for GPU-direct streaming from GCS/S3
• Operate distributed training with Ray Train: configure TorchTrainer + DDP/NCCL for multi-node H100 clusters, manage checkpoint lifecycle, implement spot-preemption recovery, and integrate warm-start fine-tuning for retrain pipelines
• Build and operate the LLM inference mesh with Ray Serve: compose vLLM (PagedAttention), SGLang (RadixAttention), and NVIDIA Triton (TensorRT/ONNX) as a unified deployment graph with Plasma zero-copy memory sharing
• Optimise inference performance: configure fractional GPU allocation, enable continuous batching, implement per-engine autoscaling based on request queue depth, and tune KV-cache block sizes
• Design and operate the model routing layer: capability-based, version-based, and tenant-based routing with cost-aware fallback between self-hosted SLMs and cloud LLMs
• Build RL training infrastructure: define Flyte workflows for RL pipelines (rollout, reward shaping, policy update, evaluation), integrate Ray RLlib or custom PPO/GRPO loops with Ray Train, and manage replay buffer persistence on GCS

• Operate the full model promotion lifecycle: quality gate → integration tests → load tests (k6) → shadow mode → A/B gate → canary (10%→100%) with golden-signal auto-rollback

• Operate the retrain pipeline: drift detection triggers, warm-start retraining, relative quality gates (V2 >= V1 − 2%), and automated Flyte DAG through to canary
• Integrate RAG retrieval into the inference mesh: vector similarity search, context assembly, and prompt construction before LLM inference

 

What You Bring


• Experience in ML engineering with time in an ML platform or MLOps role
• Production Ray depth: Ray Train, Serve, Core, and Data — debugged real production failures including NCCL timeouts, Plasma OOM, and Serve autoscaling lag
• LLM serving engines: hands-on with vLLM, SGLang, or NVIDIA Triton — PagedAttention, prefix caching, and continuous batching tuned for latency/throughput targets
• Distributed training: DDP, FSDP, NCCL collectives, gradient checkpointing, and mixed precision (BF16/FP8)
• RL working knowledge: PPO, policy gradient, or RLHF — able to translate an algorithm into distributed compute primitives

• Model lifecycle operations: MLflow registry, shadow/A/B/canary patterns, and auto-
rollback on golden signal degradation

• Vector databases: Pgvector or Qdrant — ANN index strategies, embedding upsert, and query latency tuning under inference load
• Strong Python and PyTorch; Flyte or equivalent ML orchestrator
• Quantization (nice to have): INT8/INT4/FP8 post-training quantization (GPTQ, AWQ, or bitsandbytes)
• Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent
practical experience or equivalent military experience

We offer you a competitive total rewards package, learning and tremendous opportunities to grow and advance in your career. At Saviynt, it is not typical for an individual to be hired at or near the top of the range for their role and final compensation decisions are dependent on many factors including, but not limited to location; skill sets; experience and training; licensure and certifications; and other relevant business and organizational needs.

You may also be eligible to participate in a Saviynt discretionary bonus plan, subject to the rules governing the program, whereby an award, if any, depends on various factors, including, without limitation, individual and organizational performance.

We offer you a competitive total rewards package, learning and tremendous opportunities to grow and advance in your career. At Saviynt, it is not typical for an individual to be hired at or near the top of the range for their role and final compensation decisions are dependent on many factors including but are not limited to location; skill sets; experience and training; licensure and certifications; and other relevant business and organizational needs. A reasonable estimate of the current range is $240,000 - $260,000 annually.

We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.

Vacancy posted 12 days ago
Similar jobs that could be interesting for youBased on the AI Platform Engineer, Training and Inference in Milpitas, CA vacancy
  • $229.9k - $262.4k

    Senior Lead AI Engineer (FM Hosting, LLM Inference) Overview: At Capital One, we are creating responsible and...  ...of customers. Our AI models and platforms empower teams across Capital One to...  ...components including foundation model training, large language model inference,... 
    Training
    Full time
    Part time
    Local area

    Capital One

    San Jose, CA
    1 day ago
  • $240k - $260k

     ...leader in identity security, delivering an AI-powered platform that governs and secures access to...  ...the architectural direction for how training data flows, evolves, and is governed across...  ...Platform. You define the standards ML engineers and scientists build on, and ensure... 
    Training

    Saviynt

    Milpitas, CA
    18 days ago
  • $229.9k - $262.4k

     ...Overview Sr. Lead AI Engineer (Inference Optimization, FM hosting, AI Platform) Overview: At Capital One, we are creating responsible and reliable AI...  ...AI software components including foundation model training, large language model inference, similarity search... 
    Training
    Full time
    Part time
    Local area

    Capital One

    San Jose, CA
    a month ago
  •  ...The Role We're looking for an AI Engineer to join the Fury Team with a...  ...engine of our robotic platforms. Your work will run on robots...  ...Responsibilities Design, train, and evaluate state-of-the-art...  ..., simulation, and real-world inference Conduct experiments to benchmark... 
    Training
    Full time
    Relocation package

    Scout Ai

    Sunnyvale, CA
    18 hours ago
  • $200k - $400k

     ...machines at scale. At Scout AI, we’re developing Fury, the first...  ...for a Senior or Staff AI Engineer to join the Fury Orchestration...  ...contribute across the stack: model training and evaluation, reinforcement...  ...agent coordination, and edge inference optimization. Expect to... 
    Training
    Full time
    Relocation package

    Scout Ai

    Sunnyvale, CA
    18 hours ago
  • $2,000 per month

     ...Etched Etched is building AI chips that are hard-...  ...Summary Etched’s Inference SW team enables optimal...  ...skilled and motivated engineer to join our team as we...  ...or accelerator hardware platforms. Solid systems knowledge...  ...using more FLOPs to train and run models, and the... 
    Training
    Full time
    Work at office
    Relocation package

    Etched

    San Jose, CA
    18 hours ago
  •  ...AI Engineer Opportunity Hope you are doing well Number of Position: 2 Only W2 I Abhishek...  ..., PyTorch ). Experience with model training, tuning, and evaluation. Knowledge of NVIDIA...  .... Understanding of generative AI and inference engines. Responsibilities: Preparing and... 
    Training
    Work visa

    Syntricate Technologies

    Santa Clara, CA
    3 days ago
  • $180k

     ...xAI’s mission is to create AI systems that can accurately understand...  ...motivated, and focused on engineering excellence. This organization...  ...xAI builds an end to end RL training framework to enable pretrain...  ...frameworks Experience in inference systems Proficiency in... 
    Training
    Full time
    Work at office
    Work from home
    Relocation

    Xai

    Palo Alto, CA
    18 hours ago
  •  ...AI Engineer — Customer Success & Services (F5) Location: Hybrid...  ...Compliance, and external vendor platforms to ship production-grade...  ...lifecycle: data pipelines, training, evaluation, fine-tuning, validation...  ...services and APIs (scalable inference, caching, batching, latency... 
    Training

    F5

    San Jose, CA
    18 hours ago
  • $159.5k - $271.2k

     ...expert teams of physicists, engineers, data scientists and problem-...  ...passionate and motivated Senior AI Engineer with experience...  ...~ Experience with LLM pre-training is optional, but a significant...  ...~ Understanding of cloud platforms and MLOps for scalable AI deployment... 
    Training
    Minimum wage
    Work experience placement
    Flexible hours

    KLA

    Milpitas, CA
    2 days ago
  •  ..., we build multi-agent AI systems that can automate...  ...workflows across platforms like SAP, Salesforce, Workday...  ...for a top-quality AI engineer with a strong focus on...  ...processing pipelines to train and fine-tune custom ML...  ...Azure, or GCP. Deploy inference endpoints and serve AI... 
    Training

    Tessera Labs

    San Jose, CA
    18 hours ago
  •  ...driving the transformation to AI-enabled software-defined...  ...a highly motivated Staff AI Engineer with expertise in data analytics...  ...Proven experience in designing, training, tuning, and evaluating...  ...techniques (e.g., A/B testing, causal inference). Preferred: Master... 
    Training
    Full time
    Work at office
    Worldwide
    Flexible hours
    Shift work

    The Posted Salary Range

    Sunnyvale, CA
    18 hours ago
  • $130k - $220k

     ...AIDC (AI Data Center) PM Engineer Location: Hybrid Remote – Milpitas, CA (3 days onsite per week) Term: Perm, Full-Time, Direct Hire Salary...  ...cost, and risk. Manage the onsite PM team, with proper training to improve the capability of the team members.... 
    Training
    Permanent employment
    Full time
    Remote work
    Visa sponsorship
    Free visa
    3 days per week

    Comrise

    Milpitas, CA
    1 day ago
  • $144.25k - $256.25k

     ...Staff AI Engineering - Agentic AI New York, NY, United States Phoenix...  ...Technology, we are building platforms, products, and governance...  ...LLM infrastructure, inference, and model gateways Evaluation...  ...~ Career development and training opportunities For a full... 
    Training
    Full time
    Work at office
    Local area
    Remote work
    Visa sponsorship
    Flexible hours
    3 days per week

    American Express

    Palo Alto, CA
    3 days ago
  •  ...Description We are seeking a motivated AI / Machine Learning Engineer with hands-on experience in...  ...data pipelines for model training and inference Collaborate with data scientists...  ...Implement AI solutions using cloud platforms and modern ML frameworks Document... 
    Training
    Full time
    Internship
    Relocation

    Hudson Manpower

    San Jose, CA
    18 hours ago
  • $136.3k - $231.7k

     ...expert teams of physicists, engineers, data scientists and...  ...is seeking a motivated AI Engineer with a growth...  ...innovative inspection platforms. Key Responsibilities...  ...Vision Design, train, and deploy deep learning...  ...accuracy; optimize models for inference throughput, including... 
    Training
    Minimum wage
    Full time
    Work experience placement
    Flexible hours

    KLA

    Milpitas, CA
    22 days ago
  • $197.3k - $225.1k

     ...Lead AI Engineer (Vision model customization, VLM) Overview At Capital One...  ...millions of customers. Our AI models and platforms empower teams across Capital One...  ...including foundation model training, large language model inference, similarity search, guardrails, model... 
    Training
    Full time
    Part time
    Local area

    Capital One

    San Jose, CA
    4 days ago
  • $229.9k - $262.4k

     ...Overview Senior Lead AI Engineer (Gen AI Platform Services, Agentic AI) Overview: At Capital One, we are creating responsible...  ...AI software components including foundation model training, large language model inference, similarity search, guardrails, model evaluation,... 
    Training
    Full time
    Part time
    Local area

    Capital One

    San Jose, CA
    more than 2 months ago
  • $229.9k - $262.4k

     ...responsible and reliable AI systems, changing banking...  ...class applied science and engineering teams to deliver our...  ...customers. Our AI models and platforms empower teams across Capital...  ...foundation model training, large language model inference, similarity search, guardrails... 
    Training
    Full time
    Part time
    Local area

    Information Technology Senior Management Forum

    San Jose, CA
    6 days ago
  • $144k - $236k

     ...responsible for scaling LinkedIn’s AI model training, feature engineering and serving with hundreds of billions...  ...with the state-of-the-art Feature Platform, which empowers AI Users to...  ...of native cloud, enable GPU based inference for a large variety of use cases, cuda... 
    Training
    Full time
    For contractors
    Work experience placement
    Work at office
    Flexible hours

    LinkedIn

    Mountain View, CA
    1 day ago
  • Figure is an AI robotics company developing autonomous general-purpose humanoid robots...  ...autonomy. We are looking for a Helix AI Engineer, Modeling to design and advance the core...  ...from initial research and prototyping to training and deployment Collaborate closely... 
    Training
    Full time
    Work at office

    Figure

    San Jose, CA
    18 hours ago
  •  ...Figure is an AI robotics company developing autonomous general-purpose humanoid robots...  ...autonomy. We are looking for a Helix AI Engineer, Pretraining to build large-scale foundation...  .... Responsibilities Design and train large-scale foundation models across multimodal... 
    Training
    Full time
    Work at office

    Figure

    San Jose, CA
    18 hours ago
  • Figure is an AI robotics company developing autonomous general-purpose humanoid robots....  ...autonomy. We are looking for a Helix AI Engineer, Reinforcement Learning to develop learning...  ...real-world and simulated environments Train policies that learn from interaction, feedback... 
    Training
    Full time
    Work at office

    Figure

    San Jose, CA
    18 hours ago
  • Figure is an AI robotics company developing autonomous general-purpose humanoid robots...  ...autonomy. We are looking for a Helix AI Engineer, Video Pretraining to lead the...  ...development of large-scale video foundation models trained on diverse real-world and robot-collected... 
    Training
    Full time
    Work at office

    Figure

    San Jose, CA
    18 hours ago
  • Figure is an AI robotics company developing autonomous general-purpose humanoid robots...  ...autonomy. We are looking for a Helix AI Engineer, Generative AI to build and scale...  ...the physical world. This role focuses on training and deploying diffusion and generative models... 
    Training
    Full time
    Work at office

    Figure

    San Jose, CA
    18 hours ago
  •  ...Helix AI Engineer, Robot Learning Figure is an AI robotics company developing autonomous...  .... Responsibilities Design, train, evaluate, and deploy learning-based visuomotor...  ...humanoids or highly dexterous robotic platforms Publication record in robot learning,... 
    Training
    Full time

    Figureai

    San Jose, CA
    18 hours ago
  • $126k - $248k

     ...About the Role We’re looking for a Senior Engineer to help build the next-generation inference platform that supports embedding models used for semantic search, retrieval, and AI-native experiences in the company Atlas. You’ll join the broader Search and AI Platform organization... 
    Local area
    Flexible hours

    United States Digital Space LLC

    Palo Alto, CA
    4 days ago
  • $227.5k - $300k

     ...driving the transformation to AI-enabled software-defined vehicles...  ...are seeking a Senior Staff AI Engineer with a combination of...  ...and design of an extensible AI platform. Drive architectural consensus...  ...development, including modeling, training, tuning, validating, deploying... 
    Training
    Full time
    Work at office
    Worldwide
    Flexible hours
    Shift work

    Sonatus

    Sunnyvale, CA
    18 hours ago
  •  ...driver, combining cutting-edge AI with automotive-grade...  ...the automakers and mobility platforms a clear path to AVs at...  ...the Role As a software engineering intern, you will work closely...  ...and optimize on-cloud training and onboard inference. Our solutions include a distributed... 
    Training
    Internship

    Nuro

    Mountain View, CA
    18 hours ago
  • $250k - $300k

     ...intelligence . As the only vertically integrated AI infrastructure company built from the...  ...in production. That means owning the inference stack end to end: profiling where time...  ...will also work directly with customer engineering teams to tailor deployments to their needs... 
    Temporary work

    Crusoe

    Sunnyvale, CA
    8 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Platform Engineer, Training and Inference. Be the first to apply!