Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Platform Engineer, Training and Inference

$274k - $300k

Saviynt

Job Description

Job Description

Saviynt's AI-powered identity platform manages and governs human and non-human access to all of an organization's applications, data, and business processes. Customers trust Saviynt to safeguard their digital assets, drive operational efficiency, and reduce compliance costs. Built for the AI age, Saviynt is today helping organizations safely accelerate their deployment and usage of AI. Saviynt is recognized as the leader in identity security, with solutions that protect and empower the world’s leading brands, Fortune 500 companies and government institutions. For more information, please visit 

AI Platform Engineer – Training & Inference


Saviynt's AI-powered identity platform manages and governs human and non-human access to all of an organization's applications, data, and business processes. Customers trust Saviynt to safeguard their digital assets, drive operational efficiency, and reduce compliance costs. Built for the AI age, Saviynt is today helping organizations safely accelerate their deployment and usage of AI. Saviynt is recognized as the leader in identity security, with solutions that protect and empower the world's leading brands, Fortune 500 companies and government institutions. For more information, please visit

The AI Platform team is building the compute layer that trains, evaluates, and serves every AI model at Saviynt. We need an ML Platform Engineer to own distributed training on Ray + H100s, the multi-engine LLM inference mesh (vLLM, SGLang, NVIDIA Triton), and the full model promotion lifecycle — from shadow mode through canary rollout to GA.

The AI Platform team's mission is to build a secure, scalable, product-agnostic AI foundation that enables Saviynt's identity products to deliver measurable AI-powered outcomes. Training & Inference is the engine — it turns data into deployed models that make Saviynt's products smarter.

 

What You Will Be Doing


• Own the Ray ecosystem end-to-end: manage KubeRay on GKE, tune Ray Core Task/Actor scheduling, operate the Plasma distributed object store, and configure Ray Data for GPU-direct streaming from GCS/S3
• Operate distributed training with Ray Train: configure TorchTrainer + DDP/NCCL for multi-node H100 clusters, manage checkpoint lifecycle, implement spot-preemption recovery, and integrate warm-start fine-tuning for retrain pipelines
• Build and operate the LLM inference mesh with Ray Serve: compose vLLM (PagedAttention), SGLang (RadixAttention), and NVIDIA Triton (TensorRT/ONNX) as a unified deployment graph with Plasma zero-copy memory sharing
• Optimise inference performance: configure fractional GPU allocation, enable continuous batching, implement per-engine autoscaling based on request queue depth, and tune KV-cache block sizes
• Design and operate the model routing layer: capability-based, version-based, and tenant-based routing with cost-aware fallback between self-hosted SLMs and cloud LLMs
• Build RL training infrastructure: define Flyte workflows for RL pipelines (rollout, reward shaping, policy update, evaluation), integrate Ray RLlib or custom PPO/GRPO loops with Ray Train, and manage replay buffer persistence on GCS

• Operate the full model promotion lifecycle: quality gate → integration tests → load tests (k6) → shadow mode → A/B gate → canary (10%→100%) with golden-signal auto-rollback

• Operate the retrain pipeline: drift detection triggers, warm-start retraining, relative quality gates (V2 >= V1 − 2%), and automated Flyte DAG through to canary
• Integrate RAG retrieval into the inference mesh: vector similarity search, context assembly, and prompt construction before LLM inference

 

What You Bring


• Experience in ML engineering with time in an ML platform or MLOps role
• Production Ray depth: Ray Train, Serve, Core, and Data — debugged real production failures including NCCL timeouts, Plasma OOM, and Serve autoscaling lag
• LLM serving engines: hands-on with vLLM, SGLang, or NVIDIA Triton — PagedAttention, prefix caching, and continuous batching tuned for latency/throughput targets
• Distributed training: DDP, FSDP, NCCL collectives, gradient checkpointing, and mixed precision (BF16/FP8)
• RL working knowledge: PPO, policy gradient, or RLHF — able to translate an algorithm into distributed compute primitives

• Model lifecycle operations: MLflow registry, shadow/A/B/canary patterns, and auto-
rollback on golden signal degradation

• Vector databases: Pgvector or Qdrant — ANN index strategies, embedding upsert, and query latency tuning under inference load
• Strong Python and PyTorch; Flyte or equivalent ML orchestrator
• Quantization (nice to have): INT8/INT4/FP8 post-training quantization (GPTQ, AWQ, or bitsandbytes)
• Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent
practical experience or equivalent military experience

We offer you a competitive total rewards package, learning and tremendous opportunities to grow and advance in your career. At Saviynt, it is not typical for an individual to be hired at or near the top of the range for their role and final compensation decisions are dependent on many factors including, but not limited to location; skill sets; experience and training; licensure and certifications; and other relevant business and organizational needs.

You may also be eligible to participate in a Saviynt discretionary bonus plan, subject to the rules governing the program, whereby an award, if any, depends on various factors, including, without limitation, individual and organizational performance.

We offer you a competitive total rewards package, learning and tremendous opportunities to grow and advance in your career. At Saviynt, it is not typical for an individual to be hired at or near the top of the range for their role and final compensation decisions are dependent on many factors including but are not limited to location; skill sets; experience and training; licensure and certifications; and other relevant business and organizational needs. A reasonable estimate of the current range is $274,000 - $300,000 annually.

Saviynt is an amazing place to work. We are a high-growth, Platform as a Service company focused on Identity Authority to power and protect the world at work. You will experience tremendous growth and learning opportunities through challenging yet rewarding work which directly impacts our customers, all within a welcoming and positive work environment. If you're resilient and enjoy working in a dynamic environment you belong with us!

Security & Compliance
This role requires adherence to Saviynt’s information security and privacy policies and procedures, including annual security training.

Saviynt is an equal opportunity employer and we welcome everyone to our team.  All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, or veteran status.

We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.

Vacancy posted 28 days ago
Similar jobs that could be interesting for youBased on the AI Platform Engineer, Training and Inference in Milpitas, CA vacancy
  •  ...individual can thrive.Job DescriptionThe AI Inference Engineer plays a critical role in the AI...  ...including Docker, Kubernetes, and cloud platforms such as AWS, GCP, and Azure. Hardware...  ...assignment, compensation, promotion, benefits, training, discipline, and termination. F5... 
    Training
    Full time
    Local area
    Immediate start

    F5 Networks

    San Jose, CA
    5 days ago
  •  ...enabling human life on Mars.SOFTWARE ENGINEER, INFERENCE (AI DATA ENGINEERING)The application software...  ...a high-performance AI inference platform that serves the best models internally...  ...as we'll also be providing support for training workloads.RESPONSIBILITIES:Develop highly... 
    Training
    Permanent employment
    Temporary work
    Remote work
    Worldwide
    Weekend work

    SpaceX

    Palo Alto, CA
    4 days ago
  •  ...generation computing experiences—from AI and data centers, to PCs, gaming and...  ...career. THE ROLEWe are hiring AI / ML Platform Engineers to build the platform layer that...  ...large-scale agent execution, distributed training and inference, experiment tracking, benchmark automation... 
    Training

    AMD

    Santa Clara, CA
    6 days ago
  • $229.9k - $262.4k

    Sr. Lead AI Engineer (Inference Optimization, FM hosting, AI Platform) Overview: At Capital One, we are creating responsible and reliable AI systems, changing banking...  ...AI software components including foundation model training, large language model inference, similarity search... 
    Training
    Full time
    Part time
    Local area

    Capital One Financial Corp

    San Jose, CA
    5 days ago
  • $136.3k - $231.7k

     ...expert teams of physicists, engineers, data scientists and...  ...is seeking a motivated AI Engineer with a growth...  ...innovative inspection platforms.Key ResponsibilitiesDeep...  ...Computer VisionDesign, train, and deploy deep...  ...accuracy; optimize models for inference throughput, including... 
    Training
    Minimum wage
    Full time
    Work experience placement
    Flexible hours

    KLA-Tencor

    Milpitas, CA
    5 days ago
  •  ...Systems builds the world's largest AI chip, 56 times larger than...  ...to deliver industry-leading training and inference speeds; over 10 times faster...  ...We're hiring a Software Engineer to help contribute to projects on our Inference Platform team. Our team primarily owns... 
    Training
    Full time

    Cerebras Systems

    Sunnyvale, CA
    1 day ago
  •  ...builds the world's largest AI chip, 56 times larger than...  ...to deliver industry-leading training and inference speeds; over 10 times faster...  ...systems that power engineering workflows across Cerebras.Our...  ...tools, and reusable software platforms that allow engineers to build... 
    Training

    Cerebras Systems

    Sunnyvale, CA
    4 days ago
  • $229.9k - $262.4k

    Senior Lead AI Engineer (GenAI Platform Services) Overview At Capital One, we are creating responsible and reliable AI systems,...  ...AI software components including foundation model training, large language model inference, similarity search, guardrails, model evaluation,... 
    Training
    Full time
    Part time
    Local area

    Capital One Financial Corp

    San Jose, CA
    3 days ago
  • $240k - $260k

     ...leader in identity security, delivering an AI-powered platform that governs and secures access to...  ...the architectural direction for how training data flows, evolves, and is governed across...  ...Platform. You define the standards ML engineers and scientists build on, and ensure... 
    Training

    Saviynt

    Milpitas, CA
    4 days ago
  • $40 - $85 per hour

     ...genres and languages. The AI Platform team builds the...  ...systems run on, from large-scale training platforms and post-training/...  ...infrastructure to GPU-optimized inference and serving. You'll work...  ...Machine Learning, Computer Engineering, or a related field Research... 
    Training
    Hourly pay
    Full time
    Internship
    Immediate start
    Remote work
    Flexible hours

    Netflix

    Los Gatos, CA
    2 days ago
  • $50k - $120k

     ...Mission Altimate AI, founded in 2022 in San Francisco...  ...forefront of the AI-powered data engineering revolution. You can read more...  ..., Prompt Tuning, and Adapter Training Why you should join...  ...Kubernetes) for large-scale training, inference, and multi-agent... 
    Training
    Full time
    Worldwide

    Pa Early Stage Partners

    Sunnyvale, CA
    1 day ago
  • $200k - $400k

     ...machines at scale. At Scout AI, we’re developing Fury, the first...  ...for a Senior or Staff AI Engineer to join the Fury Orchestration...  ...contribute across the stack: model training and evaluation, reinforcement...  ...agent coordination, and edge inference optimization. Expect to... 
    Training
    Full time
    Relocation package

    Scout Ai

    Sunnyvale, CA
    1 day ago
  • $147k - $237.5k

     ...Integrity, and Inclusion. We weave AI into the fabric of everything...  .... As a Principal Software Engineer, you will own the technical vision for our AI-powered platform, shaping how agentic workflows...  ...optimize LLM performance and reduce inference costSolid skills in multi-... 
    Full time
    Work at office

    Palo Alto Networks

    Santa Clara, CA
    5 days ago
  • $152k - $241.5k

    We are seeking a Senior AI/ML Performance and Efficiency Engineer, GPU Clusters at NVIDIA to join our AI Efficiency...  ...investigating, and resolving, training & inference performance end to endDebugging...  ...with cloud computing platforms (e.g., AWS, GCP, Azure) in addition... 
    Training
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    5 days ago
  •  ...builds the world's largest AI chip, 56 times larger than...  ...to deliver industry-leading training and inference speeds; over 10 times faster...  ...." You'll sit between engineering, product, and customer-facing...  ...# Build a breakthrough AI platform beyond the constraints of the... 
    Training
    Full time

    Cerebras Systems

    Sunnyvale, CA
    1 day ago
  • $100k

     ...the industry on cutting-edge AI technology, revolutionizing...  ...software models, compilers, platforms, networking, and semiconductors...  .../ Signal Integrity Engineer to design and validate high-...  ...technologies for next-generation AI inference and training clusters. This role is on-... 
    Training
    Permanent employment

    Tenstorrent

    Santa Clara, CA
    4 days ago
  • $120k - $220k

     ...the Content Intelligence platform shaping the future...  ...information powered by advanced AI, recommendation systems...  ...cycle.We're hiring the engineer who owns this agent end...  ...checkpoint — Train and ship on past CPI-winning...  ...a LoRA, optimized inference, built a ComfyUI workflow... 
    Training
    Full time
    Local area
    Work from home

    News Break

    Mountain View, CA
    5 days ago
  • $184k - $287.5k

    We're looking for outstanding AI systems engineers to develop groundbreaking technologies in the inference systems software stack! We build innovative AI systems software...  ...and library solutions for LLM inference and training (e.g. FlashInfer, Flash Attention)Expertise in... 
    Training
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $150k - $250k

     ...build state-of-the-art AI capabilities for Cylake...  ...generation cybersecurity platform. Working at the...  ...learning solutions, and train deep learning and LLM models...  ...researchers and software engineers to develop, productize,...  ..., AI agents, model inference, and serving optimization... 
    Training
    Full time

    Cylake, Inc

    Sunnyvale, CA
    1 day ago
  •  ...computing experiences—from AI and data centers, to...  ...ROLEWe are hiring AI Engineers to build recursive...  ...compute workloads and platforms.The work requires turning...  ...learning, post-training, reward modeling, and...  ...distributed training/inference systems.Experience with... 
    Training

    AMD

    Santa Clara, CA
    5 days ago
  •  ...computing experiences—from AI and data centers, to...  ...Deployed Research Engineer to build, evaluate, and...  ...agentic workflows, and post-training techniques.THE PERSON:...  ...the evolution of our AI platform.What Makes This Role...  ...EvaluationAI Training or Inference InfrastructureDeep understanding... 
    Training

    AMD

    Santa Clara, CA
    6 days ago
  • $255.65k - $299k

     ...positionWhat you can expect:As a Senior AI Software Engineer, you will collaborate to design,...  ...software applications. You will ensure AI training, inference, deployment, and operation are...  ...out to build the best collaboration platform for the enterprise, and today help people... 
    Training
    Full time
    Work at office
    Remote work

    Zoom

    San Jose, CA
    4 days ago
  • $144k - $236k

     ...the team.Responsibilities: AI is at the core of how...  ...marketing, content, and trust platforms. As a Senior AI Software Engineer you will own end-to-end...  ...scale. You won't just train models, you will own the...  ...quality improvement (i.e. inference/training efficiency, engineer... 
    Training
    For contractors
    Work at office
    Immediate start
    Flexible hours

    Linkedin

    Mountain View, CA
    7 days ago
  •  ...IT Consulting services in the US. We are actively seeking AI DevOps Engineer for one of our client, Please share your resume with...  ...TFX) • Solid understanding of computer algorithms, AI training, inference, and AI powered use cases • Good to have infrastructure... 
    Training

    Rootshell Enterprise Technologies

    Santa Clara, CA
    3 days ago
  •  ...is building the best AI systems for heavy industries...  ...for Backend AI Engineers to design, build, and...  ...model-serving pipelines, inference and orchestration layers...  ...build the reusable platform capabilities that our...  ...models, even without training them yourself. Experience... 
    Training
    Full time

    Nexxa.ai

    Sunnyvale, CA
    1 day ago
  • $175k - $287k

     ...the team.Responsibilities: AI is at the core of how...  ...marketing, content, and trust platforms. As a Staff AI Software Engineer you will own end-to-end machine...  ...scale. You won't just train models, you will own the...  ...quality improvement (i.e. inference/training efficiency, engineer... 
    Training
    For contractors
    Work at office
    Immediate start
    Flexible hours

    Linkedin

    Mountain View, CA
    7 days ago
  • $170.5k - $315.49k

     ...Description: We are looking for a performance-obsessed AI Infrastructure Engineer to push LLM inference to its absolute limits on Intel's next-generation GPU...  ...skills, experience, and relevant education or training. Your recruiter can share more about the specific compensation... 
    Training
    Full time
    Local area
    Immediate start
    Shift work

    Intel

    Santa Clara, CA
    5 days ago
  • $189.3k - $290.7k

    Job DescriptionAs an AI Engineer on the team, you will build and deploy applied...  ...robust, high-performance inference pipelines. This role is not focused on training foundation models, but rather on...  ...Familiarity with GCP or other cloud platforms, Kubernetes, Docker, and... 
    Training
    Full time
    Local area
    Remote work
    Work from home
    Flexible hours

    General Motors

    Sunnyvale, CA
    7 days ago
  • $184k - $287.5k

    Joining NVIDIA's DGX Cloud AI Efficiency Team means...  ...resiliency of AI workloads - pre-training, post-training, inference. Our objective is to...  ...AI infrastructure software engineer to join our team. You'll be...  ...underpinning NVIDIA's AI platforms.Define meaningful and actionable... 
    Training
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    5 days ago
  •  ...builds the world's largest AI chip, 56 times larger than GPUs...  ...to deliver industry-leading training and inference speeds; over 10 times faster...  ...We're hiring a Staff Engineer to own major areas of the architecture...  ...of our Inference Cloud Platform. This team owns the cloud layer... 
    Training
    Full time

    Cerebras Systems

    Sunnyvale, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Platform Engineer, Training and Inference. Be the first to apply!