AI Platform Engineer, Training and Inference
$274k - $304kSaviynt
Saviynt's AI-powered identity platform manages and governs human and non-human access to all of an organization's applications, data, and business processes. Customers trust Saviynt to safeguard their digital assets, drive operational efficiency, and reduce compliance costs. Built for the AI age, Saviynt is today helping organizations safely accelerate their deployment and usage of AI. Saviynt is recognized as the leader in identity security, with solutions that protect and empower the world's leading brands, Fortune 500 companies and government institutions. For more information, please visit AI Platform Engineer - Training & Inference Saviynt's AI-powered identity platform manages and governs human and non-human access to all of an organization's applications, data, and business processes. Customers trust Saviynt to safeguard their digital assets, drive operational efficiency, and reduce compliance costs. Built for the AI age, Saviynt is today helping organizations safely accelerate their deployment and usage of AI. Saviynt is recognized as the leader in identity security, with solutions that protect and empower the world's leading brands, Fortune 500 companies and government institutions. For more information, please visit The AI Platform team is building the compute layer that trains, evaluates, and serves every AI model at Saviynt. We need an ML Platform Engineer to own distributed training on Ray + H100s, the multi-engine LLM inference mesh (vLLM, SGLang, NVIDIA Triton), and the full model promotion lifecycle - from shadow mode through canary rollout to GA. The AI Platform team's mission is to build a secure, scalable, product-agnostic AI foundation that enables Saviynt's identity products to deliver measurable AI-powered outcomes. Training & Inference is the engine - it turns data into deployed models that make Saviynt's products smarter. What You Will Be Doing
- Own the Ray ecosystem end-to-end: manage KubeRay on GKE, tune Ray Core Task/Actor scheduling, operate the Plasma distributed object store, and configure Ray Data for GPU-direct streaming from GCS/S3
- Operate distributed training with Ray Train: configure TorchTrainer + DDP/NCCL for multi-node H100 clusters, manage checkpoint lifecycle, implement spot-preemption recovery, and integrate warm-start fine-tuning for retrain pipelines
- Build and operate the LLM inference mesh with Ray Serve: compose vLLM (PagedAttention), SGLang (RadixAttention), and NVIDIA Triton (TensorRT/ONNX) as a unified deployment graph with Plasma zero-copy memory sharing
- Optimise inference performance: configure fractional GPU allocation, enable continuous batching, implement per-engine autoscaling based on request queue depth, and tune KV-cache block sizes
- Design and operate the model routing layer: capability-based, version-based, and tenant-based routing with cost-aware fallback between self-hosted SLMs and cloud LLMs
- Build RL training infrastructure: define Flyte workflows for RL pipelines (rollout, reward shaping, policy update, evaluation), integrate Ray RLlib or custom PPO/GRPO loops with Ray Train, and manage replay buffer persistence on GCS
- Operate the full model promotion lifecycle: quality gate - integration tests - load tests (k6) - shadow mode - A/B gate - canary (10%-100%) with golden-signal auto-rollback
- Operate the retrain pipeline: drift detection triggers, warm-start retraining, relative quality gates (V2 >= V1 - 2%), and automated Flyte DAG through to canary
- Integrate RAG retrieval into the inference mesh: vector similarity search, context assembly, and prompt construction before LLM inference
- Experience in ML engineering with time in an ML platform or MLOps role
- Production Ray depth: Ray Train, Serve, Core, and Data - debugged real production failures including NCCL timeouts, Plasma OOM, and Serve autoscaling lag
- LLM serving engines: hands-on with vLLM, SGLang, or NVIDIA Triton - PagedAttention, prefix caching, and continuous batching tuned for latency/throughput targets
- Distributed training: DDP, FSDP, NCCL collectives, gradient checkpointing, and mixed precision (BF16/FP8)
- RL working knowledge: PPO, policy gradient, or RLHF - able to translate an algorithm into distributed compute primitives
- Model lifecycle operations: MLflow registry, shadow/A/B/canary patterns, and auto-
- Vector databases: Pgvector or Qdrant - ANN index strategies, embedding upsert, and query latency tuning under inference load
- Strong Python and PyTorch; Flyte or equivalent ML orchestrator
- Quantization (nice to have): INT8/INT4/FP8 post-training quantization (GPTQ, AWQ, or bitsandbytes)
- Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent
- ...individual can thrive.Job DescriptionThe AI Inference Engineer plays a critical role in the AI... ...including Docker, Kubernetes, and cloud platforms such as AWS, GCP, and Azure. Hardware... ...assignment, compensation, promotion, benefits, training, discipline, and termination. F5...TrainingFull timeLocal areaImmediate start
- ...generation computing experiences—from AI and data centers, to PCs, gaming and... ...career. THE ROLEWe are hiring AI / ML Platform Engineers to build the platform layer that... ...large-scale agent execution, distributed training and inference, experiment tracking, benchmark automation...Training
- ...enabling human life on Mars.SOFTWARE ENGINEER, INFERENCE (AI DATA ENGINEERING)The application software... ...a high-performance AI inference platform that serves the best models internally... ...as we'll also be providing support for training workloads.RESPONSIBILITIES:Develop highly...TrainingPermanent employmentTemporary workRemote workWorldwideWeekend work
- ...Senior Lead Software Engineer Be an integral part... ...Sector, Infrastructure Platforms team, you are an integral... ...optimized for AI/ML workloads. Partner... ...and skills Formal training or certification on software... ..., ML training, and inference. Experience with...TrainingFor contractors
$229.9k - $262.4k
Sr. Lead AI Engineer (Inference Optimization, FM hosting, AI Platform) Overview: At Capital One, we are creating responsible and reliable AI systems, changing banking... ...AI software components including foundation model training, large language model inference, similarity search...TrainingFull timePart timeLocal area$229.9k - $286.2k
AI Engineer 5 (FM Hosting, LLM Inference) At Capital One, we are creating responsible and reliable AI systems,... ...millions of customers. Our AI models and platforms empower teams across Capital One to... ...including foundation model training, large language model inference, agents...TrainingFull timePart timeLocal area$107.5k - $204.5k
...team: We are seeking an experienced AI Platform Engineer to design, build, and operate the... ...experience with AI gateways, model serving, inference platforms, model routing, agent... ...work experience, location, education/training, and key skills.Hired applicants may be...TrainingPermanent employmentFull timeTemporary workWork experience placementWork at officeRemote workFlexible hours$117.7k - $221.4k
...important scenarios, prepare training-ready data, and support fast... ...cost efficient for embodied AI systems. We believe the next... ...operating model reflects how Cola engineers think: build durable... ...processing, featurization, and inference foundations that power scalable...TrainingFull timeLocal areaRemote workWork from homeRelocation packageFlexible hours- Senior Lead Software Engineer Be an integral part of... ...Sector, Infrastructure Platforms team, you are an integral... ...optimized for AI/ML workloads. Partner... ...capabilities, and skills Formal training or certification on... ..., ML training, and inference. Experience with Infrastructure...TrainingFor contractors
$240k - $260k
...leader in identity security, delivering an AI-powered platform that governs and secures access to... ...the architectural direction for how training data flows, evolves, and is governed across... ...Platform. You define the standards ML engineers and scientists build on, and ensure...Training$136.3k - $231.7k
...expert teams of physicists, engineers, data scientists and... ...is seeking a motivated AI Engineer with a growth... ...innovative inspection platforms.Key ResponsibilitiesDeep... ...Computer VisionDesign, train, and deploy deep... ...accuracy; optimize models for inference throughput, including...TrainingMinimum wageFull timeWork experience placementFlexible hours- ...Systems builds the world's largest AI chip, 56 times larger than... ...to deliver industry-leading training and inference speeds; over 10 times faster... ...RoleWe're hiring a Software Engineer to help contribute to projects on our Inference Platform team. Our team primarily owns...Training
- ...builds the world's largest AI chip, 56 times larger than... ...to deliver industry-leading training and inference speeds; over 10 times faster... ...systems that power engineering workflows across Cerebras.Our... ...tools, and reusable software platforms that allow engineers to build...Training
$250.8k - $286.2k
AI Engineer 5 (Gen AI Platform Services: Agentic AI, Guardrails, Evaluation) At Capital One, we are creating responsible and reliable... ...software components including foundation model training, large language model inference, agents and multi-agent workflows, similarity...TrainingFull timePart timeLocal area$229.9k - $286.2k
AI Engineer 5 (Gen AI Platform Services) Overview: At Capital One, we are creating responsible and reliable AI systems, changing... ...AI software components including foundation model training, large language model inference, agents and multi-agent workflows, similarity...TrainingFull timePart timeLocal area$229.9k - $286.2k
AI Engineer 5 (Gen AI Platform Services - Agentic AI) Overview: At Capital One, we are creating responsible and reliable AI systems... ...software components including foundation model training, large language model inference, agents and multi-agent workflows, similarity...TrainingFull timePart timeLocal area$151.8k - $332.2k
What you can expect We are looking for an AI Inference Engineer with a solid background in speech recognition and model inference. In this role... ...is to deliver the most unique AI-powered collaboration platform to users across the globe.ResponsibilitiesDeveloping state-...Full timeWork at officeRemote work- Lead AI Native Engineer RoboForce is an AI robotics company developing Physical AI-powered... ...- Experience with post-training, model evaluation, inference, or research infrastructure at a... ...with MCP infrastructure, developer platforms, knowledge graphs, RAG, or secure...TrainingWork at officeVisa sponsorship
- ...worry. Arlo's deep expertise in AI- and CV-powered analytics,... ...vehicle/package recognition, custom-trained detections, video captioning... ...About the role As a Staff AI Engineer for Applied AI, you'll be a... ...for agent runs. Keep production inference fast and economical - serving...TrainingOdd jobNight shift
$200k
...universal interface between humans and machines. While today's AI largely operates through chat boxes and decade-old devices, Hark... ...understand attention, KV-cache behavior, and where transformer inference actually spends its time and memory bandwidth. ~ You reason...Full time$230k - $350k
...an experienced Member of Technical Staff, Inference Systems, to build and optimize a high-performance AI inference platform from the ground up. This role is focused... ...inference systems and strong systems engineering skills, with Rust experience highly valued....Full timeWork at office$183.6k - $297k
...Integrity, and Inclusion. We weave AI into the fabric of everything... .... As a Principal Software Engineer, you will own the technical vision for our AI-powered platform, shaping how agentic workflows... ...optimize LLM performance and reduce inference costSolid skills in multi-...Full timeWork at office$100k
...the industry on cutting-edge AI technology, revolutionizing... ...software models, compilers, platforms, networking, and semiconductors... .../ Signal Integrity Engineer to design and validate high-... ...technologies for next-generation AI inference and training clusters. This role is on-...TrainingPermanent employment- ...computing experiences—from AI and data centers, to... ...Deployed Research Engineer to build, evaluate, and... ...agentic workflows, and post-training techniques.THE PERSON:... ...the evolution of our AI platform.What Makes This Role... ...EvaluationAI Training or Inference InfrastructureDeep understanding...Training
- ...builds the world's largest AI chip, 56 times larger than GPUs... ...to deliver industry-leading training and inference speeds; over 10 times faster... ...Data Center Infrastructure Engineer to turn cluster architecture... ...:Build a breakthrough AI platform beyond the constraints of the...TrainingContract workRemote work
- ...computing experiences—from AI and data centers, to... ...ROLEWe are hiring AI Engineers to build recursive... ...compute workloads and platforms.The work requires turning... ...learning, post-training, reward modeling, and... ...distributed training/inference systems.Experience with...Training
$152k - $241.5k
We are seeking a Senior AI/ML Performance and Efficiency Engineer, GPU Clusters at NVIDIA to join our AI Efficiency... ...investigating, and resolving, training & inference performance end to endDebugging... ...with cloud computing platforms (e.g., AWS, GCP, Azure) in addition...TrainingFull timeRemote work$152k - $241.5k
...outstanding Senior High Performance AI Engineers to build the next generation... ...full agentic AI stack—from training and improving models, to... ...stack, from models and inference through compilers, runtimes,... ...CUDA or comparable accelerator platforms, and the ability to work...TrainingFull timeRemote work$120k - $220k
...the Content Intelligence platform shaping the future... ...information powered by advanced AI, recommendation systems... ...cycle.We're hiring the engineer who owns this agent end... ...checkpoint — Train and ship on past CPI-winning... ...a LoRA, optimized inference, built a ComfyUI workflow...TrainingFull timeLocal areaWork from home$144k - $236k
...the team.Responsibilities: AI is at the core of how... ...marketing, content, and trust platforms. As a Senior AI Software Engineer you will own end-to-end... ...scale. You won't just train models, you will own the... ...quality improvement (i.e. inference/training efficiency, engineer...TrainingFor contractorsWork at officeImmediate startFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Platform Engineer, Training and Inference. Be the first to apply!


