ML Engineer, Inference & Optimization
$250k - $350kPika
About the RoleWe are seeking Senior/Staff level Inference Engineers to accelerate the performance of Pika's AI-driven products. In this highly technical role, you will operate at the intersection of cutting-edge inference acceleration, GPU parallelism, advanced model deployment, and video generation technologies. Your expertise will drive significant improvements to model speed and efficiency, ensuring our creative AI systems deliver industry-leading user experiences at scale.You will design and optimize inference pipelines, implement state-of-the-art acceleration techniques, and work closely with researchers and engineers across the team to push the boundaries of what’s possible in real-time AI deployment. Your efforts will play a foundational role in powering the next generation of Pika’s video and language models.What You’ll DoAccelerate Inference: Lead and implement advanced inference acceleration techniques, including attention optimization and quantization for efficient model serving.Maximize GPU Parallelism: Engineer and optimize GPU strategies across tensor, sequence, and pipeline parallelism (TP, SP, PP) for maximal efficiency and scalability.Programming for Performance: Develop and optimize high-performance computing kernels and distributed workloads using CUDA and NCCL.Advance AI Deployment: Collaborate with research and engineering teams to bring state-of-the-art videogen and large language models into production.Improve Training Efficiency: (Bonus) Contribute to improvements in model training speed, stability, and resource utilization as part of our deployment lifecycle.Technical Excellence: Drive rigorous code reviews, participate in technical discussions, and mentor fellow engineers on best practices in inference and GPU programming.What We’re Looking ForExperience: 5+ years engineering experience, with a strong track record in inference acceleration and model deployment at scale.Inference Mastery: Proven expertise in inference optimization, including quantization, attention acceleration, and deep learning compiler stacks.GPU & Parallelism: Deep knowledge of GPU programming (CUDA, NCCL) and experience with SP, TP, PP, and other forms of parallelism for distributed inference.AI Domain Knowledge: Familiarity with video generation (videogen) models and large language models (LLMs).Collaboration: Strong cross-discipline communication skills; able to drive shared goals across research and engineering functions.Ownership Mindset: Self-driven, solutions-oriented, and capable of managing ambiguity in a fast-paced startup environment.Bonus: Experience in enhancing training efficiency, stability, or resource optimization for large models.Nice to HaveExperience with high-throughput video or real-time streaming model deploymentFamiliarity with distributed training and optimization toolkitsContributions to open source projects in AI infrastructure or deep learning compilersStartup or rapid prototyping experienceWhat We OfferCompetitive salary in the AI industryEquity in a fast-growing startup shaping the future of AIComprehensive health benefits, monthly stipends, company retreatsA supportive and collaborative office culture—we’re all building and launching togetherAbout PikaAt Pika, we're crafting a future where video creation is seamless, intuitive, and universally accessible. Our mission is to empower creativity by breaking down technical barriers using the transformative power of AI. We’re a tight-knit, energetic team based in Palo Alto, CA, valuing efficiency, curiosity, and the ambition to make a meaningful impact on the world.We work from our Palo Alto office 3–5 days a week and welcome applicants who are eager to contribute onsite.Compensation Range: $250K - $350KLocationPalo Alto HQEmployment TypeFull timeLocation TypeOn-siteDepartmentResearchCompensationbase salary $250K – $350K
$278.1k - $347.6k
...within that runtime. As our Principal Engineer for On-Device AI Inference & Systems, you will be the foremost... ...leaves research, through export, optimization, and kernel-level tuning, to a shipped... ...and own the integration between the ML runtime and the game engine: real-time...SuggestedWork at officeWorldwideRelocation package$209k - $313k
...other digital services.Snap Engineering teams build fun and technically... ...that quantify causal impact, optimize decision-making, and drive... ...Strong understanding of causal inference and modern approaches to estimating... ...tests) and leveraging causal ML in production...SuggestedFull timeLive inWork at officeLocal area- ...About the Role We’re looking for an Applied ML Engineer to design, evaluate, and scale recommendation and ranking systems... ...skips, conversions) and sparse explicit feedback. Optimize systems for real-time inference, scalability, and robustness under non-stationary user...SuggestedFull time
$117.7k - $221.4k
...This operating model reflects how Cola engineers think: build durable intermediate artifacts... ...quality, speed, and cost instead of optimizing any one of them in isolation.The RoleWe... ...data processing, featurization, and inference foundations that power scalable world understanding...SuggestedFull timeLocal areaRemote workWork from homeRelocation packageFlexible hours- ...About the job ML Engineer Our Client Is a rapidly growing Tier 1 VC backed... ...redefining how businesses learn from and optimize their in-person customer experiences.... ...preprocessing, model training, deployment, inference, and monitoring in production...SuggestedFull time
$213k - $263k
...Waymo Ml Software Engineer Waymo is an autonomous driving technology company with the mission to be the world's most trusted driver. Since... ...feature and experiment management, model development, optimization and monitoring. These efforts have resulted in making machine...Full timeRemote work$195.2k - $361.2k
...future of AI should belong to the people it servesRole SummaryMake models fast on the hardware people actually own. You optimize inference engines (llama.cpp, vLLM) for constrained local and edge environments — GPU/iGPUs, Vulkan backends — not datacenter H100 environment...Full timeInternshipLocal areaImmediate startShift work$184k - $287.5k
We're now looking for a Sr. Inference Engineer, for GPU Kernel Optimization! What does it take to push every LLM inference operation to its performance ceiling? Our LLM Inference Performance Analysis and Optimization team builds the answer from the ground up. We develop...Full time- ...world.Role OverviewAs our Staff Software Engineer, ML infra Engineer for Search & Discovery... ...Discovery organization is responsible for optimizing customers' navigation experience and... ...ML based ranking system and online ML inference servicesBuild strong cross-functional partnerships...Temporary work
$207k - $300k
...generated content through advanced context engineering and agentic feedback loops to identify... ...and serving AIGC content to optimize efficiency for low-latency surfaces.Navigate... ...Experience optimizing machine learning inference, resolving system-level bottlenecks, and...$250k - $350k
...alongside some of the world's leading ML systems engineers, including leaders behind Megatron-LM... ...infrastructure powering our large-scale training, inference, and reinforcement learning. You'll... .... What You'll Do Build and optimize large-scale training and reinforcement...Visa sponsorship- ...per day.Your Role and ResponsibilitiesWe are seeking to hire an ML Engineer to join our growing team! This is a hands-on technical role... ...improve existing ML models, training pipelines, and run-time inference performance• Expand the scope and capabilities of our models to...Full timeTemporary workPart time
$207k - $301k
...architecture in different modeling domains.5 years of experience with ML design and ML infrastructure (e.g., model deployment, model... ...language models, as well as building real-time and large scale inference infrastructure.Personalizing our Ads has proven to be extremely...$196k - $221k
...our core AI team responsible for ML and work alongside industry-veteran scientists and engineers. As a Machine Learning Engineer... ...learning in order to scale and optimize our ML systems—creating and... ...fine-tuning, post-training, and inference strategies for large language and...Permanent employment$300k - $400k
...frontier model training and inference fast, efficient, and tightly... ...translate findings into actionable optimizations Implement direct S3... ...and benchmarking distributed ML systems to identify and eliminate... ...world’s best — the scientists, engineers, and problem-solvers who don’...Visa sponsorshipFlexible hoursShift work$175k - $275k
...datasets, or full-cycle data engineering, Abaka AI provides the foundation... ...Abaka builds, trains, and optimizes multimodal AI systems. You... ...applied machine learning or ML engineering, with a demonstrated... ...scale distributed training and inference systems. ~ Familiarity with...Full timeImmediate startFlexible hours$188.5k - $282.7k
...Rubrik's Semantic AI Governance Engine, which is the first system... ...lead even further.As an Applied ML Engineer on the SAGE team,... ...supervised fine-tuning, preference optimization (DPO/RLAIF), and distillation... ...Model Serving and Inference Infrastructure (25% of time)Designing...Permanent employmentLocal area$165.2k - $223.6k
...performance for AWS's custom ML accelerators. Working at the... ...hardware-software boundary, our engineers craft high-performance... ...every FLOP counts in delivering optimal performance for our customers... ...PyTorch, enabling unparalleled ML inference and training performance.As...InternshipLocal areaWork from homeFlexible hours$195k - $230k
...for a Senior Machine Learning Engineer to help evolve our large-... ...ranking, and multi-objective optimization to balance engagement, retention... ...from offline training online inference A/B experimentation metric analysis... ...with large-scale data and ML systems (e.g., Spark,...Full timeLocal areaWork from home$193.3k - $261.5k
...performance for AWS's custom ML accelerators. Working at the... ...hardware-software boundary, our engineers craft high-performance... ...every FLOP counts in delivering optimal performance for our customers... ...PyTorch, enabling unparalleled ML inference and training performance.As...InternshipLocal areaWork from homeFlexible hours$190k - $202k
...Machine Learning Engineer, Perception Palo Alto, California Wing offers... ...and deploy highly reliable ML solutions both for the backend as well as optimized for resource-constrained, real-time... ...TensorFlow, building training/inference/validation pipelines. ~ Familiarity...Full timeLocal area$160k - $225k
...Machine Learning Engineer Location: Mountain View, CA Company Stage... ...AI agents that continuously optimize campaigns, make intelligent... ...architectures, streaming pipelines, and ML serving infrastructure.... ...support model training and inference. Build customer-facing AI...H1bWork at officeVisa sponsorship$170k - $190k
...focus on outcomes. ASAPP’s AI Engineering team is seeking an... ...a highly experienced Lead AI/ML Engineer to join our Core GenerativeAgent... ...handling, streaming inference, and audio quality, and can translate... ...Adapt, evaluate, and optimize LLMs for domain-specific enterprise...Full time- ...democratize AI through high-performance, optimized, open-source and cutting-edge models,... ...Mistral AI is seeking a Applied AI Engineer to facilitate the adoption of its products... ...open source codebases for tasks such as inference and fine-tuning. • You’ll be involved...Full timeWork at officeVisa sponsorship
- ...automation with Moveworks’ Reasoning Engine and natural language... ...Engineer to help build cutting edge ML infrastructure for building... ...be critical in building, optimizing and scaling end-to-end machine... ...distributed training and inference pipeline for large language models...Work at officeRemote workFlexible hours
$124k - $250k
...support of others.A Day in the LifeAs a member of our software engineering infra team, you'll solve technical challenges, including... ...pipeline of model delivery, including training, serving, and optimizations, etc.The Impact You'll MakeDesign, develop, and maintain large...- ...a Principal Machine Learning Engineer, you will drive the development... ...guidance to emerging ML engineers. Your role is pivotal... ...retrieval, LLM-based systems) to optimize user experience and achieve business... ...offline training and online inference at scaleOversee end-to-end...Local areaRemote work
$260k - $330k
...Experience building large-scale prediction or optimization systemsPubMatic is the leading AI-... ...for a Senior Principal Machine Learning Engineer to help build the next generation of performance... ...ROAS. The ideal candidate has strong ML fundamentals and experience building...Work at officeRemote work$230k - $260k
...You’ll Do As a Principal Machine Learning Engineer, you will operate at the company level—... ...You will lead the design of large-scale ML systems and shared platforms that power all... ...scalable ML platforms (training, evaluation, inference, safety) used across multiple teams Drive...Work at officeImmediate start3 days per week- Job Title: ML Engineer What You Will Own End‑to‑End ML Lifecycle across real products: data ingestion, feature design, model selection,... ...and unlimited PTO. Learning budget and a hybrid Bay Area setup optimized for collaboration without dogma. Work that compounds: systems...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to ML Engineer, Inference & Optimization. Be the first to apply!
- machine learning engineer Palo Alto, CA
- computer vision machine learning engineer Palo Alto, CA
- machine learning research scientist Palo Alto, CA
- data engineer machine learning Palo Alto, CA
- internship machine learning Palo Alto, CA
- machine learning scientist Palo Alto, CA
- machine learning remote Palo Alto, CA
- machine learning intern Palo Alto, CA
- machine learning Palo Alto, CA
- artificial intelligence - machine learning intern Palo Alto, CA


