Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

ML Engineer, Inference & Optimization

$250k - $350k

Pika

About the RoleWe are seeking Senior/Staff level Inference Engineers to accelerate the performance of Pika's AI-driven products. In this highly technical role, you will operate at the intersection of cutting-edge inference acceleration, GPU parallelism, advanced model deployment, and video generation technologies. Your expertise will drive significant improvements to model speed and efficiency, ensuring our creative AI systems deliver industry-leading user experiences at scale.You will design and optimize inference pipelines, implement state-of-the-art acceleration techniques, and work closely with researchers and engineers across the team to push the boundaries of what’s possible in real-time AI deployment. Your efforts will play a foundational role in powering the next generation of Pika’s video and language models.What You’ll DoAccelerate Inference: Lead and implement advanced inference acceleration techniques, including attention optimization and quantization for efficient model serving.Maximize GPU Parallelism: Engineer and optimize GPU strategies across tensor, sequence, and pipeline parallelism (TP, SP, PP) for maximal efficiency and scalability.Programming for Performance: Develop and optimize high-performance computing kernels and distributed workloads using CUDA and NCCL.Advance AI Deployment: Collaborate with research and engineering teams to bring state-of-the-art videogen and large language models into production.Improve Training Efficiency: (Bonus) Contribute to improvements in model training speed, stability, and resource utilization as part of our deployment lifecycle.Technical Excellence: Drive rigorous code reviews, participate in technical discussions, and mentor fellow engineers on best practices in inference and GPU programming.What We’re Looking ForExperience: 5+ years engineering experience, with a strong track record in inference acceleration and model deployment at scale.Inference Mastery: Proven expertise in inference optimization, including quantization, attention acceleration, and deep learning compiler stacks.GPU & Parallelism: Deep knowledge of GPU programming (CUDA, NCCL) and experience with SP, TP, PP, and other forms of parallelism for distributed inference.AI Domain Knowledge: Familiarity with video generation (videogen) models and large language models (LLMs).Collaboration: Strong cross-discipline communication skills; able to drive shared goals across research and engineering functions.Ownership Mindset: Self-driven, solutions-oriented, and capable of managing ambiguity in a fast-paced startup environment.Bonus: Experience in enhancing training efficiency, stability, or resource optimization for large models.Nice to HaveExperience with high-throughput video or real-time streaming model deploymentFamiliarity with distributed training and optimization toolkitsContributions to open source projects in AI infrastructure or deep learning compilersStartup or rapid prototyping experienceWhat We OfferCompetitive salary in the AI industryEquity in a fast-growing startup shaping the future of AIComprehensive health benefits, monthly stipends, company retreatsA supportive and collaborative office culture—we’re all building and launching togetherAbout PikaAt Pika, we're crafting a future where video creation is seamless, intuitive, and universally accessible. Our mission is to empower creativity by breaking down technical barriers using the transformative power of AI. We’re a tight-knit, energetic team based in Palo Alto, CA, valuing efficiency, curiosity, and the ambition to make a meaningful impact on the world.We work from our Palo Alto office 3–5 days a week and welcome applicants who are eager to contribute onsite.Compensation Range: $250K - $350KLocationPalo Alto HQEmployment TypeFull timeLocation TypeOn-siteDepartmentResearchCompensationbase salary $250K – $350K

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the ML Engineer, Inference & Optimization in Palo Alto, CA vacancy
  • $278.1k - $347.6k

     ...within that runtime. As our Principal Engineer for On-Device AI Inference & Systems, you will be the foremost...  ...leaves research, through export, optimization, and kernel-level tuning, to a shipped...  ...and own the integration between the ML runtime and the game engine: real-time... 
    Suggested
    Work at office
    Worldwide
    Relocation package

    Unity

    Mountain View, CA
    2 days ago
  • $209k - $313k

     ...other digital services.Snap Engineering teams build fun and technically...  ...that quantify causal impact, optimize decision-making, and drive...  ...Strong understanding of causal inference and modern approaches to estimating...  ...tests) and leveraging causal ML in production... 
    Suggested
    Full time
    Live in
    Work at office
    Local area

    Snap

    Palo Alto, CA
    9 hours ago
  •  ...About the Role We’re looking for an Applied ML Engineer to design, evaluate, and scale recommendation and ranking systems...  ...skips, conversions) and sparse explicit feedback. Optimize systems for real-time inference, scalability, and robustness under non-stationary user... 
    Suggested
    Full time

    Darwin

    Palo Alto, CA
    13 hours ago
  • $117.7k - $221.4k

     ...This operating model reflects how Cola engineers think: build durable intermediate artifacts...  ...quality, speed, and cost instead of optimizing any one of them in isolation.The RoleWe...  ...data processing, featurization, and inference foundations that power scalable world understanding... 
    Suggested
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    2 days ago
  •  ...About the job ML Engineer Our Client Is a rapidly growing Tier 1 VC backed...  ...redefining how businesses learn from and optimize their in-person customer experiences....  ...preprocessing, model training, deployment, inference, and monitoring in production... 
    Suggested
    Full time

    Catalyst Labs, LLC

    Mountain View, CA
    4 days ago
  • $213k - $263k

     ...Waymo Ml Software Engineer Waymo is an autonomous driving technology company with the mission to be the world's most trusted driver. Since...  ...feature and experiment management, model development, optimization and monitoring. These efforts have resulted in making machine... 
    Full time
    Remote work

    Latent Logic

    Mountain View, CA
    1 hour ago
  • $195.2k - $361.2k

     ...future of AI should belong to the people it servesRole SummaryMake models fast on the hardware people actually own. You optimize inference engines (llama.cpp, vLLM) for constrained local and edge environments — GPU/iGPUs, Vulkan backends — not datacenter H100 environment... 
    Full time
    Internship
    Local area
    Immediate start
    Shift work

    Intel

    Santa Clara, CA
    1 day ago
  • $184k - $287.5k

    We're now looking for a Sr. Inference Engineer, for GPU Kernel Optimization! What does it take to push every LLM inference operation to its performance ceiling? Our LLM Inference Performance Analysis and Optimization team builds the answer from the ground up. We develop... 
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  •  ...world.Role OverviewAs our Staff Software Engineer, ML infra Engineer for Search & Discovery...  ...Discovery organization is responsible for optimizing customers' navigation experience and...  ...ML based ranking system and online ML inference servicesBuild strong cross-functional partnerships... 
    Temporary work

    Coupang

    Mountain View, CA
    2 days ago
  • $207k - $300k

     ...generated content through advanced context engineering and agentic feedback loops to identify...  ...and serving AIGC content to optimize efficiency for low-latency surfaces.Navigate...  ...Experience optimizing machine learning inference, resolving system-level bottlenecks, and... 

    Google

    Mountain View, CA
    1 day ago
  • $250k - $350k

     ...alongside some of the world's leading ML systems engineers, including leaders behind Megatron-LM...  ...infrastructure powering our large-scale training, inference, and reinforcement learning. You'll...  .... What You'll Do Build and optimize large-scale training and reinforcement... 
    Visa sponsorship

    Periodic Labs

    Menlo Park, CA
    2 days ago
  •  ...per day.Your Role and ResponsibilitiesWe are seeking to hire an ML Engineer to join our growing team! This is a hands-on technical role...  ...improve existing ML models, training pipelines, and run-time inference performance• Expand the scope and capabilities of our models to... 
    Full time
    Temporary work
    Part time

    IBM

    Mountain View, CA
    3 days ago
  • $207k - $301k

     ...architecture in different modeling domains.5 years of experience with ML design and ML infrastructure (e.g., model deployment, model...  ...language models, as well as building real-time and large scale inference infrastructure.Personalizing our Ads has proven to be extremely... 

    Google

    Mountain View, CA
    1 day ago
  • $196k - $221k

     ...our core AI team responsible for ML and work alongside industry-veteran scientists and engineers. As a Machine Learning Engineer...  ...learning in order to scale and optimize our ML systems—creating and...  ...fine-tuning, post-training, and inference strategies for large language and... 
    Permanent employment

    Otter.ai

    Mountain View, CA
    2 days ago
  • $300k - $400k

     ...frontier model training and inference fast, efficient, and tightly...  ...translate findings into actionable optimizations Implement direct S3...  ...and benchmarking distributed ML systems to identify and eliminate...  ...world’s best — the scientists, engineers, and problem-solvers who don’... 
    Visa sponsorship
    Flexible hours
    Shift work

    Periodic Labs

    Menlo Park, CA
    4 days ago
  • $175k - $275k

     ...datasets, or full-cycle data engineering, Abaka AI provides the foundation...  ...Abaka builds, trains, and optimizes multimodal AI systems. You...  ...applied machine learning or ML engineering, with a demonstrated...  ...scale distributed training and inference systems. ~ Familiarity with... 
    Full time
    Immediate start
    Flexible hours

    Abaka Ai

    Palo Alto, CA
    13 hours ago
  • $188.5k - $282.7k

     ...Rubrik's Semantic AI Governance Engine, which is the first system...  ...lead even further.As an Applied ML Engineer on the SAGE team,...  ...supervised fine-tuning, preference optimization (DPO/RLAIF), and distillation...  ...Model Serving and Inference Infrastructure (25% of time)Designing... 
    Permanent employment
    Local area

    Rubrik

    Palo Alto, CA
    1 day ago
  • $165.2k - $223.6k

     ...performance for AWS's custom ML accelerators. Working at the...  ...hardware-software boundary, our engineers craft high-performance...  ...every FLOP counts in delivering optimal performance for our customers...  ...PyTorch, enabling unparalleled ML inference and training performance.As... 
    Internship
    Local area
    Work from home
    Flexible hours

    Amazon

    Cupertino, CA
    2 days ago
  • $195k - $230k

     ...for a Senior Machine Learning Engineer to help evolve our large-...  ...ranking, and multi-objective optimization to balance engagement, retention...  ...from offline training online inference A/B experimentation metric analysis...  ...with large-scale data and ML systems (e.g., Spark,... 
    Full time
    Local area
    Work from home

    News Break

    Mountain View, CA
    2 days ago
  • $193.3k - $261.5k

     ...performance for AWS's custom ML accelerators. Working at the...  ...hardware-software boundary, our engineers craft high-performance...  ...every FLOP counts in delivering optimal performance for our customers...  ...PyTorch, enabling unparalleled ML inference and training performance.As... 
    Internship
    Local area
    Work from home
    Flexible hours

    Amazon

    Cupertino, CA
    2 days ago
  • $190k - $202k

     ...Machine Learning Engineer, Perception Palo Alto, California Wing offers...  ...and deploy highly reliable ML solutions both for the backend as well as optimized for resource-constrained, real-time...  ...TensorFlow, building training/inference/validation pipelines. ~ Familiarity... 
    Full time
    Local area

    Wing

    Palo Alto, CA
    13 hours ago
  • $160k - $225k

     ...Machine Learning Engineer Location: Mountain View, CA Company Stage...  ...AI agents that continuously optimize campaigns, make intelligent...  ...architectures, streaming pipelines, and ML serving infrastructure....  ...support model training and inference. Build customer-facing AI... 
    H1b
    Work at office
    Visa sponsorship

    Recruiting from Scratch

    Mountain View, CA
    2 days ago
  • $170k - $190k

     ...focus on outcomes. ASAPP’s AI Engineering team is seeking an...  ...a highly experienced Lead AI/ML Engineer to join our Core GenerativeAgent...  ...handling, streaming inference, and audio quality, and can translate...  ...Adapt, evaluate, and optimize LLMs for domain-specific enterprise... 
    Full time

    Asapp

    Mountain View, CA
    13 hours ago
  •  ...democratize AI through high-performance, optimized, open-source and cutting-edge models,...  ...Mistral AI is seeking a Applied AI Engineer to facilitate the adoption of its products...  ...open source codebases for tasks such as inference and fine-tuning. • You’ll be involved... 
    Full time
    Work at office
    Visa sponsorship

    Mistral Ai

    Palo Alto, CA
    13 hours ago
  •  ...automation with Moveworks’ Reasoning Engine and natural language...  ...Engineer to help build cutting edge ML infrastructure for building...  ...be critical in building, optimizing and scaling end-to-end machine...  ...distributed training and inference pipeline for large language models... 
    Work at office
    Remote work
    Flexible hours

    Moveworks

    Mountain View, CA
    2 days ago
  • $124k - $250k

     ...support of others.A Day in the LifeAs a member of our software engineering infra team, you'll solve technical challenges, including...  ...pipeline of model delivery, including training, serving, and optimizations, etc.The Impact You'll MakeDesign, develop, and maintain large... 

    AppLovin

    Palo Alto, CA
    2 days ago
  •  ...a Principal Machine Learning Engineer, you will drive the development...  ...guidance to emerging ML engineers. Your role is pivotal...  ...retrieval, LLM-based systems) to optimize user experience and achieve business...  ...offline training and online inference at scaleOversee end-to-end... 
    Local area
    Remote work

    Atlassian

    Mountain View, CA
    3 days ago
  • $260k - $330k

     ...Experience building large-scale prediction or optimization systemsPubMatic is the leading AI-...  ...for a Senior Principal Machine Learning Engineer to help build the next generation of performance...  ...ROAS. The ideal candidate has strong ML fundamentals and experience building... 
    Work at office
    Remote work

    PubMatic

    Redwood City, CA
    3 days ago
  • $230k - $260k

     ...You’ll Do As a Principal Machine Learning Engineer, you will operate at the company level—...  ...You will lead the design of large-scale ML systems and shared platforms that power all...  ...scalable ML platforms (training, evaluation, inference, safety) used across multiple teams Drive... 
    Work at office
    Immediate start
    3 days per week

    Typeface

    Palo Alto, CA
    2 days ago
  • Job Title: ML Engineer What You Will Own End‑to‑End ML Lifecycle across real products: data ingestion, feature design, model selection,...  ...and unlimited PTO. Learning budget and a hybrid Bay Area setup optimized for collaboration without dogma. Work that compounds: systems... 

    MetAntz

    Palo Alto, CA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to ML Engineer, Inference & Optimization. Be the first to apply!