Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

ML Engineer, Inference & Optimization

Full-time

Pika

About the Role

We are seeking Senior/Staff level Inference Engineers to accelerate the performance of Pika's AI-driven products. In this highly technical role, you will operate at the intersection of cutting-edge inference acceleration, GPU parallelism, advanced model deployment, and video generation technologies. Your expertise will drive significant improvements to model speed and efficiency, ensuring our creative AI systems deliver industry-leading user experiences at scale.

 

You will design and optimize inference pipelines, implement state-of-the-art acceleration techniques, and work closely with researchers and engineers across the team to push the boundaries of what’s possible in real-time AI deployment. Your efforts will play a foundational role in powering the next generation of Pika’s video and language models.

 

What You’ll Do

  • Accelerate Inference : Lead and implement advanced inference acceleration techniques, including attention optimization and quantization for efficient model serving.

  • Maximize GPU Parallelism : Engineer and optimize GPU strategies across tensor, sequence, and pipeline parallelism (TP, SP, PP) for maximal efficiency and scalability.

  • Programming for Performance : Develop and optimize high-performance computing kernels and distributed workloads using CUDA and NCCL.

  • Advance AI Deployment : Collaborate with research and engineering teams to bring state-of-the-art videogen and large language models into production.

  • Improve Training Efficiency : (Bonus) Contribute to improvements in model training speed, stability, and resource utilization as part of our deployment lifecycle.

  • Technical Excellence : Drive rigorous code reviews, participate in technical discussions, and mentor fellow engineers on best practices in inference and GPU programming.

 

What We’re Looking For

  • Experience : 5+ years engineering experience, with a strong track record in inference acceleration and model deployment at scale.

  • Inference Mastery : Proven expertise in inference optimization, including quantization, attention acceleration, and deep learning compiler stacks.

  • GPU & Parallelism : Deep knowledge of GPU programming (CUDA, NCCL) and experience with SP, TP, PP, and other forms of parallelism for distributed inference.

  • AI Domain Knowledge : Familiarity with video generation (videogen) models and large language models (LLMs).

  • Collaboration : Strong cross-discipline communication skills; able to drive shared goals across research and engineering functions.

  • Ownership Mindset : Self-driven, solutions-oriented, and capable of managing ambiguity in a fast-paced startup environment.

  • Bonus : Experience in enhancing training efficiency, stability, or resource optimization for large models.

 

Nice to Have

  • Experience with high-throughput video or real-time streaming model deployment

  • Familiarity with distributed training and optimization toolkits

  • Contributions to open source projects in AI infrastructure or deep learning compilers

  • Startup or rapid prototyping experience

 

What We Offer

  • Competitive salary in the AI industry

  • Equity in a fast-growing startup shaping the future of AI

  • Comprehensive health benefits, monthly stipends, company retreats

  • A supportive and collaborative office culture—we’re all building and launching together

 

About Pika

At Pika, we're crafting a future where video creation is seamless, intuitive, and universally accessible. Our mission is to empower creativity by breaking down technical barriers using the transformative power of AI. We’re a tight-knit, energetic team based in Palo Alto, CA, valuing efficiency, curiosity, and the ambition to make a meaningful impact on the world.

 

We work from our Palo Alto office 3–5 days a week and welcome applicants who are eager to contribute onsite.

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the ML Engineer, Inference & Optimization in Remote vacancy
  • $100.4k - $180.7k

    Posting TitleML and Optimization Engineer.LocationCO - Golden.Position TypeLimited Term (Fixed Term)...  ...strengths in high‑performance computing, AI/ML, modeling and simulation, and...  ...learning, probabilistic modeling, large-scale inference, foundation models, and applied AI in a... 
    Suggested
    Full time
    Fixed term contract
    Live in
    Local area
    Remote work
    Relocation
    Shift work

    National Renewable Energy Laboratory

    Golden, CO
    2 days ago
  • $159.05k - $199.3k

     .... About The Role We are looking for a software engineer with deep experience in optimizing ML models and deploying them on production-grade embedded...  ...to optimize efficiency and latency of model inference for compute boards selected by our customers Work... 
    Suggested
    Full time
    For contractors
    For subcontractor
    Casual work
    Work at office
    Remote work
    Day shift

    Applied Intuition

    California
    8 days ago
  • $141k - $249k

     ...Collaborate closely with autonomy and algorithm engineers to scale safe self-driving systems using...  ...such as TensorRT and modelopt to optimize the models running on the truck. Create and benchmark new CUDA kernels for inference. Comprehensively profile model... 
    Suggested
    Work at office
    Work from home
    Flexible hours

    Waabi

    San Francisco, CA
    2 days ago
  • We are looking for a Machine Learning Runtime Optimization Engineer to work on an innovative project that redefines...  ...role, you will focus on backend optimizations for ML runtimes, including hardware acceleration, and inference speed improvements. Experience with ML... 
    Suggested
    Remote job
    Full time
    Worldwide

    Interop Labs

    Remote
    1 day ago
  • $209k - $313k

     ...other digital services.Snap Engineering teams build fun and technically...  ...that quantify causal impact, optimize decision-making, and drive...  ...Strong understanding of causal inference and modern approaches to estimating...  ...tests) and leveraging causal ML in production... 
    Suggested
    Full time
    Live in
    Work at office
    Local area

    Snap

    Palo Alto, CA
    3 days ago
  •  ...deeply engaged audiences.Our TeamThe Decisioning & Optimization engineering team owns the systems that determine which ad...  ...surfaces. Our work spans three platform areas:ML infrastructure for model serving: real-time inference at 1M+ QPS, multi-model parallel evaluation,... 
    Hourly pay
    Full time
    Immediate start
    Flexible hours
    Shift work

    Netflix

    New York, NY
    2 days ago
  •  ...connect a billion people with optimism and civility, and looking for...  ...experiences for everyone. Our engine’s resource management and...  ...optimization. You will establish the ML framework for predictive...  ...roadmap. Design ML models that infer player and interaction... 
    Full time

    Roblox

    Remote
    1 day ago
  •  ...Technologies is seeking a Machine Learning Senior Software Engineer to join the Akamai Inference Cloud Team. This remote US-friendly role focuses on...  ...TensorFlow, or JAX, develop pipelines for model scanning, optimize latency, and implement guardrails for safety and... 
    Remote work

    Akamai Technologies GmbH

    Cambridge, MA
    1 day ago
  • $117.7k - $221.4k

     ...This operating model reflects how Cola engineers think: build durable intermediate artifacts...  ...quality, speed, and cost instead of optimizing any one of them in isolation.The RoleWe...  ...data processing, featurization, and inference foundations that power scalable world understanding... 
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    8 hours ago
  •  ...Service to Marketing we rely on ML to ensure that guests and...  ...LLM fine-tuning, alignment and optimization, RAG/Search, LLM evaluation...  ...a principal machine learning engineer, you will be responsible for...  ...tuning, optimizing models and inference run-time ~ Post-training experience... 
    Remote job
    Full time
    Casual work
    Live in
    Work at office

    Airbnb, Inc.

    United States
    1 day ago
  • $170k - $216k

     ...with downstream teams on the optimization and integration into the Waymo...  ...set of sensors, enabling engineers like you to (1) develop methods...  ...in model training and model inference through model architecture/ hardware...  ...Python ~ Experience with ML frameworks like PyTorch or... 
    Full time
    Remote work

    Waymo

    Remote
    1 day ago
  • $298k - $368k

     ...work jointly with downstream teams on the optimization and integration into the Waymo Driver....  ...a diverse set of sensors, enabling engineers like you to (1) develop methods for efficiently...  ...expertise in low-latency on-device inference techniques and a deep understanding of... 
    Full time
    Remote work

    Waymo

    San Francisco, CA
    1 day ago
  •  ...Machine Learning Engineer - Inference / Serving Join to apply for the Machine Learning Engineer...  ...inference and serving at Yobi, you’ll design, optimize, and operate the systems that bring our...  ...CTV products. This is an applied ML systems role—equal parts engineering... 
    Full time
    Remote work

    Yobi AI

    New York, NY
    1 day ago
  •  ...are looking for a performance engineer who specializes in making...  ...nodes, and high-throughput batch inference sweeping petabytes of real-...  ...generation, and evaluation. The optimization target here is not tail...  ...intersection of accelerators, ML frameworks, and large-scale data... 
    Full time
    For contractors
    For subcontractor
    Casual work
    Work at office
    Remote work
    Day shift

    Applied Intuition

    Sunnyvale, CA
    1 day ago
  • $130k

     ...experience in machine learning engineering or a comparable role. We...  ...building or supporting real-time inference systems for recommendation,...  ...learning, or other adaptive ML applications. We require...  ...experience with graph-level optimization and low-level inference performance... 
    Full time
    Temporary work
    Relocation

    Bitus Labs

    Irvine, CA
    9 days ago
  • $198k - $286k

    Role Description At Modular, we optimize inference from kernel to cloud on one unified stack. We are building a differentiated cloud platform...  ...apply the latest optimizations across kernels, the inference engine, and distributed systems so that customer workloads stay on the... 
    Full time
    Local area
    Flexible hours

    Modular

    Remote
    1 day ago
  • $137.1k - $201.6k

     ...Team The mission of the Marketplace Optimization team is to ensure we maintain a healthy...  ...leverage artificial intelligence and advanced ML, deep learning techniques to power...  ...We’re looking for a Machine Learning Engineer to help design, build, optimize and scale... 
    Hourly pay
    Work at office
    Local area
    Remote work
    Flexible hours

    Doordash Usa

    Remote
    1 day ago
  •  ...Machine Learning Engineer (Llama AI Platform) Location: Remote (...  ...to help design, deploy, and optimize AI solutions built on Llama models...  ...architectures. Optimize inference performance, latency, and...  ...tuning open-source LLMs. ML Engineering and MLOps practices... 
    Full time
    Remote work

    Performacentric

    Denver, CO
    1 day ago
  •  ...About the Role We’re looking for an Applied ML Engineer to design, evaluate, and scale recommendation and ranking systems...  ...skips, conversions) and sparse explicit feedback. Optimize systems for real-time inference, scalability, and robustness under non-stationary user... 
    Full time

    Darwin

    Palo Alto, CA
    1 day ago
  • $213k - $263k

     ...simulation across 15+ U.S. states. The ML Platform team at Waymo provides a set of...  ...management, model development, optimization and monitoring. These efforts have resulted...  ...Research and Simulation. We are looking for engineers with ML software or ML systems expertise... 
    Full time
    Remote work

    Waymo

    Remote
    1 day ago
  •  ...the TeamThe mission of the Marketplace Optimization team is to ensure we maintain a healthy...  ...leverage artificial intelligence and advanced ML, deep learning techniques to power...  ...RoleWe’re looking for a Machine Learning Engineer to help design, build, optimize and scale... 
    Hourly pay
    Work at office
    Local area
    Remote work
    Flexible hours

    Doordash

    San Francisco, CA
    8 hours ago
  •  ...perform real-time multi-objective optimization across distributed systems at...  ...is our Machine Learning and Inference Platform that powers the...  ...leader with deep experience in ML serving, high-performance...  ...- someone excited to mentor engineers, innovate at scale, and shape... 
    Work at office
    Local area
    Remote work
    Monday to Thursday
    Flexible hours

    Roku

    Austin, TX
    8 hours ago
  •  ...growth and superior returns, as we deliver rare value and impact across our businesses.  The Role As a Senior ML Engineer for AWS and Real-Time Inference, you'll own the fast path: ingesting live trading data and scoring it in near real time. It's a systems-heavy... 
    Full time

    Twg Global Ai

    Remote
    1 day ago
  • $185.8k - $303.4k

     ...Description This role sits in the Ads Optimization and Ads Marketplace Quality (AMQ)...  ...platform. You’ll join a set of tight-knit engineers working on high-impact, internet-scale...  ..., and production monitoring for ML systems. Experience collaborating with... 
    Full time
    For contractors
    Work experience placement
    Flexible hours

    Reddit

    United States
    1 day ago
  •  ...advance the frontier of intelligence. About the role: As a ML Engineer, you’ll build and operate the infrastructure that makes...  ..., and control. What you might work on: Optimize inference and training throughput for novel model architectures Build... 
    Full time
    Internship

    Tilde Research

    Remote
    1 day ago
  • $213k - $263k

     ..., learning from demonstration, generative modeling, Bayesian inference, hierarchical learning, and robust evaluation. This role follows...  ...models. Implement efficient training & inference optimization methods to maximize throughput Hands-on experience with Transformer... 
    Full time
    Temporary work
    Remote work

    Waymo

    New York, NY
    1 day ago
  • $116.1k - $170.28k

     ...skilled and experienced Senior Big Data Engineer to join our dynamic team. The ideal candidate...  ...pipelines for data collection or batch inference. This is a remote position, requiring...  ...and managing those solutions, and optimizing returns into the future. Named a best place... 
    Remote job
    Full time

    Rackspace

    Remote
    1 day ago
  •  ...OVERVIEW: We are seeking an ML Ops Engineer to own the machine-learning lifecycle in production. You will be responsible...  ...across the five models. Serving models for real-time inference - package and optimize PyTorch models, run them in the low-latency inference workers... 

    Bana Solutions, LLC

    Chantilly, Loudoun County, VA
    3 days ago
  •  ...shapes what gets built. About the Role As a ML engineer at Wispr, you'll play a crucial role in building the first...  ...Previous founding or startup experience Experience optimizing ML inference or engineering systems for research teams Fluency in Python... 
    H1b
    Work at office
    Remote work
    Relocation
    Visa sponsorship
    Flexible hours

    Wispr Flow

    San Francisco, CA
    2 days ago
  •  ...Join our San Francisco office as an ML Engineer focused on Data Engineering. Visa sponsorship...  ...high availability and reliability. Optimize input/output operations to accelerate...  ...retrieval and processing during training and inference phases. Build and manage backend... 
    Full time
    H1b
    Work at office
    Immediate start
    Visa sponsorship

    Dfbooking Recruitment Services

    Remote
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to ML Engineer, Inference & Optimization. Be the first to apply!