Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

ML Engineer, Inference & Optimization

Full-time

Pika

About the Role

We are seeking Senior/Staff level Inference Engineers to accelerate the performance of Pika's AI-driven products. In this highly technical role, you will operate at the intersection of cutting-edge inference acceleration, GPU parallelism, advanced model deployment, and video generation technologies. Your expertise will drive significant improvements to model speed and efficiency, ensuring our creative AI systems deliver industry-leading user experiences at scale.

 

You will design and optimize inference pipelines, implement state-of-the-art acceleration techniques, and work closely with researchers and engineers across the team to push the boundaries of what’s possible in real-time AI deployment. Your efforts will play a foundational role in powering the next generation of Pika’s video and language models.

 

What You’ll Do

  • Accelerate Inference : Lead and implement advanced inference acceleration techniques, including attention optimization and quantization for efficient model serving.

  • Maximize GPU Parallelism : Engineer and optimize GPU strategies across tensor, sequence, and pipeline parallelism (TP, SP, PP) for maximal efficiency and scalability.

  • Programming for Performance : Develop and optimize high-performance computing kernels and distributed workloads using CUDA and NCCL.

  • Advance AI Deployment : Collaborate with research and engineering teams to bring state-of-the-art videogen and large language models into production.

  • Improve Training Efficiency : (Bonus) Contribute to improvements in model training speed, stability, and resource utilization as part of our deployment lifecycle.

  • Technical Excellence : Drive rigorous code reviews, participate in technical discussions, and mentor fellow engineers on best practices in inference and GPU programming.

 

What We’re Looking For

  • Experience : 5+ years engineering experience, with a strong track record in inference acceleration and model deployment at scale.

  • Inference Mastery : Proven expertise in inference optimization, including quantization, attention acceleration, and deep learning compiler stacks.

  • GPU & Parallelism : Deep knowledge of GPU programming (CUDA, NCCL) and experience with SP, TP, PP, and other forms of parallelism for distributed inference.

  • AI Domain Knowledge : Familiarity with video generation (videogen) models and large language models (LLMs).

  • Collaboration : Strong cross-discipline communication skills; able to drive shared goals across research and engineering functions.

  • Ownership Mindset : Self-driven, solutions-oriented, and capable of managing ambiguity in a fast-paced startup environment.

  • Bonus : Experience in enhancing training efficiency, stability, or resource optimization for large models.

 

Nice to Have

  • Experience with high-throughput video or real-time streaming model deployment

  • Familiarity with distributed training and optimization toolkits

  • Contributions to open source projects in AI infrastructure or deep learning compilers

  • Startup or rapid prototyping experience

 

What We Offer

  • Competitive salary in the AI industry

  • Equity in a fast-growing startup shaping the future of AI

  • Comprehensive health benefits, monthly stipends, company retreats

  • A supportive and collaborative office culture—we’re all building and launching together

 

About Pika

At Pika, we're crafting a future where video creation is seamless, intuitive, and universally accessible. Our mission is to empower creativity by breaking down technical barriers using the transformative power of AI. We’re a tight-knit, energetic team based in Palo Alto, CA, valuing efficiency, curiosity, and the ambition to make a meaningful impact on the world.

 

We work from our Palo Alto office 3–5 days a week and welcome applicants who are eager to contribute onsite.

Vacancy posted 16 hours ago
Similar jobs that could be interesting for youBased on the ML Engineer, Inference & Optimization in Remote vacancy
  • $100.4k - $180.7k

    Posting TitleML and Optimization Engineer.LocationCO - Golden.Position TypeLimited Term (Fixed Term)...  ...strengths in high‑performance computing, AI/ML, modeling and simulation, and...  ...learning, probabilistic modeling, large-scale inference, foundation models, and applied AI in a... 
    Suggested
    Full time
    Fixed term contract
    Live in
    Local area
    Remote work
    Relocation
    Shift work

    National Renewable Energy Laboratory

    Golden, CO
    3 days ago
  • $159.05k - $199.3k

     ...commitments. About the role We are looking for a software engineer with deep experience in optimizing ML models and deploying them on production‑grade...  ...strategies to optimize efficiency and latency of model inference for compute boards selected by our customers Work on... 
    Suggested
    Full time
    For contractors
    For subcontractor
    Casual work
    Work at office
    Remote work
    Day shift

    Applied Intuition

    Sunnyvale, CA
    1 day ago
  • $209k - $313k

     ...other digital services.Snap Engineering teams build fun and technically...  ...that quantify causal impact, optimize decision-making, and drive...  ...Strong understanding of causal inference and modern approaches to estimating...  ...tests) and leveraging causal ML in production... 
    Suggested
    Full time
    Live in
    Work at office
    Local area

    Snap

    Palo Alto, CA
    5 days ago
  •  ...deeply engaged audiences.Our TeamThe Decisioning & Optimization engineering team owns the systems that determine which ad...  ...surfaces. Our work spans three platform areas:ML infrastructure for model serving: real-time inference at 1M+ QPS, multi-model parallel evaluation,... 
    Suggested
    Hourly pay
    Full time
    Immediate start
    Flexible hours
    Shift work

    Netflix

    New York, NY
    3 days ago
  • $141k - $249k

     ...Collaborate closely with autonomy and algorithm engineers to scale safe self-driving systems using...  ...such as TensorRT and modelopt to optimize the models running on the truck. - Create and benchmark new CUDA kernels for inference. - Comprehensively profile model... 
    Suggested
    Work at office
    Work from home
    Flexible hours

    Waabi

    Pittsburgh, PA
    8 days ago
  •  ...connect a billion people with optimism and civility, and looking for...  ...experiences for everyone. Our engine’s resource management and...  ...optimization. You will establish the ML framework for predictive...  ...roadmap. Design ML models that infer player and interaction... 
    Full time

    Roblox

    Remote
    16 hours ago
  • Jaide Health is seeking an engineer for their Model Efficiency team...  ...focuses on building reliable ML systems while enhancing core...  ...techniques such as GPU/CUDA optimizations and collaborate closely with...  ...Python and insights into the LLM inference ecosystem. A commitment to... 
    Remote job

    Jaide Health

    San Francisco, CA
    1 day ago
  • $117.7k - $221.4k

     ...This operating model reflects how Cola engineers think: build durable intermediate artifacts...  ...quality, speed, and cost instead of optimizing any one of them in isolation.The RoleWe...  ...data processing, featurization, and inference foundations that power scalable world understanding... 
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    1 day ago
  •  ...Service to Marketing we rely on ML to ensure that guests and...  ...LLM fine-tuning, alignment and optimization, RAG/Search, LLM evaluation...  ...a principal machine learning engineer, you will be responsible for...  ...tuning, optimizing models and inference run-time ~ Post-training experience... 
    Remote job
    Full time
    Casual work
    Live in
    Work at office

    Airbnb, Inc.

    United States
    16 hours ago
  • $170k - $216k

     ...with downstream teams on the optimization and integration into the Waymo...  ...set of sensors, enabling engineers like you to (1) develop methods...  ...in model training and model inference through model architecture/ hardware...  ...Python ~ Experience with ML frameworks like PyTorch or... 
    Full time
    Remote work

    Waymo

    San Francisco, CA
    16 hours ago
  • $298k - $368k

     ...work jointly with downstream teams on the optimization and integration into the Waymo Driver....  ...a diverse set of sensors, enabling engineers like you to (1) develop methods for efficiently...  ...expertise in low-latency on-device inference techniques and a deep understanding of... 
    Full time
    Remote work

    Waymo

    San Francisco, CA
    16 hours ago
  • $215k - $285k

     ...are looking for a performance engineer who specializes in making...  ...nodes, and high-throughput batch inference sweeping petabytes of real-...  ...generation, and evaluation. The optimization target here is not tail...  ...intersection of accelerators, ML frameworks, and large-scale data... 
    Full time
    For contractors
    For subcontractor
    Casual work
    Work at office
    Remote work
    Day shift

    NLP PEOPLE

    Sunnyvale, CA
    2 days ago
  • Bright Vision Technologies is seeking a Machine Learning Infrastructure Engineer to design, build, and operate high-performance inference platforms for serving large ML models in production. This remote U.S. role focuses on the systems engineering side of AI deployment... 
    Remote job

    Bright Vision Technologies

    Mountain View, CA
    4 days ago
  •  ...Machine Learning Engineer (Llama AI Platform) Location: Remote (...  ...to help design, deploy, and optimize AI solutions built on Llama models...  ...architectures. Optimize inference performance, latency, and...  ...tuning open-source LLMs. ML Engineering and MLOps practices... 
    Full time
    Remote work

    Performacentric

    Indianapolis, IN
    16 hours ago
  • $213k - $263k

     ...simulation across 15+ U.S. states. The ML Platform team at Waymo provides a set of...  ...management, model development, optimization and monitoring. These efforts have resulted...  ...Research and Simulation. We are looking for engineers with ML software or ML systems expertise... 
    Full time
    Remote work

    Waymo

    Remote
    16 hours ago
  • $137.1k - $201.6k

     ...Team The mission of the Marketplace Optimization team is to ensure we maintain a healthy...  ...leverage artificial intelligence and advanced ML, deep learning techniques to power...  ...We’re looking for a Machine Learning Engineer to help design, build, optimize and scale... 
    Hourly pay
    Work at office
    Local area
    Remote work
    Flexible hours

    Doordash Usa

    Remote
    16 hours ago
  •  ...are hiring a Causal Machine Learning Engineer to help build the causal ML foundation behind how DoorDash...  ...counterfactual policy evaluation, promotion optimization, or marketplace decisioning systems...  ...practical experience with causal inference, econometrics, experimentation, or... 
    Hourly pay
    Work at office
    Local area
    Remote work
    Flexible hours

    DoorDash

    New York, NY
    16 hours ago
  •  ...perform real-time multi-objective optimization across distributed systems at...  ...is our Machine Learning and Inference Platform that powers the...  ...leader with deep experience in ML serving, high-performance...  ...- someone excited to mentor engineers, innovate at scale, and shape... 
    Work at office
    Local area
    Remote work
    Monday to Thursday
    Flexible hours

    Roku

    Austin, TX
    1 day ago
  • $175k - $200k

     ...Senior Machine Learning Engineer Truveta is the world’s first health...  ...engineering. You’ll blend ML craftsmanship with platform...  ...expertise in applied AI, model optimization, and agentic intelligence to...  ...efficiency, interpretability, and inference performance. Think and... 
    Full time
    For contractors
    Visa sponsorship
    Work visa
    Flexible hours

    Truveta

    Remote
    7 hours ago
  • $213k - $263k

     ..., learning from demonstration, generative modeling, Bayesian inference, hierarchical learning, and robust evaluation. This role follows...  ...models. Implement efficient training & inference optimization methods to maximize throughput Hands-on experience with Transformer... 
    Full time
    Temporary work
    Remote work

    Waymo

    New York, NY
    16 hours ago
  •  ...growth and superior returns, as we deliver rare value and impact across our businesses.  The Role As a Senior ML Engineer for AWS and Real-Time Inference, you'll own the fast path: ingesting live trading data and scoring it in near real time. It's a systems-heavy... 
    Full time

    Twg Global Ai

    Remote
    16 hours ago
  • $185.8k - $303.4k

     ...Description This role sits in the Ads Optimization and Ads Marketplace Quality (AMQ)...  ...platform. You’ll join a set of tight-knit engineers working on high-impact, internet-scale...  ..., and production monitoring for ML systems. Experience collaborating with... 
    Full time
    For contractors
    Work experience placement
    Flexible hours

    Reddit

    United States
    16 hours ago
  •  ...advance the frontier of intelligence. About the role: As a ML Engineer, you’ll build and operate the infrastructure that makes...  ..., and control. What you might work on: Optimize inference and training throughput for novel model architectures Build... 
    Full time
    Internship

    Tilde Research

    Remote
    16 hours ago
  • $160.5k - $240.7k

     ...Technologies, Inc. Job Area: Engineering Group, Engineering Group >...  ...and supportcutting-edgemodel optimization workflows — pushing the boundary...  ...AIMET workflows with popular ML frameworks —PyTorchand ONNX...  ...Falcon, or similar families) for inference optimization Familiarity... 
    Work experience placement
    Immediate start
    Work from home

    Socket.dev

    Santa Clara, CA
    2 days ago
  •  ...protected veteran.**Job Description:****ML Engineer****3M Health Care is now Solventum****...  ...Integration*** **Data Pipelines:** Develop and optimize ETL processes to transform healthcare...  ...usable datasets for model training and inference.* **Feature Management:** Help build... 
    H1b
    Remote work

    Solventum

    New York, NY
    4 days ago
  •  ...every person shapes what gets built. About the Role As a ML engineer at Wispr, you’ll play a crucial role in building the first...  ...looking for? Previous founding or startup experience Experience optimizing ML inference or engineering systems for research teams Fluency in... 
    H1b
    Work at office
    Remote work
    Relocation
    Visa sponsorship
    Flexible hours

    Visa Hunt

    San Francisco, CA
    1 day ago
  • $116.1k - $170.28k

     ...skilled and experienced Senior Big Data Engineer to join our dynamic team. The ideal candidate...  ...pipelines for data collection or batch inference. This is a remote position, requiring...  ...and managing those solutions, and optimizing returns into the future. Named a best place... 
    Remote job
    Full time

    Rackspace

    Remote
    16 hours ago
  • $170.1k - $258.3k

     ...export, kernel development, and performance engineering so that every cycle on our accelerators...  ...sit at the heart of our on‑vehicle ML inference for ADAS and autonomous driving. We own...  ...benchmarking, profiling, debugging and optimizing accelerator libraries and kernels to... 
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    San Francisco, CA
    2 days ago
  • $193.3k - $261.5k

     ...performance for AWS's custom ML accelerators. Working at the...  ...hardware-software boundary, our engineers craft high-performance...  ...every FLOP counts in delivering optimal performance for our customers...  ...PyTorch, enabling unparalleled ML inference and training performance.As... 
    Internship
    Local area
    Work from home
    Flexible hours

    Amazon

    Cupertino, CA
    1 day ago
  • $128.7k - $261.3k

     ...development, and performance engineering so that every cycle on our accelerators...  ...models into fast, reliable inference across GPUs powering GM’s...  ...and turns them into highly optimized inference artifacts running...  ...reliable, and effortless for ML engineers across the AV organization... 
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    San Francisco, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to ML Engineer, Inference & Optimization. Be the first to apply!