Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

ML Engineer, Inference & Optimization

Full-time

Pika

About the Role

We are seeking Senior/Staff level Inference Engineers to accelerate the performance of Pika's AI-driven products. In this highly technical role, you will operate at the intersection of cutting-edge inference acceleration, GPU parallelism, advanced model deployment, and video generation technologies. Your expertise will drive significant improvements to model speed and efficiency, ensuring our creative AI systems deliver industry-leading user experiences at scale.

 

You will design and optimize inference pipelines, implement state-of-the-art acceleration techniques, and work closely with researchers and engineers across the team to push the boundaries of what’s possible in real-time AI deployment. Your efforts will play a foundational role in powering the next generation of Pika’s video and language models.

 

What You’ll Do

  • Accelerate Inference : Lead and implement advanced inference acceleration techniques, including attention optimization and quantization for efficient model serving.

  • Maximize GPU Parallelism : Engineer and optimize GPU strategies across tensor, sequence, and pipeline parallelism (TP, SP, PP) for maximal efficiency and scalability.

  • Programming for Performance : Develop and optimize high-performance computing kernels and distributed workloads using CUDA and NCCL.

  • Advance AI Deployment : Collaborate with research and engineering teams to bring state-of-the-art videogen and large language models into production.

  • Improve Training Efficiency : (Bonus) Contribute to improvements in model training speed, stability, and resource utilization as part of our deployment lifecycle.

  • Technical Excellence : Drive rigorous code reviews, participate in technical discussions, and mentor fellow engineers on best practices in inference and GPU programming.

 

What We’re Looking For

  • Experience : 5+ years engineering experience, with a strong track record in inference acceleration and model deployment at scale.

  • Inference Mastery : Proven expertise in inference optimization, including quantization, attention acceleration, and deep learning compiler stacks.

  • GPU & Parallelism : Deep knowledge of GPU programming (CUDA, NCCL) and experience with SP, TP, PP, and other forms of parallelism for distributed inference.

  • AI Domain Knowledge : Familiarity with video generation (videogen) models and large language models (LLMs).

  • Collaboration : Strong cross-discipline communication skills; able to drive shared goals across research and engineering functions.

  • Ownership Mindset : Self-driven, solutions-oriented, and capable of managing ambiguity in a fast-paced startup environment.

  • Bonus : Experience in enhancing training efficiency, stability, or resource optimization for large models.

 

Nice to Have

  • Experience with high-throughput video or real-time streaming model deployment

  • Familiarity with distributed training and optimization toolkits

  • Contributions to open source projects in AI infrastructure or deep learning compilers

  • Startup or rapid prototyping experience

 

What We Offer

  • Competitive salary in the AI industry

  • Equity in a fast-growing startup shaping the future of AI

  • Comprehensive health benefits, monthly stipends, company retreats

  • A supportive and collaborative office culture—we’re all building and launching together

 

About Pika

At Pika, we're crafting a future where video creation is seamless, intuitive, and universally accessible. Our mission is to empower creativity by breaking down technical barriers using the transformative power of AI. We’re a tight-knit, energetic team based in Palo Alto, CA, valuing efficiency, curiosity, and the ambition to make a meaningful impact on the world.

 

We work from our Palo Alto office 3–5 days a week and welcome applicants who are eager to contribute onsite.

Vacancy posted more than 2 months ago
Similar jobs that could be interesting for youBased on the ML Engineer, Inference & Optimization in Remote vacancy
  • $100.4k - $180.7k

    Posting TitleML and Optimization Engineer.LocationCO - Golden.Position TypeLimited Term (Fixed Term)...  ...strengths in high‑performance computing, AI/ML, modeling and simulation, and...  ...learning, probabilistic modeling, large-scale inference, foundation models, and applied AI in a... 
    Suggested
    Full time
    Fixed term contract
    Live in
    Local area
    Remote work
    Relocation
    Shift work

    National Renewable Energy Laboratory

    Golden, CO
    4 days ago
  • $159.05k - $199.3k

     .... About The Role We are looking for a software engineer with deep experience in optimizing ML models and deploying them on production-grade embedded...  ...to optimize efficiency and latency of model inference for compute boards selected by our customers Work... 
    Suggested
    Full time
    For contractors
    For subcontractor
    Casual work
    Work at office
    Remote work
    Day shift

    Applied Intuition

    California
    20 days ago
  • $141k - $249k

     ...Collaborate closely with autonomy and algorithm engineers to scale safe self-driving systems using...  ...such as TensorRT and modelopt to optimize the models running on the truck. - Create and benchmark new CUDA kernels for inference. - Comprehensively profile model... 
    Suggested
    Work at office
    Work from home
    Flexible hours

    Waabi

    Dallas, TX
    1 day ago
  •  ...production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI.... 
    Suggested
    Full time

    Nebius

    Remote
    1 day ago
  • $209k - $313k

     ...other digital services.Snap Engineering teams build fun and technically...  ...that quantify causal impact, optimize decision-making, and drive...  ...Strong understanding of causal inference and modern approaches to estimating...  ...tests) and leveraging causal ML in production... 
    Suggested
    Full time
    Live in
    Work at office
    Local area

    Snap

    Seattle, WA
    2 days ago
  •  ...deeply engaged audiences.Our TeamThe Decisioning & Optimization engineering team owns the systems that determine which ad...  ...surfaces. Our work spans three platform areas:ML infrastructure for model serving: real-time inference at 1M+ QPS, multi-model parallel evaluation,... 
    Hourly pay
    Full time
    Immediate start
    Flexible hours
    Shift work

    Netflix

    New York, NY
    4 days ago
  • $117.7k - $221.4k

     ...This operating model reflects how Cola engineers think: build durable intermediate artifacts...  ...quality, speed, and cost instead of optimizing any one of them in isolation.The RoleWe...  ...data processing, featurization, and inference foundations that power scalable world understanding... 
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, TX
    4 days ago
  • $203.5k - $299.3k

     ...are hiring a Causal Machine Learning Engineer to help build the causal ML foundation behind how DoorDash...  ...counterfactual policy evaluation, promotion optimization, or marketplace decisioning systems...  ...practical experience with causal inference, econometrics, experimentation, or... 
    Hourly pay
    Work at office
    Local area
    Remote work
    Flexible hours

    Doordash

    Seattle, WA
    2 days ago
  •  ...impact. Make Wayve the experience that defines your career! The role As a Staff ML Performance Engineer, you’ll play a key role in high-impact projects, optimising ML inference for edge accelerators and GPUs. The focus of this team is to run large transformer-... 
    Full time
    Work at office
    Work from home

    Wayve

    United Kingdom
    15 days ago
  •  ...perform real-time multi-objective optimization across distributed systems at...  ...is our Machine Learning and Inference Platform that powers the...  ...leader with deep experience in ML serving, high-performance...  ...- someone excited to mentor engineers, innovate at scale, and shape... 
    Work at office
    Local area
    Remote work
    Monday to Thursday
    Flexible hours

    Roku

    Austin, TX
    2 days ago
  • $147k - $268.4k

     ...never been done. If you're an engineer, scientist, or builder who...  ...What You'll Be Doing As an ML Ops Engineer, you build and operate...  ...at scale. You optimize infrastructure and GPU resources...  ...Operate and optimize large-scale inference platforms that support scientific... 
    Remote work
    Flexible hours
    2 days per week

    Eli Lilly

    San Francisco, CA
    3 days ago
  • $147k - $268.4k

     ...never been done. If you're an engineer, scientist, or builder who...  ...What You'll Be Doing As an ML Ops Engineer, you build and...  ...reproducibility at scale. You optimize infrastructure and GPU resources...  ...and optimize large-scale inference platforms that support scientific... 
    Full time
    Remote work
    Flexible hours
    2 days per week

    Eli Lilly

    South San Francisco, CA
    4 days ago
  •  ...Job Title : ML Engineer Location : Minnetonka, MN FULLTIME ONLY...  ...and pipelines. • Build training and inference code with reproducibility, versioning,...  ...offline), batching, and latency/throughput optimization. • Integrate model lifecycle tooling... 
    Full time

    AceStack LLC

    Minnetonka, MN
    2 days ago
  •  ...dedicated owner of Confido's ML platform. Our AI/ML team already...  ...behind training, inference, and agentic workloads...  ...ML safely ~ Serve and optimize inference and forecasting workloads...  ...the data interface with data engineering: serve the right data to models... 
    Local area
    Relocation package
    Weekend work

    Confido

    New York, NY
    3 days ago
  • $128.7k - $261.3k

     ...development, and performance engineering so that every cycle on our accelerators...  ...models into fast, reliable inference across GPUs powering GM’s...  ...and turns them into highly optimized inference artifacts running...  ...reliable, and effortless for ML engineers across the AV organization... 
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Austin, TX
    2 days ago
  • $173k - $253k

    Matterport - Senior ML Ops Engineer Job Description CoStar Group is a leading global provider...  ...bottlenecks, applying advanced optimization techniques, and deploying highly efficient...  ...analyze model performance, optimize inference speed and resource utilization, and... 
    Full time
    Work at office
    Work from home

    Matterport

    Sunnyvale, CA
    1 day ago
  • $170.1k - $258.3k

     ...export, kernel development, and performance engineering so that every cycle on our accelerators...  ...sit at the heart of our on‑vehicle ML inference for ADAS and autonomous driving. We own...  ...benchmarking, profiling, debugging and optimizing accelerator libraries and kernels to... 
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Austin, TX
    3 days ago
  • $165.2k - $223.6k

     ...performance for AWS's custom ML accelerators. Working at the...  ...hardware-software boundary, our engineers craft high-performance...  ...every FLOP counts in delivering optimal performance for our customers...  ...PyTorch, enabling unparalleled ML inference and training performance.As... 
    Internship
    Local area
    Work from home
    Flexible hours

    Amazon

    Cupertino, CA
    2 days ago
  • AI/ML Ops EngineerLocation: Remote / Hybrid (Client-Facing Consulting...  ...seeking an experienced AI/ML Engineer to design, deploy, and...  ...involving forecasting, optimization, decision support, and advanced...  ...including real-time APIs, batch inference pipelines, and feature stores... 
    Full time
    Contract work
    Local area
    Remote work
    Flexible hours

    Slalom

    Houston, TX
    2 days ago
  •  ...career growth potential. Job Title: Edge ML Engineer Location: 100% Remote (U.S.) Position...  ...looking for an Edge ML Engineer to design, optimize, and deploy machine learning models...  ...Experience with at least one major edge inference framework. Solid understanding of... 
    Full time
    Local area
    Remote work

    Bright Vision Technologies

    Dublin, OH
    4 days ago
  • $180k - $260k

    You will work on model architecture, distributed training, inference optimization, and evaluation systems that power Tenzin and the broader Octave-X platform. The role bridges applied ML engineering and production reliability for regulated and high-trust environments.... 
    Full time
    Remote work

    Octave X

    Chicago, IL
    2 days ago
  •  ...every person shapes what gets built. About the Role As a ML engineer at Wispr, you’ll play a crucial role in building the first...  ...for? Previous founding or startup experience Experience optimizing ML inference or engineering systems for research teams Fluency in Python... 
    H1b
    Work at office
    Remote work
    Relocation
    Visa sponsorship
    Flexible hours

    Visa Hunt

    Brooklyn, NY
    3 days ago
  •  ...a protected veteran. Job Description ML Engineer 3M Health Care is now Solventum At Solventum...  ...Data Pipelines: Develop and optimize ETL processes to transform healthcare data...  ...usable datasets for model training and inference. Feature Management: Help build and maintain... 
    H1b
    Remote work

    Solventum

    Austin, TX
    5 days ago
  •  ...experienced Machine Learning Engineers (MLE Bench) to contribute to...  ...on work with production-grade ML codebases, model training and...  ...training, evaluation, and inference pipelines . Prepare datasets...  ..., evaluation metrics, optimization). Experience working with... 
    Contract work
    Temporary work
    For contractors
    Freelance
    Remote work

    Turing

    Remote
    14 days ago
  • $80 - $150 per hour

     ...Role Overview Contribute expert machine learning engineering work to projects that develop, train, evaluate,...  ...tasks spanning model development, training and inference systems, numerical computing, and performance optimization. Key Responsibilities Develop and... 
    Hourly pay
    For contractors
    Remote work
    Flexible hours

    SaidGig

    Remote
    13 days ago
  • $80 - $150 per hour

     ...ML Engineer Pay: $80, $150/hour Location: Global, fully remote Job Type: Contractor (~15 hours per...  ...project involving model development, training and inference systems, numerical computing, performance optimization, and Python. The work involves creating,... 
    Temporary work
    For contractors
    Immediate start
    Remote work
    Flexible hours

    micro1

    Remote
    6 days ago
  • $200k - $240k

     ...Senior Machine Learning Engineer (Computer Vision / Vision-Language Models) Remote...  ...evaluation, and monitoring frameworks Optimize inference pipelines for performance, accuracy, and...  ...skills ~ Proven record of deploying ML systems into production ~ Experience... 
    Remote work

    Harnham

    New York, NY
    2 days ago
  • $260k - $330k

     ...Experience building large-scale prediction or optimization systemsPubMatic is the leading AI-...  ...for a Senior Principal Machine Learning Engineer to help build the next generation of performance...  ...ROAS. The ideal candidate has strong ML fundamentals and experience building... 
    Work at office
    Remote work

    PubMatic

    Redwood City, CA
    3 days ago
  • $180k - $280k

     ...across real-world scenarios.As a Staff AI/ML Engineer within the Onboard Embodied AI...  ...-to-end solutions capable of real-time inference and robust autonomous driving performance...  ...training methodologies, and inference optimization strategies suited for real-time onboard... 
    Full time
    Local area
    Work from home
    Relocation
    Relocation package

    General Motors

    Sunnyvale, CA
    3 days ago
  •  ...world.Role OverviewAs our Staff Software Engineer, ML infra Engineer for Search & Discovery...  ...Discovery organization is responsible for optimizing customers' navigation experience and...  ...ML based ranking system and online ML inference servicesBuild strong cross-functional partnerships... 
    Temporary work

    Coupang

    Mountain View, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to ML Engineer, Inference & Optimization. Be the first to apply!