ML Engineer, Inference & Optimization
Pika
About the Role
We are seeking Senior/Staff level Inference Engineers to accelerate the performance of Pika's AI-driven products. In this highly technical role, you will operate at the intersection of cutting-edge inference acceleration, GPU parallelism, advanced model deployment, and video generation technologies. Your expertise will drive significant improvements to model speed and efficiency, ensuring our creative AI systems deliver industry-leading user experiences at scale.
You will design and optimize inference pipelines, implement state-of-the-art acceleration techniques, and work closely with researchers and engineers across the team to push the boundaries of what’s possible in real-time AI deployment. Your efforts will play a foundational role in powering the next generation of Pika’s video and language models.
What You’ll Do
Accelerate Inference : Lead and implement advanced inference acceleration techniques, including attention optimization and quantization for efficient model serving.
Maximize GPU Parallelism : Engineer and optimize GPU strategies across tensor, sequence, and pipeline parallelism (TP, SP, PP) for maximal efficiency and scalability.
Programming for Performance : Develop and optimize high-performance computing kernels and distributed workloads using CUDA and NCCL.
Advance AI Deployment : Collaborate with research and engineering teams to bring state-of-the-art videogen and large language models into production.
Improve Training Efficiency : (Bonus) Contribute to improvements in model training speed, stability, and resource utilization as part of our deployment lifecycle.
Technical Excellence : Drive rigorous code reviews, participate in technical discussions, and mentor fellow engineers on best practices in inference and GPU programming.
What We’re Looking For
Experience : 5+ years engineering experience, with a strong track record in inference acceleration and model deployment at scale.
Inference Mastery : Proven expertise in inference optimization, including quantization, attention acceleration, and deep learning compiler stacks.
GPU & Parallelism : Deep knowledge of GPU programming (CUDA, NCCL) and experience with SP, TP, PP, and other forms of parallelism for distributed inference.
AI Domain Knowledge : Familiarity with video generation (videogen) models and large language models (LLMs).
Collaboration : Strong cross-discipline communication skills; able to drive shared goals across research and engineering functions.
Ownership Mindset : Self-driven, solutions-oriented, and capable of managing ambiguity in a fast-paced startup environment.
Bonus : Experience in enhancing training efficiency, stability, or resource optimization for large models.
Nice to Have
Experience with high-throughput video or real-time streaming model deployment
Familiarity with distributed training and optimization toolkits
Contributions to open source projects in AI infrastructure or deep learning compilers
Startup or rapid prototyping experience
What We Offer
Competitive salary in the AI industry
Equity in a fast-growing startup shaping the future of AI
Comprehensive health benefits, monthly stipends, company retreats
A supportive and collaborative office culture—we’re all building and launching together
About Pika
At Pika, we're crafting a future where video creation is seamless, intuitive, and universally accessible. Our mission is to empower creativity by breaking down technical barriers using the transformative power of AI. We’re a tight-knit, energetic team based in Palo Alto, CA, valuing efficiency, curiosity, and the ambition to make a meaningful impact on the world.
We work from our Palo Alto office 3–5 days a week and welcome applicants who are eager to contribute onsite.
- ...next steps. Our partner is looking for a Machine Learning Engineer — Inference Optimization based in Australia. This role offers the opportunity to... ...complex technical challenges and building high-performance ML systems from the ground up. Accountabilities As a...SuggestedRemote jobFull timeFlexible hours
$141k - $249k
...technologies that can be adopted into Waabi’s training and inference frameworks. Examples include designing new CUDA... .../deployment techniques. ~Work with researchers and ML engineers on best-practices for optimal resource usage. ~Create and improve tooling and dashboards...SuggestedFull timeWork at officeWork from homeFlexible hours- We are looking for a Machine Learning Runtime Optimization Engineer to work on an innovative project that redefines... ...role, you will focus on backend optimizations for ML runtimes, including hardware acceleration, and inference speed improvements. Experience with ML...SuggestedRemote jobFull timeWorldwide
- ...connect a billion people with optimism and civility, and looking for... ...experiences for everyone. Our engine’s resource management and... ...optimization. You will establish the ML framework for predictive... ...roadmap. Design ML models that infer player and interaction...SuggestedFull time
$198k - $286k
Role Description At Modular, we optimize inference from kernel to cloud on one unified stack. We are building a differentiated cloud platform... ...apply the latest optimizations across kernels, the inference engine, and distributed systems so that customer workloads stay on the...SuggestedFull timeLocal areaFlexible hours- ...Service to Marketing we rely on ML to ensure that guests and... ...LLM fine-tuning, alignment and optimization, RAG/Search, LLM evaluation... ...a principal machine learning engineer, you will be responsible for... ...tuning, optimizing models and inference run-time ~ Post-training experience...Remote jobFull timeCasual workLive inWork at office
$159.05k - $199.3k
.... About The Role We are looking for a software engineer with deep experience in optimizing ML models and deploying them on production-grade embedded... ...to optimize efficiency and latency of model inference for compute boards selected by our customers Work...Full timeFor contractorsFor subcontractorCasual workWork at officeRemote workDay shift$170k - $216k
...with downstream teams on the optimization and integration into the Waymo... ...set of sensors, enabling engineers like you to (1) develop methods... ...in model training and model inference through model architecture/ hardware... ...Python ~ Experience with ML frameworks like PyTorch or...Full timeRemote work$298k - $368k
...work jointly with downstream teams on the optimization and integration into the Waymo Driver.... ...a diverse set of sensors, enabling engineers like you to (1) develop methods for efficiently... ...expertise in low-latency on-device inference techniques and a deep understanding of...Full timeRemote work$213k - $263k
...simulation across 15+ U.S. states. The ML Platform team at Waymo provides a set of... ...management, model development, optimization and monitoring. These efforts have resulted... ...Research and Simulation. We are looking for engineers with ML software or ML systems expertise...Full timeRemote work$137.1k - $201.6k
...Team The mission of the Marketplace Optimization team is to ensure we maintain a healthy... ...leverage artificial intelligence and advanced ML, deep learning techniques to power... ...We’re looking for a Machine Learning Engineer to help design, build, optimize and scale...Hourly payWork at officeLocal areaRemote workFlexible hours- ...Machine Learning Engineer (Llama AI Platform) Location: Remote (... ...to help design, deploy, and optimize AI solutions built on Llama models... ...architectures. Optimize inference performance, latency, and... ...tuning open-source LLMs. ML Engineering and MLOps practices...Full timeRemote work
- ...About the Role We’re looking for an Applied ML Engineer to design, evaluate, and scale recommendation and ranking systems... ...skips, conversions) and sparse explicit feedback. Optimize systems for real-time inference, scalability, and robustness under non-stationary user...Full time
$216.7k - $303.4k
...Description We are hiring Machine Learning Engineers (IC4) to build and evolve the auction,... ...role, you will: ~Design and implement optimization algorithms for auctions, bidding... ...optimization, and production monitoring for ML systems. ~Experience collaborating with...Full timeWork experience placementFlexible hours- ...Description Incubate and validate new ML initiatives end-to-end. On... ...connect ingest/processing to inference and visible product outcomes... ...with Research and Data Engineering on dataset curation, annotation... ...Nice‑to‑haves ~GPU serving/optimization experience (Triton/KServe,...Full timeRemote work
$180k - $260k
...You will work on model architecture, distributed training, inference optimization, and evaluation systems that power Tenzin and the broader Octave-X platform. The role bridges applied ML engineering and production reliability for regulated and high-trust environments....Full timeRemote work$185.8k - $303.4k
...Description This role sits in the Ads Optimization and Ads Marketplace Quality (AMQ)... ...platform. You’ll join a set of tight-knit engineers working on high-impact, internet-scale... ..., and production monitoring for ML systems. Experience collaborating with...Full timeFor contractorsWork experience placementFlexible hours- ...advance the frontier of intelligence. About the role: As a ML Engineer, you’ll build and operate the infrastructure that makes... ..., and control. What you might work on: Optimize inference and training throughput for novel model architectures Build...Full timeInternship
$152k - $228k
Role Description We're hiring a Senior ML Engineer to own the productionization layer of Invoca's ML stack — model serving, inference optimization, fine-tuning, and the APIs and pipelines that tie it all together. You'll be a primary driver of the infrastructure powering...Full timeCurrently hiringRemote workFlexible hours$177.3k - $212.8k
...Description The Auto Tagger team is the engine behind our data flywheel,... ...testing. ~Architect and optimize distributed data pipelines to... ...tune both heuristic-based and ML-assisted algorithms (including... ...-throughput model serving and inference. ~Experience with semantic...Full timeWork at officeImmediate start$143k - $197k
...are looking for an experienced Senior AI/ML Engineer to spearhead the design, development,... ...training, evaluation, and real-time/batch inference using modern MLOps frameworks. ~Collaborate... ...low latency, high availability, and optimal resource utilization. ~Architect...Full time- ...Join our San Francisco office as an ML Engineer focused on Data Engineering. Visa sponsorship... ...high availability and reliability. Optimize input/output operations to accelerate... ...retrieval and processing during training and inference phases. Build and manage backend...Full timeH1bWork at officeImmediate startVisa sponsorship
$116.1k - $170.28k
...skilled and experienced Senior Big Data Engineer to join our dynamic team. The ideal candidate... ...pipelines for data collection or batch inference. This is a remote position, requiring... ...and managing those solutions, and optimizing returns into the future. Named a best place...Remote jobFull time- ...protected veteran.**Job Description:****ML Engineer****3M Health Care is now Solventum****... ...Integration*** **Data Pipelines:** Develop and optimize ETL processes to transform healthcare... ...usable datasets for model training and inference.* **Feature Management:** Help build...H1bRemote work
- ...industries, including healthcare, to optimize processes, improve customer... ..., real-time and batch inference systems, and model serving.... ...GCP. ~Experience in data engineering in Big Data systems including... ...Responsibilities ~Build, refine, and use ML Engineering platforms and...Full timeFlexible hours
- ...unexpectedly, or need to be improved, engineers rely on data to understand what actually... ...the Role We're looking for an applied ML engineer with deep infrastructure... ...that makes ML work in production: from optimizing inference pipeline throughput to standing up training...Full timeRemote work
$150k - $220k
Role Description As an Applied ML Engineer, you will own and streamline the research-to-production... ...infrastructure — a hybrid training and inference stack spanning our own GPU data centers... ...whether a model is ready to ship. ~Optimize models and serving for production:...Full time- ...the Role We are seeking a highly experienced Principal ML Engineer (Applied / Systems) to join our engineering team and... ...results and productionize successful approaches. Build, optimize, and scale inference pipelines and model serving infrastructure on AWS. Collaborate...Full time
$251k - $310k
...collaborative group of machine learning (ML) engineers, software engineers, and ML research... ...and identify bottlenecks in training and inference performance (e.g., memory bandwidth,... ...and efficient attention mechanisms. Optimize model code for specific hardware accelerators...Full timeRemote work$151k - $177.5k
...battery from the inside out today. We engineer and manufacture ground-breaking battery... ...engineering logic, statistical methods, optimization, machine learning, and foundation models... ...workflows for feature generation, inference, event detection, and feedback into operational...Full time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to ML Engineer, Inference & Optimization. Be the first to apply!
- senior ml engineer Remote
- machine learning engineer Remote
- computer vision machine learning engineer Remote
- ai ml engineer Remote
- junior machine learning research engineer Remote
- machine learning software engineer Remote
- machine learning ai engineer Remote
- machine learning intern Remote
- machine learning part time Remote
- machine learning scientist Remote










