ML Engineer, Inference & Optimization
Pika
About the Role
We are seeking Senior/Staff level Inference Engineers to accelerate the performance of Pika's AI-driven products. In this highly technical role, you will operate at the intersection of cutting-edge inference acceleration, GPU parallelism, advanced model deployment, and video generation technologies. Your expertise will drive significant improvements to model speed and efficiency, ensuring our creative AI systems deliver industry-leading user experiences at scale.
You will design and optimize inference pipelines, implement state-of-the-art acceleration techniques, and work closely with researchers and engineers across the team to push the boundaries of what’s possible in real-time AI deployment. Your efforts will play a foundational role in powering the next generation of Pika’s video and language models.
What You’ll Do
Accelerate Inference : Lead and implement advanced inference acceleration techniques, including attention optimization and quantization for efficient model serving.
Maximize GPU Parallelism : Engineer and optimize GPU strategies across tensor, sequence, and pipeline parallelism (TP, SP, PP) for maximal efficiency and scalability.
Programming for Performance : Develop and optimize high-performance computing kernels and distributed workloads using CUDA and NCCL.
Advance AI Deployment : Collaborate with research and engineering teams to bring state-of-the-art videogen and large language models into production.
Improve Training Efficiency : (Bonus) Contribute to improvements in model training speed, stability, and resource utilization as part of our deployment lifecycle.
Technical Excellence : Drive rigorous code reviews, participate in technical discussions, and mentor fellow engineers on best practices in inference and GPU programming.
What We’re Looking For
Experience : 5+ years engineering experience, with a strong track record in inference acceleration and model deployment at scale.
Inference Mastery : Proven expertise in inference optimization, including quantization, attention acceleration, and deep learning compiler stacks.
GPU & Parallelism : Deep knowledge of GPU programming (CUDA, NCCL) and experience with SP, TP, PP, and other forms of parallelism for distributed inference.
AI Domain Knowledge : Familiarity with video generation (videogen) models and large language models (LLMs).
Collaboration : Strong cross-discipline communication skills; able to drive shared goals across research and engineering functions.
Ownership Mindset : Self-driven, solutions-oriented, and capable of managing ambiguity in a fast-paced startup environment.
Bonus : Experience in enhancing training efficiency, stability, or resource optimization for large models.
Nice to Have
Experience with high-throughput video or real-time streaming model deployment
Familiarity with distributed training and optimization toolkits
Contributions to open source projects in AI infrastructure or deep learning compilers
Startup or rapid prototyping experience
What We Offer
Competitive salary in the AI industry
Equity in a fast-growing startup shaping the future of AI
Comprehensive health benefits, monthly stipends, company retreats
A supportive and collaborative office culture—we’re all building and launching together
About Pika
At Pika, we're crafting a future where video creation is seamless, intuitive, and universally accessible. Our mission is to empower creativity by breaking down technical barriers using the transformative power of AI. We’re a tight-knit, energetic team based in Palo Alto, CA, valuing efficiency, curiosity, and the ambition to make a meaningful impact on the world.
We work from our Palo Alto office 3–5 days a week and welcome applicants who are eager to contribute onsite.
$100.4k - $180.7k
Posting TitleML and Optimization Engineer.LocationCO - Golden.Position TypeLimited Term (Fixed Term)... ...strengths in high‑performance computing, AI/ML, modeling and simulation, and... ...learning, probabilistic modeling, large-scale inference, foundation models, and applied AI in a...SuggestedFull timeFixed term contractLive inLocal areaRemote workRelocationShift work$159.05k - $199.3k
.... About The Role We are looking for a software engineer with deep experience in optimizing ML models and deploying them on production-grade embedded... ...to optimize efficiency and latency of model inference for compute boards selected by our customers Work...SuggestedFull timeFor contractorsFor subcontractorCasual workWork at officeRemote workDay shift$141k - $249k
...Collaborate closely with autonomy and algorithm engineers to scale safe self-driving systems using... ...such as TensorRT and modelopt to optimize the models running on the truck. - Create and benchmark new CUDA kernels for inference. - Comprehensively profile model...SuggestedWork at officeWork from homeFlexible hours- ...production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI....SuggestedFull time
$209k - $313k
...other digital services.Snap Engineering teams build fun and technically... ...that quantify causal impact, optimize decision-making, and drive... ...Strong understanding of causal inference and modern approaches to estimating... ...tests) and leveraging causal ML in production...SuggestedFull timeLive inWork at officeLocal area- ...deeply engaged audiences.Our TeamThe Decisioning & Optimization engineering team owns the systems that determine which ad... ...surfaces. Our work spans three platform areas:ML infrastructure for model serving: real-time inference at 1M+ QPS, multi-model parallel evaluation,...Hourly payFull timeImmediate startFlexible hoursShift work
$117.7k - $221.4k
...This operating model reflects how Cola engineers think: build durable intermediate artifacts... ...quality, speed, and cost instead of optimizing any one of them in isolation.The RoleWe... ...data processing, featurization, and inference foundations that power scalable world understanding...Full timeLocal areaRemote workWork from homeRelocation packageFlexible hours$203.5k - $299.3k
...are hiring a Causal Machine Learning Engineer to help build the causal ML foundation behind how DoorDash... ...counterfactual policy evaluation, promotion optimization, or marketplace decisioning systems... ...practical experience with causal inference, econometrics, experimentation, or...Hourly payWork at officeLocal areaRemote workFlexible hours- ...impact. Make Wayve the experience that defines your career! The role As a Staff ML Performance Engineer, you’ll play a key role in high-impact projects, optimising ML inference for edge accelerators and GPUs. The focus of this team is to run large transformer-...Full timeWork at officeWork from home
- ...perform real-time multi-objective optimization across distributed systems at... ...is our Machine Learning and Inference Platform that powers the... ...leader with deep experience in ML serving, high-performance... ...- someone excited to mentor engineers, innovate at scale, and shape...Work at officeLocal areaRemote workMonday to ThursdayFlexible hours
$147k - $268.4k
...never been done. If you're an engineer, scientist, or builder who... ...What You'll Be Doing As an ML Ops Engineer, you build and operate... ...at scale. You optimize infrastructure and GPU resources... ...Operate and optimize large-scale inference platforms that support scientific...Remote workFlexible hours2 days per week$147k - $268.4k
...never been done. If you're an engineer, scientist, or builder who... ...What You'll Be Doing As an ML Ops Engineer, you build and... ...reproducibility at scale. You optimize infrastructure and GPU resources... ...and optimize large-scale inference platforms that support scientific...Full timeRemote workFlexible hours2 days per week- ...Job Title : ML Engineer Location : Minnetonka, MN FULLTIME ONLY... ...and pipelines. • Build training and inference code with reproducibility, versioning,... ...offline), batching, and latency/throughput optimization. • Integrate model lifecycle tooling...Full time
- ...dedicated owner of Confido's ML platform. Our AI/ML team already... ...behind training, inference, and agentic workloads... ...ML safely ~ Serve and optimize inference and forecasting workloads... ...the data interface with data engineering: serve the right data to models...Local areaRelocation packageWeekend work
$128.7k - $261.3k
...development, and performance engineering so that every cycle on our accelerators... ...models into fast, reliable inference across GPUs powering GM’s... ...and turns them into highly optimized inference artifacts running... ...reliable, and effortless for ML engineers across the AV organization...Full timeLocal areaRemote workWork from homeRelocation packageFlexible hours$173k - $253k
Matterport - Senior ML Ops Engineer Job Description CoStar Group is a leading global provider... ...bottlenecks, applying advanced optimization techniques, and deploying highly efficient... ...analyze model performance, optimize inference speed and resource utilization, and...Full timeWork at officeWork from home$170.1k - $258.3k
...export, kernel development, and performance engineering so that every cycle on our accelerators... ...sit at the heart of our on‑vehicle ML inference for ADAS and autonomous driving. We own... ...benchmarking, profiling, debugging and optimizing accelerator libraries and kernels to...Full timeLocal areaRemote workWork from homeRelocation packageFlexible hours$165.2k - $223.6k
...performance for AWS's custom ML accelerators. Working at the... ...hardware-software boundary, our engineers craft high-performance... ...every FLOP counts in delivering optimal performance for our customers... ...PyTorch, enabling unparalleled ML inference and training performance.As...InternshipLocal areaWork from homeFlexible hours- AI/ML Ops EngineerLocation: Remote / Hybrid (Client-Facing Consulting... ...seeking an experienced AI/ML Engineer to design, deploy, and... ...involving forecasting, optimization, decision support, and advanced... ...including real-time APIs, batch inference pipelines, and feature stores...Full timeContract workLocal areaRemote workFlexible hours
- ...career growth potential. Job Title: Edge ML Engineer Location: 100% Remote (U.S.) Position... ...looking for an Edge ML Engineer to design, optimize, and deploy machine learning models... ...Experience with at least one major edge inference framework. Solid understanding of...Full timeLocal areaRemote work
$180k - $260k
You will work on model architecture, distributed training, inference optimization, and evaluation systems that power Tenzin and the broader Octave-X platform. The role bridges applied ML engineering and production reliability for regulated and high-trust environments....Full timeRemote work- ...every person shapes what gets built. About the Role As a ML engineer at Wispr, you’ll play a crucial role in building the first... ...for? Previous founding or startup experience Experience optimizing ML inference or engineering systems for research teams Fluency in Python...H1bWork at officeRemote workRelocationVisa sponsorshipFlexible hours
- ...a protected veteran. Job Description ML Engineer 3M Health Care is now Solventum At Solventum... ...Data Pipelines: Develop and optimize ETL processes to transform healthcare data... ...usable datasets for model training and inference. Feature Management: Help build and maintain...H1bRemote work
- ...experienced Machine Learning Engineers (MLE Bench) to contribute to... ...on work with production-grade ML codebases, model training and... ...training, evaluation, and inference pipelines . Prepare datasets... ..., evaluation metrics, optimization). Experience working with...Contract workTemporary workFor contractorsFreelanceRemote work
$80 - $150 per hour
...Role Overview Contribute expert machine learning engineering work to projects that develop, train, evaluate,... ...tasks spanning model development, training and inference systems, numerical computing, and performance optimization. Key Responsibilities Develop and...Hourly payFor contractorsRemote workFlexible hours$80 - $150 per hour
...ML Engineer Pay: $80, $150/hour Location: Global, fully remote Job Type: Contractor (~15 hours per... ...project involving model development, training and inference systems, numerical computing, performance optimization, and Python. The work involves creating,...Temporary workFor contractorsImmediate startRemote workFlexible hours$200k - $240k
...Senior Machine Learning Engineer (Computer Vision / Vision-Language Models) Remote... ...evaluation, and monitoring frameworks Optimize inference pipelines for performance, accuracy, and... ...skills ~ Proven record of deploying ML systems into production ~ Experience...Remote work$260k - $330k
...Experience building large-scale prediction or optimization systemsPubMatic is the leading AI-... ...for a Senior Principal Machine Learning Engineer to help build the next generation of performance... ...ROAS. The ideal candidate has strong ML fundamentals and experience building...Work at officeRemote work$180k - $280k
...across real-world scenarios.As a Staff AI/ML Engineer within the Onboard Embodied AI... ...-to-end solutions capable of real-time inference and robust autonomous driving performance... ...training methodologies, and inference optimization strategies suited for real-time onboard...Full timeLocal areaWork from homeRelocationRelocation package- ...world.Role OverviewAs our Staff Software Engineer, ML infra Engineer for Search & Discovery... ...Discovery organization is responsible for optimizing customers' navigation experience and... ...ML based ranking system and online ML inference servicesBuild strong cross-functional partnerships...Temporary work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to ML Engineer, Inference & Optimization. Be the first to apply!
- junior machine learning research engineer Remote
- data scientist machine learning engineer Remote
- graduate machine learning engineer Remote
- junior machine learning engineer Remote
- computer vision machine learning engineer Remote
- machine learning software engineer Remote
- machine learning ai engineer Remote
- senior ml engineer Remote
- machine learning engineer Remote
- ai ml engineer Remote


