Sr. Inference Optimization Engineer (local / edge runtime)
$195.2k - $361.2kIntel
Job Details:Job Description: Job DescriptionOur MissionAt Intel, our journey is to transform AI into something safer, more trustworthy, and respectful of human privacy by design. We believe transformative AI should have a positive impact on people—powerful in capability, yet honest about its limits and protective of the data and resources it touches.To get there, we build agentic AI that combines the best of local and cloud intelligence — private, affordable, and sustainable by design. Small, efficient models run directly on the user's machine (AI PC, edge, on-prem, and beyond), keeping data private and token costs low, while powerful cloud models handle the hardest work: planning, reasoning, and complex problem-solving. Today, neither approach can deliver this alone. Together, they give people real capability without compromise—data stays private, spend stays predictable, and energy use stays in check.We're building intelligence that scales without sacrificing trust, cost, or the planet—because the future of AI should belong to the people it servesRole SummaryMake models fast on the hardware people actually own. You optimize inference engines (llama.cpp, vLLM) for constrained local and edge environments — GPU/iGPUs, Vulkan backends — not datacenter H100 environment, mostly PC/edge. KV cache, batching, quantization, scheduling, and CPU-overhead reduction are your daily tools.This is the rare skill that makes a hybrid, low-cost agent product viable.What you’ll doProfile and optimize local inference (llama.cpp-vulkan and vLLM) for latency, throughput, and memory on edge hardwareTune KV cache, continuous batching, and scheduling for interactive agent workloadsDrive quantization strategy (GGUF / AWQ / GPTQ) and validate quality impact with the Post-Training teamCut CPU overhead and improve engine startup, model load, and lifecycle (start / stop / health)Benchmark across hardware tiers and publish honest performance comparisonsUpstream fixes and patches to open-source engines where it helps usWhat you’ll learn / grow intoCuriosity is required. You will develop:The internals of modern inference engines and where the milliseconds actually goHardware-aware optimization across iGPU / CPU paths (Vulkan, SYCL, oneAPI, CUDA where relevant)The quality-vs-speed-vs-memory trade space for small modelsInterest in local / edge AI and squeezing hardwareIMPORTANT:Please be informed that Intel is proactively tryingto find candidates for this position which is frequently availableat Intel.Please note that the position may not be availableat this time. If you would be interested in this position should itbecome available, we would encourage you to apply, and ourhiring team will be glad to contact you when/if relevant. Qualifications:QualificationsMinimum qualifications are required to be initially considered for this position. Preferred qualifications are in addition to the minimum requirements and are considered a plus factor in identifying top candidates.You must possess the minimum qualifications to be initially considered for this position. Preferred qualifications are in addition to the minimum requirements and are considered a plus factor in identifying top candidates.Required QualificationsBS/MS in CS, EE, Math or related STEM field8+ years software development backgroundStrong in C++ and/or Python; comfortable reading systems-level codeExperience with LLM inference. (attention, KV cache, decoding)Experience profiling and optimizing real performance problems (CPU or GPU) and can prove the speedupLinux, build systems, and low-level debugging expertisePreferred QualificationsHands-on with llama.cpp, vLLM, ggml, or similar enginesExperience with GPU / accelerator programming (Vulkan, CUDA, SYCL, Metal) or SIMD / CPU kernelsFamiliarity with quantization formats and their quality trade-offsOpen-source contributions to inference enginesRequirements listed would be obtained through a combination of industry relevant job experience, internship experiences and or schoolwork/classes/research.Benefits at IntelOur total rewards package goes above and beyond just a paycheck. Whether you're looking to build your career, improve your health, or protect your wealth, we offer generous benefits to help you achieve your goals. Go to Intel Benefits | Intel Careers for details of benefits available to you. Intel reserves the right to modify, change or discontinue benefit plans at any time in its sole discretion.Job Type:Shift:Shift 1 (United States of America)Primary Location: US, California, Santa ClaraAdditional Locations:US, Arizona, Phoenix, US, California, Folsom, US, Oregon, HillsboroPosting Statement:All qualified applicants will receive consideration for employment without regard to race, color, religion, religious creed, sex, national origin, ancestry, age, physical or mental disability, medical condition, genetic information, military and veteran status, marital status, pregnancy, gender, gender expression, gender identity, sexual orientation, or any other characteristic protected by local law, regulation, or ordinance.Position of TrustN/ABenefitsWe offer a total compensation package that ranks among the best in the industry. It consists of competitive pay, stock bonuses, and benefit programs which include health, retirement, and vacation. Find out more about the benefits of working at Intel. Annual Salary Range for jobs which could be performed in the US: $195,200.00-361,200.00 USDThe range displayed on this job posting reflects the minimum and maximum target compensation for the position across all US locations. Within the range, individual pay is determined by work location and additional factors, including job-related skills, experience, and relevant education or training. Your recruiter can share more about the specific compensation range for your preferred location during the hiring process.Work Model for this RoleThis role will be eligible for our hybrid work model which allows employees to split their time between working on-site at their assigned Intel site and off-site. * Job posting details (such as work model, location or time type) are subject to change.*ADDITIONAL INFORMATION: Intel is committed to Responsible Business Alliance (RBA) compliance and ethical hiring practices. We do not charge any fees during our hiring process. Candidates should never be required to pay recruitment fees, medical examination fees, or any other charges as a condition of employment. If you are asked to pay any fees during our hiring process, please report this immediately to your recruiter.SummaryLocation: US, California, Santa Clara; US, Oregon, Hillsboro; US, California, Folsom; US, Arizona, PhoenixType: Full time
$152k - $241.5k
...Developer Technology Engineer, you will be at... ...workflows at the edge powered by NVIDIAs... ...enterprise ISVs on solving local end-to-end agentic... ...in suboptimal runtime performance.... ...deployment targeting optimal runtime... ...GPU-accelerated AI inference driven by NVIDIA APIs...Local areaSeniorFull time$184k - $287.5k
We're now looking for a Sr. Inference Engineer, for GPU Kernel Optimization! What does it take to push every LLM inference operation to its performance ceiling... ...across kernel execution, compiler decisions, and runtime scheduling.Direct experience with LLM inference frameworks...SeniorFull time$140k - $215k
...This is a Software Development Engineer (SDE) role in the engineering... ...agent) on various container optimized Linux distros. This role will... ...and developing container runtime engines, software that monitors... ...membership or activity in a local human rights commission, status...Local areaSeniorFull timeWork experience placementWork at officeRemote work$224k - $356.5k
...are looking for a Senior Software Engineer to build and optimize the local runtimes and agent frameworks that bring autonomous... ...combining high-performance local inference (Nemotron models) with robust... ...constraints and advantages for edge AI.Experience building virtualization...Local areaSeniorFull timeWorldwide- ...the Senior Robotics Engineer role at Knightscope... ...combines robotics, edge AI, and cloud services... ...2–based perception, localization, planning, and cloud... ...TensorFlow Lite / ONNX Runtime‑based inference, anomaly detection,... ...DDS tuning and QoS optimization. Experience with...Local areaSeniorInternship
$182.5k - $260.5k
...One platform, its Zero Trust Engine, and the powerful NewEdge... ...Learning Scientist, you own the inference and optimization layer that makes AI in... ...real hardware, and build the runtime that executes bounded AI... ...economics of agentic AI.Cutting-edge, unusual stack. The hard,...Senior- ...looking for a specialized Sr. Staff or Principal level engineer who is passionate about... ...on scaling training and inference for the latest Generative... ...model training and inference optimizations across a variety of... ...best in the field.Cutting-edge Technology: Work with state...Senior
$152k - $241.5k
...eager to work on cutting-edge AI technology for... ...as a Senior Software Engineer, and be at the forefront... ...enabling high-performance AI inference solutions for... ...TensorRT's compiler and runtime for specialized and constrained... ...to performance optimization and benchmarking...SeniorFull time$184k - $287.5k
...globally. We seek a Senior Engineer to lead technical efforts in... ...advanced AI agent frameworks and local runtimes on Windows and NVIDIA... ...By combining powerful local inference (Nemotron models) with strong... ...the engineering efforts to optimize the agent runtimes for Windows...Local areaSeniorFull timeShift work- ...industry-leading training and inference speeds; over 10 times faster... ...global enterprises, and cutting-edge AI-native startups. OpenAI... ...hiring a Senior Performance Engineer to join our Product team. You... ...TensorRT-LLM), GPU kernel-level optimization toolchains (CUDA, Triton),...SeniorContract workShift work
- ...ROLE:We are looking for a Senior GPU Inference Performance Engineer to own end-to-end performance analysis... ...GPU silicon through the software runtime, and drive competitive positioning against... ...framework performance: Profile and optimize inference engines including vLLM, SGLang...Senior
$183k - $247.6k
...software, hardware, and network engineers, supply chain specialists,... ...creates compute, accelerator, edge and storage server designs... ...building software services for optimizing SSD performance and health at... ...follow all federal, state, and local laws and Company policies....Local areaSeniorFlexible hours$170.6k - $261.3k
...Senior Machine Learning Engineer on the State Estimation and... ..., robustness under edge cases) and drive systematic... ...Implement efficient training and inference pipelines, including model optimization techniques (e.g., pruning... ...with federal, state and local laws. We encourage...Local areaSeniorFull timeRemote workWork from homeRelocation packageFlexible hours$160k - $198k
...DoAs a Senior AI Systems Engineer, you will architect,... ...AI model training and inference. You will ensure our machine... ...—to streamline and optimize the AI development lifecycle... ...productionize cutting-edge models, establish... ...protected by federal, state or local laws.Archer is...Local areaSenior$164k - $313.3k
...Systems & Efficiency Engineer to join our R&D team focused... ...ready improvements in inference performance, latency,... ..., systems, inference runtimes, and services, with a... ..., and performance optimization. You will work closely... ...accordance with state and local laws and “fair chance”...Local areaSeniorFull timeTemporary workWorldwide$166k - $244k
...Role:We are looking for a Software Engineer, Edge Systems & Runtime to join our team, focusing on high-performance... ...GPUs.Your primary focus will be optimizing our production C++ runtime... ...routing multi-sensor inputs into model inference loops.Collaborate with machine learning...Full time$124k - $195.5k
NVIDIA is recruiting a Senior Inference Performance Engineer to push NVIDIA's performance limits on large-scale AI inference benchmarks. This position... ...provides an outstanding opportunity to employ your optimization knowledge in an autonomous optimization framework. AI...Full time$153.6k - $234.1k
.... As a Senior Systems Engineer, you will be at the heart... ...subsystem design, optimize performance, and contribute... ...and safety of cutting-edge autonomous technologies... ...middleware, orchestration, and runtime; low-level SoC/MCU... ...federal, state and local laws. We encourage interested...Local areaSeniorFull timeWork experience placementWork from homeRelocation packageFlexible hours- ...ROLEWe are hiring a senior engineering leader to define,... ...processing dataflow and runtime, distributed compute/... ...pre/post-processing, inference, tracking, streaming,... ...Heterogeneous compute optimization across CPU/GPU/NPU (kernel... ...SDKs for embedded/edge markets: robotics, industrial...Senior
- ...Packard Enterprise is the global edge-to-cloud company advancing... ...high‑impact systems, mentors engineering teams, and helps define long‑... .... Experience designing and optimizing distributed query engines and... ...report the incident to your local authorities immediately.SummaryLocation...Local areaFull timeWork experience placementWork at officeImmediate start2 days per week
$193.3k - $261.5k
The AWS BIOS Engineering team creates and maintains custom firmware solutions... ...to maintain its competitive edge in cloud computing.The ideal... ...-level expertise to find optimal solutions to complex boot-... ...follow all federal, state, and local laws and Company policies. Criminal...Local areaSeniorInternshipWorldwideFlexible hours$140k - $180k
...model architectureCollaborate with Systems Engineering to properly decompose higher level... ...generationStrong knowledge of vehicle dynamics, optimization, and both classical and modern control... ...protected by federal, state or local laws. For this position we are targeting...Local areaSenior$229.9k - $262.4k
...Overview Sr. Lead AI Engineer (Inference Optimization, FM hosting, AI Platform) Overview: At Capital One, we are creating responsible and reliable... ...in compliance with applicable federal, state, and local laws. Capital One promotes a drug-free workplace. Capital...Local areaSeniorFull timePart time$184.5k - $249.6k
...is seeking a 'Principal FAE - Edge AI' to support strategic... ...relationships with architects, engineering leaders, and key decision-makers... ..., integration, and optimization to meet power, performance, scalability... ...we can offer is limited by local legal, regulatory, tax, or other...Local areaWork at office- ...of patients worldwide. Were a team of engineers, clinicians, and innovators united by one... ...with supplier engineers and vendors to optimize component DFM (injection molded and machined... ...protected under federal, state, or local applicable laws. Mandatory Notices U.S...Local areaSeniorFull timeWorldwideFlexible hoursShift work
- ...advance your career. THE ROLEWe are seeking a Principal GenAI Inference Optimization Engineer to join our Models and Applications team. This role... ...comfortable working across multiple layers—from kernels and runtimes to frameworks and serving systems—and can independently...
- ...patients worldwide.We’re a team of engineers, clinicians, and innovators... ...design requirements into optimized processes, tooling, fixturing... ...techniquesKnowledge of cutting-edge technologies used in metal... ...protected under federal, state, or local applicable laws.Mandatory...Local areaSeniorWorldwideFlexible hours
$170.6k - $261.3k
...contributor driving cutting-edge sensor calibration and state... ...of all sensors to ensure they optimally support many other vehicle subsystems... ...-functional teams, mentor engineers across the organization, and... ...the purposes of real-time localization, static and dynamic sensor calibration...Local areaSeniorFull timeRemote workWork from homeRelocation packageFlexible hours- ...millions of patients worldwide.We’re a team of engineers, clinicians, and innovators united by... ...process changes, production optimization, facility changes, new equipment qualification... ...status protected under federal, state, or local applicable laws.Mandatory NoticesU.S. Export...Local areaSeniorContract workWorldwideFlexible hours
$120.8k - $193.3k
DescriptionJob Title: Sr. Engineer, AI/ML Apps Engineering... ..., drones, and edge AI solutions. This is... ...architecture, AI software optimization, and technical leadership... ...PyTorch, TensorFlow, ONNX Runtime, OpenCV, or similar.... ...understanding of AI inference optimization, model deployment...SeniorFull timeWork at office
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Sr. Inference Optimization Engineer (local / edge runtime). Be the first to apply!
- senior manager tax Santa Clara, CA
- senior devops Santa Clara, CA
- senior associate vice president Santa Clara, CA
- senior director digital marketing Santa Clara, CA
- senior director global marketing Santa Clara, CA
- senior international accountant Santa Clara, CA
- senior vmware engineer Santa Clara, CA
- sr marketing manager Santa Clara, CA
- sr technical product manager Santa Clara, CA
- senior resident engineer Santa Clara, CA

