Sr. Inference Optimization Engineer (local / edge runtime)
$195.2k - $361.2kIntel
Job Details:Job Description: Job DescriptionOur MissionAt Intel, our journey is to transform AI into something safer, more trustworthy, and respectful of human privacy by design. We believe transformative AI should have a positive impact on people—powerful in capability, yet honest about its limits and protective of the data and resources it touches.To get there, we build agentic AI that combines the best of local and cloud intelligence — private, affordable, and sustainable by design. Small, efficient models run directly on the user's machine (AI PC, edge, on-prem, and beyond), keeping data private and token costs low, while powerful cloud models handle the hardest work: planning, reasoning, and complex problem-solving. Today, neither approach can deliver this alone. Together, they give people real capability without compromise—data stays private, spend stays predictable, and energy use stays in check.We're building intelligence that scales without sacrificing trust, cost, or the planet—because the future of AI should belong to the people it servesRole SummaryMake models fast on the hardware people actually own. You optimize inference engines (llama.cpp, vLLM) for constrained local and edge environments — GPU/iGPUs, Vulkan backends — not datacenter H100 environment, mostly PC/edge. KV cache, batching, quantization, scheduling, and CPU-overhead reduction are your daily tools.This is the rare skill that makes a hybrid, low-cost agent product viable.What you’ll doProfile and optimize local inference (llama.cpp-vulkan and vLLM) for latency, throughput, and memory on edge hardwareTune KV cache, continuous batching, and scheduling for interactive agent workloadsDrive quantization strategy (GGUF / AWQ / GPTQ) and validate quality impact with the Post-Training teamCut CPU overhead and improve engine startup, model load, and lifecycle (start / stop / health)Benchmark across hardware tiers and publish honest performance comparisonsUpstream fixes and patches to open-source engines where it helps usWhat you’ll learn / grow intoCuriosity is required. You will develop:The internals of modern inference engines and where the milliseconds actually goHardware-aware optimization across iGPU / CPU paths (Vulkan, SYCL, oneAPI, CUDA where relevant)The quality-vs-speed-vs-memory trade space for small modelsInterest in local / edge AI and squeezing hardwareIMPORTANT:Please be informed that Intel is proactively tryingto find candidates for this position which is frequently availableat Intel.Please note that the position may not be availableat this time. If you would be interested in this position should itbecome available, we would encourage you to apply, and ourhiring team will be glad to contact you when/if relevant. Qualifications:QualificationsMinimum qualifications are required to be initially considered for this position. Preferred qualifications are in addition to the minimum requirements and are considered a plus factor in identifying top candidates.You must possess the minimum qualifications to be initially considered for this position. Preferred qualifications are in addition to the minimum requirements and are considered a plus factor in identifying top candidates.Required QualificationsBS/MS in CS, EE, Math or related STEM field8+ years software development backgroundStrong in C++ and/or Python; comfortable reading systems-level codeExperience with LLM inference. (attention, KV cache, decoding)Experience profiling and optimizing real performance problems (CPU or GPU) and can prove the speedupLinux, build systems, and low-level debugging expertisePreferred QualificationsHands-on with llama.cpp, vLLM, ggml, or similar enginesExperience with GPU / accelerator programming (Vulkan, CUDA, SYCL, Metal) or SIMD / CPU kernelsFamiliarity with quantization formats and their quality trade-offsOpen-source contributions to inference enginesRequirements listed would be obtained through a combination of industry relevant job experience, internship experiences and or schoolwork/classes/research.Benefits at IntelOur total rewards package goes above and beyond just a paycheck. Whether you're looking to build your career, improve your health, or protect your wealth, we offer generous benefits to help you achieve your goals. Go to Intel Benefits | Intel Careers for details of benefits available to you. Intel reserves the right to modify, change or discontinue benefit plans at any time in its sole discretion.Job Type:Shift:Shift 1 (United States of America)Primary Location: US, California, Santa ClaraAdditional Locations:US, Arizona, Phoenix, US, California, Folsom, US, Oregon, HillsboroPosting Statement:All qualified applicants will receive consideration for employment without regard to race, color, religion, religious creed, sex, national origin, ancestry, age, physical or mental disability, medical condition, genetic information, military and veteran status, marital status, pregnancy, gender, gender expression, gender identity, sexual orientation, or any other characteristic protected by local law, regulation, or ordinance.Position of TrustN/ABenefitsWe offer a total compensation package that ranks among the best in the industry. It consists of competitive pay, stock bonuses, and benefit programs which include health, retirement, and vacation. Find out more about the benefits of working at Intel. Annual Salary Range for jobs which could be performed in the US: $195,200.00-361,200.00 USDThe range displayed on this job posting reflects the minimum and maximum target compensation for the position across all US locations. Within the range, individual pay is determined by work location and additional factors, including job-related skills, experience, and relevant education or training. Your recruiter can share more about the specific compensation range for your preferred location during the hiring process.Work Model for this RoleThis role will be eligible for our hybrid work model which allows employees to split their time between working on-site at their assigned Intel site and off-site. * Job posting details (such as work model, location or time type) are subject to change.*ADDITIONAL INFORMATION: Intel is committed to Responsible Business Alliance (RBA) compliance and ethical hiring practices. We do not charge any fees during our hiring process. Candidates should never be required to pay recruitment fees, medical examination fees, or any other charges as a condition of employment. If you are asked to pay any fees during our hiring process, please report this immediately to your recruiter.SummaryLocation: US, California, Santa Clara; US, Oregon, Hillsboro; US, California, Folsom; US, Arizona, PhoenixType: Full time
- ...agentic AI that combines the best of local and cloud intelligence - private, affordable... ...on the user's machine (AI PC, edge, on-prem, and beyond), keeping data private... ...the hardware people actually own. You optimize inference engines (llama.cpp, vLLM) for constrained...Local areaSeniorShift work
$193.3k - $261.5k
...looking for a Senior Inference Engineer to own inference for real... ...building thereal-time runtime that serves it within... ...in• Implement and optimize the inference path for... ...AWS Neuron/Trainium, edge accelerators) and how... ...all federal, state, and local laws and Company policies...Local areaSeniorInternshipFlexible hours$152k - $241.5k
...Developer Technology Engineer, you will be at... ...workflows at the edge powered by NVIDIAs... ...enterprise ISVs on solving local end-to-end agentic... ...in suboptimal runtime performance.... ...deployment targeting optimal runtime... ...GPU-accelerated AI inference driven by NVIDIA APIs...Local areaSeniorFull time- ...seeking outstanding AI systems engineers to advance the inference software stack. You will design and optimize kernels, build new... ...contribute to accelerators and runtimes that power large language models... ...CUDA C/C++, Triton, and cutting-edge MLIR-based tooling, and participate...Senior
$184k - $287.5k
We're now looking for a Sr. Inference Engineer, for GPU Kernel Optimization! What does it take to push every LLM inference operation to its performance ceiling... ...across kernel execution, compiler decisions, and runtime scheduling.Direct experience with LLM inference frameworks...SeniorFull time$182.5k - $260.5k
...One platform, its Zero Trust Engine, and the powerful NewEdge... ...Learning Scientist, you own the inference and optimization layer that makes AI in... ...real hardware, and build the runtime that executes bounded AI... ...economics of agentic AI.Cutting-edge, unusual stack. The hard,...Senior$229.9k - $262.4k
Sr. Lead AI Engineer (Inference Optimization, FM hosting, AI Platform) Overview: At Capital One, we are creating responsible and reliable AI systems, changing... ...in compliance with applicable federal, state, and local laws. Capital One promotes a drug-free workplace....Local areaSeniorFull timePart time- ...the Senior Robotics Engineer role at Knightscope... ...combines robotics, edge AI, and cloud services... ...2-based perception, localization, planning, and cloud... ...TensorFlow Lite / ONNX Runtime‑based inference, anomaly detection,... ...DDS tuning and QoS optimization. Experience with...Local areaSeniorInternship
$197.53k - $276.54k
...blending proven and cutting edge manufacturing... ...highly motivated Software Engineers (AI/ML Ops) to join... ..., model training, and inference at scale. Set up and... ...troubleshooting, and optimization of AI and ML models to... ...federal, state, and/or local law. Blue Origin will...Local areaSeniorPermanent employmentFull timeTemporary work$152k - $241.5k
...eager to work on cutting-edge AI technology for... ...as a Senior Software Engineer, and be at the forefront... ...enabling high-performance AI inference solutions for... ...TensorRT's compiler and runtime for specialized and constrained... ...to performance optimization and benchmarking...SeniorFull time$184k - $287.5k
...globally. We seek a Senior Engineer to lead technical efforts in... ...advanced AI agent frameworks and local runtimes on Windows and NVIDIA... ...By combining powerful local inference (Nemotron models) with strong... ...the engineering efforts to optimize the agent runtimes for Windows...Local areaSeniorFull timeShift work$193.3k - $261.5k
...accelerators. Join us to optimize the latest models to run... ...Trainium hardware.As a Sr. Software Development Engineer on the Inference Model Enablement team, you... ...with the compiler and runtime teams* Mentoring junior... ...all federal, state, and local laws and Company policies...Local areaSeniorInternshipFlexible hours- ...ROLE:We are looking for a Senior GPU Inference Performance Engineer to own end-to-end performance analysis... ...GPU silicon through the software runtime, and drive competitive positioning against... ...framework performance: Profile and optimize inference engines including vLLM, SGLang...Senior
$286.2k - $326.7k
...Senior Distinguished Engineer, AI Compute (... ...scale developer and runtime environments required... ...training, model inference and feature... ...Capital One to help optimize business outcomes... ...00 - $326,700 for Sr. Distinguished AI... ...federal, state, and local laws. Capital One...Local areaSeniorFull timePart timeRemote work$183k - $247.6k
...software, hardware, and network engineers, supply chain specialists,... ...creates compute, accelerator, edge and storage server designs... ...building software services for optimizing SSD performance and health at... ...follow all federal, state, and local laws and Company policies....Local areaSeniorFlexible hours$174.05k - $278.45k
Distinguished Technologist, Edge AI Architect... ...of AI will be built local, more secure, more mobile... ...OS trust foundation, inference serving, agentic runtimes and creation tooling,... ...and establishes the engineering standards and best practices that optimize development. Working...Local areaFull timeContract workTemporary workWork experience placementFlexible hoursShift work$170.6k - $261.3k
...Senior Machine Learning Engineer on the State Estimation and... ..., robustness under edge cases) and drive systematic... ...Implement efficient training and inference pipelines, including model optimization techniques (e.g., pruning... ...with federal, state and local laws. We encourage...Local areaSeniorFull timeRemote workWork from homeRelocation packageFlexible hours$160k - $198k
...DoAs a Senior AI Systems Engineer, you will architect,... ...AI model training and inference. You will ensure our machine... ...—to streamline and optimize the AI development lifecycle... ...productionize cutting-edge models, establish... ...protected by federal, state or local laws.Archer is...Local areaSenior$166k - $244k
...Role:We are looking for a Software Engineer, Edge Systems & Runtime to join our team, focusing on high-performance... ...GPUs.Your primary focus will be optimizing our production C++ runtime... ...routing multi-sensor inputs into model inference loops.Collaborate with machine learning...Full time- ...industry-leading training and inference speeds; over 10 times faster... ...global enterprises, and cutting-edge AI-native startups. OpenAI... ...hiring a Senior Performance Engineer to join our Product team. You... ...TensorRT-LLM), GPU kernel-level optimization toolchains (CUDA, Triton),...SeniorContract workShift work
- ...Packard Enterprise is the global edge-to-cloud company advancing... ...high‑impact systems, mentors engineering teams, and helps define long‑... .... Experience designing and optimizing distributed query engines and... ...report the incident to your local authorities immediately.SummaryLocation...Local areaFull timeWork experience placementWork at officeImmediate start2 days per week
$153.6k - $234.1k
.... As a Senior Systems Engineer, you will be at the heart... ...subsystem design, optimize performance, and contribute... ...and safety of cutting-edge autonomous technologies... ...middleware, orchestration, and runtime; low-level SoC/MCU... ...federal, state and local laws. We encourage interested...Local areaSeniorFull timeWork experience placementWork from homeRelocation packageFlexible hours- ...ROLEWe are hiring a senior engineering leader to define,... ...processing dataflow and runtime, distributed compute/... ...pre/post-processing, inference, tracking, streaming,... ...Heterogeneous compute optimization across CPU/GPU/NPU (kernel... ...SDKs for embedded/edge markets: robotics, industrial...Senior
$149.1k - $215.93k
...cloud, networking, and edge. As an independent... ...You will be one of the engineers who turns that raw delivery... ...flows, including runtime and resource prediction... ...debug, runtime/resource optimization, and knowledge capture... ...protected by local law, regulation, or ordinance...Local areaSeniorFull timeShift work$193.3k - $261.5k
The AWS BIOS Engineering team creates and maintains custom firmware solutions... ...to maintain its competitive edge in cloud computing.The ideal... ...-level expertise to find optimal solutions to complex boot-... ...follow all federal, state, and local laws and Company policies. Criminal...Local areaSeniorInternshipWorldwideFlexible hours$106k - $205k
Title: Staff System Engineers Engineer, AI/... ...study, model, and optimize AI/ML workloads for... ...teams (compilers, runtimes, ML frameworks), serving... ...ML acceleration on edge devices — NPUs, dedicated inference accelerators, DSP-based... ...specified by local, state or federal law...Local areaFull timeWork at office$184.5k - $249.6k
...is seeking a 'Principal FAE - Edge AI' to support strategic... ...relationships with architects, engineering leaders, and key decision-makers... ..., integration, and optimization to meet power, performance, scalability... ...we can offer is limited by local legal, regulatory, tax, or other...Local areaWork at office- ...seeking a Senior Agentic AI Software Engineer (Finance) to build agentic systems and advance AI inference workloads in real-world finance contexts... ...that drives experimental agents, optimize performance, and contribute to cutting-edge AI research at scale. This role collaborates...Senior
- ...advance your career. THE ROLEWe are seeking a Principal GenAI Inference Optimization Engineer to join our Models and Applications team. This role... ...comfortable working across multiple layers—from kernels and runtimes to frameworks and serving systems—and can independently...
$170.6k - $261.3k
...As a Senior State Estimation Engineer within the SEAM Embodied AI organization... ...contributor driving cutting-edge sensor calibration and state... ...all sensors to ensure they optimally support many other vehicle... ...the purposes of real-time localization, static and dynamic sensor...Local areaSeniorFull timeWork at officeWork from homeFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Sr. Inference Optimization Engineer (local / edge runtime). Be the first to apply!
- senior associate architect Santa Clara, CA
- senior dynamics crm developer Santa Clara, CA
- senior application security Santa Clara, CA
- senior account director Santa Clara, CA
- senior supervisor Santa Clara, CA
- senior plumbing designer Santa Clara, CA
- senior advisor Santa Clara, CA
- senior cloud data engineer Santa Clara, CA
- senior customer service manager Santa Clara, CA
- senior marketing coordinator Santa Clara, CA

