Sr. Inference Optimization Engineer (local / edge runtime) (Santa Clara)
$195.2k - $361.2kIntel
Job Details:Job Description: Job DescriptionOur MissionAt Intel, our journey is to transform AI into something safer, more trustworthy, and respectful of human privacy by design. We believe transformative AI should have a positive impact on people—powerful in capability, yet honest about its limits and protective of the data and resources it touches.To get there, we build agentic AI that combines the best of local and cloud intelligence — private, affordable, and sustainable by design. Small, efficient models run directly on the user's machine (AI PC, edge, on-prem, and beyond), keeping data private and token costs low, while powerful cloud models handle the hardest work: planning, reasoning, and complex problem-solving. Today, neither approach can deliver this alone. Together, they give people real capability without compromise—data stays private, spend stays predictable, and energy use stays in check.We're building intelligence that scales without sacrificing trust, cost, or the planet—because the future of AI should belong to the people it servesRole SummaryMake models fast on the hardware people actually own. You optimize inference engines (llama.cpp, vLLM) for constrained local and edge environments — GPU/iGPUs, Vulkan backends — not datacenter H100 environment, mostly PC/edge. KV cache, batching, quantization, scheduling, and CPU-overhead reduction are your daily tools.This is the rare skill that makes a hybrid, low-cost agent product viable.What you’ll doProfile and optimize local inference (llama.cpp-vulkan and vLLM) for latency, throughput, and memory on edge hardwareTune KV cache, continuous batching, and scheduling for interactive agent workloadsDrive quantization strategy (GGUF / AWQ / GPTQ) and validate quality impact with the Post-Training teamCut CPU overhead and improve engine startup, model load, and lifecycle (start / stop / health)Benchmark across hardware tiers and publish honest performance comparisonsUpstream fixes and patches to open-source engines where it helps usWhat you’ll learn / grow intoCuriosity is required. You will develop:The internals of modern inference engines and where the milliseconds actually goHardware-aware optimization across iGPU / CPU paths (Vulkan, SYCL, oneAPI, CUDA where relevant)The quality-vs-speed-vs-memory trade space for small modelsInterest in local / edge AI and squeezing hardwareIMPORTANT:Please be informed that Intel is proactively tryingto find candidates for this position which is frequently availableat Intel.Please note that the position may not be availableat this time. If you would be interested in this position should itbecome available, we would encourage you to apply, and ourhiring team will be glad to contact you when/if relevant. Qualifications:QualificationsMinimum qualifications are required to be initially considered for this position. Preferred qualifications are in addition to the minimum requirements and are considered a plus factor in identifying top candidates.You must possess the minimum qualifications to be initially considered for this position. Preferred qualifications are in addition to the minimum requirements and are considered a plus factor in identifying top candidates.Required QualificationsBS/MS in CS, EE, Math or related STEM field8+ years software development backgroundStrong in C++ and/or Python; comfortable reading systems-level codeExperience with LLM inference. (attention, KV cache, decoding)Experience profiling and optimizing real performance problems (CPU or GPU) and can prove the speedupLinux, build systems, and low-level debugging expertisePreferred QualificationsHands-on with llama.cpp, vLLM, ggml, or similar enginesExperience with GPU / accelerator programming (Vulkan, CUDA, SYCL, Metal) or SIMD / CPU kernelsFamiliarity with quantization formats and their quality trade-offsOpen-source contributions to inference enginesRequirements listed would be obtained through a combination of industry relevant job experience, internship experiences and or schoolwork/classes/research.Benefits at IntelOur total rewards package goes above and beyond just a paycheck. Whether you're looking to build your career, improve your health, or protect your wealth, we offer generous benefits to help you achieve your goals. Go to Intel Benefits | Intel Careers for details of benefits available to you. Intel reserves the right to modify, change or discontinue benefit plans at any time in its sole discretion.Job Type:Shift:Shift 1 (United States of America)Primary Location: US, California, Santa ClaraAdditional Locations:US, Arizona, Phoenix, US, California, Folsom, US, Oregon, HillsboroPosting Statement:All qualified applicants will receive consideration for employment without regard to race, color, religion, religious creed, sex, national origin, ancestry, age, physical or mental disability, medical condition, genetic information, military and veteran status, marital status, pregnancy, gender, gender expression, gender identity, sexual orientation, or any other characteristic protected by local law, regulation, or ordinance.Position of TrustN/ABenefitsWe offer a total compensation package that ranks among the best in the industry. It consists of competitive pay, stock bonuses, and benefit programs which include health, retirement, and vacation. Find out more about the benefits of working at Intel. Annual Salary Range for jobs which could be performed in the US: $195,200.00-361,200.00 USDThe range displayed on this job posting reflects the minimum and maximum target compensation for the position across all US locations. Within the range, individual pay is determined by work location and additional factors, including job-related skills, experience, and relevant education or training. Your recruiter can share more about the specific compensation range for your preferred location during the hiring process.Work Model for this RoleThis role will be eligible for our hybrid work model which allows employees to split their time between working on-site at their assigned Intel site and off-site. * Job posting details (such as work model, location or time type) are subject to change.*ADDITIONAL INFORMATION: Intel is committed to Responsible Business Alliance (RBA) compliance and ethical hiring practices. We do not charge any fees during our hiring process. Candidates should never be required to pay recruitment fees, medical examination fees, or any other charges as a condition of employment. If you are asked to pay any fees during our hiring process, please report this immediately to your recruiter.SummaryLocation: US, California, Santa Clara; US, Oregon, Hillsboro; US, California, Folsom; US, Arizona, PhoenixType: Full time
$170.5k - $315.49k
## Inference Optimization Engineer (local / edge runtime)Applylocations: US, California, Santa Clara: US, Oregon, Hillsboro: US, California, Folsom: US, Arizona, Phoenixtime type: Full timeposted on: Posted Yesterdayjob requisition id: JR0284871# **Job Details:**## Job...Local areaInternshipImmediate startShift work$182.5k - $260.5k
...spread across offices in Santa Clara, St. Louis,... ...Scientist, you own the inference and optimization layer that makes AI... ...hardware, and build the runtime that executes... ...agentic AI.Cutting-edge, unusual stack. The... ...systems and backend engineers to ship capabilities...SeniorPart timeWork at office$184k - $287.5k
...globally. We seek a Senior Engineer to lead technical... ...agent frameworks and local runtimes on Windows and NVIDIA... ...powerful local inference (Nemotron models) with... ...engineering efforts to optimize the agent runtimes for... ...SummaryLocation: US, CA, Santa Clara; US, WA, RedmondType:...Local areaSeniorFull timePart timeShift work$195.2k - $361.2k
...execution stack targeting edge and robotic... ...of the firmware, runtime, and performance infrastructure... ..., develop, and optimize firmware and... ...by establishing engineering practices, driving... ...:US, California, Santa Clara, US, Texas,... ...characteristic protected by local law, regulation,...Local areaSeniorFull timePart timeWork at officeImmediate startShift work$184k - $287.5k
...for a Senior DL Algorithms Engineer! NVIDIA is seeking senior engineers... ...of performance analysis and optimization to help us squeeze every... ...and multimodal model inference as part of NVIDIA Inference... ...law.SummaryLocation: US, CA, Santa Clara; US, CA, RemoteType: Full time...SeniorFull timePart time- ...ROLE:We are looking for a Senior GPU Inference Performance Engineer to own end-to-end performance analysis... ...GPU silicon through the software runtime, and drive competitive positioning against... ...framework performance: Profile and optimize inference engines including vLLM, SGLang...SeniorPart time
- ...Hybrid, working onsite at our Santa Clara, CA, headquarters 3 days per... ...week.The role: Senior toStaff Runtime Systems EngineerWhat You Will... ...on in-memory compute for AI inference in datacenters.This position is for runtime software engineering, working on the architecture,...SeniorPart time3 days per week
- ...working on-site at our Santa Clara, CA, headquarters 3... ...per week.The role: Sr. Staff, ML... ...that will be used to optimize large language model inference on DNN accelerators... ...researchers, and ML engineers who create and apply... ...to the most cutting-edge and high-impact research...SeniorPart time3 days per week
$195.2k - $361.2k
...combines the best of local and cloud... ...machine (AI PC, edge, on-prem, and... ...Machine Learning Engineer / Data... ...tasks.Debug and optimize training runs —... ...interact with runtime constraints on... ...US, California, Santa ClaraAdditional... ...California, Santa Clara; US, Oregon, Hillsboro...Local areaSeniorFull timePart timeInternshipImmediate startShift work$224k - $356.5k
...performance across various inference frameworks.... ...includes choosing GPUs, optimizing costs, reducing... ..., you will lead the engineering team within NVIDIA’s... ...tool for datacenter, local, and edge use cases. This span... ...SummaryLocation: US, CA, Santa Clara; US, TX, Austin; US,...Local areaFull timePart timeRemote workWorldwide- ...play a pivotal role in optimizing and developing deep... ...SOTA LLM and Multimodal inference at scale across multi-... ...and optimize cutting-edge compiler technologies... ...THE PERSON: Skilled engineer with strong technical... ...Integrate and optimize runtime execution through graph...SeniorPart time
$224k - $356.5k
...world.NVIDIA's Local AI team is... ...efficiency on NVIDIA edge AI hardware.... ...-source LLM inference frameworks — identify... ...paths, and optimization... ...Science, Computer Engineering, Electrical Engineering... ...optimization, runtime configuration,... ...: US, CA, Santa Clara; US, MA, Westford...Local areaSeniorFull timePart time$100k
...leading the industry on cutting-edge AI technology,... ...seniorities.As a Software Engineer on the Metal Runtime team at Tenstorrent, you’ll... ...accelerators. You’ll build and optimize high-performance runtime systems... ...is hybrid, based out of Santa Clara, CA; Austin, TX; Toronto,...Permanent employmentPart time$100k
...leading the industry on cutting-edge AI technology,... ...We are looking for a talented engineer to join our CPU design team to... ..., based out of Austin, TX or Santa Clara, CA.We welcome candidates at... ...Use innovative techniques to optimize power, performance, and area...SeniorPermanent employmentPart time- ...sits at the leading edge of what’s possible with LLM inference on heterogeneous... ...patterns to deep optimization of inference kernels... ...research and engineering team that moves fast... ...build the tools, runtimes, and frameworks that... ..., and benefits in Santa Clara, CA.Equal Opportunity...Part time
$100k
...is leading the industry on cutting-edge AI technology, revolutionizing performance... ....We are seeking an Senior Engineer to develop and optimize the software stack for our IP customers... ...out of Toronto, ON, Austin, TX, or Santa Clara, CA.We welcome candidates at various...SeniorPermanent employmentPart time$164.47k - $269.1k
...Analog Circuit Design Engineer, you will be at the forefront... ...developing cutting-edge analog circuits in... ...play a critical role in optimizing performance, power,... ...Phoenix, US, California, Santa Clara, US, Texas,... ...characteristic protected by local law, regulation, or ordinance...Local areaFull timePart timeInternshipImmediate startShift work$256.05k - $361.48k
...Physical Design Integration Engineer performs physical design implementation... ...lowpower synthesizable CPU. Optimizes CPU design to improve... ...FolsomAdditional Locations:US, California, Santa Clara, US, Texas, AustinBusiness... ...characteristic protected by local law, regulation, or ordinance...Local areaSeniorFull timePart timeWork experience placementImmediate startFlexible hoursShift work$184k - $230k
...customers to eliminate security blind spots, optimize network traffic, and dramatically... ...and educational organizations.As a Sr. Staff Hardware Engineer, you will drive the design and... ...under applicable federal, state, and/or local law. For more information, please refer...Local areaSeniorContract workPart timeWorldwide$168k - $258.75k
...product efforts for local AI on Linux and... ...and the edge. Developers want... ...data on-prem. Inference stacks like vLLM... ...becoming the default runtime for these... ...functional teams — engineering, DevRel,... ...Experience with model optimization — quantization,... ...: US, CA, Santa ClaraType: Full...Local areaSeniorFull timePart timeShift work$184k - $287.5k
...and low-level hardware optimization has never been more... ...in both training and inference pipelines.Collaborate... ...exploratory tools and runtime systems to profile... ...Computer Science, Computer Engineering, Electrical... ...SummaryLocation: US, CA, Santa Clara; US, TX, Austin; US,...SeniorFull timePart time$184k - $287.5k
...highly skilled and motivated software engineers to join us and build AI inference systems that serve large-scale... ...high-performance inference stacks, optimize GPU kernels and compilers, drive industry... ...protected by law.SummaryLocation: US, CA, Santa ClaraType: Full time...SeniorFull timePart time$134.39k - $201.3k
...mentoring junior verification engineers Education Bachelor's degree... ...processes across the team Build and optimize verification infrastructure... ...tapeout success Use leading edge AI tools to develop... ...to employment.#LI-SA1SummaryLocation: Santa Clara, CAType: Full time...SeniorPermanent employmentFull timePart timeInternshipWork from homeShift workNight shift$184k - $287.5k
...it. The TensorRT inference platform is the backbone... ...of cutting-edge deep learning... ...skilled and driven Engineering Manager to take the... ...LLM inference runtime. Your work will be... ...development, runtime optimizations, and frameworks... ...: US, CA, Santa Clara; US, CA, RemoteType...Full timePart time$173.66k - $330.34k
...analytics, and cloud-to-edge technology is at the... ...a Senior EDA Software Engineer for the Silicon Chassis... ...Performance, and Area (PPA) optimization loops by building... ...Location: US, California, Santa ClaraAdditional Locations... ...protected by local law, regulation, or ordinance...Local areaSeniorFull timePart timeInternshipImmediate startShift workNight shift$141.91k - $269.1k
...Intel's mission to engineer world-changing technology... ...efficiency and optimize power and... ...Location: US, California, Santa ClaraAdditional Locations... ...delivering cutting-edge silicon process and... ...protected by local law, regulation, or... ...California, Santa Clara; US, Oregon, Hillsboro...Local areaFull timePart timeImmediate startShift work$173.66k - $284.58k
...technologies and cutting-edge products to the... ...EDA Tools Software Engineer, you will play a... ...requirements and implement optimal solutions.Automate... ..., US, California, Santa ClaraBusiness group... ...protected by local law, regulation, or... ...California, Santa Clara; US, Arizona, PhoenixType...Local areaFull timePart timeImmediate startShift work$195.2k - $275.58k
...Software Development Engineer to contribute... ...and optimization of oneDNN, a complex... ...PyTorch, ONNX Runtime, and many others... ...engineering, and cutting‑edge Intel hardware,... ...deep‑learning inference and training... ..., California, Santa ClaraBusiness... ...protected by local law, regulation...Local areaSeniorFull timePart timeImmediate startRemote workWorldwideFlexible hoursShift work$272k - $431.25k
...NVIDIA is the engine of modern AI, and robotics is where AI... ...simulation, training, and on-robot inference on Jetson and edge platforms.Deliver. Drive... ...(across both numerical optimization and sampling-based... ....SummaryLocation: US, CA, Santa Clara; US, TX, Austin; US, VA, Charlottesville...SeniorFull timePart time$100k
...Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing... ...house designs to architect, enable, and optimize the firmware that powers next-generation... ...systems, and software teams to solve frontier engineering challenges.How to help define the...SeniorPermanent employmentPart time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Sr. Inference Optimization Engineer (local / edge runtime) (Santa Clara). Be the first to apply!
- senior lighting artist Santa Clara, CA
- senior hvac project manager Santa Clara, CA
- senior technical product manager Santa Clara, CA
- senior medical science liaison Santa Clara, CA
- senior accountant remote Santa Clara, CA
- senior app developer Santa Clara, CA
- senior marketing account manager Santa Clara, CA
- senior robotics software engineer Santa Clara, CA
- sr project manager Santa Clara, CA
- senior compensation manager Santa Clara, CA









