Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Sr. Inference Optimization Engineer (local / edge runtime)

$195.2k - $361.2k

Intel

Job Details:Job Description: Job DescriptionOur MissionAt Intel, our journey is to transform AI into something safer, more trustworthy, and respectful of human privacy by design. We believe transformative AI should have a positive impact on people—powerful in capability, yet honest about its limits and protective of the data and resources it touches.To get there, we build agentic AI that combines the best of local and cloud intelligence — private, affordable, and sustainable by design. Small, efficient models run directly on the user's machine (AI PC, edge, on-prem, and beyond), keeping data private and token costs low, while powerful cloud models handle the hardest work: planning, reasoning, and complex problem-solving. Today, neither approach can deliver this alone. Together, they give people real capability without compromise—data stays private, spend stays predictable, and energy use stays in check.We're building intelligence that scales without sacrificing trust, cost, or the planet—because the future of AI should belong to the people it servesRole SummaryMake models fast on the hardware people actually own. You optimize inference engines (llama.cpp, vLLM) for constrained local and edge environments — GPU/iGPUs, Vulkan backends — not datacenter H100 environment, mostly PC/edge. KV cache, batching, quantization, scheduling, and CPU-overhead reduction are your daily tools.This is the rare skill that makes a hybrid, low-cost agent product viable.What you’ll doProfile and optimize local inference (llama.cpp-vulkan and vLLM) for latency, throughput, and memory on edge hardwareTune KV cache, continuous batching, and scheduling for interactive agent workloadsDrive quantization strategy (GGUF / AWQ / GPTQ) and validate quality impact with the Post-Training teamCut CPU overhead and improve engine startup, model load, and lifecycle (start / stop / health)Benchmark across hardware tiers and publish honest performance comparisonsUpstream fixes and patches to open-source engines where it helps usWhat you’ll learn / grow intoCuriosity is required. You will develop:The internals of modern inference engines and where the milliseconds actually goHardware-aware optimization across iGPU / CPU paths (Vulkan, SYCL, oneAPI, CUDA where relevant)The quality-vs-speed-vs-memory trade space for small modelsInterest in local / edge AI and squeezing hardwareIMPORTANT:Please be informed that Intel is proactively tryingto find candidates for this position which is frequently availableat Intel.Please note that the position may not be availableat this time. If you would be interested in this position should itbecome available, we would encourage you to apply, and ourhiring team will be glad to contact you when/if relevant. Qualifications:QualificationsMinimum qualifications are required to be initially considered for this position. Preferred qualifications are in addition to the minimum requirements and are considered a plus factor in identifying top candidates.You must possess the minimum qualifications to be initially considered for this position. Preferred qualifications are in addition to the minimum requirements and are considered a plus factor in identifying top candidates.Required QualificationsBS/MS in CS, EE, Math or related STEM field8+ years software development backgroundStrong in C++ and/or Python; comfortable reading systems-level codeExperience with LLM inference. (attention, KV cache, decoding)Experience profiling and optimizing real performance problems (CPU or GPU) and can prove the speedupLinux, build systems, and low-level debugging expertisePreferred QualificationsHands-on with llama.cpp, vLLM, ggml, or similar enginesExperience with GPU / accelerator programming (Vulkan, CUDA, SYCL, Metal) or SIMD / CPU kernelsFamiliarity with quantization formats and their quality trade-offsOpen-source contributions to inference enginesRequirements listed would be obtained through a combination of industry relevant job experience, internship experiences and or schoolwork/classes/research.Benefits at IntelOur total rewards package goes above and beyond just a paycheck. Whether you're looking to build your career, improve your health, or protect your wealth, we offer generous benefits to help you achieve your goals. Go to Intel Benefits | Intel Careers for details of benefits available to you. Intel reserves the right to modify, change or discontinue benefit plans at any time in its sole discretion.Job Type:Shift:Shift 1 (United States of America)Primary Location: US, California, Santa ClaraAdditional Locations:US, Arizona, Phoenix, US, California, Folsom, US, Oregon, HillsboroPosting Statement:All qualified applicants will receive consideration for employment without regard to race, color, religion, religious creed, sex, national origin, ancestry, age, physical or mental disability, medical condition, genetic information, military and veteran status, marital status, pregnancy, gender, gender expression, gender identity, sexual orientation, or any other characteristic protected by local law, regulation, or ordinance.Position of TrustN/ABenefitsWe offer a total compensation package that ranks among the best in the industry. It consists of competitive pay, stock bonuses, and benefit programs which include health, retirement, and vacation. Find out more about the benefits of working at Intel. Annual Salary Range for jobs which could be performed in the US: $195,200.00-361,200.00 USDThe range displayed on this job posting reflects the minimum and maximum target compensation for the position across all US locations. Within the range, individual pay is determined by work location and additional factors, including job-related skills, experience, and relevant education or training. Your recruiter can share more about the specific compensation range for your preferred location during the hiring process.Work Model for this RoleThis role will be eligible for our hybrid work model which allows employees to split their time between working on-site at their assigned Intel site and off-site. * Job posting details (such as work model, location or time type) are subject to change.*ADDITIONAL INFORMATION: Intel is committed to Responsible Business Alliance (RBA) compliance and ethical hiring practices. We do not charge any fees during our hiring process. Candidates should never be required to pay recruitment fees, medical examination fees, or any other charges as a condition of employment. If you are asked to pay any fees during our hiring process, please report this immediately to your recruiter.SummaryLocation: US, California, Santa Clara; US, Oregon, Hillsboro; US, California, Folsom; US, Arizona, PhoenixType: Full time

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Sr. Inference Optimization Engineer (local / edge runtime) in Santa Clara, CA vacancy
  • $152k - $241.5k

     ...Developer Technology Engineer, you will be at...  ...workflows at the edge powered by NVIDIAs...  ...enterprise ISVs on solving local end-to-end agentic...  ...in suboptimal runtime performance....  ...deployment targeting optimal runtime...  ...GPU-accelerated AI inference driven by NVIDIA APIs... 
    Local area
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $184k - $287.5k

    We're now looking for a Sr. Inference Engineer, for GPU Kernel Optimization! What does it take to push every LLM inference operation to its performance ceiling...  ...across kernel execution, compiler decisions, and runtime scheduling.Direct experience with LLM inference frameworks... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    7 hours ago
  • $140k - $215k

     ...This is a Software Development Engineer (SDE) role in the engineering...  ...agent) on various container optimized Linux distros. This role will...  ...and developing container runtime engines, software that monitors...  ...membership or activity in a local human rights commission, status... 
    Local area
    Senior
    Full time
    Work experience placement
    Work at office
    Remote work

    CrowdStrike

    Sunnyvale, CA
    4 days ago
  • $224k - $356.5k

     ...are looking for a Senior Software Engineer to build and optimize the local runtimes and agent frameworks that bring autonomous...  ...combining high-performance local inference (Nemotron models) with robust...  ...constraints and advantages for edge AI.Experience building virtualization... 
    Local area
    Senior
    Full time
    Worldwide

    Nvidia

    Santa Clara, CA
    3 days ago
  •  ...the Senior Robotics Engineer role at Knightscope...  ...combines robotics, edge AI, and cloud services...  ...2–based perception, localization, planning, and cloud...  ...TensorFlow Lite / ONNX Runtime‑based inference, anomaly detection,...  ...DDS tuning and QoS optimization. Experience with... 
    Local area
    Senior
    Internship

    Knightscope

    Sunnyvale, CA
    3 days ago
  • $182.5k - $260.5k

     ...One platform, its Zero Trust Engine, and the powerful NewEdge...  ...Learning Scientist, you own the inference and optimization layer that makes AI in...  ...real hardware, and build the runtime that executes bounded AI...  ...economics of agentic AI.Cutting-edge, unusual stack. The hard,... 
    Senior

    Netskope

    Santa Clara, CA
    7 hours ago
  •  ...looking for a specialized Sr. Staff or Principal level engineer who is passionate about...  ...on scaling training and inference for the latest Generative...  ...model training and inference optimizations across a variety of...  ...best in the field.Cutting-edge Technology: Work with state... 
    Senior

    AMD

    San Jose, CA
    4 days ago
  • $152k - $241.5k

     ...eager to work on cutting-edge AI technology for...  ...as a Senior Software Engineer, and be at the forefront...  ...enabling high-performance AI inference solutions for...  ...TensorRT's compiler and runtime for specialized and constrained...  ...to performance optimization and benchmarking... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

     ...globally. We seek a Senior Engineer to lead technical efforts in...  ...advanced AI agent frameworks and local runtimes on Windows and NVIDIA...  ...By combining powerful local inference (Nemotron models) with strong...  ...the engineering efforts to optimize the agent runtimes for Windows... 
    Local area
    Senior
    Full time
    Shift work

    Nvidia

    Santa Clara, CA
    3 days ago
  •  ...industry-leading training and inference speeds; over 10 times faster...  ...global enterprises, and cutting-edge AI-native startups. OpenAI...  ...hiring a Senior Performance Engineer to join our Product team. You...  ...TensorRT-LLM), GPU kernel-level optimization toolchains (CUDA, Triton),... 
    Senior
    Contract work
    Shift work

    Cerebras Systems

    Sunnyvale, CA
    4 days ago
  •  ...ROLE:We are looking for a Senior GPU Inference Performance Engineer to own end-to-end performance analysis...  ...GPU silicon through the software runtime, and drive competitive positioning against...  ...framework performance: Profile and optimize inference engines including vLLM, SGLang... 
    Senior

    AMD

    Santa Clara, CA
    4 days ago
  • $183k - $247.6k

     ...software, hardware, and network engineers, supply chain specialists,...  ...creates compute, accelerator, edge and storage server designs...  ...building software services for optimizing SSD performance and health at...  ...follow all federal, state, and local laws and Company policies.... 
    Local area
    Senior
    Flexible hours

    AmazonWebServices

    Cupertino, CA
    4 days ago
  • $170.6k - $261.3k

     ...Senior Machine Learning Engineer on the State Estimation and...  ..., robustness under edge cases) and drive systematic...  ...Implement efficient training and inference pipelines, including model optimization techniques (e.g., pruning...  ...with federal, state and local laws. We encourage... 
    Local area
    Senior
    Full time
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    3 days ago
  • $160k - $198k

     ...DoAs a Senior AI Systems Engineer, you will architect,...  ...AI model training and inference. You will ensure our machine...  ...—to streamline and optimize the AI development lifecycle...  ...productionize cutting-edge models, establish...  ...protected by federal, state or local laws.Archer is... 
    Local area
    Senior

    Archer Aviation

    San Jose, CA
    1 day ago
  • $164k - $313.3k

     ...Systems & Efficiency Engineer to join our R&D team focused...  ...ready improvements in inference performance, latency,...  ..., systems, inference runtimes, and services, with a...  ..., and performance optimization. You will work closely...  ...accordance with state and local laws and “fair chance”... 
    Local area
    Senior
    Full time
    Temporary work
    Worldwide

    Adobe Systems

    San Jose, CA
    3 days ago
  • $166k - $244k

     ...Role:We are looking for a Software Engineer, Edge Systems & Runtime to join our team, focusing on high-performance...  ...GPUs.Your primary focus will be optimizing our production C++ runtime...  ...routing multi-sensor inputs into model inference loops.Collaborate with machine learning... 
    Full time

    X Company

    Mountain View, CA
    4 days ago
  • $124k - $195.5k

    NVIDIA is recruiting a Senior Inference Performance Engineer to push NVIDIA's performance limits on large-scale AI inference benchmarks. This position...  ...provides an outstanding opportunity to employ your optimization knowledge in an autonomous optimization framework. AI... 
    Full time

    Nvidia

    Santa Clara, CA
    6 hours ago
  • $153.6k - $234.1k

     ....  As a Senior Systems Engineer, you will be at the heart...  ...subsystem design, optimize performance, and contribute...  ...and safety of cutting-edge autonomous technologies...  ...middleware, orchestration, and runtime; low-level SoC/MCU...  ...federal, state and local laws. We encourage interested... 
    Local area
    Senior
    Full time
    Work experience placement
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    2 days ago
  •  ...ROLEWe are hiring a senior engineering leader to define,...  ...processing dataflow and runtime, distributed compute/...  ...pre/post-processing, inference, tracking, streaming,...  ...Heterogeneous compute optimization across CPU/GPU/NPU (kernel...  ...SDKs for embedded/edge markets: robotics, industrial... 
    Senior

    AMD

    San Jose, CA
    3 days ago
  •  ...Packard Enterprise is the global edge-to-cloud company advancing...  ...high‑impact systems, mentors engineering teams, and helps define long‑...  .... Experience designing and optimizing distributed query engines and...  ...report the incident to your local authorities immediately.SummaryLocation... 
    Local area
    Full time
    Work experience placement
    Work at office
    Immediate start
    2 days per week

    Hewlett Packard Enterprise

    San Jose, CA
    6 hours ago
  • $193.3k - $261.5k

    The AWS BIOS Engineering team creates and maintains custom firmware solutions...  ...to maintain its competitive edge in cloud computing.The ideal...  ...-level expertise to find optimal solutions to complex boot-...  ...follow all federal, state, and local laws and Company policies. Criminal... 
    Local area
    Senior
    Internship
    Worldwide
    Flexible hours

    Amazon

    Cupertino, CA
    2 days ago
  • $140k - $180k

     ...model architectureCollaborate with Systems Engineering to properly decompose higher level...  ...generationStrong knowledge of vehicle dynamics, optimization, and both classical and modern control...  ...protected by federal, state or local laws. For this position we are targeting... 
    Local area
    Senior

    Archer Aviation

    San Jose, CA
    1 day ago
  • $229.9k - $262.4k

     ...Overview Sr. Lead AI Engineer (Inference Optimization, FM hosting, AI Platform) Overview: At Capital One, we are creating responsible and reliable...  ...in compliance with applicable federal, state, and local laws. Capital One promotes a drug-free workplace. Capital... 
    Local area
    Senior
    Full time
    Part time

    Capital One

    San Jose, CA
    more than 2 months ago
  • $184.5k - $249.6k

     ...is seeking a 'Principal FAE - Edge AI' to support strategic...  ...relationships with architects, engineering leaders, and key decision-makers...  ..., integration, and optimization to meet power, performance, scalability...  ...we can offer is limited by local legal, regulatory, tax, or other... 
    Local area
    Work at office

    ARM

    San Jose, CA
    3 days ago
  •  ...of patients worldwide. Were a team of engineers, clinicians, and innovators united by one...  ...with supplier engineers and vendors to optimize component DFM (injection molded and machined...  ...protected under federal, state, or local applicable laws. Mandatory Notices U.S... 
    Local area
    Senior
    Full time
    Worldwide
    Flexible hours
    Shift work

    Intuitive

    Sunnyvale, CA
    2 days ago
  •  ...advance your career. THE ROLEWe are seeking a Principal GenAI Inference Optimization Engineer to join our Models and Applications team. This role...  ...comfortable working across multiple layers—from kernels and runtimes to frameworks and serving systems—and can independently... 

    AMD

    San Jose, CA
    3 days ago
  •  ...patients worldwide.We’re a team of engineers, clinicians, and innovators...  ...design requirements into optimized processes, tooling, fixturing...  ...techniquesKnowledge of cutting-edge technologies used in metal...  ...protected under federal, state, or local applicable laws.Mandatory... 
    Local area
    Senior
    Worldwide
    Flexible hours

    Intuitive Surgical

    Santa Clara, CA
    7 hours ago
  • $170.6k - $261.3k

     ...contributor driving cutting-edge sensor calibration and state...  ...of all sensors to ensure they optimally support many other vehicle subsystems...  ...-functional teams, mentor engineers across the organization, and...  ...the purposes of real-time localization, static and dynamic sensor calibration... 
    Local area
    Senior
    Full time
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    1 day ago
  •  ...millions of patients worldwide.We’re a team of engineers, clinicians, and innovators united by...  ...process changes, production optimization, facility changes, new equipment qualification...  ...status protected under federal, state, or local applicable laws.Mandatory NoticesU.S. Export... 
    Local area
    Senior
    Contract work
    Worldwide
    Flexible hours

    Intuitive Surgical

    Sunnyvale, CA
    2 days ago
  • $120.8k - $193.3k

    DescriptionJob Title: Sr. Engineer, AI/ML Apps Engineering...  ..., drones, and edge AI solutions. This is...  ...architecture, AI software optimization, and technical leadership...  ...PyTorch, TensorFlow, ONNX Runtime, OpenCV, or similar....  ...understanding of AI inference optimization, model deployment... 
    Senior
    Full time
    Work at office

    SiMa Technologies

    San Jose, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Sr. Inference Optimization Engineer (local / edge runtime). Be the first to apply!