Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Sr. Inference Optimization Engineer (local / edge runtime) (Santa Clara)

$195.2k - $361.2k
Part-time

Intel

Job Details:Job Description: Job DescriptionOur MissionAt Intel, our journey is to transform AI into something safer, more trustworthy, and respectful of human privacy by design. We believe transformative AI should have a positive impact on people—powerful in capability, yet honest about its limits and protective of the data and resources it touches.To get there, we build agentic AI that combines the best of local and cloud intelligence — private, affordable, and sustainable by design. Small, efficient models run directly on the user's machine (AI PC, edge, on-prem, and beyond), keeping data private and token costs low, while powerful cloud models handle the hardest work: planning, reasoning, and complex problem-solving. Today, neither approach can deliver this alone. Together, they give people real capability without compromise—data stays private, spend stays predictable, and energy use stays in check.We're building intelligence that scales without sacrificing trust, cost, or the planet—because the future of AI should belong to the people it servesRole SummaryMake models fast on the hardware people actually own. You optimize inference engines (llama.cpp, vLLM) for constrained local and edge environments — GPU/iGPUs, Vulkan backends — not datacenter H100 environment, mostly PC/edge. KV cache, batching, quantization, scheduling, and CPU-overhead reduction are your daily tools.This is the rare skill that makes a hybrid, low-cost agent product viable.What you’ll doProfile and optimize local inference (llama.cpp-vulkan and vLLM) for latency, throughput, and memory on edge hardwareTune KV cache, continuous batching, and scheduling for interactive agent workloadsDrive quantization strategy (GGUF / AWQ / GPTQ) and validate quality impact with the Post-Training teamCut CPU overhead and improve engine startup, model load, and lifecycle (start / stop / health)Benchmark across hardware tiers and publish honest performance comparisonsUpstream fixes and patches to open-source engines where it helps usWhat you’ll learn / grow intoCuriosity is required. You will develop:The internals of modern inference engines and where the milliseconds actually goHardware-aware optimization across iGPU / CPU paths (Vulkan, SYCL, oneAPI, CUDA where relevant)The quality-vs-speed-vs-memory trade space for small modelsInterest in local / edge AI and squeezing hardwareIMPORTANT:Please be informed that Intel is proactively tryingto find candidates for this position which is frequently availableat Intel.Please note that the position may not be availableat this time. If you would be interested in this position should itbecome available, we would encourage you to apply, and ourhiring team will be glad to contact you when/if relevant. Qualifications:QualificationsMinimum qualifications are required to be initially considered for this position. Preferred qualifications are in addition to the minimum requirements and are considered a plus factor in identifying top candidates.You must possess the minimum qualifications to be initially considered for this position. Preferred qualifications are in addition to the minimum requirements and are considered a plus factor in identifying top candidates.Required QualificationsBS/MS in CS, EE, Math or related STEM field8+ years software development backgroundStrong in C++ and/or Python; comfortable reading systems-level codeExperience with LLM inference. (attention, KV cache, decoding)Experience profiling and optimizing real performance problems (CPU or GPU) and can prove the speedupLinux, build systems, and low-level debugging expertisePreferred QualificationsHands-on with llama.cpp, vLLM, ggml, or similar enginesExperience with GPU / accelerator programming (Vulkan, CUDA, SYCL, Metal) or SIMD / CPU kernelsFamiliarity with quantization formats and their quality trade-offsOpen-source contributions to inference enginesRequirements listed would be obtained through a combination of industry relevant job experience, internship experiences and or schoolwork/classes/research.Benefits at IntelOur total rewards package goes above and beyond just a paycheck. Whether you're looking to build your career, improve your health, or protect your wealth, we offer generous benefits to help you achieve your goals. Go to Intel Benefits | Intel Careers for details of benefits available to you. Intel reserves the right to modify, change or discontinue benefit plans at any time in its sole discretion.Job Type:Shift:Shift 1 (United States of America)Primary Location: US, California, Santa ClaraAdditional Locations:US, Arizona, Phoenix, US, California, Folsom, US, Oregon, HillsboroPosting Statement:All qualified applicants will receive consideration for employment without regard to race, color, religion, religious creed, sex, national origin, ancestry, age, physical or mental disability, medical condition, genetic information, military and veteran status, marital status, pregnancy, gender, gender expression, gender identity, sexual orientation, or any other characteristic protected by local law, regulation, or ordinance.Position of TrustN/ABenefitsWe offer a total compensation package that ranks among the best in the industry. It consists of competitive pay, stock bonuses, and benefit programs which include health, retirement, and vacation. Find out more about the benefits of working at Intel. Annual Salary Range for jobs which could be performed in the US: $195,200.00-361,200.00 USDThe range displayed on this job posting reflects the minimum and maximum target compensation for the position across all US locations. Within the range, individual pay is determined by work location and additional factors, including job-related skills, experience, and relevant education or training. Your recruiter can share more about the specific compensation range for your preferred location during the hiring process.Work Model for this RoleThis role will be eligible for our hybrid work model which allows employees to split their time between working on-site at their assigned Intel site and off-site. * Job posting details (such as work model, location or time type) are subject to change.*ADDITIONAL INFORMATION: Intel is committed to Responsible Business Alliance (RBA) compliance and ethical hiring practices. We do not charge any fees during our hiring process. Candidates should never be required to pay recruitment fees, medical examination fees, or any other charges as a condition of employment. If you are asked to pay any fees during our hiring process, please report this immediately to your recruiter.SummaryLocation: US, California, Santa Clara; US, Oregon, Hillsboro; US, California, Folsom; US, Arizona, PhoenixType: Full time

Vacancy posted 2 hours ago
Similar jobs that could be interesting for youBased on the Sr. Inference Optimization Engineer (local / edge runtime) (Santa Clara) in Santa Clara, CA vacancy
  • $170.5k - $315.49k

    ## Inference Optimization Engineer (local / edge runtime)Applylocations: US, California, Santa Clara: US, Oregon, Hillsboro: US, California, Folsom: US, Arizona, Phoenixtime type: Full timeposted on: Posted Yesterdayjob requisition id: JR0284871# **Job Details:**## Job... 
    Local area
    Internship
    Immediate start
    Shift work

    Intel

    Santa Clara, CA
    3 days ago
  • $182.5k - $260.5k

     ...spread across offices in Santa Clara, St. Louis,...  ...Scientist, you own the inference and optimization layer that makes AI...  ...hardware, and build the runtime that executes...  ...agentic AI.Cutting-edge, unusual stack. The...  ...systems and backend engineers to ship capabilities... 
    Senior
    Part time
    Work at office

    Netskope

    Santa Clara, CA
    2 hours ago
  • $184k - $287.5k

     ...globally. We seek a Senior Engineer to lead technical...  ...agent frameworks and local runtimes on Windows and NVIDIA...  ...powerful local inference (Nemotron models) with...  ...engineering efforts to optimize the agent runtimes for...  ...SummaryLocation: US, CA, Santa Clara; US, WA, RedmondType:... 
    Local area
    Senior
    Full time
    Part time
    Shift work

    Nvidia

    Santa Clara, CA
    2 hours ago
  • $195.2k - $361.2k

     ...execution stack targeting edge and robotic...  ...of the firmware, runtime, and performance infrastructure...  ..., develop, and optimize firmware and...  ...by establishing engineering practices, driving...  ...:US, California, Santa Clara, US, Texas,...  ...characteristic protected by local law, regulation,... 
    Local area
    Senior
    Full time
    Part time
    Work at office
    Immediate start
    Shift work

    Intel

    Santa Clara, CA
    2 hours ago
  • $184k - $287.5k

     ...for a Senior DL Algorithms Engineer! NVIDIA is seeking senior engineers...  ...of performance analysis and optimization to help us squeeze every...  ...and multimodal model inference as part of NVIDIA Inference...  ...law.SummaryLocation: US, CA, Santa Clara; US, CA, RemoteType: Full time... 
    Senior
    Full time
    Part time

    Nvidia

    Santa Clara, CA
    2 hours ago
  •  ...ROLE:We are looking for a Senior GPU Inference Performance Engineer to own end-to-end performance analysis...  ...GPU silicon through the software runtime, and drive competitive positioning against...  ...framework performance: Profile and optimize inference engines including vLLM, SGLang... 
    Senior
    Part time

    AMD

    Santa Clara, CA
    2 hours ago
  •  ...Hybrid, working onsite at our Santa Clara, CA, headquarters 3 days per...  ...week.The role: Senior toStaff Runtime Systems EngineerWhat You Will...  ...on in-memory compute for AI inference in datacenters.This position is for runtime software engineering, working on the architecture,... 
    Senior
    Part time
    3 days per week

    d-Matrix

    Santa Clara, CA
    2 hours ago
  •  ...working on-site at our Santa Clara, CA, headquarters 3...  ...per week.The role: Sr. Staff, ML...  ...that will be used to optimize large language model inference on DNN accelerators...  ...researchers, and ML engineers who create and apply...  ...to the most cutting-edge and high-impact research... 
    Senior
    Part time
    3 days per week

    d-Matrix

    Santa Clara, CA
    2 hours ago
  • $195.2k - $361.2k

     ...combines the best of local and cloud...  ...machine (AI PC, edge, on-prem, and...  ...Machine Learning Engineer / Data...  ...tasks.Debug and optimize training runs —...  ...interact with runtime constraints on...  ...US, California, Santa ClaraAdditional...  ...California, Santa Clara; US, Oregon, Hillsboro... 
    Local area
    Senior
    Full time
    Part time
    Internship
    Immediate start
    Shift work

    Intel

    Santa Clara, CA
    2 hours ago
  • $224k - $356.5k

     ...performance across various inference frameworks....  ...includes choosing GPUs, optimizing costs, reducing...  ..., you will lead the engineering team within NVIDIA’s...  ...tool for datacenter, local, and edge use cases. This span...  ...SummaryLocation: US, CA, Santa Clara; US, TX, Austin; US,... 
    Local area
    Full time
    Part time
    Remote work
    Worldwide

    Nvidia

    Santa Clara, CA
    2 hours ago
  •  ...play a pivotal role in optimizing and developing deep...  ...SOTA LLM and Multimodal inference at scale across multi-...  ...and optimize cutting-edge compiler technologies...  ...THE PERSON:   Skilled engineer with strong technical...  ...Integrate and optimize runtime execution through graph... 
    Senior
    Part time

    AMD

    Santa Clara, CA
    2 hours ago
  • $224k - $356.5k

     ...world.NVIDIA's Local AI team is...  ...efficiency on NVIDIA edge AI hardware....  ...-source LLM inference frameworks — identify...  ...paths, and optimization...  ...Science, Computer Engineering, Electrical Engineering...  ...optimization, runtime configuration,...  ...: US, CA, Santa Clara; US, MA, Westford... 
    Local area
    Senior
    Full time
    Part time

    Nvidia

    Santa Clara, CA
    2 hours ago
  • $100k

     ...leading the industry on cutting-edge AI technology,...  ...seniorities.As a Software Engineer on the Metal Runtime team at Tenstorrent, you’ll...  ...accelerators. You’ll build and optimize high-performance runtime systems...  ...is hybrid, based out of Santa Clara, CA; Austin, TX; Toronto,... 
    Permanent employment
    Part time

    Tenstorrent

    Santa Clara, CA
    2 hours ago
  • $100k

     ...leading the industry on cutting-edge AI technology,...  ...We are looking for a talented engineer to join our CPU design team to...  ..., based out of Austin, TX or Santa Clara, CA.We welcome candidates at...  ...Use innovative techniques to optimize power, performance, and area... 
    Senior
    Permanent employment
    Part time

    Tenstorrent

    Santa Clara, CA
    2 hours ago
  •  ...sits at the leading edge of what’s possible with LLM inference on heterogeneous...  ...patterns to deep optimization of inference kernels...  ...research and engineering team that moves fast...  ...build the tools, runtimes, and frameworks that...  ..., and benefits in Santa Clara, CA.Equal Opportunity... 
    Part time

    d-Matrix

    Santa Clara, CA
    2 hours ago
  • $100k

     ...is leading the industry on cutting-edge AI technology, revolutionizing performance...  ....We are seeking an Senior Engineer to develop and optimize the software stack for our IP customers...  ...out of Toronto, ON, Austin, TX, or Santa Clara, CA.We welcome candidates at various... 
    Senior
    Permanent employment
    Part time

    Tenstorrent

    Santa Clara, CA
    2 hours ago
  • $164.47k - $269.1k

     ...Analog Circuit Design Engineer, you will be at the forefront...  ...developing cutting-edge analog circuits in...  ...play a critical role in optimizing performance, power,...  ...Phoenix, US, California, Santa Clara, US, Texas,...  ...characteristic protected by local law, regulation, or ordinance... 
    Local area
    Full time
    Part time
    Internship
    Immediate start
    Shift work

    Intel

    Santa Clara, CA
    2 hours ago
  • $256.05k - $361.48k

     ...Physical Design Integration Engineer performs physical design implementation...  ...lowpower synthesizable CPU. Optimizes CPU design to improve...  ...FolsomAdditional Locations:US, California, Santa Clara, US, Texas, AustinBusiness...  ...characteristic protected by local law, regulation, or ordinance... 
    Local area
    Senior
    Full time
    Part time
    Work experience placement
    Immediate start
    Flexible hours
    Shift work

    Intel

    Santa Clara, CA
    7 hours ago
  • $184k - $230k

     ...customers to eliminate security blind spots, optimize network traffic, and dramatically...  ...and educational organizations.As a Sr. Staff Hardware Engineer, you will drive the design and...  ...under applicable federal, state, and/or local law. For more information, please refer... 
    Local area
    Senior
    Contract work
    Part time
    Worldwide

    Gigamon

    Santa Clara, CA
    2 hours ago
  • $168k - $258.75k

     ...product efforts for local AI on Linux and...  ...and the edge. Developers want...  ...data on-prem. Inference stacks like vLLM...  ...becoming the default runtime for these...  ...functional teams — engineering, DevRel,...  ...Experience with model optimization — quantization,...  ...: US, CA, Santa ClaraType: Full... 
    Local area
    Senior
    Full time
    Part time
    Shift work

    Nvidia

    Santa Clara, CA
    2 hours ago
  • $184k - $287.5k

     ...and low-level hardware optimization has never been more...  ...in both training and inference pipelines.Collaborate...  ...exploratory tools and runtime systems to profile...  ...Computer Science, Computer Engineering, Electrical...  ...SummaryLocation: US, CA, Santa Clara; US, TX, Austin; US,... 
    Senior
    Full time
    Part time

    Nvidia

    Santa Clara, CA
    2 hours ago
  • $184k - $287.5k

     ...highly skilled and motivated software engineers to join us and build AI inference systems that serve large-scale...  ...high-performance inference stacks, optimize GPU kernels and compilers, drive industry...  ...protected by law.SummaryLocation: US, CA, Santa ClaraType: Full time... 
    Senior
    Full time
    Part time

    Nvidia

    Santa Clara, CA
    2 hours ago
  • $134.39k - $201.3k

     ...mentoring junior verification engineers Education Bachelor's degree...  ...processes across the team Build and optimize verification infrastructure...  ...tapeout success Use leading edge AI tools to develop...  ...to employment.#LI-SA1SummaryLocation: Santa Clara, CAType: Full time... 
    Senior
    Permanent employment
    Full time
    Part time
    Internship
    Work from home
    Shift work
    Night shift

    Marvell

    Santa Clara, CA
    2 hours ago
  • $184k - $287.5k

     ...it. The TensorRT inference platform is the backbone...  ...of cutting-edge deep learning...  ...skilled and driven Engineering Manager to take the...  ...LLM inference runtime. Your work will be...  ...development, runtime optimizations, and frameworks...  ...: US, CA, Santa Clara; US, CA, RemoteType... 
    Full time
    Part time

    Nvidia

    Santa Clara, CA
    2 hours ago
  • $173.66k - $330.34k

     ...analytics, and cloud-to-edge technology is at the...  ...a Senior EDA Software Engineer for the Silicon Chassis...  ...Performance, and Area (PPA) optimization loops by building...  ...Location: US, California, Santa ClaraAdditional Locations...  ...protected by local law, regulation, or ordinance... 
    Local area
    Senior
    Full time
    Part time
    Internship
    Immediate start
    Shift work
    Night shift

    Intel

    Santa Clara, CA
    2 hours ago
  • $141.91k - $269.1k

     ...Intel's mission to engineer world-changing technology...  ...efficiency and optimize power and...  ...Location: US, California, Santa ClaraAdditional Locations...  ...delivering cutting-edge silicon process and...  ...protected by local law, regulation, or...  ...California, Santa Clara; US, Oregon, Hillsboro... 
    Local area
    Full time
    Part time
    Immediate start
    Shift work

    Intel

    Santa Clara, CA
    2 hours ago
  • $173.66k - $284.58k

     ...technologies and cutting-edge products to the...  ...EDA Tools Software Engineer, you will play a...  ...requirements and implement optimal solutions.Automate...  ..., US, California, Santa ClaraBusiness group...  ...protected by local law, regulation, or...  ...California, Santa Clara; US, Arizona, PhoenixType... 
    Local area
    Full time
    Part time
    Immediate start
    Shift work

    Intel

    Santa Clara, CA
    2 hours ago
  • $195.2k - $275.58k

     ...Software Development Engineer to contribute...  ...and optimization of oneDNN, a complex...  ...PyTorch, ONNX Runtime, and many others...  ...engineering, and cutting‑edge Intel hardware,...  ...deep‑learning inference and training...  ..., California, Santa ClaraBusiness...  ...protected by local law, regulation... 
    Local area
    Senior
    Full time
    Part time
    Immediate start
    Remote work
    Worldwide
    Flexible hours
    Shift work

    Intel

    Santa Clara, CA
    2 hours ago
  • $272k - $431.25k

     ...NVIDIA is the engine of modern AI, and robotics is where AI...  ...simulation, training, and on-robot inference on Jetson and edge platforms.Deliver. Drive...  ...(across both numerical optimization and sampling-based...  ....SummaryLocation: US, CA, Santa Clara; US, TX, Austin; US, VA, Charlottesville... 
    Senior
    Full time
    Part time

    Nvidia

    Santa Clara, CA
    7 hours ago
  • $100k

     ...Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing...  ...house designs to architect, enable, and optimize the firmware that powers next-generation...  ...systems, and software teams to solve frontier engineering challenges.How to help define the... 
    Senior
    Permanent employment
    Part time

    Tenstorrent

    Santa Clara, CA
    2 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Sr. Inference Optimization Engineer (local / edge runtime) (Santa Clara). Be the first to apply!