Sr. Inference Optimization Engineer (local / edge runtime)
PVH (Tommy Hilfiger/Calvin Klein)
Job Details: Job Description Our MissionAt Intel, our journey is to transform AI into something safer, more trustworthy, and respectful of human privacy by design. We believe transformative AI should have a positive impact on people-powerful in capability, yet honest about its limits and protective of the data and resources it touches.To get there, we build agentic AI that combines the best of local and cloud intelligence - private, affordable, and sustainable by design. Small, efficient models run directly on the user's machine (AI PC, edge, on-prem, and beyond), keeping data private and token costs low, while powerful cloud models handle the hardest work: planning, reasoning, and complex problem-solving. Today, neither approach can deliver this alone. Together, they give people real capability without compromise-data stays private, spend stays predictable, and energy use stays in check.We're building intelligence that scales without sacrificing trust, cost, or the planet-because the future of AI should belong to the people it serves Role Summary Make models fast on the hardware people actually own. You optimize inference engines (llama.cpp, vLLM) for constrained local and edge environments - GPU/iGPUs, Vulkan backends - not datacenter H100 environment, mostly PC/edge. KV cache, batching, quantization, scheduling, and CPU-overhead reduction are your daily tools.This is the rare skill that makes a hybrid, low-cost agent product viable. What you’ll do Profile and optimize local inference (llama.cpp-vulkan and vLLM) for latency, throughput, and memory on edge hardware Tune KV cache, continuous batching, and scheduling for interactive agent workloads Drive quantization strategy (GGUF / AWQ / GPTQ) and validate quality impact with the Post-Training team Cut CPU overhead and improve engine startup, model load, and lifecycle (start / stop / health) Benchmark across hardware tiers and publish honest performance comparisons Upstream fixes and patches to open-source engines where it helps us What you’ll learn / grow into The internals of modern inference engines and where the milliseconds actually go Hardware‑aware optimization across iGPU / CPU paths (Vulkan, SYCL, oneAPI, CUDA where relevant) The quality-vs-speed-vs-memory trade space for small models Interest in local / edge AI and squeezing hardware Qualifications Minimum qualifications are required to be initially considered for this position. Preferred qualifications are in addition to the minimum requirements and are considered a plus factor in identifying top candidates.You must possess the minimum qualifications to be initially considered for this position. Preferred qualifications are in addition to the minimum requirements and are considered a plus factor in identifying top candidates. Required Qualifications BS/MS in CS, EE, Math or related STEM field 8+ years software development background Strong in C++ and/or Python; comfortable reading systems‑level code Experience with LLM inference. (attention, KV cache, decoding) Experience profiling and optimizing real performance problems (CPU or GPU) and can prove the speedup Linux, build systems, and low-level debugging expertise Preferred Qualifications Hands‑on with llama.cpp, vLLM, ggml, or similar engines Experience with GPU / accelerator programming (Vulkan, CUDA, SYCL, Metal) or SIMD / CPU kernels Familiarity with quantization formats and their quality trade‑offs Open‑source contributions to inference engines Benefits at Intel Our total rewards package goes above and beyond just a paycheck. Whether you're looking to build your career, improve your health, or protect your wealth, we offer generous benefits to help you achieve your goals. Go to Intel Benefits | Intel Careers for details of benefits available to you. Intel reserves the right to modify, change or discontinue benefit plans at any time in its sole discretion. Job Type Shift:Shift 1 (United States of America) Primary Location US, California, Santa Clara Additional Locations US, Arizona, Phoenix, US, California, Folsom, US, Oregon, Hillsboro Posting Statement All qualified applicants will receive consideration for employment without regard to race, color, religion, religious creed, sex, national origin, ancestry, age, physical or mental disability, medical condition, genetic information, military and veteran #J-18808-Ljbffr PVH (Tommy Hilfiger/Calvin Klein)
$195.2k - $361.2k
...agentic AI that combines the best of local and cloud intelligence — private, affordable... ...on the user's machine (AI PC, edge, on-prem, and beyond), keeping data private... ...the hardware people actually own. You optimize inference engines (llama.cpp, vLLM) for constrained...Local areaSeniorFull timeInternshipImmediate startShift work- NVIDIA seeks a Senior Systems Software Engineer to tackle client-side AI challenges on Windows and Linux PCs... ...align AI strategies for RTX and DGX ecosystems, while optimizing AI models, data pipelines, and inference runtimes for performance on next-generation GPUs. The role...Local areaSenior
$152k - $241.5k
...Developer Technology Engineer, you will be at... ...workflows at the edge powered by NVIDIAs... ...enterprise ISVs on solving local end-to-end agentic... ...in suboptimal runtime performance.... ...deployment targeting optimal runtime... ...GPU-accelerated AI inference driven by NVIDIA APIs...Local areaSeniorFull time- ...seeking outstanding AI systems engineers to advance the inference software stack. You will design and optimize kernels, build new... ...contribute to accelerators and runtimes that power large language models... ...CUDA C/C++, Triton, and cutting-edge MLIR-based tooling, and participate...Senior
- NVIDIA seeks a Senior Software Engineer for Local AI on Windows and Linux PCs. You will partner with software, research,... ...strategies for RTX and DGX PCs, and work on end-to-end optimization of inference runtimes and edge deployment. The role requires a strong background...Local areaSenior
$184k - $287.5k
We're now looking for a Sr. Inference Engineer, for GPU Kernel Optimization! What does it take to push every LLM inference operation to its performance ceiling... ...across kernel execution, compiler decisions, and runtime scheduling.Direct experience with LLM inference frameworks...SeniorFull time$140k - $215k
...This is a Software Development Engineer (SDE) role in the engineering... ...agent) on various container optimized Linux distros. This role will... ...and developing container runtime engines, software that monitors... ...membership or activity in a local human rights commission, status...Local areaSeniorFull timeWork experience placementWork at officeRemote work- Intel in Santa Clara, CA, seeks an experienced software engineer to optimize local inference for edge devices. You will work on llama.cpp, vLLM, and quantization, improving latency, memory usage, and startup times across hardware tiers. You will profile performance, collaborate...Local areaSenior
- Intel is seeking a seasoned software engineer to accelerate AI inference on edge hardware. You will optimize llama.cpp/vLLM, tune KV cache, batching and scheduling, and push quantization strategies to balance speed and quality. This role focuses on low-latency, privacy-...Senior
- ...the Senior Robotics Engineer role at Knightscope... ...combines robotics, edge AI, and cloud services... ...2–based perception, localization, planning, and cloud... ...TensorFlow Lite / ONNX Runtime‑based inference, anomaly detection,... ...DDS tuning and QoS optimization. Experience with...Local areaSeniorInternship
$182.5k - $260.5k
...One platform, its Zero Trust Engine, and the powerful NewEdge... ...Learning Scientist, you own the inference and optimization layer that makes AI in... ...real hardware, and build the runtime that executes bounded AI... ...economics of agentic AI.Cutting-edge, unusual stack. The hard,...Senior$229.9k - $262.4k
...Overview Sr. Lead AI Engineer (Inference Optimization, FM hosting, AI Platform) Overview: At Capital One, we are creating responsible and reliable... ...in compliance with applicable federal, state, and local laws. Capital One promotes a drug-free workplace. Capital...Local areaSeniorFull timePart time- NVIDIA is seeking a Senior DL Algorithms Engineer to optimize LLM/Omni models and enhance performance across its software stack. The ideal... ...and 3+ years of experience in deep learning, specifically in inference. This role involves profiling, analyzing bottlenecks, and...Senior
- ...looking for a specialized Sr. Staff or Principal level engineer who is passionate about... ...on scaling training and inference for the latest Generative... ...model training and inference optimizations across a variety of... ...best in the field.Cutting-edge Technology: Work with state...Senior
$152k - $241.5k
...eager to work on cutting-edge AI technology for... ...as a Senior Software Engineer, and be at the forefront... ...enabling high-performance AI inference solutions for... ...TensorRT's compiler and runtime for specialized and constrained... ...to performance optimization and benchmarking...SeniorFull time$184k - $287.5k
...globally. We seek a Senior Engineer to lead technical efforts in... ...advanced AI agent frameworks and local runtimes on Windows and NVIDIA... ...By combining powerful local inference (Nemotron models) with strong... ...the engineering efforts to optimize the agent runtimes for Windows...Local areaSeniorFull timeShift work- ...ROLE:We are looking for a Senior GPU Inference Performance Engineer to own end-to-end performance analysis... ...GPU silicon through the software runtime, and drive competitive positioning against... ...framework performance: Profile and optimize inference engines including vLLM, SGLang...Senior
- ...industry-leading training and inference speeds; over 10 times faster... ...global enterprises, and cutting-edge AI-native startups. OpenAI... ...hiring a Senior Performance Engineer to join our Product team. You... ...TensorRT-LLM), GPU kernel-level optimization toolchains (CUDA, Triton),...SeniorContract workShift work
$193.3k - $261.5k
...accelerators. Join us to optimize the latest models to run... ...Trainium hardware.As a Sr. Software Development Engineer on the Inference Model Enablement team, you... ...with the compiler and runtime teams* Mentoring junior... ...all federal, state, and local laws and Company policies...Local areaSeniorInternshipFlexible hours$183k - $247.6k
...software, hardware, and network engineers, supply chain specialists,... ...creates compute, accelerator, edge and storage server designs... ...building software services for optimizing SSD performance and health at... ...follow all federal, state, and local laws and Company policies....Local areaSeniorFlexible hours$160k - $198k
...DoAs a Senior AI Systems Engineer, you will architect,... ...AI model training and inference. You will ensure our machine... ...—to streamline and optimize the AI development lifecycle... ...productionize cutting-edge models, establish... ...protected by federal, state or local laws.Archer is...Local areaSenior$170.6k - $261.3k
...Senior Machine Learning Engineer on the State Estimation and... ..., robustness under edge cases) and drive systematic... ...Implement efficient training and inference pipelines, including model optimization techniques (e.g., pruning... ...with federal, state and local laws. We encourage...Local areaSeniorFull timeRemote workWork from homeRelocation packageFlexible hours$166k - $244k
...Role:We are looking for a Software Engineer, Edge Systems & Runtime to join our team, focusing on high-performance... ...GPUs.Your primary focus will be optimizing our production C++ runtime... ...routing multi-sensor inputs into model inference loops.Collaborate with machine learning...Full time$124k - $195.5k
NVIDIA is recruiting a Senior Inference Performance Engineer to push NVIDIA's performance limits on large-scale AI inference benchmarks. This position... ...provides an outstanding opportunity to employ your optimization knowledge in an autonomous optimization framework. AI...Full time$153.6k - $234.1k
.... As a Senior Systems Engineer, you will be at the heart... ...subsystem design, optimize performance, and contribute... ...and safety of cutting-edge autonomous technologies... ...middleware, orchestration, and runtime; low-level SoC/MCU... ...federal, state and local laws. We encourage interested...Local areaSeniorFull timeWork experience placementWork from homeRelocation packageFlexible hours- ...Packard Enterprise is the global edge-to-cloud company advancing... ...high‑impact systems, mentors engineering teams, and helps define long‑... .... Experience designing and optimizing distributed query engines and... ...report the incident to your local authorities immediately.SummaryLocation...Local areaFull timeWork experience placementWork at officeImmediate start2 days per week
- ...ROLEWe are hiring a senior engineering leader to define,... ...processing dataflow and runtime, distributed compute/... ...pre/post-processing, inference, tracking, streaming,... ...Heterogeneous compute optimization across CPU/GPU/NPU (kernel... ...SDKs for embedded/edge markets: robotics, industrial...Senior
$193.3k - $261.5k
The AWS BIOS Engineering team creates and maintains custom firmware solutions... ...to maintain its competitive edge in cloud computing.The ideal... ...-level expertise to find optimal solutions to complex boot-... ...follow all federal, state, and local laws and Company policies. Criminal...Local areaSeniorInternshipWorldwideFlexible hours$184.5k - $249.6k
...is seeking a 'Principal FAE - Edge AI' to support strategic... ...relationships with architects, engineering leaders, and key decision-makers... ..., integration, and optimization to meet power, performance, scalability... ...we can offer is limited by local legal, regulatory, tax, or other...Local areaWork at office- ...advance your career. THE ROLEWe are seeking a Principal GenAI Inference Optimization Engineer to join our Models and Applications team. This role... ...comfortable working across multiple layers—from kernels and runtimes to frameworks and serving systems—and can independently...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Sr. Inference Optimization Engineer (local / edge runtime). Be the first to apply!
- senior technology project manager Santa Clara, CA
- senior c++ developer Santa Clara, CA
- remote senior business analyst Santa Clara, CA
- senior manager clinical operations Santa Clara, CA
- senior supervisor Santa Clara, CA
- senior researcher Santa Clara, CA
- senior leadership Santa Clara, CA
- sr. process development engineer Santa Clara, CA
- senior manager data science Santa Clara, CA
- sr field service engineer Santa Clara, CA


