Senior Edge Inference Optimization Engineer
Intel Corporation
Intel Corporation is seeking a software engineer to make models fast on hardware people own, optimizing inference engines for edge environments. You will work with llama.cpp and vLLM, tuning KV cache, batching, and quantization while reducing CPU overhead and startup costs. You’ll benchmark across hardware tiers and contribute patches to open‑source engines, aligning with Intel’s mission to improve AI safety, privacy, and efficiency on local devices. #J-18808-Ljbffr Intel Corporation
$195.2k - $361.2k
...efficient models run directly on the user's machine (AI PC, edge, on-prem, and beyond), keeping data private and token costs... ...models fast on the hardware people actually own. You optimize inference engines (llama.cpp, vLLM) for constrained local and edge environments...SeniorFull timeInternshipLocal areaImmediate startShift work- ...efficient models run directly on the user's machine (AI PC, edge, on-prem, and beyond), keeping data private and token costs... ...Make models fast on the hardware people actually own. You optimize inference engines (llama.cpp, vLLM) for constrained local and edge environments...SeniorLocal areaShift work
$184k - $287.5k
We're now looking for a Sr. Inference Engineer, for GPU Kernel Optimization! What does it take to push every LLM inference operation to its performance ceiling? Our LLM Inference Performance Analysis and Optimization team builds the answer from the ground up. We develop...SeniorFull time- Intel is seeking a seasoned software engineer to accelerate AI inference on edge hardware. You will optimize llama.cpp/vLLM, tune KV cache, batching and scheduling, and push quantization strategies to balance speed and quality. This role focuses on low-latency, privacy-...Senior
- NVIDIA is seeking a Senior DL Algorithms Engineer to optimize LLM/Omni models and enhance performance across its software stack. The ideal candidate will... ...years of experience in deep learning, specifically in inference. This role involves profiling, analyzing bottlenecks,...Senior
- NVIDIA seeks a Senior Systems Software Engineer to tackle client-side AI challenges on Windows and Linux PCs with limited resources... ...strategies for RTX and DGX ecosystems, while optimizing AI models, data pipelines, and inference runtimes for performance on next-generation...SeniorLocal area
$193.3k - $261.5k
We are looking for a Senior Inference Engineer to own inference for real-time multimodalconversational... ...before they arelocked in• Implement and optimize the inference path for large-scale... ...backends (NVIDIA GPU, AWS Neuron/Trainium, edge accelerators) and how architecture...SeniorInternshipLocal areaFlexible hours- ...industry-leading training and inference speeds; over 10 times... ...enterprises, and cutting-edge AI-native startups. OpenAI... ...About The RoleWe are hiring a Senior Performance Engineer to join our Product team.... ...-LLM), GPU kernel-level optimization toolchains (CUDA, Triton),...SeniorContract workShift work
$152k - $241.5k
...world.As a Developer Technology Engineer, you will be at the forefront... ...agentic AI workflows at the edge powered by NVIDIAs RTX and... ...agentic AI deployment targeting optimal runtime performance.Improve... ...Experience with GPU-accelerated AI inference driven by NVIDIA APIs and...SeniorFull timeLocal area$182.5k - $260.5k
...One platform, its Zero Trust Engine, and the powerful NewEdge... ....Positions are available at Senior Staff and above. Candidates... ...Learning Scientist, you own the inference and optimization layer that makes AI in... ...economics of agentic AI.Cutting-edge, unusual stack. The hard, interesting...Senior$152k - $241.5k
...and eager to work on cutting-edge AI technology for safety-... ...NVIDIA's TensorRT team as a Senior Software Engineer, and be at the forefront of... ...high-performance AI inference solutions for automotive safety... ...requirementsContribute to performance optimization and benchmarking efforts...SeniorFull time- NVIDIA is seeking a Senior Agentic AI Software Engineer (Finance) to build agentic systems and advance AI inference workloads in real-world finance contexts... ...that drives experimental agents, optimize performance, and contribute to cutting-edge AI research at scale. This...Senior
- ...Santa Clara, CA is seeking outstanding AI systems engineers to advance the inference software stack. You will design and optimize kernels, build new abstractions for LLM... ...teams, work on CUDA C/C++, Triton, and cutting-edge MLIR-based tooling, and participate in open source...Senior
- Intel in Santa Clara, CA, seeks an experienced software engineer to optimize local inference for edge devices. You will work on llama.cpp, vLLM, and quantization, improving latency, memory usage, and startup times across hardware tiers. You will profile performance, collaborate...SeniorLocal area
- NVIDIA Corporation is seeking a Senior Software Engineer for the TensorRT Edge-LLM team in the US. You will develop a high-performance inference framework in modern C++ that extends TensorRT... ...across CUDA and robotics teams, optimize transformer components, and contribute...Senior
$184k - $287.5k
We are now looking for a Senior DL Algorithms Engineer! NVIDIA is seeking senior engineers who are mindful of performance analysis and optimization to help us squeeze every last clock cycle out... ...language and multimodal model inference as part of NVIDIA Inference Microservices...SeniorFull time- ...we advance your career. THE ROLE:We are looking for a Senior GPU Inference Performance Engineer to own end-to-end performance analysis of GPU-accelerated... ....AI serving framework performance: Profile and optimize inference engines including vLLM, SGLang, and emerging...Senior
- CoreWeave is seeking a Senior Engineer for its Benchmarking & Performance team to own kernel-level optimization for LLM inference and end-to-end model serving, focusing on CUDA kernels and throughput/latency improvements. You will lead kernel design reviews, mentor engineers...Senior
$184k - $287.5k
We are seeking highly skilled and motivated software engineers to join us and build AI inference systems that serve large-scale models with extreme efficiency... ...and implement high-performance inference stacks, optimize GPU kernels and compilers, drive industry benchmarks,...SeniorFull time$184k - $356.5k
...leading technology company in California is seeking a Senior DL Algorithms Engineer to drive inference performance for Deep Learning workloads. The role... ...inference and collaborating with co-design teams to optimize performance across hardware and software interfaces....Senior$166k - $244k
Google is looking for a Senior Software Engineer in Sunnyvale, CA to lead GPU performance optimizations for cutting-edge AI and machine learning technologies. This role offers the opportunity to work on innovative projects that impact billions of users around the globe...Senior- NVIDIA Corporation in Santa Clara, CA seeks a Senior Software Engineer to advance Deep Learning Inference within TensorRT. You will build scalable inferencing software... ...DL inference, leveraging C++, Python, and cutting-edge NVIDIA GPUs. This role offers impact, collaboration...Senior
- ...Staff or Principal level engineer who is passionate... ...scaling training and inference for the latest Generative... ...Exciting Opportunities: As a Senior member on the team,... ...and inference optimizations across a variety of applications... ...in the field.Cutting-edge Technology: Work with...Senior
- NVIDIA is seeking a Senior Agentic AI Software Engineer to advance agentic AI systems and workloads from scalable... ...build agentic components, analyze inference dynamics, and collaborate with teams... ...agents is essential, with a strong edge in inference or agent architectures...Senior
$184k - $287.5k
...best work. Come join the team and see how you can make a lasting impact on the worldWe are looking for an experienced Compiler Optimization Engineer for an exciting role in our Compute Compiler Team. We deliver features and improvements to CUDA and other compute compilers...SeniorFull time- Applied Intuition's Axion Data team builds the data engine to train perception models, facilitating edge and cloud runs and rapid model iteration. The engineer will optimize pipelines, scale data labels, and implement MLOps tooling to provide visibility to stakeholders....Senior
- ...strategy, consulting, and customer experience with agile engineering and problem-solving creativity. United by our core... ...their customers truly value.OverviewWe’re looking for a Senior Adobe Journey Optimizer engineer who can take a business flow from a Miro board...Senior
- Lemurian Labs is building a hardware-agnostic AI compiler stack and seeks a Graph Optimization Compiler Engineer to own the middle tier, transforming high-level graphs into efficient code. You will influence fusion, layout, and IR design that directly impacts performance...Senior
- ...seeking a highly skilled individual to develop and optimize containerized inference execution for their cutting-edge AI models. The role involves collaborating... ...advanced degrees in Computer Science or Electrical Engineering, coupled with extensive experience in AI systems...Senior
$138k - $206k
...hardware and software engineers to identify and address... ..., prototype, and optimize next-generation AI systems... ...hands-on with cutting-edge accelerator hardware,... ...workloads.We are seeking a Senior LLM Systems... ...reasoning, disaggregated inference, and Mixture-of-Experts...SeniorWork experience placementWork at officeFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Edge Inference Optimization Engineer. Be the first to apply!
- senior lead project manager Santa Clara, CA
- senior robotics software engineer Santa Clara, CA
- senior devops engineer remote Santa Clara, CA
- senior IT manager Santa Clara, CA
- sr project manager Santa Clara, CA
- senior windows systems engineer Santa Clara, CA
- senior researcher Santa Clara, CA
- senior manager data science Santa Clara, CA
- senior principal engineer Santa Clara, CA
- senior application developer Santa Clara, CA
