Inference Engineer
Designworks Talent LLC
Inference Engineer Location: Hybrid | Bellevue, WA Area Titles: Senior and Staff (multiple roles available) Build the Inference Platform Powering Next-Generation AI Applications About the Opportunity A well-funded, rapidly growing AI infrastructure company is building a next-generation cloud platform designed to power the full lifecycle of artificial intelligence. The organization is developing a comprehensive AI infrastructure, platform, and services portfolio that supports the full spectrum of AI workloads—including large‑scale compute, model training, fine‑tuning, inference, and emerging agentic AI applications. Backed by significant long-term investment, the company combines the speed, ownership, and innovation of a startup with the stability and resources of an established parent organization. Engineering teams are intentionally lean, highly collaborative, and AI‑native, leveraging modern tooling and automation to build infrastructure capable of supporting the industry's most demanding AI workloads. We're seeking Inference Engineers to build and operate the model‑serving systems behind a next‑generation AI inference platform. This team focuses on delivering high‑throughput, low‑latency, reliable inference experiences that enable customers to consume advanced AI capabilities through production‑scale APIs. The Opportunity This is a foundational engineering role focused on building the systems that bring AI models from research environments into reliable production services. You'll work on the infrastructure layer responsible for serving large models efficiently, optimizing performance, and ensuring reliability as usage scales. You'll collaborate closely with GPU performance, AI training infrastructure, platform engineering, and operations teams to solve complex challenges around model serving, latency optimization, resource efficiency, and production reliability. This opportunity is ideal for engineers who enjoy working at the intersection of distributed systems, machine learning infrastructure, GPU computing, and large‑scale production systems. What You'll Do Build and operate production‑grade model‑serving and inference systems supporting high‑throughput, low‑latency AI workloads. Optimize inference infrastructure for token throughput, latency, scalability, and cost efficiency across different model architectures and workloads. Design systems that maximize GPU utilization while maintaining predictable performance and reliability. Improve the scalability and operational maturity of inference platforms as customer demand grows. Partner with AI training, GPU performance, orchestration, and infrastructure teams to ensure smooth transitions from model development to production serving. Develop monitoring, alerting, and operational practices to maintain reliable inference services. Investigate and resolve performance, reliability, and capacity challenges across inference workloads. Contribute to architecture decisions and engineering standards as the platform evolves. What We're Looking For Experience building and operating production machine learning inference or model‑serving systems at scale. Strong understanding of the performance trade‑offs involved in serving large AI models, including latency, throughput, memory utilization, and cost efficiency. Experience designing reliable distributed systems or production infrastructure. Understanding of GPU‑backed AI workloads and the challenges of scaling inference systems. Strong engineering fundamentals and the ability to independently own complex technical problems. Comfortable working in a fast‑moving environment where systems and processes are being built from the ground up. Preferred Qualifications Experience with modern inference‑serving frameworks such as vLLM, SGLang, TensorRT‑LLM, Triton Inference Server, or similar technologies. Experience optimizing LLM inference workloads or large‑scale AI serving platforms. Background operating API‑based AI products or high‑volume production services. Experience with GPU scheduling, distributed systems, Kubernetes, or cloud infrastructure platforms. Familiarity with model optimization techniques such as quantization, batching, caching, or performance tuning. Experience working at a hyperscaler, AI lab, GPU cloud provider, or large‑scale ML infrastructure organization. Compensation Competitive base pay for Bellevue market Certain roles are eligible for additional rewards, including merit increases, annual bonus, and long‑term incentives. These awards are allocated based on individual performance U.S. based employees have access to medical, dental, and vision insurance, a 401(k) plan and company match, employees also receive per calendar year, paid holidays. Location Hybrid role based in the Bellevue, WA area. Approximately three days per week in the office. Candidates elsewhere in the U.S. who are open to relocation are encouraged to apply. U.S. work authorization is required. Visa sponsorship is not currently available. Why Join? Build the inference platform powering the next generation of AI applications. Work directly on large‑scale model serving, GPU optimization, and production AI systems. Solve complex challenges around latency, throughput, reliability, and cost efficiency. Join early enough to influence architecture, tooling, and engineering practices. Collaborate with a highly experienced team building critical AI infrastructure from the ground up. Enjoy the ownership and technical impact of a startup environment backed by significant long‑term investment. #J-18808-Ljbffr Designworks Talent LLC
- CoreWeave is seeking a Senior Engineer for its Benchmarking & Performance team to own kernel authoring and optimization for LLM inference. You will write and tune CUDA kernels, optimize tensor cores, and push end-to-end latency down while maintaining accuracy. You will...Suggested
$184k - $287.5k
We're now looking for a Sr. Inference Engineer, for GPU Kernel Optimization! What does it take to push every LLM inference operation to its performance ceiling? Our LLM Inference Performance Analysis and Optimization team builds the answer from the ground up. We develop...SuggestedFull time$165k - $242k
CoreWeave in Bellevue, WA is seeking an experienced engineer to lead design reviews and optimize distributed systems. The role requires 5-8 years in cloud services, strong skills in Python or Go, and hands-on experience with Kubernetes. Key responsibilities include defining...SuggestedFlexible hours$92k - $135k
...intelligence that drives innovation. What You’ll Do: Join the Inference team to ship production features that improve latency,... ...practices, and grow quickly with mentorship from experienced engineers. About the role: Implement well-scoped features and...SuggestedPermanent employmentFull timeTemporary workCasual workInternshipWork at officeRemote workFlexible hours$130.4k - $195.6k
...: adapting PolySpatial to stream Unity content into other game engines and 3D environments — in-process, cross-process, and over the network... ...with MCP-style tool interfaces, LLM agents, multimodal models, inference pipelines, embeddings, vector search, or AI systems that...SuggestedFull timeWork at officeRemote workWorldwide- NVIDIA in Redmond, WA is seeking a Senior Software Engineer for Quantized Inference to accelerate the development of efficient inference recipes for LLMs. You will translate recipe specifications into high-performance code, including Triton kernels and quantize/dequantize...
- NVIDIA is seeking a Senior Software Engineer for Quantized Inference to accelerate the development of efficient inference recipes for LLMs. You will implement quantized and sparse recipes, work on kernel and model-level implementations, and collaborate with partner teams...
- ...Computer Vision Research Engineer The Cloud AI team at the Huawei Research America is looking for highly motivated and qualified... ...segmentation algorithms, knowledge graph engine and efficient inference algorithm, large scale data analysis framework, etc. We are leading...Worldwide
- NVIDIA Gruppe is seeking a Senior Software Engineer for Quantized Inference in Redmond, Washington. You will be responsible for implementing quantized and sparse recipes in inference engines and enhancing developer productivity across the team. Successful candidates will...
$81.8k - $118.6k
...top design firms while building our clean energy future.The OpportunityWe are seeking an Automation / Instrumentation & Controls Engineer to join our Energy and Resources consulting team. This position is intended for Engineers looking to demonstrate and grow their skills...Full timeTemporary workPart timeFor contractorsCasual workLocal areaRemote workFlexible hours$129.2k - $174.8k
...an innovative, systems-oriented Computer Vision & Automation Engineer to help design and deploy next-generation intelligent automation... ...reliable data capture, processing pipelines, and low-latency inference on industrial equipment - Own hardware-software integration, including...Flexible hours$206.4k - $379.1k
...’s Generative AI Services team is seeking a Principal Service Engineer to serve as the technical lead for our GenAI Services domain.... ...generative models into Adobe’s flagship products.Design and architect inference infrastructure for enterprise-scale model customization,...Full timeTemporary workLocal areaWorldwide- ...further to learn how you could help make great things possible not only in your community, but around the world.We believe building engineering is more than systems and structures, it’s about powering progress and enabling innovation. As part of HDR’s Building Engineering...For contractorsWork at office
$137.3k - $185.7k
...., current Green Card holder), or lawfully admitted into the U.S. as a refugee or granted asylum.As a Sr. PCBA Manufacturing Test Engineer, you will be responsible for deployment, qualification, continuous improvement of test systems, processes, and procedures at Contract...Permanent employmentContract workTemporary workFlexible hours$177k - $239.4k
...including hardware, software, supply chain, and manufacturing.This role includes direct responsibility for program management of various engineering work-streams under manufacturing, system integration, and production test systems. This is a highly technical role requiring...Permanent employmentContract workFlexible hours$115.8k - $160k
...regular updates on test progress, issues, and resolutions.Basic qualifications- Bachelor's degree in Computer Science, Software Engineering, MIS, or a related field- 3 years of experience in software testing, with at least 1 years on managing test suites for ERP systems...Flexible hours- ByteDance is seeking a Research Engineer - LLM/VLM Inference Optimization in Seattle. The role involves designing and optimizing high-performance inference systems for large-scale LLMs and VLMs, requiring expertise in C/C++ and Python, and familiarity with GPU optimization...
$118.3k - $207.1k
...good, then Jacobs is where you belong. We are looking for a passionate and dedicated Instrumentation & Control Systems (I&CS) Design Engineer. As part of our Northwest team in Bellevue, you’ll have the chance to work on projects that bring innovative solutions to our...Full timeWork at officeRemote workWorldwide$165k - $242k
...degradation, rollback/traffic-shift strategies. Mentor IC1/IC2 engineers; review cross-team designs and elevate coding/testing... ...(Prometheus, Grafana, OpenTelemetry). Practical knowledge of inference internals: batching, caching, mixed precision (BF16/FP8), streaming...Permanent employmentTemporary workWork at officeRemote workFlexible hoursShift work$127.4k - $191.2k
...— across billions of monthly users on the world's leading game engine. Recommendation and ranking systems are the core of this work:... ...time ad delivery.Design and run rigorous experiments using causal inference, A/B testing, and offline evaluation frameworks to measure and...Full timeInternshipWork at officeWorldwideShift work$142.3k - $192.4k
We are seeking Controls System Development Engineers to join our In-House Controls (IHC) team within World-Wide Central Engineering. This global role requires a dynamic individual eager to travel and lead the development and deployment of advanced automation solutions across...InternshipWorldwideFlexible hours$139k - $204k
What You’ll Do Senior engineers are area owners who lead designs, raise engineering standards, and deliver measurable improvements to... ...orchestration, and hardware teams to evolve our Kubernetes‑native inference platform and meet strict P99 SLAs at scale. About The Role...Permanent employmentTemporary workCasual workWork at officeRemote workFlexible hoursShift work- Job Description goes here!! US Salary Range $1 - $100,000 USD The salary range for this role is an estimate based on a wide range of compensation factors, inclusive of base salary only. Actual salary offer may vary based on (but not limited to) work experience...Full timeWork experience placementImmediate start
$90 - $140 per hour
We’re ALTEN Technology USA, an engineering company helping clients bring groundbreaking ideas to life—from advancing space exploration and life-saving medical devices to building autonomous electric vehicles. With 3,000+ experts across North America, we partner with leading...Permanent employmentFor contractorsRemote work- Designworks Talent is seeking an Inference Engineer to build and operate production-grade model-serving and inference systems in a hybrid Bellevue, WA setting. You’ll work across distributed systems, GPU optimization, and production AI infrastructure to deliver high-throughput...
- ...to learn how you could help make great things possible not only in your community, but around the world. BES:We believe building engineering is more than systems and structures, it’s about powering progress and enabling innovation. As part of HDR’s Building Engineering...Contract work
$94.21k - $141.31k
...Systems (AES) LocationRedmond, WA DescriptionAstronics Advanced Electronic Systems (AES) is seeking a Senior or Staff LevelQuality Engineer - New Product Introduction to join our cohesive and diverse team of professional problem solvers in Redmond, WA.This is a hybrid /...Part timeFor contractorsLocal areaWork from homeRelocation packageFlexible hours2 days per week- ...better. Read further to learn how you could help make great things possible not only in your community, but around the world.HDR Engineering is currently seeking a Civil/Senior Site Civil Engineer to join one of the largest, fastest growing, and comprehensive TMT (Tech,...Local area
- ...projects that will benefit future generations. Grow with us, H2O+U.Your OpportunityStantec is seeking an experienced and versatile civil engineer with design and construction experience of gravity sewer pipes, force mains, water distribution systems, and water transmission...Full timeTemporary workPart timeFor contractorsFor subcontractorCasual workLocal areaFlexible hours
$38.46 - $45.67 per hour
A leading global real estate firm is seeking an Operating Engineer to support maintenance processes at their Bellevue, WA location. The ideal candidate will have over 3 years of experience in repair and maintenance, particularly in HVAC and plumbing systems. Responsibilities...Hourly pay
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Inference Engineer. Be the first to apply!


