Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Inference Engineer

Designworks Talent LLC

Inference Engineer Location: Hybrid | Bellevue, WA Area Titles: Senior and Staff (multiple roles available) Build the Inference Platform Powering Next-Generation AI Applications About the Opportunity A well-funded, rapidly growing AI infrastructure company is building a next-generation cloud platform designed to power the full lifecycle of artificial intelligence. The organization is developing a comprehensive AI infrastructure, platform, and services portfolio that supports the full spectrum of AI workloads—including large‑scale compute, model training, fine‑tuning, inference, and emerging agentic AI applications. Backed by significant long-term investment, the company combines the speed, ownership, and innovation of a startup with the stability and resources of an established parent organization. Engineering teams are intentionally lean, highly collaborative, and AI‑native, leveraging modern tooling and automation to build infrastructure capable of supporting the industry's most demanding AI workloads. We're seeking Inference Engineers to build and operate the model‑serving systems behind a next‑generation AI inference platform. This team focuses on delivering high‑throughput, low‑latency, reliable inference experiences that enable customers to consume advanced AI capabilities through production‑scale APIs. The Opportunity This is a foundational engineering role focused on building the systems that bring AI models from research environments into reliable production services. You'll work on the infrastructure layer responsible for serving large models efficiently, optimizing performance, and ensuring reliability as usage scales. You'll collaborate closely with GPU performance, AI training infrastructure, platform engineering, and operations teams to solve complex challenges around model serving, latency optimization, resource efficiency, and production reliability. This opportunity is ideal for engineers who enjoy working at the intersection of distributed systems, machine learning infrastructure, GPU computing, and large‑scale production systems. What You'll Do Build and operate production‑grade model‑serving and inference systems supporting high‑throughput, low‑latency AI workloads. Optimize inference infrastructure for token throughput, latency, scalability, and cost efficiency across different model architectures and workloads. Design systems that maximize GPU utilization while maintaining predictable performance and reliability. Improve the scalability and operational maturity of inference platforms as customer demand grows. Partner with AI training, GPU performance, orchestration, and infrastructure teams to ensure smooth transitions from model development to production serving. Develop monitoring, alerting, and operational practices to maintain reliable inference services. Investigate and resolve performance, reliability, and capacity challenges across inference workloads. Contribute to architecture decisions and engineering standards as the platform evolves. What We're Looking For Experience building and operating production machine learning inference or model‑serving systems at scale. Strong understanding of the performance trade‑offs involved in serving large AI models, including latency, throughput, memory utilization, and cost efficiency. Experience designing reliable distributed systems or production infrastructure. Understanding of GPU‑backed AI workloads and the challenges of scaling inference systems. Strong engineering fundamentals and the ability to independently own complex technical problems. Comfortable working in a fast‑moving environment where systems and processes are being built from the ground up. Preferred Qualifications Experience with modern inference‑serving frameworks such as vLLM, SGLang, TensorRT‑LLM, Triton Inference Server, or similar technologies. Experience optimizing LLM inference workloads or large‑scale AI serving platforms. Background operating API‑based AI products or high‑volume production services. Experience with GPU scheduling, distributed systems, Kubernetes, or cloud infrastructure platforms. Familiarity with model optimization techniques such as quantization, batching, caching, or performance tuning. Experience working at a hyperscaler, AI lab, GPU cloud provider, or large‑scale ML infrastructure organization. Compensation Competitive base pay for Bellevue market Certain roles are eligible for additional rewards, including merit increases, annual bonus, and long‑term incentives. These awards are allocated based on individual performance U.S. based employees have access to medical, dental, and vision insurance, a 401(k) plan and company match, employees also receive per calendar year, paid holidays. Location Hybrid role based in the Bellevue, WA area. Approximately three days per week in the office. Candidates elsewhere in the U.S. who are open to relocation are encouraged to apply. U.S. work authorization is required. Visa sponsorship is not currently available. Why Join? Build the inference platform powering the next generation of AI applications. Work directly on large‑scale model serving, GPU optimization, and production AI systems. Solve complex challenges around latency, throughput, reliability, and cost efficiency. Join early enough to influence architecture, tooling, and engineering practices. Collaborate with a highly experienced team building critical AI infrastructure from the ground up. Enjoy the ownership and technical impact of a startup environment backed by significant long‑term investment. #J-18808-Ljbffr Designworks Talent LLC

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Inference Engineer in Bellevue, WA vacancy
  • CoreWeave is seeking a Senior Engineer for its Benchmarking & Performance team to own kernel authoring and optimization for LLM inference. You will write and tune CUDA kernels, optimize tensor cores, and push end-to-end latency down while maintaining accuracy. You will... 
    Suggested

    CoreWeave

    Bellevue, WA
    3 days ago
  • $184k - $287.5k

    We're now looking for a Sr. Inference Engineer, for GPU Kernel Optimization! What does it take to push every LLM inference operation to its performance ceiling? Our LLM Inference Performance Analysis and Optimization team builds the answer from the ground up. We develop... 
    Suggested
    Full time

    Nvidia

    Seattle, WA
    2 days ago
  • $165k - $242k

    CoreWeave in Bellevue, WA is seeking an experienced engineer to lead design reviews and optimize distributed systems. The role requires 5-8 years in cloud services, strong skills in Python or Go, and hands-on experience with Kubernetes. Key responsibilities include defining... 
    Suggested
    Flexible hours

    CoreWeave

    Bellevue, WA
    16 hours ago
  • $92k - $135k

     ...intelligence that drives innovation.  What You’ll Do: Join the Inference team to ship production features that improve latency,...  ...practices, and grow quickly with mentorship from experienced engineers. About the role: Implement well-scoped features and... 
    Suggested
    Permanent employment
    Full time
    Temporary work
    Casual work
    Internship
    Work at office
    Remote work
    Flexible hours

    Coreweave

    Bellevue, WA
    1 day ago
  • $130.4k - $195.6k

     ...: adapting PolySpatial to stream Unity content into other game engines and 3D environments — in-process, cross-process, and over the network...  ...with MCP-style tool interfaces, LLM agents, multimodal models, inference pipelines, embeddings, vector search, or AI systems that... 
    Suggested
    Full time
    Work at office
    Remote work
    Worldwide

    Unity Technologies

    Bellevue, WA
    5 days ago
  • NVIDIA in Redmond, WA is seeking a Senior Software Engineer for Quantized Inference to accelerate the development of efficient inference recipes for LLMs. You will translate recipe specifications into high-performance code, including Triton kernels and quantize/dequantize... 

    NVIDIA

    Redmond, WA
    3 days ago
  • NVIDIA is seeking a Senior Software Engineer for Quantized Inference to accelerate the development of efficient inference recipes for LLMs. You will implement quantized and sparse recipes, work on kernel and model-level implementations, and collaborate with partner teams... 

    NVIDIA AI

    Redmond, WA
    2 days ago
  •  ...Computer Vision Research Engineer The Cloud AI team at the Huawei Research America is looking for highly motivated and qualified...  ...segmentation algorithms, knowledge graph engine and efficient inference algorithm, large scale data analysis framework, etc. We are leading... 
    Worldwide

    Netpace

    Bellevue, WA
    5 days ago
  • NVIDIA Gruppe is seeking a Senior Software Engineer for Quantized Inference in Redmond, Washington. You will be responsible for implementing quantized and sparse recipes in inference engines and enhancing developer productivity across the team. Successful candidates will... 

    NVIDIA Gruppe

    Redmond, WA
    16 hours ago
  • $81.8k - $118.6k

     ...top design firms while building our clean energy future.The OpportunityWe are seeking an Automation / Instrumentation & Controls Engineer to join our Energy and Resources consulting team. This position is intended for Engineers looking to demonstrate and grow their skills... 
    Full time
    Temporary work
    Part time
    For contractors
    Casual work
    Local area
    Remote work
    Flexible hours

    Stantec

    Bellevue, WA
    3 days ago
  • $129.2k - $174.8k

     ...an innovative, systems-oriented Computer Vision & Automation Engineer to help design and deploy next-generation intelligent automation...  ...reliable data capture, processing pipelines, and low-latency inference on industrial equipment - Own hardware-software integration, including... 
    Flexible hours

    Amazon

    Bellevue, WA
    5 days ago
  • $206.4k - $379.1k

     ...’s Generative AI Services team is seeking a Principal Service Engineer to serve as the technical lead for our GenAI Services domain....  ...generative models into Adobe’s flagship products.Design and architect inference infrastructure for enterprise-scale model customization,... 
    Full time
    Temporary work
    Local area
    Worldwide

    Adobe Systems

    Seattle, WA
    1 day ago
  •  ...further to learn how you could help make great things possible not only in your community, but around the world.We believe building engineering is more than systems and structures, it’s about powering progress and enabling innovation. As part of HDR’s Building Engineering... 
    For contractors
    Work at office

    HDR

    Bellevue, WA
    1 day ago
  • $137.3k - $185.7k

     ...., current Green Card holder), or lawfully admitted into the U.S. as a refugee or granted asylum.As a Sr. PCBA Manufacturing Test Engineer, you will be responsible for deployment, qualification, continuous improvement of test systems, processes, and procedures at Contract... 
    Permanent employment
    Contract work
    Temporary work
    Flexible hours

    Amazon

    Bellevue, WA
    2 days ago
  • $177k - $239.4k

     ...including hardware, software, supply chain, and manufacturing.This role includes direct responsibility for program management of various engineering work-streams under manufacturing, system integration, and production test systems. This is a highly technical role requiring... 
    Permanent employment
    Contract work
    Flexible hours

    Amazon

    Bellevue, WA
    2 days ago
  • $115.8k - $160k

     ...regular updates on test progress, issues, and resolutions.Basic qualifications- Bachelor's degree in Computer Science, Software Engineering, MIS, or a related field- 3 years of experience in software testing, with at least 1 years on managing test suites for ERP systems... 
    Flexible hours

    Amazon

    Bellevue, WA
    2 days ago
  • ByteDance is seeking a Research Engineer - LLM/VLM Inference Optimization in Seattle. The role involves designing and optimizing high-performance inference systems for large-scale LLMs and VLMs, requiring expertise in C/C++ and Python, and familiarity with GPU optimization... 

    ByteDance

    Seattle, WA
    16 hours ago
  • $118.3k - $207.1k

     ...good, then Jacobs is where you belong. We are looking for a passionate and dedicated Instrumentation & Control Systems (I&CS) Design Engineer. As part of our Northwest team in Bellevue, you’ll have the chance to work on projects that bring innovative solutions to our... 
    Full time
    Work at office
    Remote work
    Worldwide

    Jacobs

    Bellevue, WA
    1 day ago
  • $165k - $242k

     ...degradation, rollback/traffic-shift strategies. Mentor IC1/IC2 engineers; review cross-team designs and elevate coding/testing...  ...(Prometheus, Grafana, OpenTelemetry). Practical knowledge of inference internals: batching, caching, mixed precision (BF16/FP8), streaming... 
    Permanent employment
    Temporary work
    Work at office
    Remote work
    Flexible hours
    Shift work

    CoreWeave

    Bellevue, WA
    16 hours ago
  • $127.4k - $191.2k

     ...— across billions of monthly users on the world's leading game engine. Recommendation and ranking systems are the core of this work:...  ...time ad delivery.Design and run rigorous experiments using causal inference, A/B testing, and offline evaluation frameworks to measure and... 
    Full time
    Internship
    Work at office
    Worldwide
    Shift work

    Unity Technologies

    Bellevue, WA
    5 days ago
  • $142.3k - $192.4k

    We are seeking Controls System Development Engineers to join our In-House Controls (IHC) team within World-Wide Central Engineering. This global role requires a dynamic individual eager to travel and lead the development and deployment of advanced automation solutions across... 
    Internship
    Worldwide
    Flexible hours

    Amazon

    Bellevue, WA
    2 days ago
  • $139k - $204k

    What You’ll Do Senior engineers are area owners who lead designs, raise engineering standards, and deliver measurable improvements to...  ...orchestration, and hardware teams to evolve our Kubernetes‑native inference platform and meet strict P99 SLAs at scale. About The Role... 
    Permanent employment
    Temporary work
    Casual work
    Work at office
    Remote work
    Flexible hours
    Shift work

    CoreWeave

    Bellevue, WA
    16 hours ago
  • Job Description goes here!! US Salary Range $1 - $100,000 USD The salary range for this role is an estimate based on a wide range of compensation factors, inclusive of base salary only. Actual salary offer may vary based on (but not limited to) work experience...
    Full time
    Work experience placement
    Immediate start

    Anduril

    Bellevue, WA
    1 day ago
  • $90 - $140 per hour

    We’re ALTEN Technology USA, an engineering company helping clients bring groundbreaking ideas to life—from advancing space exploration and life-saving medical devices to building autonomous electric vehicles. With 3,000+ experts across North America, we partner with leading... 
    Permanent employment
    For contractors
    Remote work

    Alten

    Bellevue, WA
    5 days ago
  • Designworks Talent is seeking an Inference Engineer to build and operate production-grade model-serving and inference systems in a hybrid Bellevue, WA setting. You’ll work across distributed systems, GPU optimization, and production AI infrastructure to deliver high-throughput... 

    Designworks Talent LLC

    Bellevue, WA
    2 days ago
  •  ...to learn how you could help make great things possible not only in your community, but around the world. BES:We believe building engineering is more than systems and structures, it’s about powering progress and enabling innovation. As part of HDR’s Building Engineering... 
    Contract work

    HDR

    Bellevue, WA
    5 days ago
  • $94.21k - $141.31k

     ...Systems (AES) LocationRedmond, WA DescriptionAstronics Advanced Electronic Systems (AES) is seeking a Senior or Staff LevelQuality Engineer - New Product Introduction to join our cohesive and diverse team of professional problem solvers in Redmond, WA.This is a hybrid /... 
    Part time
    For contractors
    Local area
    Work from home
    Relocation package
    Flexible hours
    2 days per week

    Astronics

    Kirkland, WA
    4 days ago
  •  ...better. Read further to learn how you could help make great things possible not only in your community, but around the world.HDR Engineering is currently seeking a Civil/Senior Site Civil Engineer to join one of the largest, fastest growing, and comprehensive TMT (Tech,... 
    Local area

    HDR

    Bellevue, WA
    1 day ago
  •  ...projects that will benefit future generations. Grow with us, H2O+U.Your OpportunityStantec is seeking an experienced and versatile civil engineer with design and construction experience of gravity sewer pipes, force mains, water distribution systems, and water transmission... 
    Full time
    Temporary work
    Part time
    For contractors
    For subcontractor
    Casual work
    Local area
    Flexible hours

    Stantec

    Bellevue, WA
    2 days ago
  • $38.46 - $45.67 per hour

    A leading global real estate firm is seeking an Operating Engineer to support maintenance processes at their Bellevue, WA location. The ideal candidate will have over 3 years of experience in repair and maintenance, particularly in HVAC and plumbing systems. Responsibilities... 
    Hourly pay

    JLL

    Bellevue, WA
    5 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Inference Engineer. Be the first to apply!