Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Cloud Inference Launch Engineer: Scale & Optimize LLM Infra

Jobleads-US

United States Digital Space LLC is seeking an engineer to join the Cloud Inference team and scale Claude across AWS, GCP, Azure, and future CSPs. You will own end-to-end inference on each cloud platform, from API integration to deployment and daily operations. You will focus on fast, cost-efficient validation, performance improvements, and reliability to ensure consistent behavior across providers and accelerate model delivery. #J-18808-Ljbffr Jobleads-US

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Cloud Inference Launch Engineer: Scale & Optimize LLM Infra in Seattle, WA vacancy
  • ByteDance is seeking a Research Engineer - LLM/VLM Inference Optimization in Seattle. The role involves designing and optimizing high-performance inference systems for large-scale LLMs and VLMs, requiring expertise in C/C++ and Python, and familiarity with GPU optimization... 
    Suggested

    ByteDance

    Seattle, WA
    16 hours ago
  • $236k - $330k

     ...and advance the state of the art in LLM inference systems and optimization.Our mission is to build the next...  ...manual tuning. We embrace AI-native engineering, using AI not only as the workload...  ...cutting-edge research into production-scale AI.ResponsibilitiesDesign and... 
    Suggested

    Snowflake

    Bellevue, WA
    3 days ago
  • $320k

     ...of committed researchers, engineers, policy experts, and...  ...About the role The Cloud Inference team scales and optimizes Claude to serve the massive...  ...Inference, the model & inference launch team owns the validation...  ...a strong interest in LLM serving; prior inference... 
    Cloud
    Visa sponsorship

    Jobleads-US

    Seattle, WA
    2 days ago
  • $184k - $287.5k

    We're now looking for a Sr. Inference Engineer, for GPU Kernel Optimization! What does it take to push every LLM inference operation to its performance ceiling? Our LLM Inference Performance Analysis and Optimization team builds the answer from the ground up. We develop... 
    Suggested
    Full time

    Nvidia

    Seattle, WA
    3 days ago
  •  ...model-serving and inference systems for high-throughput...  ...AI workloads. Optimize inference...  ...orchestration, platform engineering, and operations...  ...systems at scale. Strong understanding...  ...SGLang, TensorRT-LLM, or Triton Inference...  ..., Kubernetes, or cloud infrastructure... 
    Cloud
    Full time
    Work at office
    Relocation
    3 days per week

    Designworks Talent

    Bellevue, WA
    2 days ago
  •  ...thrive.Job DescriptionThe AI Inference Engineer plays a critical role in the...  ...model development and optimized deployment environments. This...  ...ensuring high performance at scale. Handle deployment optimizations...  ...Docker, Kubernetes, and cloud platforms such as AWS, GCP,... 
    Cloud
    Full time
    Local area
    Immediate start

    F5 Networks

    Seattle, WA
    2 days ago
  • $152.2k - $205.9k

     ...building the future of cloud financial...  ...in production at scale. Spend is driven...  ...tokens, model choice, inference patterns, and...  ...cost, usage, and optimization insights that power...  ...most consequential launches, run the customer...  ...directly with engineering, and dig into consumption... 
    Cloud
    Flexible hours

    Amazon

    Seattle, WA
    4 days ago
  •  ...Senior Lead Software Engineer at JPMorgan Chase within...  ...secure, scalable cloud platforms optimized for AI/ML workloads.Partner...  ...by automation at scale.Required...  ...architecture, ML training, and inference.Experience with Infrastructure...  ...a high-performance LLM inference platform—... 
    Cloud
    For contractors

    JP Morgan Chase

    Seattle, WA
    3 days ago
  •  ...builds foundational cloud and AI computing...  ...ByteDance and Volcano Engine. Our mission is to...  ...operate at cloud scale and help shape the...  ...world‑class inference solutions using vLLM...  ...SGLang, TensorRT-LLM, and other LLM engines...  ..., performance optimization, and open-source innovation... 
    Cloud
    Summer work
    Internship

    ByteDance

    Seattle, WA
    3 days ago
  • $152k

     ...continue our growth and launch new services at...  ...Detection Engineering (DE) team within Coupang...  ...threats at scale. The team builds and...  ...develops AI Agent and LLM-based analysis capabilities...  .... Analyze and optimize automation...  ...operating systems in cloud environments, preferably... 
    Cloud
    Temporary work
    Flexible hours

    Coupang

    Seattle, WA
    1 day ago
  • $184k - $287.5k

    NVIDIA is seeking an NCX Senior Engineer to join our DSX team,...  ...custom AI solutions on NCP and Neo Cloud platforms, including distributed training, inference optimization, and MLOps pipelines constructed...  ...builds.Profile and tune large-scale training and inference workloads... 
    Cloud
    Full time
    Remote work

    Nvidia

    Seattle, WA
    3 days ago
  • $142k - $220.5k

    Job DescriptionA Senior Engineer on the AI Enablement team designs...  ...Generative AI safely and at scale. The role owns work that...  ...integrate with developer tools, cloud services, LLM providers, and observability...  ..., and processes.Develop and optimize databases and infrastructure... 
    Cloud
    Full time
    Temporary work

    Nordstrom

    Seattle, WA
    3 days ago
  • AssemblyAI is hiring a Software Engineer to turn cutting-edge AI research into products our customers rely...  ...-documented APIs serving production traffic at scale. You’ll design and ship customer-facing APIs, scale inference infrastructure for 1M+ users, and collaborate with... 

    AssemblyAI, Inc.

    Seattle, WA
    2 days ago
  •  ...analysis and AI training and inference. Designed from the...  ...data center, edge, and cloud. About the Role As a Forward Deployed Engineer (FDE), you are a core...  ...most strategic, large-scale customer environments....  ...customer's technical team to optimize infrastructure... 
    Cloud
    Remote work

    VAST Data

    Seattle, WA
    1 day ago
  • $234.4k - $296.6k

     ...for hybrid, multi-cloud environments. Join...  ...expertise with the scale and operational...  ...and Cisco’s global engineering capabilities. Our...  ...distributed training and inference pipelines to...  ...model deployment, optimization, and evaluation.Preferred...  ...in the code-gen / LLM community.Why... 
    Cloud
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    Seattle, WA
    2 days ago
  • $148.7k - $201.2k

     ...the people who keep the cloud running. We support...  ...diverse team of data engineers, business intelligence...  ...to join our Material Optimization team. This role will drive...  ...-visibility, large-scale, innovative system solutions...  ..., and post launch improvements• Manage change... 
    Cloud
    Flexible hours

    Amazon

    Bellevue, WA
    1 day ago
  • $210.2k - $284.3k

     ...commitment, aligning engineering and business...  ...available at scale.Key job...  ...KPIs (cost per inference, time from PoC...  ...input to post-launch tooling- Convert...  ...broadly adopted cloud platform. We pioneered...  ...inference cost optimization. You have built...  ...implementing enterprise LLM proxy or... 
    Cloud
    Local area
    Flexible hours
    Shift work

    AmazonWebServices

    Seattle, WA
    3 days ago
  • $152.2k - $243.7k

     ...opportunity to create impact at scale — tackling meaningful...  ...entrepreneurs and engineers in 2016, Pismo is a...  ...in the market. Pismo’s cloud-based platform empowers firms to build and launch financial products rapidly...  ...and optimize high-performance data services... 
    Cloud
    Full time
    Work experience placement
    Work at office
    Local area

    Visa

    Bellevue, WA
    2 days ago
  • $146k - $194k

     ...Platform is the internal engineering force multiplier behind...  ...is how Anduril scales its business systems with...  ...production stability.Support launch readiness, production...  ...distributed systems, CI/CD, and cloud or platform...  ...AI coding assistants, LLM-enabled applications, or... 
    Cloud
    Full time
    Work experience placement
    Immediate start

    Anduril Industries

    Seattle, WA
    3 days ago
  • $171k - $231.4k

     ...a broad set of global cloud-based services including...  ...faster, lower IT costs, and scale. You will be a part of...  ...AWS Product Compliance Engineering team within AWS Global...  ...region planning and launches. You will develop compliance...  ...Engineering, AWS Infra Service Supply Chain, Compliance... 
    Cloud
    Local area
    Flexible hours

    AmazonWebServices

    Seattle, WA
    1 day ago
  • $148.7k - $201.2k

     ...backbone of Generative AI cloud at AWS? Do you want to...  ...for AI training and inference? Want to do industry...  ...applied to those at cloud scale? If yes, then come...  ...you. The AWS Hardware Engineering team creates server designs...  ...conception, test, launch, and operations.... 
    Cloud
    Internship
    Local area
    Flexible hours

    Amazon

    Seattle, WA
    2 days ago
  • Snowflake is seeking AI-native thinkers for the AI Research team to advance LLM inference systems and optimization. You will work on high-performance, adaptive inference across distributed serving, GPU kernels, and system co-design, collaborating with researchers and product... 

    Snowflake Computing

    Bellevue, WA
    4 days ago
  • $254k - $350k

     ...Machine Learning and System Optimization Engineer, you will orchestrate and allocate...  ...allow for more efficient inference by sharing various parts of...  ..., production-ready large-scale models to our on-vehicle stack...  ...(e.g., TensorRT-LLM).$254,000 - $350,000 a yearBase... 
    Full time
    Temporary work
    Relocation package

    Zoox

    Seattle, WA
    4 days ago
  • $209.1k - $282.9k

     ...seek a Staff Robotics Engineer to help create the next...  ..., independent of cloud services.You will provide...  ...methods, and performance optimization.Demonstrated ability...  ...with AI and perception inference pipelines deployed on...  ...services are built and scaled at Arm.In addition to... 
    Cloud
    Work at office
    Local area
    Visa sponsorship
    Relocation package

    ARM

    Seattle, WA
    4 days ago
  • $143.7k - $194.4k

     ...AWS Inferentia and Trainium cloud-scale machinelearning accelerators...  ...role is for a senior software engineer in the Machine Learning Inference Applications team. This role...  ...development and performance optimization of core building blocks of LLM Inference - Attention, MLP,... 
    Cloud
    Internship
    Flexible hours

    Amazon

    Seattle, WA
    4 days ago
  • $183k - $247.6k

     ...backbone of Generative AI cloud at AWS? Do you want to...  ...for AI training and inference? Want to do industry leading...  ...hardware, and network engineers, supply chain...  ...training and inference at scale.You will define and drive...  ...corrective actions. After launch, you will oversee the... 
    Cloud
    Local area
    Flexible hours

    AmazonWebServices

    Seattle, WA
    5 days ago
  •  ...Hardware / Machine Design Engineer Location: Hybrid |...  ...a next-generation cloud platform designed to power...  ...workloads—including large-scale compute, model training, fine-tuning, inference, and emerging agentic...  ...into highly optimized, production-ready systems... 
    Cloud
    Work at office
    Relocation
    3 days per week

    Designworks Talent LLC

    Bellevue, WA
    1 day ago
  •  ...GPU Performance / Kernel Engineer Location: Hybrid |...  ...multiple roles available Optimize the Performance Layer...  ...building a next-generation cloud platform designed to...  ...workloads—including large-scale compute, model training, fine-tuning, inference, and emerging agentic AI... 
    Cloud
    Work at office
    Relocation
    3 days per week

    Jobleads-US

    Bellevue, WA
    4 days ago
  • $162.7k - $220.2k

     ...passionate about helping customers optimize price/performance for AI Inference at scale? AWS is seeking a Go-To-Market (...  ...to prospect and develop pipeline, launch new opportunities and drive usage...  ...Experience selling enterprise software or cloud-based applications- Experience... 
    Cloud
    Local area
    Worldwide
    Flexible hours

    AmazonWebServices

    Bellevue, WA
    2 days ago
  •  ...Stripe is seeking a Staff Infrastructure Engineer to lead and evolve our compute infrastructure, enabling global scale and cost efficiency. You will collaborate with product...  ...reliability challenges in a fast-paced, cloud-centric environment. #J-18808-Ljbffr Jobleads... 
    Cloud

    Jobleads-US

    Seattle, WA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Cloud Inference Launch Engineer: Scale & Optimize LLM Infra. Be the first to apply!