Cloud Inference Launch Engineer: Scale & Optimize LLM Infra
Jobleads-US
United States Digital Space LLC is seeking an engineer to join the Cloud Inference team and scale Claude across AWS, GCP, Azure, and future CSPs. You will own end-to-end inference on each cloud platform, from API integration to deployment and daily operations. You will focus on fast, cost-efficient validation, performance improvements, and reliability to ensure consistent behavior across providers and accelerate model delivery. #J-18808-Ljbffr Jobleads-US
- ByteDance is seeking a Research Engineer - LLM/VLM Inference Optimization in Seattle. The role involves designing and optimizing high-performance inference systems for large-scale LLMs and VLMs, requiring expertise in C/C++ and Python, and familiarity with GPU optimization...Suggested
$236k - $330k
...and advance the state of the art in LLM inference systems and optimization.Our mission is to build the next... ...manual tuning. We embrace AI-native engineering, using AI not only as the workload... ...cutting-edge research into production-scale AI.ResponsibilitiesDesign and...Suggested$320k
...of committed researchers, engineers, policy experts, and... ...About the role The Cloud Inference team scales and optimizes Claude to serve the massive... ...Inference, the model & inference launch team owns the validation... ...a strong interest in LLM serving; prior inference...CloudVisa sponsorship$184k - $287.5k
We're now looking for a Sr. Inference Engineer, for GPU Kernel Optimization! What does it take to push every LLM inference operation to its performance ceiling? Our LLM Inference Performance Analysis and Optimization team builds the answer from the ground up. We develop...SuggestedFull time- ...model-serving and inference systems for high-throughput... ...AI workloads. Optimize inference... ...orchestration, platform engineering, and operations... ...systems at scale. Strong understanding... ...SGLang, TensorRT-LLM, or Triton Inference... ..., Kubernetes, or cloud infrastructure...CloudFull timeWork at officeRelocation3 days per week
- ...thrive.Job DescriptionThe AI Inference Engineer plays a critical role in the... ...model development and optimized deployment environments. This... ...ensuring high performance at scale. Handle deployment optimizations... ...Docker, Kubernetes, and cloud platforms such as AWS, GCP,...CloudFull timeLocal areaImmediate start
$152.2k - $205.9k
...building the future of cloud financial... ...in production at scale. Spend is driven... ...tokens, model choice, inference patterns, and... ...cost, usage, and optimization insights that power... ...most consequential launches, run the customer... ...directly with engineering, and dig into consumption...CloudFlexible hours- ...Senior Lead Software Engineer at JPMorgan Chase within... ...secure, scalable cloud platforms optimized for AI/ML workloads.Partner... ...by automation at scale.Required... ...architecture, ML training, and inference.Experience with Infrastructure... ...a high-performance LLM inference platform—...CloudFor contractors
- ...builds foundational cloud and AI computing... ...ByteDance and Volcano Engine. Our mission is to... ...operate at cloud scale and help shape the... ...world‑class inference solutions using vLLM... ...SGLang, TensorRT-LLM, and other LLM engines... ..., performance optimization, and open-source innovation...CloudSummer workInternship
$152k
...continue our growth and launch new services at... ...Detection Engineering (DE) team within Coupang... ...threats at scale. The team builds and... ...develops AI Agent and LLM-based analysis capabilities... .... Analyze and optimize automation... ...operating systems in cloud environments, preferably...CloudTemporary workFlexible hours$184k - $287.5k
NVIDIA is seeking an NCX Senior Engineer to join our DSX team,... ...custom AI solutions on NCP and Neo Cloud platforms, including distributed training, inference optimization, and MLOps pipelines constructed... ...builds.Profile and tune large-scale training and inference workloads...CloudFull timeRemote work$142k - $220.5k
Job DescriptionA Senior Engineer on the AI Enablement team designs... ...Generative AI safely and at scale. The role owns work that... ...integrate with developer tools, cloud services, LLM providers, and observability... ..., and processes.Develop and optimize databases and infrastructure...CloudFull timeTemporary work- AssemblyAI is hiring a Software Engineer to turn cutting-edge AI research into products our customers rely... ...-documented APIs serving production traffic at scale. You’ll design and ship customer-facing APIs, scale inference infrastructure for 1M+ users, and collaborate with...
- ...analysis and AI training and inference. Designed from the... ...data center, edge, and cloud. About the Role As a Forward Deployed Engineer (FDE), you are a core... ...most strategic, large-scale customer environments.... ...customer's technical team to optimize infrastructure...CloudRemote work
$234.4k - $296.6k
...for hybrid, multi-cloud environments. Join... ...expertise with the scale and operational... ...and Cisco’s global engineering capabilities. Our... ...distributed training and inference pipelines to... ...model deployment, optimization, and evaluation.Preferred... ...in the code-gen / LLM community.Why...CloudFull timeTemporary workLocal areaFlexible hours$148.7k - $201.2k
...the people who keep the cloud running. We support... ...diverse team of data engineers, business intelligence... ...to join our Material Optimization team. This role will drive... ...-visibility, large-scale, innovative system solutions... ..., and post launch improvements• Manage change...CloudFlexible hours$210.2k - $284.3k
...commitment, aligning engineering and business... ...available at scale.Key job... ...KPIs (cost per inference, time from PoC... ...input to post-launch tooling- Convert... ...broadly adopted cloud platform. We pioneered... ...inference cost optimization. You have built... ...implementing enterprise LLM proxy or...CloudLocal areaFlexible hoursShift work$152.2k - $243.7k
...opportunity to create impact at scale — tackling meaningful... ...entrepreneurs and engineers in 2016, Pismo is a... ...in the market. Pismo’s cloud-based platform empowers firms to build and launch financial products rapidly... ...and optimize high-performance data services...CloudFull timeWork experience placementWork at officeLocal area$146k - $194k
...Platform is the internal engineering force multiplier behind... ...is how Anduril scales its business systems with... ...production stability.Support launch readiness, production... ...distributed systems, CI/CD, and cloud or platform... ...AI coding assistants, LLM-enabled applications, or...CloudFull timeWork experience placementImmediate start$171k - $231.4k
...a broad set of global cloud-based services including... ...faster, lower IT costs, and scale. You will be a part of... ...AWS Product Compliance Engineering team within AWS Global... ...region planning and launches. You will develop compliance... ...Engineering, AWS Infra Service Supply Chain, Compliance...CloudLocal areaFlexible hours$148.7k - $201.2k
...backbone of Generative AI cloud at AWS? Do you want to... ...for AI training and inference? Want to do industry... ...applied to those at cloud scale? If yes, then come... ...you. The AWS Hardware Engineering team creates server designs... ...conception, test, launch, and operations....CloudInternshipLocal areaFlexible hours- Snowflake is seeking AI-native thinkers for the AI Research team to advance LLM inference systems and optimization. You will work on high-performance, adaptive inference across distributed serving, GPU kernels, and system co-design, collaborating with researchers and product...
$254k - $350k
...Machine Learning and System Optimization Engineer, you will orchestrate and allocate... ...allow for more efficient inference by sharing various parts of... ..., production-ready large-scale models to our on-vehicle stack... ...(e.g., TensorRT-LLM).$254,000 - $350,000 a yearBase...Full timeTemporary workRelocation package$209.1k - $282.9k
...seek a Staff Robotics Engineer to help create the next... ..., independent of cloud services.You will provide... ...methods, and performance optimization.Demonstrated ability... ...with AI and perception inference pipelines deployed on... ...services are built and scaled at Arm.In addition to...CloudWork at officeLocal areaVisa sponsorshipRelocation package$143.7k - $194.4k
...AWS Inferentia and Trainium cloud-scale machinelearning accelerators... ...role is for a senior software engineer in the Machine Learning Inference Applications team. This role... ...development and performance optimization of core building blocks of LLM Inference - Attention, MLP,...CloudInternshipFlexible hours$183k - $247.6k
...backbone of Generative AI cloud at AWS? Do you want to... ...for AI training and inference? Want to do industry leading... ...hardware, and network engineers, supply chain... ...training and inference at scale.You will define and drive... ...corrective actions. After launch, you will oversee the...CloudLocal areaFlexible hours- ...Hardware / Machine Design Engineer Location: Hybrid |... ...a next-generation cloud platform designed to power... ...workloads—including large-scale compute, model training, fine-tuning, inference, and emerging agentic... ...into highly optimized, production-ready systems...CloudWork at officeRelocation3 days per week
- ...GPU Performance / Kernel Engineer Location: Hybrid |... ...multiple roles available Optimize the Performance Layer... ...building a next-generation cloud platform designed to... ...workloads—including large-scale compute, model training, fine-tuning, inference, and emerging agentic AI...CloudWork at officeRelocation3 days per week
$162.7k - $220.2k
...passionate about helping customers optimize price/performance for AI Inference at scale? AWS is seeking a Go-To-Market (... ...to prospect and develop pipeline, launch new opportunities and drive usage... ...Experience selling enterprise software or cloud-based applications- Experience...CloudLocal areaWorldwideFlexible hours- ...Stripe is seeking a Staff Infrastructure Engineer to lead and evolve our compute infrastructure, enabling global scale and cost efficiency. You will collaborate with product... ...reliability challenges in a fast-paced, cloud-centric environment. #J-18808-Ljbffr Jobleads...Cloud
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Cloud Inference Launch Engineer: Scale & Optimize LLM Infra. Be the first to apply!
- senior cloud security engineer Seattle, WA
- principal cloud computing engineer Seattle, WA
- senior principal cloud computing engineer Seattle, WA
- cloud engineering manager Seattle, WA
- big data cloud engineer Seattle, WA
- cloud security engineer Seattle, WA
- salesforce marketing cloud developer Seattle, WA
- aws cloud security engineer Seattle, WA
- cloud engineer Seattle, WA
- cloud developer Seattle, WA


