LLM Inference Engineer — Scalable AI Serving
Reducto, Inc.
A tech startup in AI model serving located in San Francisco is seeking a qualified candidate to architect scalable inference systems. The role focuses on optimizing model serving performance and integrating advanced techniques for AI deployment. Candidates should have strong expertise in Python and PyTorch, along with low-level systems knowledge. This in-person position offers a fast-paced work environment that is ideal for those eager to tackle complex challenges and shape the future of AI technology.#J-18808-Ljbffr
$160k - $230k
About the RoleAt Together.ai, we are building state... ...enable efficient and scalable inference for large language... ...Frameworks and Optimization Engineer to design, develop,... ...shape the future of LLM inference infrastructure... ...for high-performance serving.Apply CUDA graph...SuggestedFull time- Anyscale is seeking a Distributed LLM Inference Engineer in San Francisco, California. This pivotal role involves pushing the boundaries of performance for ML inference at scale. You'll work closely with product teams to deliver end-to-end solutions while leveraging open...Suggested
$170k - $245k
...creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI... ...tech stacks to accelerate the progress of AI applications out into the real world.... ...to date.About the roleAs a Distributed LLM Inference Engineer, you will help systems and...SuggestedWork at office- ...Francisco is seeking experienced backend engineers to own the systems that serve our diffusion LLMs in production.... ...that handles billions of inference requests, optimizing for latency, throughput... ..., with responsibilities spanning scalable services, model serving, load...Suggested
$300 per month
...vertically integrated AI infrastructure company... ...Crusoe, our Production Engineering team ensures the reliability and scalability of Crusoe’s AI-... ...services with a focus on serving and scaling LLM workloadsBuild automation... ...distributed AI pipelines and inference servicesDefine,...SuggestedTemporary work$175k - $250k
Global Inference Library Engineer $175000 - $250000 per year | San Francisco, CA |... ...about us: We're a well-funded AI infrastructure startup... ...who understands how modern LLM inference systems work under... ...inference frameworks and model-serving infrastructure Hands-on experience...Permanent employmentLocal area- MakerMaker.AI is looking for a Senior Machine Learning Systems Engineer in San Francisco. In this role, you will build and operate production inference systems, optimizing for performance and reliability... ...in production-grade serving infrastructure, be fluent in Python...
$300 per month
...vertically integrated AI infrastructure company... ...Crusoe, our Production Engineering team ensures the reliability and scalability of Crusoe’s AI-... ...services with a focus on serving and scaling LLM workloadsDefine, measure... ...large-scale training and inference clustersAutomate...Temporary work$91.1k - $179.5k
...success. We are hiring an AI Engineer to build and operate the data... ...reliable, secure, and scalable AI solutions. This role is... ...model training, real-time inference, and LLM applications using Claude-,... ...datasets and feature engineering/serving for ML training and real-time...Local area- Inception is seeking engineers and scientists to design, optimize, and scale the diffusion LLM serving systems powering production inference. Your work will help make inference faster, more cost-efficient, and more reliable. You will extend orchestration frameworks (Kubernetes...
- ...patients worldwide.We’re a team of engineers, clinicians, and innovators... ...Systems GPU Engineer - AI & Robotics, you will be... ...optimized, robust, validated and scalable medical device products and... ...models to real-time onboard inference—while serving as a core contributor to team...Local areaWorldwideFlexible hours
- Help build the inference stack behind the next generation... ...architectures are served at scale. You’ll be building... ...real-time multimodal AI capable of processing... ...research and product engineering, designing the... ...architectures. Design scalable, reliable distributed...Work at officeRelocation package
$286.2k - $326.7k
...Senior Distinguished Engineer, AI Compute (Remote Eligible... ...experiences and scalable, high-performance AI infrastructure... ...to reimagine how we serve our customers and... ...model training, model inference and feature generation... ...workloads from LLM pre-training and reinforcement...Full timePart timeLocal areaRemote work$170k - $200k
...around the globe. Our AI-powered labor marketplace... ...a Senior Software Engineer - Instawork Robotics to... ...requirements into clear, scalable technical solutions, leveraging... ...rolloutsExposure to LLM-based systems—includes... ...AI-powered platform serves thousands of businesses...Hourly payTemporary workLocal areaShift work- ...funded startup building AI-native solutions for... ...looking for a Founding Engineer & CTO to be the technical... ...our product vision into scalable, production-grade AI... ...from data pipelines to LLM orchestration to client... ...-tenant SaaS platforms serving enterprise clients. Familiarity...Flexible hours
- Magic AI, Inc. is seeking a Member of Technical Staff to design and operate distributed systems for serving models in production and driving large-scale post-training workflows... ...the infrastructure enabling fast inference and scalable RL iteration, balancing KV-cache strategies...
- Vast.ai Inc. is seeking a systems engineer with HPC or parallel programming experience to help scale AI inference. You will design and optimize GPU kernels and tensor libraries, leveraging... ...strong emphasis on HPC techniques and scalable workloads. You will collaborate with...
- ...powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion... ...help build the platform engineers turn to to ship AI... ...hardware. We believe that as LLM and multi-modal workloads... ...compute for Disaggregated Serving, Wide Expert Parallelism...Full timeFlexible hours
$150k - $190k
...our 3D lidar technology will serve as the foundation of tomorrow... ...Role SummaryAs a Staff Software Engineer in Test, you will be the... ...Design, develop, and maintain a scalable and modular automated test... ...and UDP.Experience leveraging AI and LLM tools to assist in analysis and...Work experience placementLocal area- ...Head of Internal Tools Engineering, Artificial Intelligence (AI) Required, Work From Home... ...productivity. - Design scalable, secure architectures using... ...unified technical vision. - Serve as the bridge between business... ...Tools Engineering, LLM, SaaS, SDLC, Software as a...Remote workWork from home
$200.8k - $251k
...A leading AI technology company in San Francisco seeks a team member to build and optimize a machine learning framework for large... ...should have system optimization experience and solid software engineering skills, particularly in tools like CUDA and Pytorch. This full-...Full time- Senior ML Systems Engineer, Frameworks & Tooling at Cohere... ...scale intelligence to serve humanity. We’re... ...enterprises who are building AI systems to power magical... ...fast, reliable, and scalable model training and build... ...responsible for large-scale LLM training. Design...Full timeWork at officeRemote workFlexible hours
$206.3k - $388k
...looking for a Principal ML Engineer to architect and scale... ...-ready data Scale up inference throughput across the... ...compute scheduling ARCHITECT SCALABLE DATA INFRASTRUCTURE... ...store, index, and serve billions of data points... ...into impact, powered by AI and driven by human ingenuity...Full timeTemporary workLocal areaWorldwide$221k - $247k
...reflects the people we serve.All full-time employees... ...ourTotal Rewards philosophy.AI is a fundamental part... ...ecosystem to improve scalability, efficiency, and... ...EnablementPartner with EAIT engineers, administrators, AI architects... ...knowledge of leading LLM platforms (e.g., OpenAI...Full timeContract workWork at officeLocal areaShift work2 days per week3 days per week- ...running the world’s best data and AI infrastructure platform so our... ...their business. Founded by engineers — and customer obsessed — we... ...metric views, and definitions that serve as the source of truth for... ...applications or agentic workflows (LLM-powered apps and automations)....Worldwide
$264.8k - $331k
...Learning Systems Research Engineer, Agent Post-training - Enterprise GenAI AI is becoming vitally... ...research and resources that serve all of our enterprise... ...optimize our training and inference framework. Post-train state... ...: At least 1-3 years of LLM training in a production...Full timeContract workFor contractorsFor subcontractorWork at office$102k - $182.71k
...Senior Search Systems Engineer to build the intelligence... ...of our marketing and AI visibility data. This role... .... This role serves as the bridge between SEO... ...teams to operationalize scalable datasets, transformation... ...frameworks related to LLM analysis, AI agents, search...Shift work$200k - $240k
...blockchain analytics and AI solutions to help law... ...world for all. The AI Engineering Team is chartered with... ...petabyte-scale pipelines, serve models with millisecond... ...-edge tools in the LLM and agent space — including... ...out a modular and scalable AI infrastructure stack...Remote workWorldwide$105.4k - $124k
...the customers and businesses we serve to make better and smarter... ...Intelligent Document Processing Engineer to join the Intelligent Document... ...team delivers enterprise-wide AI, Machine Learning, Generative... ...solutions, and deliver scalable applications that leverage Tungsten...Full timeLocal area3 days per week$165k - $206k
...we are hiring the world’s best engineers, scientists, designers,... ...Systems Engineer who leads with AI - not as a tool on the side, but... ...to let AI do the heavy lifting.Serve as a cross-functional technical... ...live supportLeverage Cursor with LLM-pair programming (Claude/...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to LLM Inference Engineer — Scalable AI Serving. Be the first to apply!




