Staff Applied AI Inference Engineer
$215k - $260kCrusoe
Job Description
Job Description
Crusoe is on a mission to accelerate the abundance of energy and intelligence . As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power the world's most ambitious AI workloads. When you join Crusoe, you join a team that is building the future, faster.
We're in the midst of the greatest industrial revolution of our time. The demand for AI compute is boundless, and power is a bottleneck. We're solving that — with an energy-first approach that makes AI infrastructure better for the world and faster for the people innovating with AI.
We're looking for problem-solving, opportunity-finding teammates with a sense of urgency, who believe in the scale of our ambition and thrive on a path not fully paved — people who want to grow their careers alongside a team of experts across energy, manufacturing, data center construction, and cloud services.
If you want to do the most meaningful work of your career, help our customers and partners advance their AI strategies, and be part of a high-performing team that believes in each other, come build with us at Crusoe.
About the Role
You will spend your time making large language models run faster, cheaper, and more reliably in production. That means owning the inference stack end to end: profiling where time and cost go, bringing modern optimization techniques into real deployments, and getting deep into the serving code when the defaults are not good enough. This is core systems and performance work on some of the most demanding models in use today.
The work is applied, not academic. The optimizations you build land in real customer deployments, each with its own models, traffic patterns, latency targets, and cost constraints. So while performance is the heart of the role, you will also work directly with customer engineering teams to tailor deployments to their needs, take a workload from an early proof of concept to a fully monitored production service, and make sure the gains you engineer actually show up for the people running the workload.
To set expectations clearly, this is a hands-on engineering role built around coding, profiling, and low-level optimization. It also carries a customer-facing side, along with elements of product and technical solutions work, because that is where the performance work gets proven.
What You'll Be Working On:
Bring current inference techniques into production and refine them.
Design and optimize serving architectures, including prefill and decode disaggregation, request routing, and related approaches.
Work down into the serving stack, from frameworks like vLLM and SGLang to the CUDA kernels underneath, profiling and running in-depth analysis to find and fix performance problems.
Adapt and scale optimization methods across many kinds of ML models, with an emphasis on large language models.
Profile and tune deployments against clear targets for latency, throughput, and cost, and keep them dependable under real traffic.
Tailor deployments to each customer's models and constraints, partnering with their engineering teams to move a workload from an early proof of concept through to a live, well-monitored production service.
Build and support the software and product features around the inference stack in a production setting, using one or more general-purpose languages, with Python preferred given how central it is to ML work.
Experiment quickly: take fuzzy goals, shape them into clear specs and focused proofs of concept, run fast experiments to find what works, and ship well-tested results without delay.
Own delivery end to end, from the first experiment through to the optimization running in production, keeping the underlying performance goals, clear specs, and follow-through front of mind, and drafting features and product requirement documents together with other engineering and product teams.
Work through ambiguity and make sound calls on tradeoffs and tooling, steering away from complexity that is not needed.
Take real pride and ownership in your work, hold yourself accountable, and look for the same from the people around you.
What You'll Bring to the Team:
A Bachelor's, Master's, or Ph.D. in Computer Science, Engineering, Mathematics, or a related field.
Hands-on experience shipping code in production with one or more general-purpose languages, such as Python or C++, with a strong preference for Python.
Familiarity with methods for optimizing LLMs for high throughput / low latency inference.
Comfort with modern LLM serving frameworks such as vLLM or SGLang, and with profiling and analyzing performance down to the kernel level.
A firm grasp of how GPUs are built and how they behave.
Clear interest and hands-on experience with large language models.
A working knowledge of AI/ML pipelines and the full path of developing and deploying ML models.
Strong communication skills, particularly when explaining hard technical topics to customers and teammates.
Bonus points:
A track record of making software systems run faster, especially for large language models.
Experience with CUDA or comparable technologies.
A strong command of software engineering fundamentals, with a record of building and shipping AI/ML inference systems.
Experience with Docker and Kubernetes.
Prior work building or tuning AI/ML projects, particularly in a customer-facing setting.
Benefits:
Competitive compensation and equity packages
Restricted Stock Units
Paid time off, paid holidays & leave of absence programs
Comprehensive health, dental & vision insurance
Employer contributions to HSA account
Paid parental leave
Paid life insurance, short-term and long-term disability
Professional development & tuition reimbursement
Mental health & wellness support
Commuter benefits (parking & transit)
Cell phone stipend
401(k) Retirement plan with company match up to 4% of salary
Volunteer time off
Global travel insurance & emergency assistance
Daily meals allowance
Additional perks & programs specific to location
Compensation Range
Compensation will be paid in the range of up to $215,000 - $260,000 + Bonus. Restricted Stock Units are included in all offers. Compensation to be determined by the applicant's knowledge, education, and abilities, as well as internal equity and alignment with market data.
Crusoe is an Equal Opportunity Employer. Employment decisions are made without regard to race, color, religion, disability, genetic information, pregnancy, citizenship, marital status, sex/gender, sexual preference/ orientation, gender identity, age, veteran status, national origin, or any other status protected by law or regulation.
$150k - $250k
...You will build state-of-the-art AI capabilities for Cylake's next-... ...researchers and software engineers to develop, productize, and deploy... ...security products. Research and apply advancements in LLM architectures, AI agents, model inference, and serving optimization....SuggestedFull time- The Staff AI Engineer - Applied Research will report to the VP Research in the Future Forward organization. This is a high-impact individual contributor... ...(PyTorch, TensorFlow), edge computing for low-latency inference, and integration of AI with physical robotic systems or...SuggestedLocal areaWorldwideFlexible hoursShift work
- ...community where each individual can thrive.Job DescriptionThe AI Inference Engineer plays a critical role in the AI lifecycle by bridging the... ...protected by applicable local, state, or federal laws. This policy applies to all aspects of employment, including, but not limited to,...SuggestedFull timeLocal areaImmediate start
$152k - $241.5k
NVIDIA's Silicon Co-Design Group is seeking an Applied AI Engineer to innovate, develop, and integrate innovative AI solutions into the design and automation infrastructure that powers our chips. Every CPU, GPU, and Tegra SoC NVIDIA has shipped in the past four years passed...SuggestedFull timeRemote work$168k - $264.5k
Nvidia's SOC Design (SOCD) team is looking for an Applied AI Engineer who is passionate about eliminating bottlenecks in SOC integration workflows through intelligent automation. If you are driven to build AI-powered tools, agents, and automation solutions to dramatically...SuggestedFull time$184k - $287.5k
...in high performance computing, gaming and AI. Our GPUs and SOCs give outstanding... ...Blackwell generation alone! Now we're hiring the engineer who will lead the rebuild of that... ...checkpoints.Lead eval-driven development for applied AI in production: error analysis on real...Full timeImmediate start- ...products that accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems.... ...beyond. Together, we advance your career. THE ROLEWe are hiring Applied AI Engineers to work directly with hardware and software engineering teams...
$135k - $185k
...news and information powered by advanced AI, recommendation systems, and adtech.Recognized... ...algorithm work as an excellent engineer to join our advertising team? In this role... ...field of ad delivery, with more than 2 years applying large-model capabilities to systems such...Full timeWork experience placementLocal areaWork from home$207k - $300k
...serving frontier models.Translate AI/ML research into production-... ...services, optimizing model inference latency, throughput, and large... ...architectures and cognitive planning engines capable of executing complex,... ...industry problems.As a Staff Software Engineer in the Research...$177k - $226k
...critical care through our rapid seizure detection technology, come join the movement!Position Overview:The Senior Manager, Applied AI Engineering is a senior individual contributor role with broad ownership across Ceribell's internal AI engineering portfolio. This person...For contractorsWork at officeLocal areaImmediate startRemote work$152k - $230k
...recently, GPU deep learning ignited modern AI — the next era of computing. NVIDIA is a... ...choice to join us today.Design-for-Test Engineering at NVIDIA works on groundbreaking innovations... ...complex datasets and explorations using Applied AI methods.In addition, you will help...Full time$229.9k - $262.4k
Senior Lead AI Engineer (FM Hosting, LLM Inference) Overview: At Capital One, we are creating responsible and reliable AI systems, changing banking... ...banking. We are committed to continuing to build world-class applied science and engineering teams to deliver our industry...Full timePart timeLocal area- ...Systems builds the world's largest AI chip, 56 times larger than... ...-leading training and inference speeds; over 10 times faster... ...visible role working directly with Engineering, Product, Infrastructure, SRE... ...to work at Cerebras here! Apply today and become part of the...Remote work
- ...analysis and optimization of LLM and VLM inference on NVIDIA GPU-accelerated systems. Develop... ...improvements to open-source inference engines to reduce latency and increase throughput... ...over 6 years of experience in full-stack AI inference performance with strong programming...
- NVIDIA is seeking a Senior Software Engineer to advance AI inference performance on GPU-accelerated systems. You will optimize LLM/VLM workloads, profile with Nsight tools, and contribute to open-source inference engines while collaborating with model, kernel, and networking...
- NVIDIA seeks a Senior Software Engineer - AI Inference Performance to push LLM/VLM workloads toward practical performance limits on NVIDIA GPUs. You will lead end-to-end analysis, build performance models, and optimize latency, throughput, and energy efficiency across models...
$2,000 per month
...are heavily focused on inference . Backed by hundreds... ...and staffed by leading engineers, Etched is redefining the... ...We are using AI to build AI chips. AI agents... ...and push past it. As an Applied AI Engineer, you will embed... ...all of our technical staff to contribute to both and...Work at officeRelocation packageNight shift$184k - $287.5k
...Manufacturing & System Co-Design Workflow Engineer to lead the methodology and infrastructure... ...workflows. This is the infrastructure that makes AI genuinely usable in a rigorous engineering... ...outputs (speed, power, binning) and apply AI with genuine judgment: reviewable...Full timeImmediate start- Applied AI Engineer, Use-case - Palo Alto/NYC Full-time Onsite 2+ years exp Mistral AI is a pioneering company focused on democratizing AI through... ...contributing to our open source codebases for tasks such as inference and fine‑tuning You’ll be involved in pre‑sales calls to...Full timeH1bWork at office
- AI Fund is seeking an Applied AI Engineer in Mountain View, CA. In this role, you will design state-of-the-art document AI solutions and engage with clients to drive measurable business outcomes. The ideal candidate will have over 5 years of experience in AI/ML, strong...
$175k - $225k
...world’s documents computable. We are an AI-native company transforming PDFs, PowerPoints... ...together some of the strongest AI Engineers and Machine Learning Engineers in the industry... ...builders of agentic AI systems and have applied that work in the real world through Agentic...Contract workWork at office$135k - $155k
...news and information powered by advanced AI, recommendation systems, and adtech. Recognized... ...Are you a recent graduate excited to apply LLMs and Agent technology to real... ...optimization platform. You'll work alongside senior engineers to develop AI advertising expert systems...Full timeWork experience placementInternshipLocal areaWork from home- DoorDash is seeking a Member of Technical Staff in Applied AI Research in the San Francisco Bay Area. You will work directly with the cofounder to direct AI research, develop agent environments, and design systems to improve agent performance at real-world scale. You will...
- ...Senior Applied AI Engineer Agentic Systems Location: Mountain View, CA Job Description Agentic Feature Development & Full Stack Delivery Design, build, and ship agentic features directly within EAS - autonomous workflow agents, multi-step task orchestration...
$130k - $220k
Santa Clara, CAData Engineering - ML Infrastructure /Full-time /HybridFinding... ...work at the intersection of applied machine learning, information... ...operate distributed mining, inference, and indexing pipelines over... ...use artificial intelligence (AI) tools to support parts of...Full time- Walmart Global Tech in Sunnyvale, CA is seeking a Group Director, Applied AI & Engineering to lead an elite team building Walmart’s next-generation intelligence layer. You will own vision, strategy, and execution for scalable foundational models trained on multi-modal...
$144k - $236k
...of the team.Responsibilities: AI is at the core of how... ...platforms. As a Senior AI Software Engineer you will own end-to-end machine... ...or quality improvement (i.e. inference/training efficiency, engineer... ...model paradigmsExperience applying AI/ML to recommender systems...For contractorsWork at officeImmediate startFlexible hours$152k - $241.5k
We are seeking a Senior AI/ML Performance and Efficiency Engineer, GPU Clusters at NVIDIA to join our AI Efficiency... ...tools, frameworks, and apply ML techniques to detect & analyze efficiency... ...investigating, and resolving, training & inference performance end to endDebugging and...Full timeRemote work- ...Systems builds the world's largest AI chip, 56 times larger than... ...-leading training and inference speeds; over 10 times faster... ...loop." You'll sit between engineering, product, and customer-facing... ...work at Cerebras here ! Apply today and become part of the...Full time
$100k
...the industry on cutting-edge AI technology, revolutionizing performance... .../ Signal Integrity Engineer to design and validate high-bandwidth... ...for next-generation AI inference and training clusters. This role... ..., and E2). These requirements apply to persons located in the U.S....Permanent employment
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Staff Applied AI Inference Engineer. Be the first to apply!
- engineering aide Sunnyvale, CA
- senior staff engineer Sunnyvale, CA
- assistant engineer Sunnyvale, CA
- software engineer staff Sunnyvale, CA
- staff engineer Sunnyvale, CA
- senior staff systems engineer Sunnyvale, CA
- staff data engineer Sunnyvale, CA
- technology administrator Sunnyvale, CA
- field applications engineer Sunnyvale, CA
- technical application engineer Sunnyvale, CA



