Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff Applied AI Inference Engineer

$215k - $260k

Crusoe

Job Description

Job Description

Crusoe is on a mission to accelerate the abundance of energy and intelligence . As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power the world's most ambitious AI workloads. When you join Crusoe, you join a team that is building the future, faster.

We're in the midst of the greatest industrial revolution of our time. The demand for AI compute is boundless, and power is a bottleneck. We're solving that — with an energy-first approach that makes AI infrastructure better for the world and faster for the people innovating with AI.

We're looking for problem-solving, opportunity-finding teammates with a sense of urgency, who believe in the scale of our ambition and thrive on a path not fully paved — people who want to grow their careers alongside a team of experts across energy, manufacturing, data center construction, and cloud services.

If you want to do the most meaningful work of your career, help our customers and partners advance their AI strategies, and be part of a high-performing team that believes in each other, come build with us at Crusoe.

About the Role

You will spend your time making large language models run faster, cheaper, and more reliably in production. That means owning the inference stack end to end: profiling where time and cost go, bringing modern optimization techniques into real deployments, and getting deep into the serving code when the defaults are not good enough. This is core systems and performance work on some of the most demanding models in use today.

The work is applied, not academic. The optimizations you build land in real customer deployments, each with its own models, traffic patterns, latency targets, and cost constraints. So while performance is the heart of the role, you will also work directly with customer engineering teams to tailor deployments to their needs, take a workload from an early proof of concept to a fully monitored production service, and make sure the gains you engineer actually show up for the people running the workload.

To set expectations clearly, this is a hands-on engineering role built around coding, profiling, and low-level optimization. It also carries a customer-facing side, along with elements of product and technical solutions work, because that is where the performance work gets proven.

What You'll Be Working On:

  • Bring current inference techniques into production and refine them.

  • Design and optimize serving architectures, including prefill and decode disaggregation, request routing, and related approaches.

  • Work down into the serving stack, from frameworks like vLLM and SGLang to the CUDA kernels underneath, profiling and running in-depth analysis to find and fix performance problems.

  • Adapt and scale optimization methods across many kinds of ML models, with an emphasis on large language models.

  • Profile and tune deployments against clear targets for latency, throughput, and cost, and keep them dependable under real traffic.

  • Tailor deployments to each customer's models and constraints, partnering with their engineering teams to move a workload from an early proof of concept through to a live, well-monitored production service.

  • Build and support the software and product features around the inference stack in a production setting, using one or more general-purpose languages, with Python preferred given how central it is to ML work.

  • Experiment quickly: take fuzzy goals, shape them into clear specs and focused proofs of concept, run fast experiments to find what works, and ship well-tested results without delay.

  • Own delivery end to end, from the first experiment through to the optimization running in production, keeping the underlying performance goals, clear specs, and follow-through front of mind, and drafting features and product requirement documents together with other engineering and product teams.

  • Work through ambiguity and make sound calls on tradeoffs and tooling, steering away from complexity that is not needed.

  • Take real pride and ownership in your work, hold yourself accountable, and look for the same from the people around you.

What You'll Bring to the Team:

  • A Bachelor's, Master's, or Ph.D. in Computer Science, Engineering, Mathematics, or a related field.

  • Hands-on experience shipping code in production with one or more general-purpose languages, such as Python or C++, with a strong preference for Python.

  • Familiarity with methods for optimizing LLMs for high throughput / low latency inference.

  • Comfort with modern LLM serving frameworks such as vLLM or SGLang, and with profiling and analyzing performance down to the kernel level.

  • A firm grasp of how GPUs are built and how they behave.

  • Clear interest and hands-on experience with large language models.

  • A working knowledge of AI/ML pipelines and the full path of developing and deploying ML models.

  • Strong communication skills, particularly when explaining hard technical topics to customers and teammates.

Bonus points:

  • A track record of making software systems run faster, especially for large language models.

  • Experience with CUDA or comparable technologies.

  • A strong command of software engineering fundamentals, with a record of building and shipping AI/ML inference systems.

  • Experience with Docker and Kubernetes.

  • Prior work building or tuning AI/ML projects, particularly in a customer-facing setting.

Benefits:

  • Competitive compensation and equity packages

  • Restricted Stock Units

  • Paid time off, paid holidays & leave of absence programs

  • Comprehensive health, dental & vision insurance

  • Employer contributions to HSA account

  • Paid parental leave

  • Paid life insurance, short-term and long-term disability

  • Professional development & tuition reimbursement

  • Mental health & wellness support

  • Commuter benefits (parking & transit)

  • Cell phone stipend

  • 401(k) Retirement plan with company match up to 4% of salary

  • Volunteer time off

  • Global travel insurance & emergency assistance

  • Daily meals allowance

  • Additional perks & programs specific to location

Compensation Range

Compensation will be paid in the range of up to $215,000 - $260,000 + Bonus. Restricted Stock Units are included in all offers. Compensation to be determined by the applicant's knowledge, education, and abilities, as well as internal equity and alignment with market data.

Crusoe is an Equal Opportunity Employer. Employment decisions are made without regard to race, color, religion, disability, genetic information, pregnancy, citizenship, marital status, sex/gender, sexual preference/ orientation, gender identity, age, veteran status, national origin, or any other status protected by law or regulation.

Vacancy posted 15 days ago
Similar jobs that could be interesting for youBased on the Staff Applied AI Inference Engineer in Sunnyvale, CA vacancy
  • $150k - $250k

     ...You will build state-of-the-art AI capabilities for Cylake's next-...  ...researchers and software engineers to develop, productize, and deploy...  ...security products. Research and apply advancements in LLM architectures, AI agents, model inference, and serving optimization.... 
    Suggested
    Full time

    Cylake, Inc

    Sunnyvale, CA
    12 hours ago
  • The Staff AI Engineer - Applied Research will report to the VP Research in the Future Forward organization. This is a high-impact individual contributor...  ...(PyTorch, TensorFlow), edge computing for low-latency inference, and integration of AI with physical robotic systems or... 
    Suggested
    Local area
    Worldwide
    Flexible hours
    Shift work

    Socket

    Sunnyvale, CA
    5 days ago
  •  ...community where each individual can thrive.Job DescriptionThe AI Inference Engineer plays a critical role in the AI lifecycle by bridging the...  ...protected by applicable local, state, or federal laws. This policy applies to all aspects of employment, including, but not limited to,... 
    Suggested
    Full time
    Local area
    Immediate start

    F5 Networks

    San Jose, CA
    2 days ago
  • $152k - $241.5k

    NVIDIA's Silicon Co-Design Group is seeking an Applied AI Engineer to innovate, develop, and integrate innovative AI solutions into the design and automation infrastructure that powers our chips. Every CPU, GPU, and Tegra SoC NVIDIA has shipped in the past four years passed... 
    Suggested
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $168k - $264.5k

    Nvidia's SOC Design (SOCD) team is looking for an Applied AI Engineer who is passionate about eliminating bottlenecks in SOC integration workflows through intelligent automation. If you are driven to build AI-powered tools, agents, and automation solutions to dramatically... 
    Suggested
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $184k - $287.5k

     ...in high performance computing, gaming and AI. Our GPUs and SOCs give outstanding...  ...Blackwell generation alone! Now we're hiring the engineer who will lead the rebuild of that...  ...checkpoints.Lead eval-driven development for applied AI in production: error analysis on real... 
    Full time
    Immediate start

    Nvidia

    Santa Clara, CA
    1 day ago
  •  ...products that accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems....  ...beyond. Together, we advance your career. THE ROLEWe are hiring Applied AI Engineers to work directly with hardware and software engineering teams... 

    AMD

    Santa Clara, CA
    12 hours ago
  • $135k - $185k

     ...news and information powered by advanced AI, recommendation systems, and adtech.Recognized...  ...algorithm work as an excellent engineer to join our advertising team? In this role...  ...field of ad delivery, with more than 2 years applying large-model capabilities to systems such... 
    Full time
    Work experience placement
    Local area
    Work from home

    News Break

    Mountain View, CA
    4 days ago
  • $207k - $300k

     ...serving frontier models.Translate AI/ML research into production-...  ...services, optimizing model inference latency, throughput, and large...  ...architectures and cognitive planning engines capable of executing complex,...  ...industry problems.As a Staff Software Engineer in the Research... 

    Google

    Sunnyvale, CA
    2 days ago
  • $177k - $226k

     ...critical care through our rapid seizure detection technology, come join the movement!Position Overview:The Senior Manager, Applied AI Engineering is a senior individual contributor role with broad ownership across Ceribell's internal AI engineering portfolio. This person... 
    For contractors
    Work at office
    Local area
    Immediate start
    Remote work

    Ceribell

    Sunnyvale, CA
    1 day ago
  • $152k - $230k

     ...recently, GPU deep learning ignited modern AI — the next era of computing. NVIDIA is a...  ...choice to join us today.Design-for-Test Engineering at NVIDIA works on groundbreaking innovations...  ...complex datasets and explorations using Applied AI methods.In addition, you will help... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $229.9k - $262.4k

    Senior Lead AI Engineer (FM Hosting, LLM Inference) Overview: At Capital One, we are creating responsible and reliable AI systems, changing banking...  ...banking. We are committed to continuing to build world-class applied science and engineering teams to deliver our industry... 
    Full time
    Part time
    Local area

    Capital One

    San Jose, CA
    4 days ago
  •  ...Systems builds the world's largest AI chip, 56 times larger than...  ...-leading training and inference speeds; over 10 times faster...  ...visible role working directly with Engineering, Product, Infrastructure, SRE...  ...to work at Cerebras here! Apply today and become part of the... 
    Remote work

    Cerebras Systems

    Sunnyvale, CA
    12 hours ago
  •  ...analysis and optimization of LLM and VLM inference on NVIDIA GPU-accelerated systems. Develop...  ...improvements to open-source inference engines to reduce latency and increase throughput...  ...over 6 years of experience in full-stack AI inference performance with strong programming... 

    NVIDIA AI

    Santa Clara, CA
    2 days ago
  • NVIDIA is seeking a Senior Software Engineer to advance AI inference performance on GPU-accelerated systems. You will optimize LLM/VLM workloads, profile with Nsight tools, and contribute to open-source inference engines while collaborating with model, kernel, and networking... 

    Nvidia Corporation in

    Santa Clara, CA
    1 day ago
  • NVIDIA seeks a Senior Software Engineer - AI Inference Performance to push LLM/VLM workloads toward practical performance limits on NVIDIA GPUs. You will lead end-to-end analysis, build performance models, and optimize latency, throughput, and energy efficiency across models... 

    NVIDIA

    Santa Clara, CA
    2 days ago
  • $2,000 per month

     ...are heavily focused on inference . Backed by hundreds...  ...and staffed by leading engineers, Etched is redefining the...  ...We are using AI to build AI chips. AI agents...  ...and push past it. As an Applied AI Engineer, you will embed...  ...all of our technical staff to contribute to both and... 
    Work at office
    Relocation package
    Night shift

    Etched

    San Jose, CA
    15 days ago
  • $184k - $287.5k

     ...Manufacturing & System Co-Design Workflow Engineer to lead the methodology and infrastructure...  ...workflows. This is the infrastructure that makes AI genuinely usable in a rigorous engineering...  ...outputs (speed, power, binning) and apply AI with genuine judgment: reviewable... 
    Full time
    Immediate start

    Nvidia

    Santa Clara, CA
    4 days ago
  • Applied AI Engineer, Use-case - Palo Alto/NYC Full-time Onsite 2+ years exp Mistral AI is a pioneering company focused on democratizing AI through...  ...contributing to our open source codebases for tasks such as inference and fine‑tuning You’ll be involved in pre‑sales calls to... 
    Full time
    H1b
    Work at office

    Jobright.ai

    Palo Alto, CA
    1 day ago
  • AI Fund is seeking an Applied AI Engineer in Mountain View, CA. In this role, you will design state-of-the-art document AI solutions and engage with clients to drive measurable business outcomes. The ideal candidate will have over 5 years of experience in AI/ML, strong... 

    AI Fund

    Mountain View, CA
    4 days ago
  • $175k - $225k

     ...world’s documents computable. We are an AI-native company transforming PDFs, PowerPoints...  ...together some of the strongest AI Engineers and Machine Learning Engineers in the industry...  ...builders of agentic AI systems and have applied that work in the real world through Agentic... 
    Contract work
    Work at office

    AI Fund

    Mountain View, CA
    1 day ago
  • $135k - $155k

     ...news and information powered by advanced AI, recommendation systems, and adtech. Recognized...  ...Are you a recent graduate excited to apply LLMs and Agent technology to real...  ...optimization platform. You'll work alongside senior engineers to develop AI advertising expert systems... 
    Full time
    Work experience placement
    Internship
    Local area
    Work from home

    NewsBreak

    Mountain View, CA
    7 days ago
  • DoorDash is seeking a Member of Technical Staff in Applied AI Research in the San Francisco Bay Area. You will work directly with the cofounder to direct AI research, develop agent environments, and design systems to improve agent performance at real-world scale. You will... 

    DoorDash

    Sunnyvale, CA
    1 day ago
  •  ...Senior Applied AI Engineer Agentic Systems Location: Mountain View, CA Job Description Agentic Feature Development & Full Stack Delivery Design, build, and ship agentic features directly within EAS - autonomous workflow agents, multi-step task orchestration... 

    Swift Hire LLC

    Mountain View, CA
    2 days ago
  • $130k - $220k

    Santa Clara, CAData Engineering - ML Infrastructure /Full-time /HybridFinding...  ...work at the intersection of applied machine learning, information...  ...operate distributed mining, inference, and indexing pipelines over...  ...use artificial intelligence (AI) tools to support parts of... 
    Full time

    Plus.ai

    Santa Clara, CA
    2 days ago
  • Walmart Global Tech in Sunnyvale, CA is seeking a Group Director, Applied AI & Engineering to lead an elite team building Walmart’s next-generation intelligence layer. You will own vision, strategy, and execution for scalable foundational models trained on multi-modal... 

    Walmart Global Tech

    Sunnyvale, CA
    2 days ago
  • $144k - $236k

     ...of the team.Responsibilities: AI is at the core of how...  ...platforms. As a Senior AI Software Engineer you will own end-to-end machine...  ...or quality improvement (i.e. inference/training efficiency, engineer...  ...model paradigmsExperience applying AI/ML to recommender systems... 
    For contractors
    Work at office
    Immediate start
    Flexible hours

    Linkedin

    Mountain View, CA
    4 days ago
  • $152k - $241.5k

    We are seeking a Senior AI/ML Performance and Efficiency Engineer, GPU Clusters at NVIDIA to join our AI Efficiency...  ...tools, frameworks, and apply ML techniques to detect & analyze efficiency...  ...investigating, and resolving, training & inference performance end to endDebugging and... 
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  •  ...Systems builds the world's largest AI chip, 56 times larger than...  ...-leading training and inference speeds; over 10 times faster...  ...loop." You'll sit between engineering, product, and customer-facing...  ...work at Cerebras here ! Apply today and become part of the... 
    Full time

    Cerebras Systems

    Sunnyvale, CA
    12 hours ago
  • $100k

     ...the industry on cutting-edge AI technology, revolutionizing performance...  .../ Signal Integrity Engineer to design and validate high-bandwidth...  ...for next-generation AI inference and training clusters. This role...  ..., and E2). These requirements apply to persons located in the U.S.... 
    Permanent employment

    Tenstorrent

    Santa Clara, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff Applied AI Inference Engineer. Be the first to apply!