Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Staff Software Engineer, Cloud Availability Platform

$250k - $300k

Jobleads-US

Crusoe is on a mission to accelerate the abundance of energy and intelligence. As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power the world’s most ambitious AI workloads. When you join Crusoe, you join a team that is building the future, faster.

We’re in the midst of the greatest industrial revolution of our time. The demand for AI compute is boundless, and power is a bottleneck. We’re solving that — with an energy-first approach that makes AI infrastructure better for the world and faster for the people innovating with AI.

We’re looking for problem-solving, opportunity-finding teammates with a sense of urgency, who believe in the scale of our ambition and thrive on a path not fully paved — people who want to grow their careers alongside a team of experts across energy, manufacturing, data center construction, and cloud services.

If you want to do the most meaningful work of your career, help our customers and partners advance their AI strategies, and be part of a high-performing team that believes in each other, come build with us at Crusoe.

Build the operating system for the AI datacenter

Crusoe operates one of the world’s largest managed GPU fleets, and it is growing fast. A fleet at this scale cannot be run the way GPU clouds have traditionally been run: runbooks, war rooms, and heroics. It has to be run by a unified platform that senses, reasons about, and acts on the entire fleet, so infrastructure that used to take a team to operate takes a service instead. That is what our cloud platform team builds. You will work on the control plane for one of the largest AI fleets in the world, at a point where fleet autonomy is still an open problem: nobody has fully solved this at this scale, in this market.

A true platform, not internal tooling: We want to be explicit about that heading, because it is the thing most infrastructure roles get wrong. Everything we build ships as a product, fleet engineers, SREs, and product teams across Crusoe build their own services and workflows on top of what we ship. A platform team does not scale by doing everyone’s work; it scales by making everyone’s work self-serve. Concretely:

  • API-first. Every capability is exposed through well-designed, versioned APIs behind a single gateway. If it isn’t an API, it doesn’t exist. No side doors, including for us.

  • SDKs and paved paths. First-class client libraries, workflow templates, and golden paths so a fleet or SRE engineer can ship a new remediation or lifecycle workflow in days without asking the platform team.

  • Micro frontends and a self-serve portal. Teams plug their own UI surfaces into one developer portal instead of building one-off dashboards. One console for the fleet, extensible by every team.

  • Platform as product. Internal teams are customers. We own contracts, versioning, deprecation policy, quotas, documentation, and support. Adoption is our success metric: the platform wins when other teams choose it because it is the fastest path, not because it is mandated.

What this platform is

Four layers, built as one system:

  • Agents on every site and host that collect telemetry and execute commands.

  • A distributed infra graph : Models system connections down to the rack, fabric, power, and cooling layers. By integrating these connections with telemetry signals, the platform can precisely trace events to identify their blast radius and root cause.

  • A reconciliation core : workflow engine, policy engine, and state reconciler that continuously close the gap between intended state and reality, exposed through the API gateway.

  • Domain services : Services spanning provisioning, firmware upgrade, validation, deployment, repair and RMA, capacity, power and thermal, and Day-2 operations. Built once, run fleet-wide, consumable by any team through APIs and SDKs.

We operate on a continuous autonomy loop—sense, correlate, reason, act, learn—incorporating guardrails that evolve from recommendation to full automation. We treat every recurring manual intervention as a signal to engineer the next automation.

You’ll thrive here if you

  • Want to build a platform, not integrate one. This is core distributed-systems engineering: event buses, graph models, reconciliation loops, policy evaluation.

  • Treat internal engineers as customers and sweat API ergonomics, docs, and onboarding the way product teams sweat UX.

  • Like owning a hard abstraction and defending it as ten teams build on top of you.

  • Believe the interesting problems are where physical infrastructure meets software: a firmware counter, a thermal event, and a scheduling decision are one problem, not three.

  • Measure yourself by what stops paging humans, and by how fast another team ships on your platform.

What you’ll do

  • Design and build core platform services: RBAC, tenancy, the workflow engine, policy engine, and state reconciler that drive fleet actions safely at scale.

  • Design the public face of the platform: the API gateway, resource model, and versioned API contracts that fleet, SRE, and product teams build against.

  • Build SDKs, workflow templates, and golden paths that make the platform self-serve, plus the developer portal and micro frontend framework that let teams bring their own UI surfaces.

  • Build the inventory and topology graph as the fleet’s source of intended truth, and the pipelines that keep it honest against reality (metadata drift is one of our top verified incident root causes; you will kill it).

  • Build site, GPU, and network agents and the event bus that moves fleet telemetry and commands reliably.

  • Deliver the platform roadmap: pilot site on the foundation layer, first site deployed entirely through the platform, zero-downtime firmware upgrades, first fully automated RMA, then 100K+ GPUs on platform with MTTD under 60 seconds and MTTR under 30 minutes.

  • Work with embedded engineers from fleet and production engineering who bring the operational scar tissue, and turn it into services other teams extend.

Requirements

  • 10+ years building distributed systems, control planes, or infrastructure platforms.

  • Strong software engineering skills in Go, Python, or Rust.

  • Experience building platforms other engineers consume: public or internal APIs, SDKs, or developer tooling with real adoption.

  • Depth in at least one of: workflow/orchestration engines (Temporal or similar), event-driven architectures, graph data models, policy/rules engines, or reconciliation-based control loops (Kubernetes operator patterns).

  • Experience running what you build: you have carried a pager for a platform other teams depend on.

  • Systems thinking across the hardware/software boundary.

  • Bachelor’s degree in Computer Science, Data Science, or a closely related technical field.

Bonus experience

  • Internal developer platforms: API gateways, service catalogs, Backstage-style portals, micro frontend architectures.

  • GPU or bare-metal fleet infrastructure: DCGM, Redfish/IPMI, firmware lifecycle.

  • High-cardinality observability platforms (per-GPU telemetry at fleet scale).

  • InfiniBand or RoCE fabrics.

  • AI agents applied to infrastructure triage and autonomous remediation.

About CAPE

Vision. Crusoe’s infrastructure runs as a self-aware, self-healing system: anticipating and auto-remediating failures, shaping its own power demand, and tuning silicon-to-orchestration as one instrument. The world’s most reliable, efficient, and sustainable AI compute platform.

Mission. We design, build, and operate the world’s most reliable and energy-efficient AI infrastructure platform by treating the physical and digital layers as one software-defined system. Every day, for every workload, we automate away the latency, waste, and fragility between stranded energy and delivered intelligence.

Benefits:

  • Competitive compensation and equity packages

  • Restricted Stock Units

  • Paid time off, paid holidays & leave of absence programs

  • Comprehensive health, dental & vision insurance

  • Employer contributions to HSA account

  • Paid parental leave

  • Paid life insurance, short-term and long-term disability

  • Professional development & tuition reimbursement

  • Mental health & wellness support

  • Commuter benefits (parking & transit)

  • Cell phone stipend

  • 401(k) Retirement plan with company match up to 4% of salary

  • Volunteer time off

  • Global travel insurance & emergency assistance

  • Daily meals allowance

  • Additional perks & programs specific to location

Compensation Range

Compensation will be paid in the range of up to $250,000 - $300,000 + Bonus. Restricted Stock Units are included in all offers. Compensation to be determined by the applicant’s knowledge, education, and abilities, as well as internal equity and alignment with market data.

Crusoe is an Equal Opportunity Employer. Employment decisions are made without regard to race, color, religion, disability, genetic information, pregnancy, citizenship, marital status, sex/gender, sexual preference/ orientation, gender identity, age, veteran status, national origin, or any other status protected by law or regulation.

#J-18808-Ljbffr Jobleads-US
Vacancy posted 2 hours ago
Similar jobs that could be interesting for youBased on the Senior Staff Software Engineer, Cloud Availability Platform in Seattle, WA vacancy
  • $262k - $364k

     ...Senior Staff Software Engineer, Cloud SDK Share Senior Staff Software Engineer, Cloud SDK Google Kirkland...  ...benefits package, which is available to all eligible US based employees....  ...programmatic front door to 150+ Google Cloud Platform (GCP) APIs for virtually every... 
    Platform
    Senior
    Temporary work

    Jobleads-US

    Seattle, WA
    2 days ago
  • $180k - $240k

     ...Senior/Staff Software Engineer (Platform and Execution Model) Seattle, WA (Preferred) or McLean, VA or Remote...  ...Go, Rust, Java, or TypeScript) and cloud-native stacks (containers, CI/CD, IaC...  ...401K, FSA, and equity incentives available. ~ Mental health benefits are available... 
    Platform
    Senior
    Remote job
    Full time
    Shift work

    Jobleads-US

    Seattle, WA
    1 day ago
  • $236k - $325k

     ...working collaboratively with senior architects and ML team...  ...services and meet reliability, availability, and performance commitmentsCollaborate...  ...and/or machine learning platforms.Preferred: Strong...  ...PythonExperience with feature engineering platforms and ML platformsSnowflake... 
    Platform
    Senior

    Snowflake

    Bellevue, WA
    4 days ago
  • $254k - $350k

     ..., WA / Remote (United States)Software - Software Systems /Full-time...  ...executing this mission. The ML Platform team at Zoox plays a crucial...  ...a team of strong software engineers and act as a force multiplier...  ...application materials based on available information. These tools assist... 
    Platform
    Senior
    Full time
    Remote work

    Zoox

    Seattle, WA
    2 days ago
  • $230k - $315k

     .../ San Diego, CASoftware - Software Platforms and Product /Full-time /HybridWe are looking for a Senior/Staff Software Engineer to lead the development and scaling of our cloud services and APIs that handle...  ...materials based on available information. These tools assist... 
    Platform
    Senior
    Full time
    Temporary work
    Relocation package

    Zoox

    Seattle, WA
    3 days ago
  • $264k - $310k

     ...Senior Staff Software Engineer, Data Platform Bellevue, WA Join Us In Building The Future Of Finance Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest... 
    Platform
    Senior
    Work at office
    Flexible hours
    Shift work
    3 days per week

    Robinhood

    Bellevue, WA
    3 days ago
  •  ...the operating system for the AI datacenter, delivering a true platform rather than internal tooling. You will design scalable platform...  ..., API gateways, and developer tooling to empower fleet engineers and product teams across Crusoe''s energy-efficient AI compute... 
    Platform
    Senior

    Jobleads-US

    Seattle, WA
    2 hours ago
  • $262k - $364k

     ...package, which is available to all eligible...  ...experience in systems software engineering, software...  ..., OS kernels, or cloud infrastructure)....  ...strategy and mentoring senior engineers....  .... As a Senior Staff Software Engineer...  ...providing the essential platforms that enable... 
    Platform
    Senior
    Temporary work
    Worldwide

    Jobleads-US

    Seattle, WA
    2 hours ago
  •  ...a financial infrastructure platform for businesses. Millions of...  ...tooling that hundreds of Stripe engineers use to build world-class...  ...changing how users interact with software, and we believe Stripe is...  ...with care.What you'll doAs a Senior Staff Engineer on the Merchant... 
    Platform
    Senior

    Stripe

    Seattle, WA
    4 days ago
  • $220.4k - $297.4k

     ...world's best data and AI infrastructure platform, so our customers can focus on the high-...  ...that are central to their missions.Our engineering teams build highly technical products that...  ...from bad actors. We are looking for senior leaders such as yourselves to create the... 
    Platform
    Senior
    Local area
    Worldwide

    DataBricks

    Seattle, WA
    2 days ago
  • $209.1k - $282.9k

    As an engineer on Arm’s AI Inference Cloud team, you will shape the technical direction and develop highly available, scalable services for running AI inference...  ...usability of Arm’s AI platform.Responsibilities:Define...  ...management.Strong software and production engineering... 
    Platform
    Work at office
    Local area
    Visa sponsorship
    Relocation package

    ARM

    Seattle, WA
    12 hours ago
  •  ...is a financial infrastructure platform for businesses. Millions of...  ...build powerful interfaces for engineers that depend on these systems—...  ...Kubernetes—while keeping them highly available and performant. We're looking...  ...leaders, you will ensure the software your team builds meets the... 
    Platform
    Remote work

    Stripe

    Seattle, WA
    12 hours ago
  • $174k - $299k

     ...of the Observability Engineering team at Coupang to build...  ...Observability Platform based on Kubernetes and...  ..., as well as building software components from scratch...  ...SLOs) related to system availability, performance, and reliability...  ...certifications in cloud platforms, monitoring... 
    Platform
    Senior
    Temporary work
    Work experience placement
    Flexible hours

    Coupang

    Seattle, WA
    12 hours ago
  • $209.1k - $282.9k

    As a software Engineer on the AI Compute Platform team, you will design and build a secure, reliable, and easy-to...  ...experience building distributed systems, cloud platforms, or production backend...  ...where equal opportunities are available to all applicants and colleagues.... 
    Platform
    Work at office
    Local area
    Visa sponsorship
    Relocation package

    ARM

    Seattle, WA
    12 hours ago
  •  ...data and AI infrastructure platform, enabling our customers to focus on innovation.Our engineering teams push the boundaries of...  ...build impactful solutions.As a Senior Staff Software Engineer, you will define...  ...RBAC, GDPR, HIPAA), as well as cloud infrastructure (AWS, Azure,... 
    Platform
    Senior
    Worldwide

    DataBricks

    Seattle, WA
    2 days ago
  • $220k - $292k

     ...years.ABOUT THE TEAMThe Developer Platform team serves as the backbone of Anduril's engineering ecosystem, building critical...  ...CD workflows powering Anduril's Software Factories across the US,...  ...competitive benefits package (available at little to no cost to employees... 
    Platform
    Full time
    Work experience placement
    Immediate start

    Anduril Industries

    Seattle, WA
    2 days ago
  • $200k - $260k

     ...Senior/Staff Software Engineer In Office - Santa Clara, CA OR Seattle, WA About Us Orkes is a platform for developers to build durable, distributed event driven applications. Based...  ...systems (Kafka, RabbitMQ, etc.), and cloud-native architectures. Demonstrated... 
    Platform
    Senior
    Full time
    Work at office
    Remote work
    Flexible hours

    Orkes

    Seattle, WA
    2 days ago
  • $165k - $215k

     ...the only complete agentic AI platform for revenue teams. Outreach...  ...RoleJoin a small, high impact engineering team building the core...  ...You:8+ years of professional software development experience.Deep...  ...application materials based on available information. These tools assist... 
    Platform
    Full time
    Work at office
    Flexible hours
    Night shift

    Outreach

    Seattle, WA
    3 days ago
  • $170k - $250k

     ...CPAs, tax attorneys, and engineers, Taxbit is the leading...  .... Taxbit's AI-enabled platform streamlines compliance...  ...is looking for a Staff Software Engineer to join our Systems...  ...take ownership of the cloud infrastructure that...  ...secure, highly available cloud infrastructure on... 
    Platform
    Work at office
    Work from home

    Taxbit

    Seattle, WA
    3 days ago
  •  ....Snowflake runs large scale cloud infrastructure to deliver its...  ...self-serve cloud efficiency platform along with AI skills and...  ...optimization of our cloud spend.AS A SOFTWARE ENGINEER AT SNOWFLAKE YOU WILL:Design...  ...monitoring.Ensure high availability, reliability, and... 
    Platform
    Senior

    Snowflake

    Bellevue, WA
    2 days ago
  •  ...is only one Data Cloud. Snowflake's founders...  ...designed a data platform built for the...  ...stop there. They engineered Snowflake to power...  ...optimizes, provides high availability and data...  ...operates on.As a Senior Distributed Systems...  ...by applying your software engineering and... 
    Platform
    Senior

    Snowflake

    Bellevue, WA
    3 days ago
  • $174k - $299k

     ...systems. Building highly available and scalable systems to...  ...experienced and passionate engineer who can design and build...  ...the core Workflow Platform Infrastructure. (e.g., Workflow...  ...experience in backend software developmentExperience working in cloud environments,... 
    Platform
    Senior
    Temporary work

    Coupang

    Seattle, WA
    1 day ago
  • $110k - $135k

    We are seeking a Senior Cloud Engineer to join our Infrastructure Engineering team and lead the management...  ...infrastructure.What You’ll Do:Cloud Platform Management and Excellence* Manage and...  .... Reasonable accommodations are available for candidates during all aspects of the... 
    Platform
    Senior
    Temporary work
    Local area

    Slalom

    Seattle, WA
    12 hours ago
  •  ...Senior Cloud Engineer (Azure) Location: Seattle, WA Visa: GC or Citizen Or H1B Note: Its a Senior...  ...build an elastic infrastructure and platform used to power provisioning, deployment...  ...Build microservices that are highly available and fault tolerant Partner with various... 
    Platform
    Senior
    H1b

    Georgia IT Inc

    Seattle, WA
    1 day ago
  • $230k - $270k

     ...Staff Software Engineer, Storage Platform Bellevue, WA Join us in building the future of finance. Our mission is to democratize finance for...  ...millions of users and critical brokerage workloads. Availability is our highest priority — our systems are designed to... 
    Platform
    Work at office
    Flexible hours
    Shift work
    3 days per week

    Robinhood

    Bellevue, WA
    1 day ago
  •  ...Claude across multiple cloud service providers....  ...with internal engineering teams and cloud...  ...versions across cloud platforms. Develop cross-...  ...Significant software engineering experience...  ...policy requiring staff to work from an office...  ...sponsorship is available with reasonable efforts... 
    Platform
    Senior
    Full time
    Work at office
    Visa sponsorship
    Flexible hours

    Anthropic

    Seattle, WA
    11 days ago
  •  ...is the largest live shopping platform in North America and Europe...  ...Whatnot updates on our news and engineering blogs and join us as we...  ...is not a role we open often. Senior Staff Engineers at Whatnot operate...  ...you should have 10+ years of software engineering experience,... 
    Platform
    Senior
    Temporary work
    Work experience placement
    Work at office
    Local area
    Remote work
    Work from home

    Jobleads-US

    Seattle, WA
    12 hours ago
  • $148.5k - $223.9k

     ...future of Salesforce.Our Public Cloud engineering teams are responsible for...  ...systems engineering platform that ships hundreds of features...  ...craft solutions that are highly available, and a proven ability to design...  ...required4+ years backend software development experienceDeep knowledge... 
    Platform
    Senior
    Full time

    Salesforce

    Bellevue, WA
    12 hours ago
  • $160k - $180k

     ...SummaryAt the heart of our Cloud team is a mission to build a...  ...scalable, and high-performance platform that powers our cybersecurity products. As a Senior Software Engineer, Cloud, you will own the...  ...designing and scaling high-availability distributed systems with automated... 
    Platform
    Senior
    Permanent employment
    Remote work

    ExtraHop Networks

    Seattle, WA
    4 hours ago
  • $110.7k - $218.3k

     ...Summary Salesforce Life Sciences Cloud Senior ConsultantDeloitte's Sales & Service...  ...builtPartnering with solution architects and platform teams to translate future state process...  ...hiring and ban-the-box laws where available. Fair Chance Hiring and Ban-the-Box Notices... 
    Platform
    Senior
    Local area

    Deloitte

    Seattle, WA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Staff Software Engineer, Cloud Availability Platform. Be the first to apply!