Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff Network Reliability Engineer - Scale & Incident Response

$195k - $235k

Crusoe Energy Systems LLC

Crusoe Energy Systems LLC is looking for a Staff Network Operations Engineer to ensure production reliability across its global network infrastructure. This role is critical in maintaining uptime and facilitating AI workloads via incident response and operational excellence. The ideal candidate has 8+ years of experience in network engineering, specializing in operations and incident response. You'll work with advanced monitoring tools and help shape the future of AI infrastructure. Compensation ranges from $195,000 to $235,000, plus bonuses and stock options. #J-18808-Ljbffr Crusoe Energy Systems LLC

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Staff Network Reliability Engineer - Scale & Incident Response in San Francisco, CA vacancy
  • $250k

     ...team of hardcore engineers, designers, and researchers...  ...the tooling and response loops that keep us safe as we scale. You’ll partner...  ...and enforce network topology standards...  ...observability, detection, and incident response for...  ...in risk, reliability, or incident outcomes... 
    Suggested

    XOXO AI

    San Francisco, CA
    2 hours ago
  • $210k - $240k

     ...to perform under real‑world scale, reliability, and security demands - and we’re looking for an engineer who wants to own the...  ...design and operate the global network and reliability layer behind...  ...monitoring, alerting, and incident response - SLOs, runbooks, and on‑call... 
    Suggested

    Alembic Technologies

    San Francisco, CA
    12 hours ago
  • Delta Dental of California is seeking an Expert (Staff) Cyber Risk Management Engineer to evolve detection, response, and purple team initiatives. You will act as Incident Commander, build detections, and lead cross‑functional IR efforts to learn from adversary behaviors... 
    Suggested

    Delta Dental of California

    San Francisco, CA
    2 days ago
  • $182k - $250k

    Grow Therapy is looking for a Senior Platform Reliability Engineer to enhance reliability across teams. This role includes defining SLOs/SLAs, improving observability, and evolving incident response procedures. You will foster a culture of reliability while collaborating... 
    Suggested

    Grow Therapy

    San Francisco, CA
    12 hours ago
  • $150k - $300k

     ...Francisco, is seeking an Infrastructure Engineer to construct the foundational...  .... This role involves significant responsibilities regarding system reliability and security, requiring expertise...  ...background and experience managing large-scale infrastructure. Compensation is... 
    Suggested

    Sieve

    San Francisco, CA
    4 days ago
  • $250k - $350k

     ...persistent, and well-resourced anywhere. We are building Detection & Response Engineering from the ground up: engineering-led, agent-first, and built to scale across IT, OT, and physical surfaces. As the Staff Incident Responder, you are the most senior incident commander in the... 
    Contract work
    Local area

    Fluidstack

    San Francisco, CA
    12 hours ago
  • $218.03k - $256.5k

    Coinbase seeks an experienced software engineer for the Infra Reliability team in San Francisco, California. You will be responsible for scaling systems and improving service reliability, collaborating closely with teams to enhance deployment systems and manage configurations... 
    Remote work

    Coinbase

    San Francisco, CA
    12 hours ago
  • $182k - $250k

    About The Role We're hiring a Senior Platform Reliability Engineer to define and scale reliability as a first-class capability at Grow. In this role...  ...standards around observability, SLOs/SLAs, and incident response—while also helping translate those standards into self... 
    Full time
    Work at office
    Remote work
    Home office
    Flexible hours
    Day shift
    3 days per week

    Grow Therapy

    San Francisco, CA
    12 hours ago
  • $180k - $250k

     ...foundational infrastructure engineers who can help scale the platform. This is a...  ...Kubernetes, containers, networking, cloud infrastructure,...  ...workloads safely, quickly, and reliably across cloud environments...  ..., monitoring, alerting, incident response, deployment safety, and... 
    Full time
    Immediate start

    Crosscheck Staffing

    San Francisco, CA
    12 hours ago
  • $186.07k - $218.9k

     ...You’ll join the IT Operations Corporate Engineering team as a Senior Site Reliability Engineer focused on building and scaling Coinbase's identity and access management...  ...end-to-end, including on-call rotation, incident response, root cause analysis, and blameless retrospectives... 
    Local area

    Coinbase

    San Francisco, CA
    12 hours ago
  •  ...looking for a Infrastructure Engineer to take the lead on scaling our operational resilience as...  ...role where you’ll shape how reliability is done - reducing incident load, building internal tooling...  ...Ownership We take full responsibility for our work, outcomes, and team... 
    Worldwide
    Shift work

    HappyRobot

    San Francisco, CA
    12 hours ago
  •  ...less than 1 second and quickly scales out to thousands of GPUs....  ...Deploy and validate data center network infrastructure (front‑end,...  ...ICT, Hardware, and Network Engineering to identify blockers early,...  ...during and after deployments: incident response, troubleshooting, and break‑... 

    BEAM inc.

    San Francisco, CA
    2 days ago
  • $190k - $260k

     ...team of lawyers, engineers and research scientists...  ...fit and are scaling our team very...  ...infrastructure. Your responsibilities will include:...  ...Establish internal incident management, on-call...  ...understanding of networking, security...  ...practices Senior or Staff level Bonus: Experience... 

    Harvey

    San Francisco, CA
    12 hours ago
  • Waymo is seeking a backend software engineer to design, develop, test, and...  ...infrastructure powering real-world event responses. You will build and evolve tools to scale services and enable new markets,...  ...focus on distributed systems and reliable, scalable design in a fast-paced... 

    Waymo

    San Francisco, CA
    2 days ago
  • $342k

     ...OpenAI is currently looking for an experienced Optical Network Engineer based in San Francisco, California. This role involves leading laser...  ...-related efforts in optical interconnect projects for large-scale compute systems, ensuring performance and manufacturability of... 

    OpenAI

    San Francisco, CA
    1 day ago
  • $225k - $250k

     ...the world's largest engineering organizations and...  ...for: You’ll be responsible for building and scaling the infrastructure...  ...design for scale, reliability, and developer velocity...  ...Azure), including networking, security, and...  ...observability, and incident response. Compensation... 

    Céline

    San Francisco, CA
    12 hours ago
  •  ...inference delivery network for high-...  ...builders, architects, engineers, and researchers...  ...re looking for a Staff Edge Network Engineer...  .... Key Responsibilities Edge Architecture...  ...on‑call and lead incident response for edge...  ...incident response at scale Salary Range Why... 
    Work at office

    AI Fabrik

    San Francisco, CA
    1 day ago
  • $320k - $405k

     ...mission is to create reliable, interpretable,...  ...researchers, engineers, policy experts,...  ...team to build and scale our next‑generation...  ...security operations. Responsibilities Build AI‑powered...  ...development to incident response Design...  ..., we expect all staff to be in one of... 
    Visa sponsorship
    Shift work

    Anthropic

    San Francisco, CA
    12 hours ago
  •  ...finance at a global scale. Attributes We...  ...IT support, engineering and business...  ...management. Network Operations builds...  ...for all staff worldwide, and ensures a reliable, stable experience...  ...team. Responsibilities Build the foundations...  ...issues, running incidents to completion... 
    Work at office
    Remote work
    Worldwide

    airwallex

    San Francisco, CA
    12 hours ago
  • A technology solutions provider is looking for a Network Engineer to enhance and maintain a large-scale network. This role involves managing both wired and wireless infrastructures, conducting assessments, and ensuring network security. Candidates should have a degree... 

    CGS Federal (Contact Government Services)

    San Francisco, CA
    2 days ago
  • $250k - $320k

     ...of AI infrastructure: large-scale AI datacenters and the...  ...Role Gimlet Labs is seeking a Network Engineer to design, build, and scale...  ...operations teams to improve network reliability, deployment velocity,...  ...deployment validation, and incident response workflows. You may be a... 

    Gimlet Labs, Inc.

    San Francisco, CA
    4 days ago
  • $350k

    Thinking Machines Lab is seeking a Network Engineer in San Francisco to manage and improve our GPU network fabric. The role requires in-depth knowledge of large-scale deployments and the ability to debug complex network issues. A collaborative environment is emphasized,... 
    Visa sponsorship

    Thinking Machines Lab

    San Francisco, CA
    2 days ago
  • $150k - $250k

    As our Founding Security Reliability Engineer at Charta Health, you'll pioneer...  ...opportunity to build and scale the foundational security...  ...mitigation, and efficient incident response. You’ll be crucial in engineering...  ...(primarily AWS), including network security, identity and... 

    Charta Health, Inc.

    San Francisco, CA
    12 hours ago
  • $150k - $170k

     ...looking for an Integration Reliability Engineer to own the...  ...warehouses. This role is responsible for making systems observable...  ...and repeatable as we scale across deployments,...  ...Define and improve incident response, severity...  ...across infrastructure, networking, and distributed... 
    Permanent employment

    Claryo, Inc.

    San Francisco, CA
    2 days ago
  • Airwallex is seeking a Senior Site Reliability Engineer based in San Francisco to architect and implement scalable cloud infrastructure. You'll drive incident response and operational excellence in a collaborative Engineering team focused on innovation in financial technology... 

    Airwallex

    San Francisco, CA
    12 hours ago
  • $150k - $170k

     ...Inc. is seeking an Integration Reliability Engineer in San Francisco, CA, responsible for ensuring the reliability of...  ...observability tools and improve incident response processes. Qualifications...  ...experience in SRE, strong Linux and networking skills, and familiarity with... 

    Claryo, Inc.

    San Francisco, CA
    2 days ago
  • $160k - $220k

     ...Role Senior Database Reliability Engineer (DBRE) Experience...  ...expertise in PostgreSQL at scale and solid experience...  ...administering it. Responsibilities Architecture,...  ...MySQL. Operations & Incident Response: Lead response...  ...with Linux systems, networking fundamentals, and... 
    Permanent employment
    Work at office
    Local area
    Flexible hours

    Okta

    San Francisco, CA
    1 day ago
  • $195k - $235k

     ...urgency, who believe in the scale of our ambition and thrive on...  ...Role Crusoe Cloud is seeking a Staff Network Operations Engineer to help own production reliability across our global network...  ...production ownership role focused on incident response, root cause analysis, and... 
    Temporary work
    Worldwide

    Crusoe Energy Systems LLC

    San Francisco, CA
    12 hours ago
  • $325k - $360k

     ..., who believe in the scale of our ambition and thrive...  ...is building the network backbone for one of the...  ...Director of Network Engineering, you will own the...  ...infrastructure fast, reliable, and ready for the next...  ...proactive monitoring, incident response, and root‑cause analysis... 
    Temporary work

    Crusoe Energy Systems LLC

    San Francisco, CA
    1 day ago
  • $350k

     ...Site Reliability Engineer (SRE) San Francisco Thinking Machines Lab's mission is to empower...  ...to production observability and incident response. Develop appropriate Service Level...  ...operating production cloud services at scale (e.g., public cloud platforms, internal... 
    Local area
    Visa sponsorship
    Work visa
    Relocation package

    Thinking Machines Lab

    San Francisco, CA
    12 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff Network Reliability Engineer - Scale & Incident Response. Be the first to apply!