Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer

$206k - $303k

CoreWeave

CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability. Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025. Learn more at What You’ll Do: The Common Services organization at CoreWeave is responsible for the shared platforms, APIs, and foundational services that power our AI cloud products and internal engineering teams. From authentication and authorization to core platform primitives and developer experience tooling, this organization ensures that the rest of CoreWeave can build, ship, and operate reliably at scale. As Reliability Lead, Common Services , you will establish and lead the Reliability Engineering and production operations practice for this organization. You’ll partner closely with engineering leaders and teams across Common Services to define how we build, release, monitor, and operate critical services—raising the bar on reliability, availability, and operational excellence across the board. About the Role: As Reliability Lead, Common Services , you will be responsible for defining the reliability strategy, processes, and standards for the Common Services portfolio and driving consistent, high-quality operational practices across multiple teams. You’ll monitor production incidents within Common Services, and work directly with your partner teams to design systems that are reliable, observable, and supportable. Your day-to-day will blend hands-on technical work and cross-functional leadership to drive continuous improvement of Common Services production operations. In this role, you will: Establish and lead the SRE / production engineering practice for the Common Services organization, including standards for reliability, incident management, and on-call, in partnership with the central Product Engineering organization. Develop an Operational Excellence strategy that focuses on not only improving system performance but also monitoring and reducing operational toil Partner with engineering and product teams to define SLOs, SLIs, and error budgets for critical Common Services, and ensure these become part of how teams plan and make tradeoffs. Own and improve the incident management lifecycle for Common Services, including on-call rotations, escalation paths, incident tooling, post-incident reviews, and follow-through on corrective actions. Drive the observability strategy (metrics, logs, traces, dashboards, alerts) for Common Services, ensuring we have actionable visibility into the health, performance, and capacity of key systems. Collaborate with engineering leads to design and review architectures for reliability, scalability, resilience, and operability, including failure modes, redundancy, and graceful degradation. Lead efforts to automate and harden operational workflows , including deployments, rollbacks, configuration management, change management, and routine maintenance tasks. Build strong, trust-based relationships with partner teams and stakeholders, becoming a go‑to leader for production readiness and operational risk within Common Services. Hire, mentor, and develop SRE and production engineering talent, fostering a culture of continuous improvement, learning from incidents, and humane on‑call . Partner with other SRE and production engineering leaders across CoreWeave to align on global practices, tools, and reliability goals , representing the needs and constraints of Common Services. Who You Are: 7+ years of experience in Site Reliability Engineering, Production Engineering, or similar roles working on distributed systems or cloud/platform services. 2+ years of technical leadership experience (team lead, staff/principal engineer, or people manager) where you drove reliability and operational improvements across multiple services or teams. Strong background in Linux-based production environments , containers, and orchestration technologies (e.g., Kubernetes), including debugging complex issues in live systems. Hands‑on experience with observability stacks (metrics, logging, tracing) and alerting systems, and a track record of designing meaningful SLIs/SLOs and alert strategies. Proven experience running on‑call rotations and incident response , including leading high‑severity incidents and driving high‑quality post‑incident reviews. Demonstrated ability to design for reliability (capacity planning, redundancy, failover, backoff, circuit breaking, graceful degradation, etc.) in large‑scale or mission‑critical systems. Comfortable working with infrastructure‑as‑code and automation tooling (e.g., Terraform, Ansible, Helm, CI/CD pipelines) to make operations repeatable, auditable, and safe. Strong cross‑functional communication skills—you can translate between engineering, product, and business stakeholders and influence without relying solely on authority. A bias toward data‑driven decision making , using production data, capacity signals, and incident trends to inform priorities and investments. Preferred: Background working with GPU workloads, high‑performance computing, or latency/throughput‑sensitive systems . Experience with multi‑tenant, multi‑region, or highly regulated environments , and the associated reliability considerations. Familiarity with service ownership models and strong opinions on how to align ownership, on‑call, and accountability in a scalable way. Experience mentoring or managing senior engineers and building high‑performing teams through coaching, feedback, and clear expectations. Wondering if you’re a good fit? We believe in investing in our people, and value candidates who can bring their own diversified experiences to our teams – even if you aren't a 100% skill or experience match. Here are a few qualities we’ve found compatible with our team. If some of this describes you, we’d love to talk. You care deeply about operational excellence and see reliability as a product feature, not an afterthought. You’re excited by the challenge of bringing order and clarity to complex, rapidly evolving systems . You’re passionate about building humane, sustainable on‑call practices and learning from incidents without blame. You enjoy partnering with multiple teams, influencing through context and clarity rather than authority alone. You’re curious about how to run large‑scale, GPU‑intensive workloads reliably and efficiently in production. Why CoreWeave? At CoreWeave, we work hard, have fun, and move fast! We’re in an exciting stage of hyper‑growth that you will not want to miss out on. We’re not afraid of a little chaos, and we’re constantly learning. Our team cares deeply about how we build our product and how we work together, which is represented through our core values: Be Curious at Your Core Act Like an Owner Empower Employees Deliver Best‑in‑Class Client Experiences Achieve More Together We support and encourage an entrepreneurial outlook and independent thinking. We foster an environment that encourages collaboration and provides the opportunity to develop innovative solutions to complex problems. As we get set for take off, the growth opportunities within the organization are constantly expanding. You will be surrounded by some of the best talent in the industry, who will want to learn from you, too. Come join us! The base salary range for this role is $206,000 to $303,000. The starting salary will be determined based on job‑related knowledge, skills, experience, and market location. We strive for both market alignment and internal equity when determining compensation. In addition to base salary, our total rewards package includes a discretionary bonus, equity awards, and a comprehensive benefits program (all based on eligibility). What We Offer The range we’ve posted represents the typical compensation range for this role. To determine actual compensation, we review the market rate for each candidate which can include a variety of factors. These include qualifications, experience, interview performance, and location. In addition to a competitive salary, we offer a variety of benefits to support your needs. The benefits below reflect our US‑based offerings for full‑time employees; for roles in other locations, benefits vary and are shared during the hiring process. These include: Medical, dental, and vision insurance - 100% paid for by CoreWeave Company‑paid Life Insurance Voluntary supplemental life insurance Short and long‑term disability insurance Flexible Spending Account Health Savings Account Tuition Reimbursement Ability to Participate in Employee Stock Purchase Program (ESPP) Mental Wellness Benefits through Spring Health Family‑Forming support provided by Carrot Paid Parental LeaveFlexible, full‑service childcare support with Kinside 401(k) with a generous employer match Flexible PTO Catered lunch each day in our office and data center locations A casual work environment A work culture focused on innovative disruption California Applicants California Consumer Privacy Act Equal Opportunity & Accommodations CoreWeave is an equal opportunity employer, committed to fostering an inclusive and supportive workplace. All qualified applicants and candidates will receive consideration for employment without regard to race, color, religion, sex, disability, age, sexual orientation, gender identity, national origin, veteran status, or genetic information. As part of this commitment and consistent with the Americans with Disabilities Act (ADA), CoreWeave will ensure that qualified applicants and candidates with disabilities are provided reasonable accommodations for the hiring process, unless such accommodation would cause an undue hardship. If reasonable accommodation is needed, please contact: View email address on click.appcast.io. Export Control Compliance This position requires access to export controlled information. To conform to U.S. Government export regulations applicable to that information, applicant must either be (A) a U.S. person, defined as a (i) U.S. citizen or national, (ii) U.S. lawful permanent resident (green card holder), (iii) refugee under 8 U.S.C. § 1157, or (iv) asylee under 8 U.S.C. § 1158, (B) eligible to access the export controlled information without a required export authorization, or (C) eligible and reasonably likely to obtain the required export authorization from the applicable U.S. government agency. CoreWeave may, for legitimate business reasons, decline to pursue any export licensing process. #J-18808-Ljbffr

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in California, MO vacancy
  •  ...Develop services for identity, secrets, tenant configuration, and customer bridging. Requirements 7+ years of production software engineering experience, including 2 or more years operating what you built (real on‑call experience, not just shipping code). Production‑... 
    Suggested
    Contract work

    Jobtailor

    California, MO
    1 hour ago
  • $149.4k - $202k

     ...Senior Software Engineer- Site Reliability Engineering (SRE) DC, MD, VA, CA The Site Reliability Engineering discipline at Noctua Technology, LLC is a strategic force driving digital transformation. We treat operations as a software engineering challenge, focusing... 
    Suggested
    Remote work

    Noctua Technology

    California, MO
    2 days ago
  •  ...Design, and GTM stakeholders to translate ambiguous business goals into a clear, phased engineering plan. Establish coding standards, CI/CD pipelines, testing strategy, reliability targets, and observability practices that future teams will inherit. Rapidly prototype with... 
    Suggested

    Jobtailor

    California, MO
    5 days ago
  • Agility Robotics in the United States is seeking a software developer to create robust software for our innovative humanoid robots. In this hybrid role, you will collaborate with multi-disciplinary teams to design and troubleshoot cutting-edge prototypes. Candidates should...
    Suggested
    Full time

    Agility Robotics

    California, MO
    1 hour ago
  • $176.4k - $226.8k

     .... The Position Saildrone is seeking a Senior Robotics Software Engineer to join our Core Vehicle Systems team. You will play a critical...  ...raises the bar on software quality for systems that must perform reliably for up to a year in the world’s harshest maritime conditions.... 
    Suggested
    Local area
    Relocation package
    Flexible hours
    3 days per week

    Saildrone Inc

    California, MO
    5 days ago
  • $170k - $250k

     ...Systems was founded by Stanford researchers and veteran systems engineers who share a vision for redefining the foundations of...  ...infrastructure struggles to meet the demands of performance, reliability, and precise coordination. Clockwork is pioneering a software-... 

    Clockwork Inc

    California, MO
    1 hour ago
  •  ...A leading Voice AI company in Missouri is seeking a Pre-Sales Solutions Engineer to drive technical pre-sales engagements. This role requires collaborating with sales teams to provide technical solutions and build proof-of-concepts. Ideal candidates will have a strong... 

    Deepgram

    California, MO
    1 hour ago
  •  ...Role: We’re looking for a driven and innovative Senior Software Engineer, experienced in distributed databases to help shape the future...  ...refining database architecture, and ensuring the system is robust and reliable for large-scale, mission-critical applications. What you'll own... 
    Remote work
    Flexible hours

    Authzed, Inc.

    California, MO
    3 days ago
  • $130.33k - $195.5k

     ...technology research, proof‑of‑concept work, and design that guides overall system and product enhancement. Contribute to software engineering best practices covering design, coding standards, performance, security, delivery, and maintainability. Qualifications Minimum... 

    E2E Alignment Healthcare USA, LLC

    California, MO
    1 hour ago
  • $168k - $270.25k

     ...premises environments. Collaborating with engineering teams across NVIDIA to deliver...  ...business outcomes. Continuously improving the reliability, security, maintainability, and operational...  ...systems, infrastructure, platform, site reliability, or cloud engineering. Deep... 

    NVIDIA Gruppe

    California, MO
    5 days ago
  • $150k - $175k

     ...converting greenhouse gas into diamond wafers using zero-emission energy. We are looking for a detail-oriented, hands-on Software Engineer to strengthen and grow our data collection, software interfaces, and infrastructure, which we use to report key business metrics and... 
    Work experience placement
    Local area
    Flexible hours

    Diamond Foundry

    California, MO
    1 hour ago
  • $10k

     ...Senior Solutions Engineer As the largest pure‑play fiber provider in the U.S., we deliver blazing‑fast broadband connectivity that unlocks the potential of millions of consumers and businesses. As a Frontier employee, you will be part of our purpose of Building Gigabit... 
    Work at office

    Frontier Communications

    California, MO
    1 day ago
  • $177.5k - $253k

     ...the Fortune 100, trust the Netskope One platform, its Zero Trust Engine, and the powerful NewEdge network to gain full visibility and...  ...Proxy, NGFW, Security Gateways, working with remote access and site to site VPN technologies, SAML/SSO, DLP, Data security and understand... 
    Remote work

    Netskope

    California, MO
    1 hour ago
  • $200k - $250k

     ...in DDoS mitigation, application security, and delivery solutions for multi-cloud environments, is seeking a highly motivated Sales Engineer (SE) to join its rapidly expanding team. As businesses face an unprecedented surge in sophisticated cyberattacks, including AI-... 
    Temporary work
    Flexible hours

    Radware

    California, MO
    18 hours ago
  • $215.69k - $287.59k

     ...Technical Solutions Engineer (AI / Cloud) – California (U.S. Preferred) Compensation: US$215,692 – US$287,590 Full‑Time or Priority Part...  ...specifications for global R&D and product teams Deliver remote or on‑site technical support for AI/LLM and cloud platform deployments... 
    Full time
    Part time
    Remote work
    Worldwide

    RGH-Global Ltd

    California, MO
    2 days ago
  • $125k - $145k

     ...scalable technology solutions? We're looking for a Solutions Engineer who can bridge business and technology by designing, building,...  ...DEVELOPMENT Implement testing methodologies to ensure system reliability and quality. Develop test plans and perform unit testing of solutions... 
    Work at office
    Local area

    Munger Tolles & Olson

    California, MO
    1 hour ago
  •  ...meaningful virtual support to patients across an expansive array of specialties, in all 50 states. About the Role As a Senior Solutions Engineer , you’ll serve as the technical bridge between our product, engineering, and customer teams. You’ll design and implement end-to-... 
    Flexible hours

    OpenLoop Health, Inc.

    California, MO
    2 days ago
  • $106.4k - $133k

     ...Thinking. It’s how we stay driven, supportive, and always one step ahead as AI reshapes our world. Why this role? As a Solutions Engineer, you'll be the technical voice of the sales team — running discovery, delivering demos, and leading proof-of-concept evaluations that... 
    Work at office
    Work from home
    Flexible hours

    Clutch Canada

    California, MO
    2 days ago
  • $215k - $323k

     ...This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Senior Solution Sales Engineer - Position Description: We believe Solutions Engineers at Okta are involved in all stages of the customer’s development lifecycle... 
    Work at office
    Local area
    Worldwide
    Flexible hours

    Empleora

    California, MO
    5 days ago
  •  ...Canary Data is seeking a versatile Software Engineer to own features end-to-end across our tech stack. You will expand our authenticated web applications, design high-throughput data processing systems, and build intuitive interfaces that present AI-driven financial insights... 

    Canary Data

    California, MO
    1 hour ago
  • Northrop Grumman seeks a Sr Principal Software Configuration Analyst to design, implement, and maintain software configuration management using IBM Rational ClearCase in a fast-paced defense environment. You will own ClearCase environments, define branching strategies,...

    Relha LLC

    California, MO
    5 days ago
  •  ...A globally leading consumer device company headquartered in Cupertino, CA is seeking a senior Keynote Shader & Effects Engineer who lives and breathes realtime graphics to own the design, implementation, and optimization of the shaders and visual effects that power our... 

    OSI Engineering

    California, MO
    2 days ago
  • $184k - $287.5k

     ...Join the NVIDIA Developer Tools team and empower engineers throughout the world developing groundbreaking products in Automotive, VR,...  ...realistic delivery schedule. Write fast, effective, maintainable, reliable and well-documented code. Provide peer reviews to other... 

    NVIDIA

    California, MO
    1 hour ago
  • $178k - $210k

     ...Webflow will need grit, because we move fast, without ever sacrificing craft or quality. We’re looking for a Senior Partner Solutions Engineer (PSE) that will be a strategic technical enablement partner focused on building, validating, and scaling the technical... 
    Ongoing contract
    Permanent employment
    Full time
    Temporary work
    Fixed term contract
    Remote work
    Flexible hours

    Socket

    California, MO
    5 days ago
  • $131.25k

     ...Remote-CA: Remote-WA: Remote-CO: Remote-NY: Remote-OR(PT)posted on: Posted 25 Days Agojob requisition id: JR28401## **Senior Software Engineer**WME is building the next generation of internal platforms that power one of the world's leading entertainment companies.We're... 
    Temporary work
    Local area
    Remote work

    IMG Live

    California, MO
    3 days ago
  • $124k - $195.5k

     ...We are now looking for a Deep Learning Software Engineer, TensorRT Performance! NVIDIA is seeking an experienced Deep Learning Engineer passionate about analyzing and improving the performance of NVIDIA’s inference ecosystem! NVIDIA is rapidly growing our research and... 

    NVIDIA

    California, MO
    1 hour ago
  •  ...Join to apply for the Senior Software Engineer role at Gigster Building the worlds most elite network of unicorn talent in tech ...AI...  ...managing databases and data pipelines , and improving system reliability and scalability through smart automation. While the primary focus... 
    Contract work

    Gigster

    California, MO
    1 hour ago
  •  ...Jobtailor is seeking an experienced AI software engineer to design, develop, and deploy AI-enabled enterprise solutions. You will work with product and data science teams to translate requirements into scalable software using Python and modern AI frameworks. The role... 

    Jobtailor

    California, MO
    1 hour ago
  •  ...Overview Junior Software Engineer role at Consertus. Responsibilities Develop and maintain responsive, intuitive web applications. Design...  .... Troubleshoot and resolve production issues to ensure system reliability. Create clear, detailed technical documentation, including... 
    Local area

    Consertus

    California, MO
    2 days ago
  • $180k - $200k

     ...the age of AI. What you\'ll do: We are hiring a Senior Software Engineer to own the design, implementation, and operation of...  ...scalable, highly available environments. Design systems for scale, reliability, and high availability, and take accountability for the operational... 
    Summer work
    Work at office
    Remote work
    Flexible hours

    ZEFR

    California, MO
    1 hour ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!