Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer

$206k - $303k

CoreWeave

CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability. Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025. Learn more at What You’ll Do: The Common Services organization at CoreWeave is responsible for the shared platforms, APIs, and foundational services that power our AI cloud products and internal engineering teams. From authentication and authorization to core platform primitives and developer experience tooling, this organization ensures that the rest of CoreWeave can build, ship, and operate reliably at scale. As Reliability Lead, Common Services , you will establish and lead the Reliability Engineering and production operations practice for this organization. You’ll partner closely with engineering leaders and teams across Common Services to define how we build, release, monitor, and operate critical services—raising the bar on reliability, availability, and operational excellence across the board. About the Role: As Reliability Lead, Common Services , you will be responsible for defining the reliability strategy, processes, and standards for the Common Services portfolio and driving consistent, high-quality operational practices across multiple teams. You’ll monitor production incidents within Common Services, and work directly with your partner teams to design systems that are reliable, observable, and supportable. Your day-to-day will blend hands-on technical work and cross-functional leadership to drive continuous improvement of Common Services production operations. In this role, you will: Establish and lead the SRE / production engineering practice for the Common Services organization, including standards for reliability, incident management, and on-call, in partnership with the central Product Engineering organization. Develop an Operational Excellence strategy that focuses on not only improving system performance but also monitoring and reducing operational toil Partner with engineering and product teams to define SLOs, SLIs, and error budgets for critical Common Services, and ensure these become part of how teams plan and make tradeoffs. Own and improve the incident management lifecycle for Common Services, including on-call rotations, escalation paths, incident tooling, post-incident reviews, and follow-through on corrective actions. Drive the observability strategy (metrics, logs, traces, dashboards, alerts) for Common Services, ensuring we have actionable visibility into the health, performance, and capacity of key systems. Collaborate with engineering leads to design and review architectures for reliability, scalability, resilience, and operability, including failure modes, redundancy, and graceful degradation. Lead efforts to automate and harden operational workflows , including deployments, rollbacks, configuration management, change management, and routine maintenance tasks. Build strong, trust-based relationships with partner teams and stakeholders, becoming a go‑to leader for production readiness and operational risk within Common Services. Hire, mentor, and develop SRE and production engineering talent, fostering a culture of continuous improvement, learning from incidents, and humane on‑call . Partner with other SRE and production engineering leaders across CoreWeave to align on global practices, tools, and reliability goals , representing the needs and constraints of Common Services. Who You Are: 7+ years of experience in Site Reliability Engineering, Production Engineering, or similar roles working on distributed systems or cloud/platform services. 2+ years of technical leadership experience (team lead, staff/principal engineer, or people manager) where you drove reliability and operational improvements across multiple services or teams. Strong background in Linux-based production environments , containers, and orchestration technologies (e.g., Kubernetes), including debugging complex issues in live systems. Hands‑on experience with observability stacks (metrics, logging, tracing) and alerting systems, and a track record of designing meaningful SLIs/SLOs and alert strategies. Proven experience running on‑call rotations and incident response , including leading high‑severity incidents and driving high‑quality post‑incident reviews. Demonstrated ability to design for reliability (capacity planning, redundancy, failover, backoff, circuit breaking, graceful degradation, etc.) in large‑scale or mission‑critical systems. Comfortable working with infrastructure‑as‑code and automation tooling (e.g., Terraform, Ansible, Helm, CI/CD pipelines) to make operations repeatable, auditable, and safe. Strong cross‑functional communication skills—you can translate between engineering, product, and business stakeholders and influence without relying solely on authority. A bias toward data‑driven decision making , using production data, capacity signals, and incident trends to inform priorities and investments. Preferred: Background working with GPU workloads, high‑performance computing, or latency/throughput‑sensitive systems . Experience with multi‑tenant, multi‑region, or highly regulated environments , and the associated reliability considerations. Familiarity with service ownership models and strong opinions on how to align ownership, on‑call, and accountability in a scalable way. Experience mentoring or managing senior engineers and building high‑performing teams through coaching, feedback, and clear expectations. Wondering if you’re a good fit? We believe in investing in our people, and value candidates who can bring their own diversified experiences to our teams – even if you aren't a 100% skill or experience match. Here are a few qualities we’ve found compatible with our team. If some of this describes you, we’d love to talk. You care deeply about operational excellence and see reliability as a product feature, not an afterthought. You’re excited by the challenge of bringing order and clarity to complex, rapidly evolving systems . You’re passionate about building humane, sustainable on‑call practices and learning from incidents without blame. You enjoy partnering with multiple teams, influencing through context and clarity rather than authority alone. You’re curious about how to run large‑scale, GPU‑intensive workloads reliably and efficiently in production. Why CoreWeave? At CoreWeave, we work hard, have fun, and move fast! We’re in an exciting stage of hyper‑growth that you will not want to miss out on. We’re not afraid of a little chaos, and we’re constantly learning. Our team cares deeply about how we build our product and how we work together, which is represented through our core values: Be Curious at Your Core Act Like an Owner Empower Employees Deliver Best‑in‑Class Client Experiences Achieve More Together We support and encourage an entrepreneurial outlook and independent thinking. We foster an environment that encourages collaboration and provides the opportunity to develop innovative solutions to complex problems. As we get set for take off, the growth opportunities within the organization are constantly expanding. You will be surrounded by some of the best talent in the industry, who will want to learn from you, too. Come join us! The base salary range for this role is $206,000 to $303,000. The starting salary will be determined based on job‑related knowledge, skills, experience, and market location. We strive for both market alignment and internal equity when determining compensation. In addition to base salary, our total rewards package includes a discretionary bonus, equity awards, and a comprehensive benefits program (all based on eligibility). What We Offer The range we’ve posted represents the typical compensation range for this role. To determine actual compensation, we review the market rate for each candidate which can include a variety of factors. These include qualifications, experience, interview performance, and location. In addition to a competitive salary, we offer a variety of benefits to support your needs. The benefits below reflect our US‑based offerings for full‑time employees; for roles in other locations, benefits vary and are shared during the hiring process. These include: Medical, dental, and vision insurance - 100% paid for by CoreWeave Company‑paid Life Insurance Voluntary supplemental life insurance Short and long‑term disability insurance Flexible Spending Account Health Savings Account Tuition Reimbursement Ability to Participate in Employee Stock Purchase Program (ESPP) Mental Wellness Benefits through Spring Health Family‑Forming support provided by Carrot Paid Parental LeaveFlexible, full‑service childcare support with Kinside 401(k) with a generous employer match Flexible PTO Catered lunch each day in our office and data center locations A casual work environment A work culture focused on innovative disruption California Applicants California Consumer Privacy Act Equal Opportunity & Accommodations CoreWeave is an equal opportunity employer, committed to fostering an inclusive and supportive workplace. All qualified applicants and candidates will receive consideration for employment without regard to race, color, religion, sex, disability, age, sexual orientation, gender identity, national origin, veteran status, or genetic information. As part of this commitment and consistent with the Americans with Disabilities Act (ADA), CoreWeave will ensure that qualified applicants and candidates with disabilities are provided reasonable accommodations for the hiring process, unless such accommodation would cause an undue hardship. If reasonable accommodation is needed, please contact: View email address on click.appcast.io. Export Control Compliance This position requires access to export controlled information. To conform to U.S. Government export regulations applicable to that information, applicant must either be (A) a U.S. person, defined as a (i) U.S. citizen or national, (ii) U.S. lawful permanent resident (green card holder), (iii) refugee under 8 U.S.C. § 1157, or (iv) asylee under 8 U.S.C. § 1158, (B) eligible to access the export controlled information without a required export authorization, or (C) eligible and reasonably likely to obtain the required export authorization from the applicable U.S. government agency. CoreWeave may, for legitimate business reasons, decline to pursue any export licensing process. #J-18808-Ljbffr

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in California, MO vacancy
  •  ...Develop services for identity, secrets, tenant configuration, and customer bridging. Requirements 7+ years of production software engineering experience, including 2 or more years operating what you built (real on‑call experience, not just shipping code). Production‑... 
    Suggested
    Contract work

    Jobtailor

    California, MO
    1 day ago
  • $149.4k - $202k

     ...Senior Software Engineer- Site Reliability Engineering (SRE) DC, MD, VA, CA The Site Reliability Engineering discipline at Noctua Technology, LLC is a strategic force driving digital transformation. We treat operations as a software engineering challenge, focusing... 
    Suggested
    Remote work

    Noctua Technology

    California, MO
    2 days ago
  •  ...Define and execute the long-term Site Reliability Engineering strategy and roadmap Establish the SRE operating model, team scope, engagement models, ownership boundaries, and success measures Build and develop a high-performing team of site reliability and operations... 
    Suggested
    Immediate start

    Jobtailor

    California, MO
    2 days ago
  • $149.4k - $202k

    Senior Software Engineer- Site Reliability Engineering (SRE) DC, MD, VA, CA The Site Reliability Engineering discipline at Noctua Technology, LLC is a strategic force driving digital transformation. We treat operations as a software engineering challenge, focusing on... 
    Suggested
    Remote work

    Noctua Technology

    California, MO
    2 days ago
  •  ...Southern California Edison is seeking a Principal Manager, System Operations to lead a large team in the safe, reliable operation of the transmission and distribution system in California. The role oversees about 150 employees, manages outages, and guides strategy for... 
    Suggested

    Thomson Reuters Markets Espana SL.

    California, MO
    2 days ago
  • $170k - $277k

     ...Job Summary We are seeking a Senior Principal Software Engineer who is first and foremost a software craftsman with a passion...  ...developer productivity, operational excellence, and platform reliability. If you thrive on solving complex technical challenges, influencing... 
    Full time
    Work at office
    Visa sponsorship
    Work visa

    Socket.dev

    California, MO
    3 days ago
  •  ...test and optimize CUDA/C++ libraries across different platforms Requirements ~ BS, MS, or PhD in Computer Science, Computer Engineering, or closely related field (or equivalent experience) ~15+ years of work experience in software development ~ Outstanding technical... 
    Work experience placement

    Jobleads-US

    California, MO
    2 days ago
  •  ...Jobtailor is seeking a senior software engineer to develop Java, Scala and CUDA/C++ libraries to accelerate DataFrames and I/O on Parquet, ORC and JSON formats. You will collaborate with open source communities and distributed systems teams to optimize performance... 

    Jobleads-US

    California, MO
    2 days ago
  •  ...distributed key-value storage Work fluidly between writing production code, designing new systems, and reviewing proposals of senior engineers Help integrate LLMs into processes and systems Deliver on key Health Mediated Deployment projects Requirements ~10+... 

    Jobleads-US

    California, MO
    2 days ago
  •  ...SAIC is seeking a Senior Principal Software Engineer to support the Space Development Agency (SDA) in delivering BMC3 software and hosting environments for the Proliferated Warfighter Space Architecture (PWSA). The role focuses on end-to-end software development, secure... 
    Remote job

    Jobleads-US

    California, MO
    2 days ago
  • $144k - $236k

     ...push the boundaries of scaling large models together. The team is responsible for scaling LinkedIn’s AI model training, feature engineering and serving with hundreds of billions of parameters models and large scale feature engineering infra for all AI use cases from recommendation... 
    For contractors
    Work experience placement
    Work at office
    Flexible hours

    LinkedIn

    California, MO
    2 days ago
  • $215.69k - $287.59k

     ...Technical Solutions Engineer (AI / Cloud) – California (U.S. Preferred) Compensation: US$215,692 – US$287,590 Full‑Time or Priority Part...  ...specifications for global R&D and product teams Deliver remote or on‑site technical support for AI/LLM and cloud platform deployments... 
    Full time
    Part time
    Remote work
    Worldwide

    RGH-Global Ltd

    California, MO
    2 days ago
  • $250k

     ...LoopNet - Vice President, Software Engineering Job Description Overview CoStar Group is a leading global provider of commercial and...  ...located in our Irvine, CA office and has a schedule of 4 days on‑site and 1 day work from home. Responsibilities Act as a... 
    Full time
    Work at office
    Work from home

    CoStar Group, Inc.

    California, MO
    3 days ago
  •  ...full product suite is essential to how this team operates. Run scoping calls to gather requirements, set timelines, and keep Sales, Engineering, and Customer Success aligned. Use AI tools to build and automate implementation workflows, helping customers go live faster and... 

    Jobtailor

    California, MO
    2 days ago
  •  ...meaningful virtual support to patients across an expansive array of specialties, in all 50 states. About the Role As a Senior Solutions Engineer , you’ll serve as the technical bridge between our product, engineering, and customer teams. You’ll design and implement end-to-... 
    Flexible hours

    OpenLoop Health, Inc.

    California, MO
    2 days ago
  • $125k - $145k

     ...scalable technology solutions? We're looking for a Solutions Engineer who can bridge business and technology by designing, building,...  ...DEVELOPMENT Implement testing methodologies to ensure system reliability and quality. Develop test plans and perform unit testing of solutions... 
    Work at office
    Local area

    Munger Tolles & Olson

    California, MO
    1 day ago
  •  ...A leading Voice AI company in Missouri is seeking a Pre-Sales Solutions Engineer to drive technical pre-sales engagements. This role requires collaborating with sales teams to provide technical solutions and build proof-of-concepts. Ideal candidates will have a strong... 

    Deepgram

    California, MO
    1 day ago
  •  ...EnerSys in the United States is seeking a Solutions Engineer – BESS to lead technical design, economic modeling, and deployment support across industrial, commercial, telecom, and motive power markets. You will bridge commercial strategy with engineering execution, delivering... 

    PowerToFly

    California, MO
    1 day ago
  •  ...About the opportunity Work Model: Hybrid Location: US – California Job Description Correct Title: Solutions Engineer, Expert Candidate must reside in California. Core Responsibilities: Support the maturation and day‑to‑day operationalization of the ransomware recovery... 

    Avenue Code

    California, MO
    5 days ago
  • $170k - $250k

     ...Systems was founded by Stanford researchers and veteran systems engineers who share a vision for redefining the foundations of...  ...infrastructure struggles to meet the demands of performance, reliability, and precise coordination. Clockwork is pioneering a software-... 

    Clockwork Inc

    California, MO
    1 day ago
  •  ...Responsibilities Develop and maintain material tracking and data collection systems Coordinate with a broad range of stakeholders, including R&D engineers and technicians, to deliver scalable and actionable software solutions Assist in database optimization and application... 

    NextGenEnergyJobs

    California, MO
    4 days ago
  • $141.3k - $226k

     ...design and development of highly scalable, reliable, and performant distributed services...  ...setting an example for code quality and engineering excellence. Identify and resolve complex...  ...audiences. This position is 100% on-site located in Palo Alto, CA. Must have... 
    Local area

    Broadcom Corporation

    California, MO
    2 days ago
  • Account Executive Drive sales efforts. Role based in California and focused on sales growth. Senior Manager of Workforce Management Primary function involves providing customer support services. Specific details regarding responsibilities and requirements are not fully...

    Oppurtuna

    California, MO
    1 day ago
  •  ...Senior Software Engineer 56103 Cephas Consultancy Services Private Limited, LATAM/North America/Europe, California, United States About this position Positions: 100 Contract Experience: 3 - 30 Years Senior Software Engineer Position Overview We are... 
    Hourly pay
    Contract work
    For contractors
    Immediate start
    Remote work
    Flexible hours

    Cephas Consultancy Services Private Limited

    California, MO
    2 days ago
  •  ...across the US and are building the world's largest network of on-demand aircraft. We’re looking for a full stack Senior Software Engineer to join our growing engineering team. You’ll work on all areas of the JetInsight product and help solve challenging problems involving... 

    JetInsight

    California, MO
    5 hours ago
  • $106.4k - $133k

     ...Thinking. It’s how we stay driven, supportive, and always one step ahead as AI reshapes our world. Why this role? As a Solutions Engineer, you'll be the technical voice of the sales team — running discovery, delivering demos, and leading proof-of-concept evaluations that... 
    Work at office
    Work from home
    Flexible hours

    Snyk, Inc.(USA)

    California, MO
    2 days ago
  •  ...converting greenhouse gas into diamond wafers using zero-emission energy. We are looking for a detail-oriented, hands-on Software Engineer to strengthen and grow our data collection, software interfaces, and infrastructure, which we use to report key business metrics and... 
    Local area
    Flexible hours

    Vrai & Oro

    California, MO
    4 days ago
  • $140k - $210k

     ...Systems was founded by Stanford researchers and veteran systems engineers who share a vision for redefining the foundations of...  ...infrastructure struggles to meet the demands of performance, reliability, and precise coordination. Clockwork is pioneering a software-... 

    Clockwork Inc

    California, MO
    2 days ago
  •  ...Software Engineer Python Cloud Services Space Operations Satellite Telemetry As a Space Operations Software Engineer at Spire, you will join a small, agile team to enhance the efficiency of our satellite fleet. You'll develop systems and tools for visualizing and analyzing... 
    Work at office
    3 days per week

    Spire

    California, MO
    2 days ago
  • $143.2k - $243.4k

     ...infrastructure and provide a seamless digital end-user experience. Our services has a strong emphasis on high-scale and high-performance engineering. We operate under a DevOps model and team is spread in US & India. In addition to core infrastructure services, the team... 
    Permanent employment
    Work at office
    Remote work
    Flexible hours

    ServiceNow

    California, MO
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!