Site Reliability Engineer
$120k - $180kCrowdstrike
SRE & DevOps EngineerAt CrowdStrike, our engineering organization depends on shared infrastructure platforms that power critical product capabilities at global scale. These platforms require dedicated engineering ownership to operate reliably, scale safely, harden for security, and mature into self-service capabilities that teams across the organization can depend on.As an SRE & DevOps Engineer, you will own production infrastructure spanning multiple cloud providers and regions, serving engineering teams across CrowdStrike. The work is equal parts reliability engineering and DevOps engineering - building automation, hardening security, establishing governance, and enabling consuming teams to adopt these platforms effectively.You will work across a rich technology landscape including Kubernetes, Kafka, Cassandra, PostgreSQL, Apache Pinot, OpenSearch, and Apache Spark — operating and scaling microservices-based distributed systems that process millions of security events per second with zero tolerance for data loss or downtime.What You'll Do:Run production infrastructure - Deploy, upgrade, and maintain platform services across multiple clouds and regions on Kubernetes, including microservices-based distributed systemsOwn delivery pipelines - Build and maintain scalable CI/CD pipelines using Jenkins, GitLab CI, and Bitbucket Pipelines with GitOps workflows via ArgoCD or FluxOwn capacity planning - Track usage, forecast growth, right-size clusters, and optimize infrastructure costs across multi-cloud environmentsBuild observability - Implement metrics, dashboards, alerts (Prometheus/Grafana), distributed tracing (Jaeger/OpenTelemetry), and actionable runbooksOwn on-call and incidents - Participate in on-call rotation, lead incident resolution, write blameless postmortems, and automate repeat problemsDrive reliability - Apply SRE principles and AI-driven automation to move from reactive firefighting to proactive operationsHarden security - Implement auth, encryption, secret rotation, and network policies.Own disaster recovery - Build and test backup strategies and failover mechanisms ensuring zero data lossOperate data infrastructure - Maintain reliability of Kafka, Cassandra, PostgreSQL, Apache Pinot, and OpenSearch in productionEnable and collaborate - Support engineering teams with templates and patterns; partner with Infrastructure, SRE, and Data Services on shared operational problemsExperience and Background:8+ years in SRE, Devops engineering, or infrastructure engineeringHands-on experience running stateful distributed systems and microservices architectures on Kubernetes in productionBachelor's degree in Computer Science or related field, or equivalent work experienceReliability Engineering:Deep understanding of SRE principles - SLOs, SLIs, error budgets applied to large-scale distributed systemsStrong incident management background - on-call ownership, blameless postmortems, and turning operational pain into automationExperience with chaos engineering and resilience validation for production systemsProven ability to build and maintain systems with zero tolerance for data loss or downtimeAdvanced observability experience including Prometheus, Grafana, distributed tracing (Jaeger/OpenTelemetry), and large-scale log aggregation (ELK/Splunk) with a focus on building custom SLO dashboards and reliability scorecardsProgramming & Automation:Proficiency in Python and/or Golang for automation, tooling, and platform servicesStrong scripting and automation skills - if you do it by hand more than once, you automate itPlatform and Delivery Engineering:CI/CD pipeline experience - Building and owning scalable delivery pipelines using Jenkins, GitLab CI, Bitbucket Pipelines, Tekton, or equivalentStrong proficiency in Infrastructure as Code (IaC) — Terraform, Ansible, Pulumi, or equivalentCloud and Big Data Exposure:Proficiency in at least one cloud environment (AWS, Azure, GCP) with emphasis on multi-region architecture, cloud-native reliability patterns, and security-first cloud designStrong experience with Kubernetes at scale - managing large cluster fleets, workload orchestration, and container lifecycle managementFamiliarity with distributed data systems including relational databases (PostgreSQL), NoSQL (Cassandra), OLAP (Pinot), Indexing(OpenSearch) and real-time streaming platforms (Kafka, Flink)Exposure to Big Data and analytics technologies like Spark, StormBenefits of Working at CrowdStrike:Market leader in compensation and equity awardsComprehensive physical and mental wellness programsCompetitive vacation and holidays for rechargePaid parental and adoption leavesProfessional development opportunities for all employees regardless of level or roleEmployee Networks, geographic neighborhood groups, and volunteer opportunities to build connectionsVibrant office culture with world class amenitiesGreat Place to Work Certified™ across the globeCrowdStrike is proud to be an equal opportunity employer. We are committed to fostering a culture of belonging where everyone is valued for who they are and empowered to succeed. We support veterans and individuals with disabilities through our affirmative action program.CrowdStrike is committed to providing equal employment opportunity for all employees and applicants for employment. The Company does not discriminate in employment opportunities or practices on the basis of race, color, creed, ethnicity, religion, sex (including pregnancy or pregnancy-related medical conditions), sexual orientation, gender identity, marital or family status, veteran status, age, national origin, ancestry, physical disability (including HIV and AIDS), mental disability, medical condition, membership or activity in a local human rights commission, status with regard to public assistance, or any other characteristic protected by law. We base all employment decisions--including recruitment, selection, training, compensation, benefits, discipline, promotions, transfers, lay-offs, return from lay-off, terminations and social/recreational programs--on valid job requirements.If you need assistance accessing or reviewing the information on this website or need help submitting an application for employment or requesting an accommodation, please contact us at View email address on click.appcast.io for further assistance.Find out more about your rights as an applicant.CrowdStrike participates in the E-Verify program.Notice of E-Verify ParticipationRight to WorkCrowdStrike, Inc. is committed to fair and equitable compensation practices. Placement within the pay range is dependent on a variety of factors including, but not limited to, relevant work experience, skills, certifications, job level, supervisory status, and location. The base salary range for this position for all U.S. candidates is $120,000 - $180,000 per year, with eligibility for bonuses, equity grants and a comprehensive benefits package that includes health insurance, 401k and paid time off.For detailed information about the U.S. benefits package, please click here.
$168k - $270.25k
NVIDIA is looking for a Senior Site Reliability Engineer (SRE) to join its GeForce Now (GFN) team. SRE at NVIDIA ensures that our internal and external-facing GPU cloud gaming services have reliability and uptime as promised to the users and at the same time enables developers...SuggestedFull time$128.6k - $184.9k
...global cloud platform. As a team of six engineers distributed across the US, Canada, and the... ...with a strong focus on automation, reliability, and operational excellence. We are one... ...Qualifications7+ years of experience in Site Reliability Engineering, DevOps, Infrastructure...SuggestedPermanent employmentFull timeTemporary workLocal areaWorldwideFlexible hours$230k - $250k
...network. It's the foundation for autonomous networking, giving engineers and AI agents the ability to know the impact of every change... ...how things have always been done.Forward is looking for a Site Reliability EngineerAbout the Role This is not a "keep the lights on"...SuggestedNight shift$170k - $200k
We are seeking a talented and motivated Site Reliability Engineer to join our engineering team. You will be responsible for building, maintaining, and troubleshooting cloud service/cluster, infrastructure, and monitoring systems to ensure high availability, performance,...SuggestedFull timeWorldwide- LeanData helps the world’s fastest-growing companies automate, simplify, and accelerate revenue.We are looking for a Senior Site Reliability Engineer to lead the strategic evolution of our cloud infrastructure. Reporting directly to the SVP of Engineering, this role is...SuggestedFull timeWork at office2 days per week
$148k - $235.75k
...see how you can make a lasting impact on the world.Join our team of innovative engineers who are building an AI Data Center AIOps platform that turns raw, high-volume telemetry into reliable, job-centric insights and automation for GPU fleets. We’re hiring a DevOps Engineer...Full time$167.7k - $245.2k
...requiring approximately 2 days per week on-site at Cisco offices in either San Francisco... ...AI agents behave as intended, improving reliability and reducing risks. This unified... ...and control.As a Senior Site Reliability Engineer (SRE), you will build, operate, and continuously...Full timeTemporary workLocal areaFlexible hours2 days per week$90k - $180k
...nutritionals and branded generic medicines. Our 115,000 colleagues serve people in more than 160 countries.About the RoleThis Senior Site Reliability Engineer position works on-site out of our Sylmar, CA or Sunnyvale, CA location in the Cardiac Rhythm Management Division.We are...Remote work$152k - $241.5k
...infrastructure platforms for automated host lifecycle management, fleet reliability/auto-healing, E2E observability or data-driven operations (... ...languages such as Python, Go, Perl, or Ruby.Mentored other engineers and influenced technical direction through design reviews,...Full time$101k - $161k
...excellence has earned us several prestigious awards, such as Best Engineering Team, Best Company for Diversity, Compensation, and Work-... ...we do.Job DescriptionWho You'll Work WithWe’re looking for Site Reliability Engineers to join our growing Arista’s CloudVision-as-a-...$168k - $270.25k
...phenomenal people like you to help us accelerate the next wave of artificial intelligence.Join our team at NVIDIA as a Senior Site reliability engineer focused on HPC storage and play a crucial role in designing, implementing, and optimizing on-prem High-Performance...Full time- ...Overview Title: Site Reliability Engineer SRE – ML platform Location: Austin, TX or Sunnyvale, CA Employment type: Full-time • Seniority: Mid-Senior level • ONLY W2 Responsibilities Continuous Deployment using GitHub Actions, Flux, Kustomize Design and implement cloud...Full time
- ...Google is seeking a Software Engineering Manager II in Site Reliability Engineering, based in Sunnyvale, California. This onsite role leads a team to ensure reliability and performance of critical systems, partnering with product and engineering teams to deliver scalable...
- Remote DevOp/SRE With AI-First MindsetInsight Global is looking for a remote, DevOp/SRE with an AI-first mindset coming from a start up background to join one of our cyber security customers in the Bay Area. This role can pay 140-160k based on years of experience and skillset...Remote work
- ...design by customizing MES tool per business needs Education Requirements, Ideal Experience: Associate’s degree in Industrial Engineering or IT related field Minimum of 0-3 years’ relevant experience Experience in C#, Delphi desired Knowledge of the...Work at office
- ...of Huobi globe spanning infrastructure. • Work with engineering teams to make sure new features and changes are deployed quickly... .... • Constantly improve our system performance and reliability through better tools, process and monitoring system. •...Worldwide
$145k - $165k
...A technology solutions firm in Sunnyvale, CA is looking for a highly experienced Site Reliability Engineer (SRE). This role involves maintaining uptime and performance across systems. Exceptional Linux expertise and automation skills in Bash and Python are crucial. Key...- ...A leading technology firm is in search of a Senior Wireless Network Site Reliability Engineer to manage and enhance their wireless network infrastructure. The ideal candidate has over 8 years of experience in wireless network operations and a strong background in wireless...
- ...Job Description Job Description About the Role Senior Site Reliability Engineer (Payments Infrastructure) Kody is seeking a Senior Site Reliability Engineer to ensure the reliability, availability, scalability, and operational excellence of our global payment platform...
- ...Job Description Job Description Site Reliability Engineer Onsite- Bay Area, CA Skills Relevant Skills and Experience What You’ll Do (Day-to-Day) Own and manage our cloud infrastructure (GCP or AWS, on-prem). Build, maintain, and optimize Kubernetes...
$150k - $195k
...customers worldwide. Our team is growing, and we are looking for engineers with passion for automation. You will help support the... ...alongside engineering/operations teams to improve the scalability and reliability of internal processes. Participate in an on‑call rotation....Full timeWorldwide- ...Site Reliability Engineer, Data Platform - USDS Responsibilities Engage in and improve the whole lifecycle of service, from inception and design, through to deployment, operation and refinement. Ensure reliable, fault-tolerant, efficiently scalable and cost-effective data...
$186k - $279k
...and instrumented end to end? At Everpure, identity is the front door to everything we build, and we're looking for an IAM Site Reliability Engineer to keep that door running smoothly — and to make it better every day. The Global Information Security Office (GISO) at...Full timeWork at officeFlexible hours$135.6k - $180k
...operations team. This role involves overseeing 24/7 operational stability, enhancing processes and systems, and mentoring a diverse engineering team. The ideal candidate will have over 8 years of technical operations experience, proficiency in infrastructure automation...- ..., and the challenges of building in a high-growth startup, we’d love to talk. This is more than a job—it’s a journey. Site Reliability Engineers (SREs) are responsible for the overall performance and reliability of ASAPP's infrastructure and products. The team owns...Remote work
$150.4k - $277.6k
...Technical Operations & Site Reliability Engineer, Customer SystemsAt Apple, Customer Experience is at the forefront of everything we do. The Customer Systems Operations team is looking for a highly skilled and motivated TechOps Engineer (Technical Operations & Site Reliability...Work experience placementRelocation$101k - $161k
...Requirements: We require a BS or MS in Computer Science, or equivalent relevant experience. We look for 5+ years of software engineering experience. We need experience building or operating distributed database systems or scale-out applications in a SaaS...Full time$186.9k - $267.7k
...requiring approximately 2 days per week on-site at Cisco offices in either San Francisco... ...AI agents behave as intended, improving reliability and reducing risks. This unified... ...and control.As a Staff Site Reliability Engineer (SRE), you will provide technical leadership...Full timeTemporary workLocal areaFlexible hours2 days per week$184k - $287.5k
At NVIDIA, Site Reliability Engineering provides a rare chance to define, develop, and support large-scale production systems with high efficiency and availability. This demanding position merges software and systems engineering efforts to guarantee flawless service operation...Full time$210.6k - $305.1k
...Minimum Qualifications: You have led a distributed team of 5+ engineers, can demonstrate strong technical vision for your team, and ensure... ..., and basic life insurance. Please see the Cisco careers site to discover more benefits and perks. Employees may be eligible...Full timeTemporary workLocal areaFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
- site reliability engineer sre Sunnyvale, CA
- site reliability engineer Sunnyvale, CA
- on-site clinical research associate (traveling/remote) Sunnyvale, CA
- junior website developer Sunnyvale, CA
- site leader Sunnyvale, CA
- historic site Sunnyvale, CA
- website content developer Sunnyvale, CA
- construction site safety Sunnyvale, CA
- official site Sunnyvale, CA
- site safety Sunnyvale, CA


