Site Reliability Engineer
$120k - $180kCrowdstrike
SRE & DevOps EngineerAs a global leader in cybersecurity, CrowdStrike protects the people, processes and technologies that drive modern organizations. Since 2011, our mission hasn't changed — we're here to stop breaches, and we've redefined modern security with the world's most advanced AI-native platform. We work on large scale distributed systems, processing almost 3 trillion events per day and this traffic is growing daily. Our customers span all industries, and they count on CrowdStrike to keep their businesses running, their communities safe and their lives moving forward. We're proud to work for a mission-driven company leveraging AI to transform the way we work. CrowdStrikers drive their careers through flexibility and autonomy while also being expected to contribute to a culture of responsible AI adoption, experimentation, and innovation. We use an AI-first mindset as a force multiplier to proactively and continuously accelerate execution, build expertise, uncover insights, and solve complex problems. We're always looking to add talented CrowdStrikers to the team who have limitless passion, a relentless focus on innovation and a fanatical commitment to our customers, our community and each other. Ready to join a mission that matters? The future of cybersecurity starts with you.About the Role:At CrowdStrike, our engineering organization depends on shared infrastructure platforms that power critical product capabilities at global scale. These platforms require dedicated engineering ownership to operate reliably, scale safely, harden for security, and mature into self-service capabilities that teams across the organization can depend on.As an SRE & DevOps Engineer, you will own production infrastructure spanning multiple cloud providers and regions, serving engineering teams across CrowdStrike. The work is equal parts reliability engineering and DevOps engineering - building automation, hardening security, establishing governance, and enabling consuming teams to adopt these platforms effectively.You will work across a rich technology landscape including Kubernetes, Kafka, Cassandra, PostgreSQL, Apache Pinot, OpenSearch, and Apache Spark — operating and scaling microservices-based distributed systems that process millions of security events per second with zero tolerance for data loss or downtime.What You'll Do:Run production infrastructure - Deploy, upgrade, and maintain platform services across multiple clouds and regions on Kubernetes, including microservices-based distributed systemsOwn delivery pipelines - Build and maintain scalable CI/CD pipelines using Jenkins, GitLab CI, and Bitbucket Pipelines with GitOps workflows via ArgoCD or FluxOwn capacity planning - Track usage, forecast growth, right-size clusters, and optimize infrastructure costs across multi-cloud environmentsBuild observability - Implement metrics, dashboards, alerts (Prometheus/Grafana), distributed tracing (Jaeger/OpenTelemetry), and actionable runbooksOwn on-call and incidents - Participate in on-call rotation, lead incident resolution, write blameless postmortems, and automate repeat problemsDrive reliability - Apply SRE principles and AI-driven automation to move from reactive firefighting to proactive operationsHarden security - Implement auth, encryption, secret rotation, and network policies.Own disaster recovery - Build and test backup strategies and failover mechanisms ensuring zero data lossOperate data infrastructure - Maintain reliability of Kafka, Cassandra, PostgreSQL, Apache Pinot, and OpenSearch in productionEnable and collaborate - Support engineering teams with templates and patterns; partner with Infrastructure, SRE, and Data Services on shared operational problemsExperience and Background:8+ years in SRE, Devops engineering, or infrastructure engineeringHands-on experience running stateful distributed systems and microservices architectures on Kubernetes in productionBachelor's degree in Computer Science or related field, or equivalent work experienceReliability Engineering:Deep understanding of SRE principles - SLOs, SLIs, error budgets applied to large-scale distributed systemsStrong incident management background - on-call ownership, blameless postmortems, and turning operational pain into automationExperience with chaos engineering and resilience validation for production systemsProven ability to build and maintain systems with zero tolerance for data loss or downtimeAdvanced observability experience including Prometheus, Grafana, distributed tracing ( Jaeger/OpenTelemetry ), and large-scale log aggregation ( ELK/Splunk ) with a focus on building custom SLO dashboards and reliability scorecardsProgramming & Automation:Proficiency in Python and/or Golang for automation, tooling, and platform servicesStrong scripting and automation skills - if you do it by hand more than once, you automate itPlatform and Delivery Engineering:CI/CD pipeline experience - Building and owning scalable delivery pipelines using Jenkins, GitLab CI, Bitbucket Pipelines, Tekton, or equivalentStrong proficiency in Infrastructure as Code (IaC) — Terraform, Ansible, Pulumi, or equivalentCloud and Big Data ExposureProficiency in at least one cloud environment (AWS, Azure, GCP) with emphasis on multi-region architecture, cloud-native reliability patterns, and security-first cloud designStrong experience with Kubernetes at scale - managing large cluster fleets, workload orchestration, and container lifecycle managementFamiliarity with distributed data systems including relational databases (PostgreSQL), NoSQL ( Cassandra ), OLAP ( Pinot ), Indexing(OpenSearch) and real-time streaming platforms ( Kafka, Flink )Exposure to Big Data and analytics technologies like Spark,StormBenefits of Working at CrowdStrike:Market leader in compensation and equity awardsComprehensive physical and mental wellness programsCompetitive vacation and holidays for rechargePaid parental and adoption leavesProfessional development opportunities for all employees regardless of level or roleEmployee Networks, geographic neighborhood groups, and volunteer opportunities to build connectionsVibrant office culture with world class amenitiesGreat Place to Work Certified™ across the globeCrowdStrike is proud to be an equal opportunity employer. We are committed to fostering a culture of belonging where everyone is valued for who they are and empowered to succeed. We support veterans and individuals with disabilities through our affirmative action program.CrowdStrike is committed to providing equal employment opportunity for all employees and applicants for employment. The Company does not discriminate in employment opportunities or practices on the basis of race, color, creed, ethnicity, religion, sex (including pregnancy or pregnancy-related medical conditions), sexual orientation, gender identity, marital or family status, veteran status, age, national origin, ancestry, physical disability (including HIV and AIDS), mental disability, medical condition, genetic information, membership or activity in a local human rights commission, status with regard to public assistance, or any other characteristic protected by law. We base all employment decisions--including recruitment, selection, training, compensation, benefits, discipline, promotions, transfers, lay-offs, return from lay-off, terminations and social/recreational programs--on valid job requirements.If you need assistance accessing or reviewing the information on this website or need help submitting an application for employment or requesting an accommodation, please contact us at View email address on click.appcast.io for further assistance.Find out more about your rights as an applicant.CrowdStrike participates in the E-Verify program.Notice of E-Verify ParticipationRight to WorkCrowdStrike, Inc. is committed to fair and equitable compensation practices. Placement within the pay range is dependent on a variety of factors including, but not limited to, relevant work experience, skills, certifications, job level, supervisory status, and location. The base salary range for this position for all U.S. candidates is $120,000 - $180,000 per year, with eligibility for bonuses, equity grants and a comprehensive benefits package that includes health insurance, 401k and paid time off.For detailed information about the U.S. benefits package, please click here.
$90k - $180k
...nutritionals and branded generic medicines. Our 122,000 colleagues serve people in more than 160 countries.About the RoleThis Senior Site Reliability Engineer position works on-site out of our Sylmar, CA or Sunnyvale, CA location in the Cardiac Rhythm Management Division.We are...SuggestedRemote work$128.6k - $184.9k
...private datacenters and AWS while ensuring reliable operations, resilience, and zero-... ...private cloud environments. Lead cloud engineering initiatives using Terraform, Ansible, and... ...insurance. Please see the Cisco careers site to discover more benefits and perks. Employees...SuggestedPermanent employmentFull timeTemporary workLocal areaFlexible hours$230k - $250k
...network. It's the foundation for autonomous networking, giving engineers and AI agents the ability to know the impact of every change... ...how things have always been done.Forward is looking for a Site Reliability EngineerAbout the Role This is not a "keep the lights on"...SuggestedNight shift$170k - $200k
We are seeking a talented and motivated Site Reliability Engineer to join our engineering team. You will be responsible for building, maintaining, and troubleshooting cloud service/cluster, infrastructure, and monitoring systems to ensure high availability, performance,...SuggestedFull timeWorldwide- LeanData helps the world’s fastest-growing companies automate, simplify, and accelerate revenue.We are looking for a Senior Site Reliability Engineer to lead the strategic evolution of our cloud infrastructure. Reporting directly to the SVP of Engineering, this role is...SuggestedFull timeWork at office2 days per week
$128k - $216k
...another millions of times a day - quickly, reliably, and securely. Any time you swipe your... ...make a difference at Fiserv.Job TitleSr. Site Reliability EngineerAbout CloverClover is... ...does a successful Senior Site Reliability Engineer do at Fiserv?As a Senior Site Reliability...Full timeWorldwide$148k - $235.75k
...see how you can make a lasting impact on the world.Join our team of innovative engineers who are building an AI Data Center AIOps platform that turns raw, high-volume telemetry into reliable, job-centric insights and automation for GPU fleets. We’re hiring a DevOps Engineer...Full time$160k - $240k
...one another millions of times a day - quickly, reliably, and securely. Any time you swipe your credit... ...come make a difference at Fiserv.Job TitleSenior Site Reliability EngineerWhat does a successful Site Reliability Engineer do at Fiserv?You will join our global team in...Full time$145k - $165k
...Your Ego : Selflessly collaborate towards our shared purpose. About the role Bolt Graphics is seeking a highly experienced Site Reliability Engineer (SRE) to design, build, and operate highly reliable developer and production systems. This role is mission-critical to...Work at officeImmediate start$101k - $161k
...excellence has earned us several prestigious awards, such as Best Engineering Team, Best Company for Diversity, Compensation, and Work-... ...we do.Job DescriptionWho You'll Work WithWe’re looking for Site Reliability Engineers to join our growing Arista’s CloudVision-as-a-...$168k - $270.25k
...phenomenal people like you to help us accelerate the next wave of artificial intelligence.Join our team at NVIDIA as a Senior Site reliability engineer focused on HPC storage and play a crucial role in designing, implementing, and optimizing on-prem High-Performance...Full time$255.7k - $300k
...designs from peers, providing feedback to ensure best practices in reliability, security, and efficiency.Triage and resolve complex system... ...execution of software development initiatives.Mentor other engineers and contribute to the engineering community through documentation...Full time- ...Google is seeking a Software Engineering Manager II in Site Reliability Engineering, based in Sunnyvale, California. This onsite role leads a team to ensure reliability and performance of critical systems, partnering with product and engineering teams to deliver scalable...
- ...Overview Title: Site Reliability Engineer SRE – ML platform Location: Austin, TX or Sunnyvale, CA Employment type: Full-time • Seniority: Mid-Senior level • ONLY W2 Responsibilities Continuous Deployment using GitHub Actions, Flux, Kustomize Design and implement cloud...Full time
$150k - $195k
...customers worldwide. Our team is growing, and we are looking for engineers with passion for automation. You will help support the... ...alongside engineering/operations teams to improve the scalability and reliability of internal processes. Participate in an on‑call rotation....Full timeWorldwide- ...Site Reliability Engineer, Data Platform - USDS Responsibilities Engage in and improve the whole lifecycle of service, from inception and design, through to deployment, operation and refinement. Ensure reliable, fault-tolerant, efficiently scalable and cost-effective data...
$276.1k - $311.4k
...Vehicle Software SRE team from the ground up — defining its charter, hiring its founding engineers, establishing the operating model, and creating the technical strategy that makes reliability a first-class property of the software running on our vehicles. You'll work in a...Permanent employmentFull timeWork at officeWork from home$145k - $165k
...A technology solutions firm in Sunnyvale, CA is looking for a highly experienced Site Reliability Engineer (SRE). This role involves maintaining uptime and performance across systems. Exceptional Linux expertise and automation skills in Bash and Python are crucial. Key...$222k - $300.5k
...possible.Job OverviewAbout the TeamIntuit's Infrastructure and Site Reliability organization owns the operational backbone that keeps... ...hundreds of millions of customers. The Fintech Platform Systems Engineering team builds and operates the AWS-based infrastructure, resiliency...WorldwideShift work$262k - $364k
...services within the AViD ecosystem have reliability and uptime appropriate to users' needs with... ...capacity and performance.Build creative engineering solutions to operations and... ...changing circumstances in a strategic way.Site Reliability Engineering (SRE) combines software...- ...design by customizing MES tool per business needs Education Requirements, Ideal Experience: Associate’s degree in Industrial Engineering or IT related field Minimum of 0-3 years’ relevant experience Experience in C#, Delphi desired Knowledge of the...Work at office
- ...Job Title : Senior Site Reliability Engineer Location : Santa Clara, CA Contract ENGAGEMENT SUMMARY The Candidate will provide SRE services for AI platforms and supporting infrastructure with emphasis on reliability engineering, incident response...Contract work
- ...Job Title: Site Reliability Engineer Location: Sunnyvale, CA Experience Required: 10+ years Employment Type: W2 Work Authorization: All visa types considered except H1B Job Summary: Seeking an experienced engineer who can analyze,...H1bWork visa
- ...of Huobi globe spanning infrastructure. • Work with engineering teams to make sure new features and changes are deployed quickly... .... • Constantly improve our system performance and reliability through better tools, process and monitoring system. •...Worldwide
$120k - $200k
Sr Site Reliability Engineer (Prisma Access) 2 days ago Be among the first 25 applicants Job Description This role requires US Citizenship. Your Career Palo Alto Networks runs a large infrastructure and is one of the biggest GCP customers. As a Principal SRE, you'll be...Rotating shift$200k - $322k
...PTP, DHCP, and LDAP. This includes building for performance and reliability at global scale, covering automation, monitoring, high... ...alerting, monitoring.Collaborate with NVIDIA leadership, senior engineers, program managers, and product managers to develop compelling...Full timeRemote work$255.7k - $300k
Lead a team of engineers to maintain service uptime while managing global on-call rotations... ...improve operational practices to drive reliability, maintainability, and stakeholder alignment... ...or in a Manager, Software Engineer, Site Reliability Engineering-related occupation...Full timeWork at office$207k - $300k
Lead a team of Software/Systems Engineers on projects for users and be directly responsible for uptime.Own end-to-end availability... ...Science or Engineering.1 year of people management experience. Site Reliability Engineering (SRE) combines software and systems engineering...$272k - $431.25k
NVIDIA is looking for a Cloud Site Reliability Engineering Architect to work in IPP's (Infrastructure, Planning and Process) Cloud Infrastructure Team. IPP is a global organization within NVIDIA. This group works with various other groups within NVIDIA such as Graphics...Full timeWork experience placementWorldwide$169k - $338k
...Regular/PermanentCompany: WalmartBusiness Segment: Home OfficePosition Summary...As a Distinguished AI/ML Engineer within Walmart Global Tech's Site Reliability Engineering organization, you will lead the technical development of next-generation agentic AI systems and...Full timeTemporary workPart time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
- site reliability engineer Sunnyvale, CA
- site reliability engineer sre Sunnyvale, CA
- on-site clinical research associate (traveling/remote) Sunnyvale, CA
- site safety Sunnyvale, CA
- junior website developer Sunnyvale, CA
- construction site safety Sunnyvale, CA
- IT site lead Sunnyvale, CA
- website content developer Sunnyvale, CA
- site recruiter Sunnyvale, CA
- historic site Sunnyvale, CA


