Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer Manager- Hybrid

Calance

Job Description

Job Description

We are hiring Site Reliability Engineer Manager- Hybrid for a Contract To Hire position in santa clara, CA

The Role

You will build and lead the Site Reliability Engineering team, owning the infrastructure that development, validation, and customer-facing deployments run on. This spans colocation facilities, on-premises lab clusters, cloud environments (AWS, Azure, GCP), and the platform services customers use to collaborate on hardware and software deployments.

You are both a people manager and a practicing engineer. You will set technical direction, hire and grow the team, own SLOs for critical systems, and be the senior escalation point when things go wrong. You will work closely with hardware and software development teams to ensure HPC infrastructure meets their workload requirements and partner with the Senior DevOps Lead whose pipelines and automation run on the infrastructure you own.

What You Will Do

Team Leadership & Strategy

• Develop and manage a team of 3 5 SRE engineers; establish a culture of operational excellence, ownership, and continuous improvement.

• Define the SRE team's technical roadmap: reliability architecture, automation priorities, capacity planning, and on-call model.

• Serve as the senior technical escalation for critical incidents guiding cross-team triage, driving RCA, and ensuring systemic fixes rather than point patches.

• Translate operational signals and infrastructure health into clear, actionable narratives for engineering leadership and executive stakeholders.

• Partner with hardware and software development teams to understand HPC workload requirements and ensure infrastructure capacity, performance, and reliability meet the needs of silicon and software development programs.

24 x7 Infrastructure Reliability & Observability

• Own 24 7 reliability across colocation, on-premises lab clusters, cloud, and customer-facing platform services designing for failure domains, progressive delivery, and strict change control at every tier.

• Own the full observability stack (metrics, traces, logs) and define SLOs/SLIs across all SRE systems; use AI-driven detection, correlation, and guided remediation to reduce time to detect, respond, and resolve.

• Evolve incident and problem management into a data-driven discipline: automated triage workflows, AI/analytics to identify recurring patterns, and every P0/P1 producing a written RCA with tracked systemic fixes.

• Lead FinOps and capacity planning: model TCO across cloud vs. on-prem vs. colo, drive workload placement decisions, and anticipate infrastructure needs for new silicon programs and customer deployments.

• Own infrastructure for customer collaboration environments where partners deploy and validate hardware and software.

Automation & Infrastructure as Code

• Drive IaC-first discipline across the team Terraform, Ansible, and production-quality automation for all infrastructure provisioning and lifecycle management.

• Build and mature self-healing infrastructure platforms: host lifecycle automation, fleet auto-remediation, and AIOps-driven alerting that reduce manual intervention across the operational lifecycle.

Documentation & Global Collaboration

• Build a documentation culture and scale a follow-the-sun on-call model as we expands globally runbooks, architecture diagrams, and operational playbooks maintained as living artifacts.

• Drive POC and POV evaluations for new infrastructure technologies, interconnect fabrics, and platform services relevant to our accelerator roadmap.

What You Will Bring

Required

• Bachelor's or Master's in Computer Science, Electrical Engineering, or related field; 12+ years in SRE, infrastructure engineering, or production engineering (8 years minimum).

• 3+ years managing SRE or infrastructure teams hiring, growing, and retaining engineers in a fast-moving environment.

• Deep Linux systems expertise: networking (TCP/IP, RDMA, bonding), storage, kernel tuning, and bare-metal operations.

• Proven experience operating colocation and on-premises hardware at scale: server lifecycle, power and cooling awareness, rack-level networking.

• IaC fluency: Terraform and Ansible at production scale module design, remote state, environment isolation, and change governance.

• Kubernetes cluster operations: lifecycle management, workload reliability, storage, and RBAC at scale.

• Full observability stack ownership: Prometheus, Grafana, and/or DataDog SLO definition, alert design, and E2E signal quality.

• Strong Python and/or Go production services, not just scripts; automation that touches real infrastructure safely.

• Track record of reducing MTTR/MTTD through automation, workflow orchestration, and AIOps tooling.

• Executive communication: translating infrastructure health and operational risk into clear narratives for senior leadership.

• Demonstrated track record of moving teams from reactive, process-heavy operations to automated, technology-focused models not just managing existing runbooks.

Strongly Preferred

• Experience operating customer-facing infrastructure or platform services reliability expectations beyond internal tooling.

• Knowledge of high-speed interconnect fabrics: InfiniBand, RoCE, or NVLink setup, troubleshooting, and performance tuning.

• HPC job scheduler experience: Slurm, LSF, or equivalent setup, tuning, and integration with infrastructure automation.

• Multi-cloud hybrid operations: AWS, Azure, GCP alongside on-prem/colo unified observability and IaC across all tiers.

• FinOps: cloud spend attribution, TCO modeling across cloud vs. on-prem vs. colo, and translating cost data into workload placement recommendations for engineering and executive audiences.

• ITIL knowledge or equivalent structured incident/problem/change management framework experience.

• Published technical writing, conference talks, or open-source contributions in reliability, observability, or HPC infrastructure.

Estimated Pay Range: 90-120/hr

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer Manager- Hybrid in Santa Clara, CA vacancy
  • $160k - $250k

     ...macOS.We're looking for a Senior Engineering Manager who brings the technical depth...  ...competing priorities, and hold the reliability bar under delivery pressure. You...  ....Location:This role is hybrid, requiring 2x days per week on-site at one of the posted locations.What... 
    Suggested
    Full time
    Work experience placement
    Work at office
    Local area
    Remote work

    CrowdStrike

    Sunnyvale, CA
    2 days ago
  • $276.1k - $311.4k

     ...your career! The role As SRE Manager, you'll build the Vehicle...  ...charter, hiring its founding engineers, establishing the operating model...  ...strategy that makes reliability a first-class property of the...  ...of all worlds so we operate a hybrid working policy that combines... 
    Suggested
    Permanent employment
    Full time
    Work at office
    Work from home

    Lindus Health

    Sunnyvale, CA
    3 days ago
  • $120k - $180k

     ...cutting-edge Falcon Exposure Management pillar, you'll be at the...  ...and posture scoring across hybrid environments—spanning hosts,...  ...cutting-edge technologies to engineer robust backend services that...  ...decision-making processesService Reliability: Ensure robust, healthy... 
    Suggested
    Full time
    Work experience placement
    Work at office
    Local area
    Worldwide

    CrowdStrike

    Sunnyvale, CA
    1 day ago
  •  ...revenue.We are looking for a Senior Site Reliability Engineer to lead the strategic evolution of our...  ...infrastructure security.Please note: This is a hybrid role based in our Santa Clara, CA...  ...with feature teams to refine Change Management and CI/CD pipelines, ensuring code... 
    Suggested
    Full time
    Work at office
    2 days per week

    LeanData

    Santa Clara, CA
    4 days ago
  •  ...is currently Tuesday.Engineering at Lambda is responsible...  ...system deployment, management and maintenance.What You...  ...teams to improve service reliability and deployment...  ...years of experience in Site Reliability Engineering...  ...multi-datacenter and hybrid cloud environmentsHave... 
    Suggested
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    1 day ago
  • $152k - $241.5k

     ...Infrastructure‑as‑Code) and config management to standardize and automate...  ...distributed, multi‑cloud hybrid environment - On‑prem, AWS,...  ...lifecycle management, fleet reliability/auto-healing, E2E...  ...Perl, or Ruby.Mentored other engineers and influenced technical direction... 
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $203k - $258.6k

     ...and operates under a hybrid work model.Meet the TeamJoin...  ...— partnering across engineering, security, compliance,...  ...in Application Reliability, you will own the reliability...  ...— deploying and managing applications on GKE (Kubernetes...  ...see the Cisco careers site to discover more... 
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    San Jose, CA
    2 days ago
  • $120k - $180k

     ...a Software Development Engineer in Test (SDET) in the Platform...  ...events per second and manage petabytes of critical...  ...environmentThis role is hybrid, requiring 2-3 days per week on-site at one of the posted locations...  ...features, focusing on reliability, accuracy, and... 
    Full time
    Work experience placement
    Work at office
    Local area
    Worldwide
    2 days per week
    3 days per week

    CrowdStrike

    Sunnyvale, CA
    2 days ago
  • $146.7k - $339.3k

     ...positionWhat you can expect As a Senior Lead Site Reliability Engineer, you can anticipate opportunities to work on our hybrid systems across the globe. You will be...  ...participate in on-call shifts and incident management and work after hours/weekends for application... 
    Full time
    Work at office
    Remote work
    Worldwide
    Shift work
    Weekend work

    Zoom

    San Jose, CA
    4 days ago
  • $122.5k - $175k

     ...future of cybersecurity.RoleWe are looking for a Staff Site Reliability Engineer to join our team. This is a hybrid role going into the San Jose, CA office 3 days a...  ...hands-on expertise in building infrastructure and managing platforms like Kubernetes using automation tools... 
    Full time
    Work at office
    Local area
    3 days per week

    Zscaler

    San Jose, CA
    17 hours ago
  • $124k - $271.2k

    What You Can ExpectAs a Lead Staff Site Reliability Engineer, you will be one of the technical leads for...  ...(e.g., Python, Go, Java)Deploy and manage CI/CD pipelines using tools like Git,...  ...09/17/26Ways of WorkingOur structured hybrid approach is centered around our offices... 
    Full time
    Work at office
    Remote work

    Zoom

    San Jose, CA
    3 days ago
  •  ...will doThe Senior Sourcing Manager (Hybrid) will play a critical leadership...  ...escalations, improving reliability, and building long-term capability...  ...and collaboration across sites to maximize value, mitigate...  ...management, business, finance, engineering or related field.10 years of... 
    Full time
    Contract work
    For contractors
    3 days per week

    Stryker

    San Jose, CA
    4 days ago
  •  ...centers. We are looking for a Senior Site Reliability Engineer to improve the reliability, scalability...  ..., scheduling, networking, resource management, upgrades, and common failure modes.Have...  ...data centers, private cloud, hybrid cloud, or environments without full reliance... 
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    4 days ago
  • $187.04k - $359.72k

     ...interaction, capital management, tax and exchange optimization...  ...changes that improve reliability and velocity....  ...Computer Science, Electrical Engineering, Computer Engineering...  ...and more. On-site presence across teams...  ...company is shifting from a hybrid work model to a fully... 
    Temporary work
    Local area
    Overseas
    Shift work

    Tik Tok

    San Jose, CA
    4 days ago
  •  ...individual can thrive. The Role This hybrid role combines the hands-on...  ...of a Technical Support Engineer within a SaaS (Software as a...  ...with a growing focus on Site Reliability Engineering (SRE). The ideal...  ...Familiarity with configuration management tools (e.g., Terraform)... 
    Work at office
    Local area
    Remote work
    Work from home

    F5 Networks

    San Jose, CA
    5 days ago
  • $255.7k - $300k

    Lead a team of engineers to maintain service uptime while managing global on-call rotations and evaluating...  ...practices to drive reliability, maintainability, and...  ...Manager, Software Engineer, Site Reliability Engineering-...  ...& may allow for a hybrid schedule as per Google policy... 
    Full time
    Work at office

    Google

    Sunnyvale, CA
    17 hours ago
  • $203k - $258.6k

    This role is hybrid. Onsite 3 days per week in Raleigh (Research Triangle...  ..." platform—the central AI engine that powers productivity and...  ...techniques. By partnering with product management and design teams, you will...  ...Please see the Cisco careers site to discover more benefits and... 
    Full time
    Temporary work
    Local area
    Flexible hours
    3 days per week

    CISCO Systems

    San Jose, CA
    1 day ago
  • $342.7k

     ...are received.This position is a hybrid role, requiring the employee...  ...coordinated with the team and manager.Meet the TeamArtificial Intelligence...  ...looking for a Distinguished Engineer with outstanding technical...  ...Please see the Cisco careers site to discover more benefits and... 
    Full time
    Temporary work
    Work at office
    Local area
    Remote work
    Flexible hours

    CISCO Systems

    San Jose, CA
    2 days ago
  • $172k - $300k

     ...Vehicle Autonomy is forming a centralized Site Reliability Engineering team to make reliability a measurable...  ...Excellence: Partner with Incident Management to operationalize severity, command,...  ...environment. Experience with hybrid cloud/on-premises environments and foundational... 
    Full time
    Work at office
    Local area
    Remote work
    Work from home
    Relocation
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    3 days ago
  • $169k - $338k

     ...Summary...As a Distinguished AI/ML Engineer within Walmart Global Tech's Site Reliability Engineering organization, you...  ...Engineering organization is built with hybrid systems and software engineers...  ...with intelligent capacity management and predictive performance optimization... 
    Full time
    Temporary work
    Part time

    Walmart

    Sunnyvale, CA
    1 day ago
  •  ...data-centric cybersecurity for hybrid multicloud environments,...  ...short, we focus on data exposure management to keep your information safe...  ...for a Staff Software Engineer to join our Confidential Computing...  ...are secure by design, highly reliable, and built to scale .... 
    Temporary work
    H1b
    Worldwide

    Fortanix

    Santa Clara, CA
    12 days ago
  •  ...Skills: 2+ years of experience in Site Reliability Engineering, DevOps, Infrastructure Engineering,...  .... Practical vulnerability-management experience; familiarity with Qualys,...  ...or other public cloud platforms and hybrid infrastructure environments. Knowledge... 
    Full time
    Worldwide

    Northern Base

    San Jose, CA
    2 days ago
  •  ...Senior Android Engineer We are looking for a senior level Android engineer with Kotlin experience. The ideal candidate will have a...  ...Published Android application is required. This role will be hybrid, with 2 days per week in office, team located in Sunnyvale.... 
    Work at office
    2 days per week

    Samprasoft

    Sunnyvale, CA
    4 days ago
  • $90k - $125k

     ...user mode and kernel mode. Engineering software at that depth and that...  ...is Sensor Performance and Reliability — CrowdStrike's team of debug...  ...a career.Location:This is a hybrid role based out of one of the...  ...architecture, memory management, concurrency, compilers and... 
    Full time
    Work experience placement
    Internship
    Work at office
    Local area
    Remote work
    Worldwide

    CrowdStrike

    Sunnyvale, CA
    4 days ago
  • $150.4k - $190.6k

     ...are received.This role will be Hybrid from our Guadalajara, Mexico...  ...the Team Partner Sourcing Managers are integral to our Supply Chain...  ...adept at influencing engineering and New Product Introduction...  ...Please see the Cisco careers site to discover more benefits and... 
    Full time
    Contract work
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    San Jose, CA
    1 day ago
  • $121.1k - $153.7k

     ...number of applications are received.Preferred hybrid role in Austin, TX, or San Jose, CAUS...  ..., and legal agreements.Act as a “general manager,” cognizant of all engagements / touch...  ...insurance. Please see the Cisco careers site to discover more benefits and perks. Employees... 
    Full time
    Temporary work
    Local area
    Remote work
    Flexible hours

    CISCO Systems

    San Jose, CA
    4 days ago
  • $197.5k - $249.8k

     ...applications are received.This is a hybrid role based out of Cisco's...  ...researchers, machine learning engineers, data engineers, and...  ...model deployment.Architect and manage human-in-the-loop labeling workflows...  ...Please see the Cisco careers site to discover more benefits and... 
    Full time
    Temporary work
    Work at office
    Local area
    Flexible hours
    Shift work

    CISCO Systems

    San Jose, CA
    3 days ago
  • $120k - $180k

     ...teamCrowdStrike is seeking a Cloud Software Engineer to join our expanding Surface...  ...You will join the External Attack Surface Management (EASM) engineering team, which develops...  ...agents or complex deployment.This is a hybrid role requiring 2-3 days in office in Sunnyvale... 
    Permanent employment
    Full time
    Work experience placement
    Work at office
    Local area

    CrowdStrike

    Sunnyvale, CA
    1 day ago
  • $96.49k - $144.74k

     ...are currently seeking a Data Engineer - Data Platform (Spark/Kafka/Flink/Scala/Java) - Onsite Hybrid to join our team in Cupertino...  ...tuning, scalability optimization, reliability improvements, and operational...  ...to NTT DATA offices or client sites. This ensures we can provide... 
    Full time
    Temporary work
    Work experience placement
    Work at office
    Remote work
    Flexible hours

    NTT DATA Services

    Cupertino, CA
    5 days ago
  •  ...Your role and responsibilities As a Site Reliability Engineer, you will work in an agile,...  ...handling day-to-day operations, alert management, incident support, migration tasks, and...  ...and deploy AI across business. IBM's hybrid cloud platform is one of the most comprehensive... 
    Full time
    Contract work
    Part time
    Fixed term contract
    Internship
    Worldwide
    Flexible hours
    Shift work

    IBM

    San Jose, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer Manager- Hybrid. Be the first to apply!