Site Reliability Engineer Manager- Hybrid
Calance
Job Description
Job Description
We are hiring Site Reliability Engineer Manager- Hybrid for a Contract To Hire position in santa clara, CA
The Role You will build and lead the Site Reliability Engineering team, owning the infrastructure that development, validation, and customer-facing deployments run on. This spans colocation facilities, on-premises lab clusters, cloud environments (AWS, Azure, GCP), and the platform services customers use to collaborate on hardware and software deployments. You are both a people manager and a practicing engineer. You will set technical direction, hire and grow the team, own SLOs for critical systems, and be the senior escalation point when things go wrong. You will work closely with hardware and software development teams to ensure HPC infrastructure meets their workload requirements and partner with the Senior DevOps Lead whose pipelines and automation run on the infrastructure you own. What You Will Do Team Leadership & Strategy • Develop and manage a team of 3 5 SRE engineers; establish a culture of operational excellence, ownership, and continuous improvement. • Define the SRE team's technical roadmap: reliability architecture, automation priorities, capacity planning, and on-call model. • Serve as the senior technical escalation for critical incidents guiding cross-team triage, driving RCA, and ensuring systemic fixes rather than point patches. • Translate operational signals and infrastructure health into clear, actionable narratives for engineering leadership and executive stakeholders. • Partner with hardware and software development teams to understand HPC workload requirements and ensure infrastructure capacity, performance, and reliability meet the needs of silicon and software development programs. 24 x7 Infrastructure Reliability & Observability • Own 24 7 reliability across colocation, on-premises lab clusters, cloud, and customer-facing platform services designing for failure domains, progressive delivery, and strict change control at every tier. • Own the full observability stack (metrics, traces, logs) and define SLOs/SLIs across all SRE systems; use AI-driven detection, correlation, and guided remediation to reduce time to detect, respond, and resolve. • Evolve incident and problem management into a data-driven discipline: automated triage workflows, AI/analytics to identify recurring patterns, and every P0/P1 producing a written RCA with tracked systemic fixes. • Lead FinOps and capacity planning: model TCO across cloud vs. on-prem vs. colo, drive workload placement decisions, and anticipate infrastructure needs for new silicon programs and customer deployments. • Own infrastructure for customer collaboration environments where partners deploy and validate hardware and software. Automation & Infrastructure as Code • Drive IaC-first discipline across the team Terraform, Ansible, and production-quality automation for all infrastructure provisioning and lifecycle management. • Build and mature self-healing infrastructure platforms: host lifecycle automation, fleet auto-remediation, and AIOps-driven alerting that reduce manual intervention across the operational lifecycle. Documentation & Global Collaboration • Build a documentation culture and scale a follow-the-sun on-call model as we expands globally runbooks, architecture diagrams, and operational playbooks maintained as living artifacts. • Drive POC and POV evaluations for new infrastructure technologies, interconnect fabrics, and platform services relevant to our accelerator roadmap. What You Will Bring Required • Bachelor's or Master's in Computer Science, Electrical Engineering, or related field; 12+ years in SRE, infrastructure engineering, or production engineering (8 years minimum). • 3+ years managing SRE or infrastructure teams hiring, growing, and retaining engineers in a fast-moving environment. • Deep Linux systems expertise: networking (TCP/IP, RDMA, bonding), storage, kernel tuning, and bare-metal operations. • Proven experience operating colocation and on-premises hardware at scale: server lifecycle, power and cooling awareness, rack-level networking. • IaC fluency: Terraform and Ansible at production scale module design, remote state, environment isolation, and change governance. • Kubernetes cluster operations: lifecycle management, workload reliability, storage, and RBAC at scale. • Full observability stack ownership: Prometheus, Grafana, and/or DataDog SLO definition, alert design, and E2E signal quality. • Strong Python and/or Go production services, not just scripts; automation that touches real infrastructure safely. • Track record of reducing MTTR/MTTD through automation, workflow orchestration, and AIOps tooling. • Executive communication: translating infrastructure health and operational risk into clear narratives for senior leadership. • Demonstrated track record of moving teams from reactive, process-heavy operations to automated, technology-focused models not just managing existing runbooks. Strongly Preferred • Experience operating customer-facing infrastructure or platform services reliability expectations beyond internal tooling. • Knowledge of high-speed interconnect fabrics: InfiniBand, RoCE, or NVLink setup, troubleshooting, and performance tuning. • HPC job scheduler experience: Slurm, LSF, or equivalent setup, tuning, and integration with infrastructure automation. • Multi-cloud hybrid operations: AWS, Azure, GCP alongside on-prem/colo unified observability and IaC across all tiers. • FinOps: cloud spend attribution, TCO modeling across cloud vs. on-prem vs. colo, and translating cost data into workload placement recommendations for engineering and executive audiences. • ITIL knowledge or equivalent structured incident/problem/change management framework experience. • Published technical writing, conference talks, or open-source contributions in reliability, observability, or HPC infrastructure. Estimated Pay Range: 90-120/hrVacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer Manager- Hybrid in Santa Clara, CA vacancy
$160k - $250k
...macOS.We're looking for a Senior Engineering Manager who brings the technical depth... ...competing priorities, and hold the reliability bar under delivery pressure. You... ....Location:This role is hybrid, requiring 2x days per week on-site at one of the posted locations.What...SuggestedFull timeWork experience placementWork at officeLocal areaRemote work$276.1k - $311.4k
...your career! The role As SRE Manager, you'll build the Vehicle... ...charter, hiring its founding engineers, establishing the operating model... ...strategy that makes reliability a first-class property of the... ...of all worlds so we operate a hybrid working policy that combines...SuggestedPermanent employmentFull timeWork at officeWork from home$120k - $180k
...cutting-edge Falcon Exposure Management pillar, you'll be at the... ...and posture scoring across hybrid environments—spanning hosts,... ...cutting-edge technologies to engineer robust backend services that... ...decision-making processesService Reliability: Ensure robust, healthy...SuggestedFull timeWork experience placementWork at officeLocal areaWorldwide- ...revenue.We are looking for a Senior Site Reliability Engineer to lead the strategic evolution of our... ...infrastructure security.Please note: This is a hybrid role based in our Santa Clara, CA... ...with feature teams to refine Change Management and CI/CD pipelines, ensuring code...SuggestedFull timeWork at office2 days per week
- ...is currently Tuesday.Engineering at Lambda is responsible... ...system deployment, management and maintenance.What You... ...teams to improve service reliability and deployment... ...years of experience in Site Reliability Engineering... ...multi-datacenter and hybrid cloud environmentsHave...SuggestedWork at officeLocal areaWork from homeFlexible hours
$152k - $241.5k
...Infrastructure‑as‑Code) and config management to standardize and automate... ...distributed, multi‑cloud hybrid environment - On‑prem, AWS,... ...lifecycle management, fleet reliability/auto-healing, E2E... ...Perl, or Ruby.Mentored other engineers and influenced technical direction...Full time$203k - $258.6k
...and operates under a hybrid work model.Meet the TeamJoin... ...— partnering across engineering, security, compliance,... ...in Application Reliability, you will own the reliability... ...— deploying and managing applications on GKE (Kubernetes... ...see the Cisco careers site to discover more...Full timeTemporary workLocal areaFlexible hours$120k - $180k
...a Software Development Engineer in Test (SDET) in the Platform... ...events per second and manage petabytes of critical... ...environmentThis role is hybrid, requiring 2-3 days per week on-site at one of the posted locations... ...features, focusing on reliability, accuracy, and...Full timeWork experience placementWork at officeLocal areaWorldwide2 days per week3 days per week$146.7k - $339.3k
...positionWhat you can expect As a Senior Lead Site Reliability Engineer, you can anticipate opportunities to work on our hybrid systems across the globe. You will be... ...participate in on-call shifts and incident management and work after hours/weekends for application...Full timeWork at officeRemote workWorldwideShift workWeekend work$122.5k - $175k
...future of cybersecurity.RoleWe are looking for a Staff Site Reliability Engineer to join our team. This is a hybrid role going into the San Jose, CA office 3 days a... ...hands-on expertise in building infrastructure and managing platforms like Kubernetes using automation tools...Full timeWork at officeLocal area3 days per week$124k - $271.2k
What You Can ExpectAs a Lead Staff Site Reliability Engineer, you will be one of the technical leads for... ...(e.g., Python, Go, Java)Deploy and manage CI/CD pipelines using tools like Git,... ...09/17/26Ways of WorkingOur structured hybrid approach is centered around our offices...Full timeWork at officeRemote work- ...will doThe Senior Sourcing Manager (Hybrid) will play a critical leadership... ...escalations, improving reliability, and building long-term capability... ...and collaboration across sites to maximize value, mitigate... ...management, business, finance, engineering or related field.10 years of...Full timeContract workFor contractors3 days per week
- ...centers. We are looking for a Senior Site Reliability Engineer to improve the reliability, scalability... ..., scheduling, networking, resource management, upgrades, and common failure modes.Have... ...data centers, private cloud, hybrid cloud, or environments without full reliance...Work at officeLocal areaWork from homeFlexible hours
$187.04k - $359.72k
...interaction, capital management, tax and exchange optimization... ...changes that improve reliability and velocity.... ...Computer Science, Electrical Engineering, Computer Engineering... ...and more. On-site presence across teams... ...company is shifting from a hybrid work model to a fully...Temporary workLocal areaOverseasShift work- ...individual can thrive. The Role This hybrid role combines the hands-on... ...of a Technical Support Engineer within a SaaS (Software as a... ...with a growing focus on Site Reliability Engineering (SRE). The ideal... ...Familiarity with configuration management tools (e.g., Terraform)...Work at officeLocal areaRemote workWork from home
$255.7k - $300k
Lead a team of engineers to maintain service uptime while managing global on-call rotations and evaluating... ...practices to drive reliability, maintainability, and... ...Manager, Software Engineer, Site Reliability Engineering-... ...& may allow for a hybrid schedule as per Google policy...Full timeWork at office$203k - $258.6k
This role is hybrid. Onsite 3 days per week in Raleigh (Research Triangle... ..." platform —the central AI engine that powers productivity and... ...techniques. By partnering with product management and design teams, you will... ...Please see the Cisco careers site to discover more benefits and...Full timeTemporary workLocal areaFlexible hours3 days per week$342.7k
...are received.This position is a hybrid role, requiring the employee... ...coordinated with the team and manager.Meet the TeamArtificial Intelligence... ...looking for a Distinguished Engineer with outstanding technical... ...Please see the Cisco careers site to discover more benefits and...Full timeTemporary workWork at officeLocal areaRemote workFlexible hours$172k - $300k
...Vehicle Autonomy is forming a centralized Site Reliability Engineering team to make reliability a measurable... ...Excellence: Partner with Incident Management to operationalize severity, command,... ...environment. Experience with hybrid cloud/on-premises environments and foundational...Full timeWork at officeLocal areaRemote workWork from homeRelocationRelocation packageFlexible hours$169k - $338k
...Summary...As a Distinguished AI/ML Engineer within Walmart Global Tech's Site Reliability Engineering organization, you... ...Engineering organization is built with hybrid systems and software engineers... ...with intelligent capacity management and predictive performance optimization...Full timeTemporary workPart time- ...data-centric cybersecurity for hybrid multicloud environments,... ...short, we focus on data exposure management to keep your information safe... ...for a Staff Software Engineer to join our Confidential Computing... ...are secure by design, highly reliable, and built to scale ....Temporary workH1bWorldwide
- ...Skills: 2+ years of experience in Site Reliability Engineering, DevOps, Infrastructure Engineering,... .... Practical vulnerability-management experience; familiarity with Qualys,... ...or other public cloud platforms and hybrid infrastructure environments. Knowledge...Full timeWorldwide
- ...Senior Android Engineer We are looking for a senior level Android engineer with Kotlin experience. The ideal candidate will have a... ...Published Android application is required. This role will be hybrid, with 2 days per week in office, team located in Sunnyvale....Work at office2 days per week
$90k - $125k
...user mode and kernel mode. Engineering software at that depth and that... ...is Sensor Performance and Reliability — CrowdStrike's team of debug... ...a career.Location:This is a hybrid role based out of one of the... ...architecture, memory management, concurrency, compilers and...Full timeWork experience placementInternshipWork at officeLocal areaRemote workWorldwide$150.4k - $190.6k
...are received.This role will be Hybrid from our Guadalajara, Mexico... ...the Team Partner Sourcing Managers are integral to our Supply Chain... ...adept at influencing engineering and New Product Introduction... ...Please see the Cisco careers site to discover more benefits and...Full timeContract workTemporary workLocal areaFlexible hours$121.1k - $153.7k
...number of applications are received.Preferred hybrid role in Austin, TX, or San Jose, CAUS... ..., and legal agreements.Act as a “general manager,” cognizant of all engagements / touch... ...insurance. Please see the Cisco careers site to discover more benefits and perks. Employees...Full timeTemporary workLocal areaRemote workFlexible hours$197.5k - $249.8k
...applications are received.This is a hybrid role based out of Cisco's... ...researchers, machine learning engineers, data engineers, and... ...model deployment.Architect and manage human-in-the-loop labeling workflows... ...Please see the Cisco careers site to discover more benefits and...Full timeTemporary workWork at officeLocal areaFlexible hoursShift work$120k - $180k
...teamCrowdStrike is seeking a Cloud Software Engineer to join our expanding Surface... ...You will join the External Attack Surface Management (EASM) engineering team, which develops... ...agents or complex deployment.This is a hybrid role requiring 2-3 days in office in Sunnyvale...Permanent employmentFull timeWork experience placementWork at officeLocal area$96.49k - $144.74k
...are currently seeking a Data Engineer - Data Platform (Spark/Kafka/Flink/Scala/Java) - Onsite Hybrid to join our team in Cupertino... ...tuning, scalability optimization, reliability improvements, and operational... ...to NTT DATA offices or client sites. This ensures we can provide...Full timeTemporary workWork experience placementWork at officeRemote workFlexible hours- ...Your role and responsibilities As a Site Reliability Engineer, you will work in an agile,... ...handling day-to-day operations, alert management, incident support, migration tasks, and... ...and deploy AI across business. IBM's hybrid cloud platform is one of the most comprehensive...Full timeContract workPart timeFixed term contractInternshipWorldwideFlexible hoursShift work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer Manager- Hybrid. Be the first to apply!
Related searches
- site reliability engineer Santa Clara, CA
- site reliability engineer sre Santa Clara, CA
- site director Santa Clara, CA
- site superintendent Santa Clara, CA
- site supervisor Santa Clara, CA
- site manager Santa Clara, CA
- site agent Santa Clara, CA
- hvac site manager Santa Clara, CA
- site construction manager Santa Clara, CA
- on-site clinical research associate (traveling/remote) Santa Clara, CA



