Site Reliability Engineer
Bitdeer Group
Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure. Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to support the AI revolution. Bitdeer handles complex processes involved in computing such as equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial intelligence. Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia. To learn more, visit Position Overview You run the control plane where AIOps meets tenants — where topology-aware scheduling, self-healing, and agent-driven remediation actually execute. NeoCloud is building an AI-operated GPU cloud. Kubernetes is where all of that lands on real customer workloads: the topology-aware scheduler places jobs on the right NVLink domain, the operator drains and reschedules around predicted faults, and the tenant boundary is enforced against a Bare-Metal-as-a-Service backend. In this role you design, deploy, and operate that control plane — and you make sure the AIOps substrate can reach in and remediate without a human on the pager. What you'll own Production Kubernetes clusters optimized for GPU workloads at scale (100–10,000 GPUs). Nvidia GPU operator, device plugin, MIG configuration, and GPU time-slicing policies. Topology-aware scheduling: GPU locality, NVLink domain awareness, network rail affinity. Custom Resource Definitions (CRDs) for GPU workload lifecycle management. AI framework integrations: Slurm on K8S, Ray on K8S, Kubeflow. Multi-tenant isolation: namespaces, network policies, resource quotas, RBAC, pod security standards. Bare-Metal as a Service (BMaaS): automated provisioning, tenant onboarding, lifecycle, reclamation. Terraform providers and modules for infrastructure-as-code across GPU clusters. SLIs/SLOs for cluster availability, job completion rates, and provisioning latency. Incident management: runbook automation, escalation, post-incident reviews. Monitoring stack: Prometheus, Grafana, Alertmanager, PagerDuty. GPU node failure handling: automated detection, drain/cordon/taint, workload rescheduling. Feed the AIOps substrate The remediation-actuator and workflow engine land here — you make the control plane safe for automated action. Your CRDs are the schema the platform's predictors and remediators write against. Every human intervention you do this quarter becomes an autonomous workflow next quarter. What success looks like in year 1 Automated drain/reschedule around predicted GPU faults, at scale, without customer impact. BMaaS live for external tenants with self-service onboarding. Cluster availability and job-completion SLOs published and met. Job Requirement 5+ years in Kubernetes operations, with at least 2 years managing GPU workloads on K8S Deep understanding of Nvidia GPU operator, device plugin, and GPU scheduling in K8S Experience with topology-aware scheduling and GPU-specific resource management Hands-on experience building multi-tenant K8S platforms with strong isolation guarantees Experience with bare-metal server provisioning and lifecycle automation (Ironic, MAAS, or custom) Proficiency in Terraform, Helm, and GitOps workflows (ArgoCD/Flux) Strong SRE background: SLI/SLO frameworks, incident management, capacity planning Experience with Prometheus, Grafana, and alerting at scale Strong programming skills in Go or Python for operator/CRD development AIOps aptitude — you think of the K8S control plane as an execution surface for automated remediation, not just a scheduler. You've either wired an autoscaler/remediator loop into K8S or you can design one. Runbook-as-code mindset — every SRE playbook you write should be executable by the platform. Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, color, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union. #J-18808-Ljbffr
$148k - $235.75k
...see how you can make a lasting impact on the world.Join our team of innovative engineers who are building an AI Data Center AIOps platform that turns raw, high-volume telemetry into reliable, job-centric insights and automation for GPU fleets. We’re hiring a DevOps Engineer...SuggestedFull time- LeanData helps the world’s fastest-growing companies automate, simplify, and accelerate revenue.We are looking for a Senior Site Reliability Engineer to lead the strategic evolution of our cloud infrastructure. Reporting directly to the SVP of Engineering, this role is...SuggestedFull timeWork at office2 days per week
$230k - $250k
...network. It's the foundation for autonomous networking, giving engineers and AI agents the ability to know the impact of every change... ...how things have always been done.Forward is looking for a Site Reliability EngineerAbout the Role This is not a "keep the lights on"...SuggestedNight shift$128.6k - $184.9k
...global cloud platform. As a team of six engineers distributed across the US, Canada, and the... ...with a strong focus on automation, reliability, and operational excellence. We are one... ...Qualifications7+ years of experience in Site Reliability Engineering, DevOps, Infrastructure...SuggestedPermanent employmentFull timeTemporary workLocal areaWorldwideFlexible hours$168k - $270.25k
NVIDIA is looking for a Senior Site Reliability Engineer (SRE) to join its GeForce Now (GFN) team. SRE at NVIDIA ensures that our internal and external-facing GPU cloud gaming services have reliability and uptime as promised to the users and at the same time enables developers...SuggestedFull time- ...work from home day is currently Tuesday.Engineering at Lambda is responsible for building and... ...and networking teams to improve service reliability and deployment workflowsDeploy and... ...rotationYouHave 5+ years of experience in Site Reliability Engineering, Production Engineering...Work at officeLocal areaWork from homeFlexible hours
- Senior Site Reliability Engineer ILocationSan Jose, Costa Rica - RemoteSummary of roleOwn availability, the most important product feature, by continually striving for sustained operational excellence of Sumo’s planet-scale observability and security products. Work with...Flexible hours
- ...Lambda’s designated work from home day is currently Tuesday.Engineering at Lambda is responsible for building and scaling our cloud offering... ...and SLIs for Kubernetes services, workloads, and platform reliability.You6+ years of experience in a SRE, operations engineer, or...Work at officeLocal areaWork from homeFlexible hours
$168k - $270.25k
...phenomenal people like you to help us accelerate the next wave of artificial intelligence.Join our team at NVIDIA as a Senior Site reliability engineer focused on HPC storage and play a crucial role in designing, implementing, and optimizing on-prem High-Performance...Full time$267k - $356k
...day is currently Tuesday.Lambda's Storage Engineering team is the backbone behind our world-... ...workloads in the industry, which means reliability and performance aren't just goals—they're... ...defined storage across new and existing sites using tools such as Ansible, Jenkins etc...Work experience placementWork at officeLocal areaWork from homeFlexible hours$101k - $161k
...excellence has earned us several prestigious awards, such as Best Engineering Team, Best Company for Diversity, Compensation, and Work-... ...we do.Job DescriptionWho You'll Work WithWe’re looking for Site Reliability Engineers to join our growing Arista’s CloudVision-as-a-...- ...We are seeking a Senior Database Reliability Engineer (DBRE) to design, operate, and improve reliable, scalable, secure, and highly available... ...and data platforms. The role combines database engineering, site reliability engineering, Linux systems administration, and infrastructure...
$141k - $208k
...be a part of our journey! About the role We are committed to providing our customers with reliable and secure services so we are expanding our central Site Reliability Engineering team. You will be responsible for building and leading processes to ensure the reliability...Local areaRemote workHome officeFlexible hours- Remote DevOp/SRE With AI-First MindsetInsight Global is looking for a remote, DevOp/SRE with an AI-first mindset coming from a start up background to join one of our cyber security customers in the Bay Area. This role can pay 140-160k based on years of experience and skillset...Remote work
- ...A leading technology firm is in search of a Senior Wireless Network Site Reliability Engineer to manage and enhance their wireless network infrastructure. The ideal candidate has over 8 years of experience in wireless network operations and a strong background in wireless...
$187.04k - $359.72k
...systems by pushing for changes that improve reliability and velocity. Qualifications Minimum... ...degree in Computer Science, Electrical Engineering, Computer Engineering or related areas.... ...Product Ops, Corporate Functions and more. On-site presence across teams allows the company...Temporary workLocal areaOverseasShift work$207k - $300k
...implementation of solutions to enhance the reliability of systems that support F1.Scale systems... ...for multiple teams.Engage in software engineering on services written in Java, C++, and Go... ...related technical field.Experience in a Site Reliability Engineering role.Experience...$262k - $364k
...automation, and evolve systems by pushing for changes that improve reliability and velocity.Practice sustainable incident response and... ...qualifications:Master's degree in Computer Science or Engineering.Site Reliability Engineering (SRE) combines software and systems...$248k - $396.75k
...environment, where NVIDIANs are inspired to excel and make a profound global impact.NVIDIA is seeking a Senior Manager of Site Reliability Engineering to lead and reshape how IT operations function at scale. This role goes beyond traditional service management to build...Full time$184k - $287.5k
At NVIDIA, Site Reliability Engineering provides a rare chance to define, develop, and support large-scale production systems with high efficiency and availability. This demanding position merges software and systems engineering efforts to guarantee flawless service operation...Full time$210.6k - $305.1k
...Minimum Qualifications: You have led a distributed team of 5+ engineers, can demonstrate strong technical vision for your team, and ensure... ..., and basic life insurance. Please see the Cisco careers site to discover more benefits and perks. Employees may be eligible...Full timeTemporary workLocal areaFlexible hours$122.5k - $175k
...we invite you to bring your talents to Zscaler and help shape the future of cybersecurity.RoleWe are looking for a Staff Site Reliability Engineer to join our team. This is a hybrid role going into the San Jose, CA office 3 days a week, reporting to the Chief Architect...Full timeWork at officeLocal area3 days per week$146.7k - $339.3k
Immigration sponsorship is not available for this positionWhat you can expect As a Senior Lead Site Reliability Engineer, you can anticipate opportunities to work on our hybrid systems across the globe. You will be responsible for installing, configuring, and monitoring...Full timeWork at officeRemote workWorldwideShift workWeekend work- ...powers compute provisioning and infrastructure orchestration across our physical data centers. We are looking for a Senior Site Reliability Engineer to improve the reliability, scalability, and operational maturity of these systems as Lambda’s fleet and customer base...Work at officeLocal areaWork from homeFlexible hours
$124k - $271.2k
What You Can ExpectAs a Lead Staff Site Reliability Engineer, you will be one of the technical leads for our DevOps Platforms organization. This group is responsible for DevOps Platforms including cloud infrastructure, physical data center orchestration, critical security...Full timeWork at officeRemote work$120k - $200k
Sr Site Reliability Engineer (Prisma Access) 2 days ago Be among the first 25 applicants Job Description This role requires US Citizenship. Your Career Palo Alto Networks runs a large infrastructure and is one of the biggest GCP customers. As a Principal SRE, you'll be...Rotating shift$168k - $270.25k
...deploy and run an AI data center. We take great pride in providing excellent, comprehensive support to our customers! Sr Site Reliability Engineer in this role will significantly impact and contribute to the overall success of both external customers running their clusters...Full timeWorldwide$272k - $431.25k
NVIDIA is looking for a Cloud Site Reliability Engineering Architect to work in IPP's (Infrastructure, Planning and Process) Cloud Infrastructure Team. IPP is a global organization within NVIDIA. This group works with various other groups within NVIDIA such as Graphics...Full timeWork experience placementWorldwide$186.9k - $267.7k
...collaborate with a global team of software engineers and SREs responsible for delivering... ...product, and operations partners to ensure reliability and performance.Webex is powering the... ...insurance. Please see the Cisco careers site to discover more benefits and perks. Employees...Full timeTemporary workLocal areaWorldwideFlexible hoursShift work- ...to join IBM in a full‑time role between December 2027 and August 2028 upon successful completion of their degree. As a Site Reliability Engineer, you will work in an agile, collaborative environment to build, deploy, configure, and maintain systems for the IBM client...Full timeContract workPart timeFixed term contractInternshipWorldwideFlexible hoursShift work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
- site reliability engineer San Jose, CA
- site reliability engineer sre San Jose, CA
- website content developer San Jose, CA
- site leader San Jose, CA
- on-site clinical research associate (traveling/remote) San Jose, CA
- on site coordinator San Jose, CA
- official site San Jose, CA
- historic site San Jose, CA
- IT site lead San Jose, CA
- junior website developer San Jose, CA


