Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer

Bitdeer Group

Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure. Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to support the AI revolution. Bitdeer handles complex processes involved in computing such as equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial intelligence. Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia. To learn more, visit Position Overview You run the control plane where AIOps meets tenants — where topology-aware scheduling, self-healing, and agent-driven remediation actually execute. NeoCloud is building an AI-operated GPU cloud. Kubernetes is where all of that lands on real customer workloads: the topology-aware scheduler places jobs on the right NVLink domain, the operator drains and reschedules around predicted faults, and the tenant boundary is enforced against a Bare-Metal-as-a-Service backend. In this role you design, deploy, and operate that control plane — and you make sure the AIOps substrate can reach in and remediate without a human on the pager. What you'll own Production Kubernetes clusters optimized for GPU workloads at scale (100–10,000 GPUs). Nvidia GPU operator, device plugin, MIG configuration, and GPU time-slicing policies. Topology-aware scheduling: GPU locality, NVLink domain awareness, network rail affinity. Custom Resource Definitions (CRDs) for GPU workload lifecycle management. AI framework integrations: Slurm on K8S, Ray on K8S, Kubeflow. Multi-tenant isolation: namespaces, network policies, resource quotas, RBAC, pod security standards. Bare-Metal as a Service (BMaaS): automated provisioning, tenant onboarding, lifecycle, reclamation. Terraform providers and modules for infrastructure-as-code across GPU clusters. SLIs/SLOs for cluster availability, job completion rates, and provisioning latency. Incident management: runbook automation, escalation, post-incident reviews. Monitoring stack: Prometheus, Grafana, Alertmanager, PagerDuty. GPU node failure handling: automated detection, drain/cordon/taint, workload rescheduling. Feed the AIOps substrate The remediation-actuator and workflow engine land here — you make the control plane safe for automated action. Your CRDs are the schema the platform's predictors and remediators write against. Every human intervention you do this quarter becomes an autonomous workflow next quarter. What success looks like in year 1 Automated drain/reschedule around predicted GPU faults, at scale, without customer impact. BMaaS live for external tenants with self-service onboarding. Cluster availability and job-completion SLOs published and met. Job Requirement 5+ years in Kubernetes operations, with at least 2 years managing GPU workloads on K8S Deep understanding of Nvidia GPU operator, device plugin, and GPU scheduling in K8S Experience with topology-aware scheduling and GPU-specific resource management Hands-on experience building multi-tenant K8S platforms with strong isolation guarantees Experience with bare-metal server provisioning and lifecycle automation (Ironic, MAAS, or custom) Proficiency in Terraform, Helm, and GitOps workflows (ArgoCD/Flux) Strong SRE background: SLI/SLO frameworks, incident management, capacity planning Experience with Prometheus, Grafana, and alerting at scale Strong programming skills in Go or Python for operator/CRD development AIOps aptitude — you think of the K8S control plane as an execution surface for automated remediation, not just a scheduler. You've either wired an autoscaler/remediator loop into K8S or you can design one. Runbook-as-code mindset — every SRE playbook you write should be executable by the platform. Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, color, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union. #J-18808-Ljbffr

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in San Jose, CA vacancy
  • $148k - $235.75k

     ...see how you can make a lasting impact on the world.Join our team of innovative engineers who are building an AI Data Center AIOps platform that turns raw, high-volume telemetry into reliable, job-centric insights and automation for GPU fleets. We’re hiring a DevOps Engineer... 
    Suggested
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • LeanData helps the world’s fastest-growing companies automate, simplify, and accelerate revenue.We are looking for a Senior Site Reliability Engineer to lead the strategic evolution of our cloud infrastructure. Reporting directly to the SVP of Engineering, this role is... 
    Suggested
    Full time
    Work at office
    2 days per week

    LeanData

    Santa Clara, CA
    8 hours ago
  • $230k - $250k

     ...network. It's the foundation for autonomous networking, giving engineers and AI agents the ability to know the impact of every change...  ...how things have always been done.Forward is looking for a Site Reliability EngineerAbout the Role This is not a "keep the lights on"... 
    Suggested
    Night shift

    Forward Networks

    Santa Clara, CA
    8 hours ago
  • $128.6k - $184.9k

     ...global cloud platform. As a team of six engineers distributed across the US, Canada, and the...  ...with a strong focus on automation, reliability, and operational excellence. We are one...  ...Qualifications7+ years of experience in Site Reliability Engineering, DevOps, Infrastructure... 
    Suggested
    Permanent employment
    Full time
    Temporary work
    Local area
    Worldwide
    Flexible hours

    CISCO Systems

    Santa Clara, CA
    4 days ago
  • $168k - $270.25k

    NVIDIA is looking for a Senior Site Reliability Engineer (SRE) to join its GeForce Now (GFN) team. SRE at NVIDIA ensures that our internal and external-facing GPU cloud gaming services have reliability and uptime as promised to the users and at the same time enables developers... 
    Suggested
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  •  ...work from home day is currently Tuesday.Engineering at Lambda is responsible for building and...  ...and networking teams to improve service reliability and deployment workflowsDeploy and...  ...rotationYouHave 5+ years of experience in Site Reliability Engineering, Production Engineering... 
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    2 days ago
  • Senior Site Reliability Engineer ILocationSan Jose, Costa Rica - RemoteSummary of roleOwn availability, the most important product feature, by continually striving for sustained operational excellence of Sumo’s planet-scale observability and security products. Work with... 
    Flexible hours

    Sumo Logic

    San Jose, CA
    3 days ago
  •  ...Lambda’s designated work from home day is currently Tuesday.Engineering at Lambda is responsible for building and scaling our cloud offering...  ...and SLIs for Kubernetes services, workloads, and platform reliability.You6+ years of experience in a SRE, operations engineer, or... 
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    3 days ago
  • $168k - $270.25k

     ...phenomenal people like you to help us accelerate the next wave of artificial intelligence.Join our team at NVIDIA as a Senior Site reliability engineer focused on HPC storage and play a crucial role in designing, implementing, and optimizing on-prem High-Performance... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $267k - $356k

     ...day is currently Tuesday.Lambda's Storage Engineering team is the backbone behind our world-...  ...workloads in the industry, which means reliability and performance aren't just goals—they're...  ...defined storage across new and existing sites using tools such as Ansible, Jenkins etc... 
    Work experience placement
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    2 days ago
  • $101k - $161k

     ...excellence has earned us several prestigious awards, such as Best Engineering Team, Best Company for Diversity, Compensation, and Work-...  ...we do.Job DescriptionWho You'll Work WithWe’re looking for Site Reliability Engineers to join our growing Arista’s CloudVision-as-a-... 

    Arista Networks

    Santa Clara, CA
    3 days ago
  •  ...We are seeking a Senior Database Reliability Engineer (DBRE) to design, operate, and improve reliable, scalable, secure, and highly available...  ...and data platforms. The role combines database engineering, site reliability engineering, Linux systems administration, and infrastructure... 

    Neshent Technologies

    Los Gatos, CA
    4 days ago
  • $141k - $208k

     ...be a part of our journey! About the role We are committed to providing our customers with reliable and secure services so we are expanding our central Site Reliability Engineering team. You will be responsible for building and leading processes to ensure the reliability... 
    Local area
    Remote work
    Home office
    Flexible hours

    GrabJobs

    San Jose, CA
    2 days ago
  • Remote DevOp/SRE With AI-First MindsetInsight Global is looking for a remote, DevOp/SRE with an AI-first mindset coming from a start up background to join one of our cyber security customers in the Bay Area. This role can pay 140-160k based on years of experience and skillset...
    Remote work

    Insight Global

    Santa Clara, CA
    1 day ago
  •  ...A leading technology firm is in search of a Senior Wireless Network Site Reliability Engineer to manage and enhance their wireless network infrastructure. The ideal candidate has over 8 years of experience in wireless network operations and a strong background in wireless... 

    TechDigital Group

    Santa Clara, CA
    5 days ago
  • $187.04k - $359.72k

     ...systems by pushing for changes that improve reliability and velocity. Qualifications Minimum...  ...degree in Computer Science, Electrical Engineering, Computer Engineering or related areas....  ...Product Ops, Corporate Functions and more. On-site presence across teams allows the company... 
    Temporary work
    Local area
    Overseas
    Shift work

    Tik Tok

    San Jose, CA
    5 days ago
  • $207k - $300k

     ...implementation of solutions to enhance the reliability of systems that support F1.Scale systems...  ...for multiple teams.Engage in software engineering on services written in Java, C++, and Go...  ...related technical field.Experience in a Site Reliability Engineering role.Experience... 

    Google

    San Jose, CA
    1 day ago
  • $262k - $364k

     ...automation, and evolve systems by pushing for changes that improve reliability and velocity.Practice sustainable incident response and...  ...qualifications:Master's degree in Computer Science or Engineering.Site Reliability Engineering (SRE) combines software and systems... 

    Google

    San Jose, CA
    3 days ago
  • $248k - $396.75k

     ...environment, where NVIDIANs are inspired to excel and make a profound global impact.NVIDIA is seeking a Senior Manager of Site Reliability Engineering to lead and reshape how IT operations function at scale. This role goes beyond traditional service management to build... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

    At NVIDIA, Site Reliability Engineering provides a rare chance to define, develop, and support large-scale production systems with high efficiency and availability. This demanding position merges software and systems engineering efforts to guarantee flawless service operation... 
    Full time

    Nvidia

    Santa Clara, CA
    8 hours ago
  • $210.6k - $305.1k

     ...Minimum Qualifications:  You have led a distributed team of 5+ engineers, can demonstrate strong technical vision for your team, and ensure...  ..., and basic life insurance. Please see the Cisco careers site to discover more benefits and perks. Employees may be eligible... 
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    San Jose, CA
    2 days ago
  • $122.5k - $175k

     ...we invite you to bring your talents to Zscaler and help shape the future of cybersecurity.RoleWe are looking for a Staff Site Reliability Engineer to join our team. This is a hybrid role going into the San Jose, CA office 3 days a week, reporting to the Chief Architect... 
    Full time
    Work at office
    Local area
    3 days per week

    Zscaler

    San Jose, CA
    1 day ago
  • $146.7k - $339.3k

    Immigration sponsorship is not available for this positionWhat you can expect As a Senior Lead Site Reliability Engineer, you can anticipate opportunities to work on our hybrid systems across the globe. You will be responsible for installing, configuring, and monitoring... 
    Full time
    Work at office
    Remote work
    Worldwide
    Shift work
    Weekend work

    Zoom

    San Jose, CA
    8 hours ago
  •  ...powers compute provisioning and infrastructure orchestration across our physical data centers. We are looking for a Senior Site Reliability Engineer to improve the reliability, scalability, and operational maturity of these systems as Lambda’s fleet and customer base... 
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    8 hours ago
  • $124k - $271.2k

    What You Can ExpectAs a Lead Staff Site Reliability Engineer, you will be one of the technical leads for our DevOps Platforms organization. This group is responsible for DevOps Platforms including cloud infrastructure, physical data center orchestration, critical security... 
    Full time
    Work at office
    Remote work

    Zoom

    San Jose, CA
    3 days ago
  • $120k - $200k

    Sr Site Reliability Engineer (Prisma Access) 2 days ago Be among the first 25 applicants Job Description This role requires US Citizenship. Your Career Palo Alto Networks runs a large infrastructure and is one of the biggest GCP customers. As a Principal SRE, you'll be... 
    Rotating shift

    Palo Alto Networks

    Santa Clara, CA
    5 days ago
  • $168k - $270.25k

     ...deploy and run an AI data center. We take great pride in providing excellent, comprehensive support to our customers! ​Sr Site Reliability Engineer in this role will significantly impact and contribute to the overall success of both external customers running their clusters... 
    Full time
    Worldwide

    Nvidia

    Santa Clara, CA
    8 hours ago
  • $272k - $431.25k

    NVIDIA is looking for a Cloud Site Reliability Engineering Architect to work in IPP's (Infrastructure, Planning and Process) Cloud Infrastructure Team. IPP is a global organization within NVIDIA. This group works with various other groups within NVIDIA such as Graphics... 
    Full time
    Work experience placement
    Worldwide

    Nvidia

    Santa Clara, CA
    3 days ago
  • $186.9k - $267.7k

     ...collaborate with a global team of software engineers and SREs responsible for delivering...  ...product, and operations partners to ensure reliability and performance.Webex is powering the...  ...insurance. Please see the Cisco careers site to discover more benefits and perks. Employees... 
    Full time
    Temporary work
    Local area
    Worldwide
    Flexible hours
    Shift work

    CISCO Systems

    Milpitas, CA
    3 days ago
  •  ...to join IBM in a full‑time role between December 2027 and August 2028 upon successful completion of their degree. As a Site Reliability Engineer, you will work in an agile, collaborative environment to build, deploy, configure, and maintain systems for the IBM client... 
    Full time
    Contract work
    Part time
    Fixed term contract
    Internship
    Worldwide
    Flexible hours
    Shift work

    IBM

    San Jose, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!