Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

K8 Site Reliability SME

Doist

Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure. Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to support the AI revolution. Bitdeer handles complex processes involved in computing such as equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial intelligence. Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia. To learn more, visit Position Overview You run the control plane where AIOps meets tenants — where topology-aware scheduling, self-healing, and agent-driven remediation actually execute. NeoCloud is building an AI-operated GPU cloud. Kubernetes is where all of that lands on real customer workloads: the topology-aware scheduler places jobs on the right NVLink domain, the operator drains and reschedules around predicted faults, and the tenant boundary is enforced against a Bare-Metal-as-a-Service backend. In this role you design, deploy, and operate that control plane — and you make sure the AIOps substrate can reach in and remediate without a human on the pager. What you'll own Production Kubernetes clusters optimized for GPU workloads at scale (100–10,000 GPUs). Nvidia GPU operator, device plugin, MIG configuration, and GPU time-slicing policies. Topology-aware scheduling: GPU locality, NVLink domain awareness, network rail affinity. Custom Resource Definitions (CRDs) for GPU workload lifecycle management. AI framework integrations: Slurm on K8S, Ray on K8S, Kubeflow. Multi-tenant isolation: namespaces, network policies, resource quotas, RBAC, pod security standards. Bare-Metal as a Service (BMaaS): automated provisioning, tenant onboarding, lifecycle, reclamation. Terraform providers and modules for infrastructure-as-code across GPU clusters. SLIs/SLOs for cluster availability, job completion rates, and provisioning latency. Incident management: runbook automation, escalation, post-incident reviews. Monitoring stack: Prometheus, Grafana, Alertmanager, PagerDuty. GPU node failure handling: automated detection, drain/cordon/taint, workload rescheduling. Feed the AIOps substrate The remediation-actuator and workflow engine land here — you make the control plane safe for automated action. Your CRDs are the schema the platform's predictors and remediators write against. Every human intervention you do this quarter becomes an autonomous workflow next quarter. What success looks like in year 1 Automated drain/reschedule around predicted GPU faults, at scale, without customer impact. BMaaS live for external tenants with self-service onboarding. Cluster availability and job-completion SLOs published and met. Job Requirement 5+ years in Kubernetes operations, with at least 2 years managing GPU workloads on K8S Deep understanding of Nvidia GPU operator, device plugin, and GPU scheduling in K8S Experience with topology-aware scheduling and GPU-specific resource management Hands-on experience building multi-tenant K8S platforms with strong isolation guarantees Experience with bare-metal server provisioning and lifecycle automation (Ironic, MAAS, or custom) Proficiency in Terraform, Helm, and GitOps workflows (ArgoCD/Flux) Strong SRE background: SLI/SLO frameworks, incident management, capacity planning Experience with Prometheus, Grafana, and alerting at scale Strong programming skills in Go or Python for operator/CRD development AIOps aptitude — you think of the K8S control plane as an execution surface for automated remediation, not just a scheduler. You've either wired an autoscaler/remediator loop into K8S or you can design one. Runbook-as-code mindset — every SRE playbook you write should be executable by the platform. Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, color, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union. #J-18808-Ljbffr Doist

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the K8 Site Reliability SME in Austin, TX vacancy
  • $98.58k - $138.02k

     ...This role requires a hybrid work schedule based out of one of our office locations: Austin, TX; Irvine, CA; or Akron, OH. The Site Reliability Engineer II will be responsible for supporting, enhancing, and maintaining Restaurant365’s cloud infrastructure and applications... 
    Website
    Full time
    Work at office

    Restaurant 365

    Austin, TX
    4 days ago
  • $127k - $249k

     ...scale solutions that have the ability to impact our customer’s most crucial workloads. Role OverviewWe are seeking a talented Site Reliability Engineer (SRE) with a strong infrastructure background. This role requires engineers to have a customer-first mindset to ensure... 
    Website
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    Austin, TX
    3 days ago
  •  ...Network Reliability Engineer Hybrid At Cloudflare, we are on a mission to help build a better Internet. Today the company runs one...  ...Skills, Knowledge, and Experience ~3 years of relevant Network/Site Reliability Engineering experience ~ BA/BS in Computer... 
    Website
    Local area

    Cloudflare Inc

    Austin, TX
    14 hours ago
  •  ...operational performance and availability of critical business platforms and cloud services. With a strong technical background in site reliability engineering, the ideal applicant will have excellent communication skills and a focus on continuous improvement through... 
    Website

    Take-Two Interactive

    Austin, TX
    14 hours ago
  •  ...motivated software engineer to join our Production Platform Organization. You will build the infrastructure to collect, store, and make reliability data accessible for monitoring needs, working with Product Managers and SREs to measure service quality for enterprise customers.... 
    Website

    WebHosting.coop

    Austin, TX
    7 hours ago
  • $74.1k - $148.3k

     ...forecasting, software performance analysis, and system tuning. As a Site Reliability Engineer, you will solve interesting technical challenges by...  .... You will usually get called in during major incidents as an SME, when the source of a problem is unclear. You will have the... 
    Website
    Temporary work
    Immediate start
    Flexible hours

    Oracle

    Austin, TX
    4 days ago
  • $161.5k - $190k

    Job Title Director, Operations SME - Data Centers & Critical Env Job Description Summary We are seeking a strategic and...  ..., ensuring best‑in‑class operational execution, site start‑up support, reliability, and continuous improvement across the data center portfolio... 
    Website
    Contract work
    Local area
    Flexible hours

    Cushman & Wakefield

    Austin, TX
    4 days ago
  • Selby Jennings is seeking a Senior Site Reliability Engineer to scale and support critical workflow orchestration and automation platforms across the organization. The role sits in Platform Engineering, delivering highly available and resilient infrastructure for business... 
    Website

    Selby Jennings

    Austin, TX
    5 days ago
  •  ...experiences. You will design scalable data pipelines, integrate LLM capabilities, and collaborate across teams to deliver reliable, production-grade AI systems on-site in Austin. You will mentor teammates and drive best practices while focusing on reliability, observability, and... 
    Website

    Charles Schwab

    Austin, TX
    2 days ago
  • City of Austin seeks a Compliance Analyst Senior to oversee Reliability Requirements for electric operations and energy markets. You will...  ...Operations, a strong knowledge of reliability standards, and the ability to travel across sites as needed. #J-18808-Ljbffr austintexas
    Website

    austintexas

    Austin, TX
    1 day ago
  •  ...hardware operations, network infrastructure, and third-party data center vendors to ensure operability and reliability. The role emphasizes driving cost efficiency, maintaining safety and environmental standards, and guiding site-level programs. #J-18808-Ljbffr Google
    Website

    Google

    Austin, TX
    5 days ago
  •  ..., Inc is seeking a Senior Electrical Maintenance Specialist to support multi-site industrial facilities in the Austin/Round Rock region. The role focuses on electrical maintenance, reliability, and high voltage system troubleshooting across 13.8kV equipment and substations... 
    Website

    Catalyst Recruiting, Inc

    Austin, TX
    4 days ago
  •  ...engage directly with the customer to optimize torque tools and fastening equipment, driving operational improvements and reliability. Reporting to the Site Manager, you will perform maintenance, calibration, troubleshooting, and network-related support on tooling, ensuring... 
    Website

    Atlas Copco Tools & Assembly Systems LLC

    Austin, TX
    1 day ago
  •  ...advance your career. THE ROLEAMD is seeking a Principal Cluster Reliability Architect to define and drive the reliability strategy for next...  ....Operational Readiness & Day-2 OperationsPartner with Site Reliability Engineering (SRE) and Platform Operations teams to... 
    Website

    AMD

    Austin, TX
    4 days ago
  •  ...DC engineering teams, hardware operations, network infrastructure, and vendors to ensure operability, maintainability, and reliability across sites. The role drives cost efficiency, process improvements, and incident response, while upholding safety and environmental... 
    Website

    Google Inc.

    Austin, TX
    4 days ago
  •  ...that matter to millions of clients, and to grow your career in one of the most exciting areas of technology today.As a Senior AI Site Reliability Engineer you will support reliability efforts for cutting-edge GenAI applications that enhance the client experience and... 
    Website
    Full time

    The Charles Schwab Corporation

    Austin, TX
    14 hours ago
  • $169.5k - $233k

     ...Protection Engineer / Subject Matter Expert (SME) to support our rapidly growing Data...  ...fire protection designs that meet stringent reliability, operational, and code requirements. You...  ...and oversee hydraulic calculations for site fire water distribution networks, fire pumps... 
    Website
    Full time
    For contractors
    Work at office
    Local area
    Remote work

    Jacobs

    Austin, TX
    3 days ago
  •  ...efficiency to support 100% uptime. You will lead daily engineering and facility operations on-site in Austin, manage PM programs, and coordinate vendors while focusing on reliability and safe operations. The role requires a bachelor’s in electrical or mechanical... 
    Website

    NextGenEnergyJobs

    Austin, TX
    4 days ago
  • Recognized as the No. 1 site trusted by real estate professionals, Realtor.com has been at the forefront of online real estate for...  ...confidence through expert guidance.We are seeking a Senior Site Reliability Engineer to join our newly formed Operations Excellence organization... 
    Website
    Work at office
    Local area

    Realtor.com

    Austin, TX
    1 day ago
  •  ...10s to 100s of GWs. Fly the plane while it's being built. Sites come online in pieces, and you keep the live ones running flawlessly...  ...you set become the standard. Role Scope Own reliability for named customer workloads: their clusters, their SLAs, their... 
    Website

    Fluidstack

    Austin, TX
    14 hours ago
  •  ...network engineers, systems architects, and game studio developers. This is an ownership role: driving technical direction, influencing reliability from architecture review through production operation, and closing the gap between what engineering ships and what players... 
    Website

    2K Games

    Austin, TX
    2 days ago
  •  ...SRE to join our Platform Engineering team as the operations owner of our observability platforms. You’ll be responsible for the reliability, scalability, and continued evolution of the tools that give our engineering organization visibility into everything they build and... 
    Website
    Full time

    Dimensional Fund Advisors

    Austin, TX
    14 hours ago
  •  ...table and foster a dynamic work environment where new ideas thrive. Are you ready to join our team and make an impact?As a Senior Site Reliability Engineer at TeamViewer, you’ll be a key player in ensuring the reliability, scalability, and performance of our Azure-based... 
    Website
    Temporary work
    Casual work
    Worldwide

    TeamViewer

    Austin, TX
    14 hours ago
  • $116k - $159.5k

     ...join our winning team that is focused on Quality Engineering and Reliability Engineering. We are working with engineers and scientists...  ...equipment failures during stress tests in the lab, or at customer sites. Our engineering judgment is in demand every day as we are striving... 
    Website
    Full time
    Worldwide

    Applied Materials

    Austin, TX
    2 days ago
  •  ...opening for SharePoint Online Migration Specialist / Techno-Functional SME rec 807555 This position is 6 months, with the option of...  ...migrating data to Sharepoint specifically from a Canvas LMS site Highly desired Prior experience migrating data to Sharepoint... 
    Website
    Local area
    Remote work

    FHR

    Austin, TX
    22 days ago
  • $109.65k - $182.76k

     ...and encrypt data to make the connected world more secure.Austin, TX - Hybrid (3 days a week)Position SummaryWe are seeking a Site Reliability Engineer to ensure the high level of service and operation excellence for the development of the innovative and ambitious Telecommunication... 
    Website
    Full time
    Local area
    3 days per week

    Thales Group

    Austin, TX
    2 days ago
  • $144k - $209k

     ...and product de-risk at an early stage of development.Lead system reliability efforts by working with other organizations to define...  ...plan.Develop mission profiles for chasis, rack from integration sites to field (data centers) that help predict field reliability.Implement... 
    Website
    Contract work
    Worldwide

    Google

    Austin, TX
    2 days ago
  •  ...ReliabilityServe as a primary escalation point for production support across each of the developer tooling ecosystemsOwn the operational reliability of developer tooling ecosystems, including Python toolchains (uv, Poetry, etc.), .NET SDKs, and IDE configurations and... 
    Website
    Full time
    Local area

    Dimensional Fund Advisors

    Austin, TX
    1 day ago
  •  ...fully intend for the selected candidate for this role to work on site in the specified location(s).Retail Web Technologies builds high...  ...millions of Schwab clients. In this role, you’ll contribute to reliable, secure, and scalable web experiences by designing, building, testing... 
    Website
    Full time
    Work at office

    The Charles Schwab Corporation

    Austin, TX
    4 days ago
  •  ...Job Description Job Description Senior Site Reliability Engineer - Developer Productivity & Tooling Location: Austin, TX Area (Remote-First) Requirement: Candidates must be within commuting distance of Austin. About the Role A leading asset management... 
    Website
    Local area
    Remote work

    Selby Jennings

    Austin, TX
    15 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to K8 Site Reliability SME. Be the first to apply!