K8 Site Reliability SME
Doist
Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure. Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to support the AI revolution. Bitdeer handles complex processes involved in computing such as equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial intelligence. Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia. To learn more, visit Position Overview You run the control plane where AIOps meets tenants — where topology-aware scheduling, self-healing, and agent-driven remediation actually execute. NeoCloud is building an AI-operated GPU cloud. Kubernetes is where all of that lands on real customer workloads: the topology-aware scheduler places jobs on the right NVLink domain, the operator drains and reschedules around predicted faults, and the tenant boundary is enforced against a Bare-Metal-as-a-Service backend. In this role you design, deploy, and operate that control plane — and you make sure the AIOps substrate can reach in and remediate without a human on the pager. What you'll own Production Kubernetes clusters optimized for GPU workloads at scale (100–10,000 GPUs). Nvidia GPU operator, device plugin, MIG configuration, and GPU time-slicing policies. Topology-aware scheduling: GPU locality, NVLink domain awareness, network rail affinity. Custom Resource Definitions (CRDs) for GPU workload lifecycle management. AI framework integrations: Slurm on K8S, Ray on K8S, Kubeflow. Multi-tenant isolation: namespaces, network policies, resource quotas, RBAC, pod security standards. Bare-Metal as a Service (BMaaS): automated provisioning, tenant onboarding, lifecycle, reclamation. Terraform providers and modules for infrastructure-as-code across GPU clusters. SLIs/SLOs for cluster availability, job completion rates, and provisioning latency. Incident management: runbook automation, escalation, post-incident reviews. Monitoring stack: Prometheus, Grafana, Alertmanager, PagerDuty. GPU node failure handling: automated detection, drain/cordon/taint, workload rescheduling. Feed the AIOps substrate The remediation-actuator and workflow engine land here — you make the control plane safe for automated action. Your CRDs are the schema the platform's predictors and remediators write against. Every human intervention you do this quarter becomes an autonomous workflow next quarter. What success looks like in year 1 Automated drain/reschedule around predicted GPU faults, at scale, without customer impact. BMaaS live for external tenants with self-service onboarding. Cluster availability and job-completion SLOs published and met. Job Requirement 5+ years in Kubernetes operations, with at least 2 years managing GPU workloads on K8S Deep understanding of Nvidia GPU operator, device plugin, and GPU scheduling in K8S Experience with topology-aware scheduling and GPU-specific resource management Hands-on experience building multi-tenant K8S platforms with strong isolation guarantees Experience with bare-metal server provisioning and lifecycle automation (Ironic, MAAS, or custom) Proficiency in Terraform, Helm, and GitOps workflows (ArgoCD/Flux) Strong SRE background: SLI/SLO frameworks, incident management, capacity planning Experience with Prometheus, Grafana, and alerting at scale Strong programming skills in Go or Python for operator/CRD development AIOps aptitude — you think of the K8S control plane as an execution surface for automated remediation, not just a scheduler. You've either wired an autoscaler/remediator loop into K8S or you can design one. Runbook-as-code mindset — every SRE playbook you write should be executable by the platform. Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, color, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union. #J-18808-Ljbffr Doist
$98.58k - $138.02k
...This role requires a hybrid work schedule based out of one of our office locations: Austin, TX; Irvine, CA; or Akron, OH. The Site Reliability Engineer II will be responsible for supporting, enhancing, and maintaining Restaurant365’s cloud infrastructure and applications...WebsiteFull timeWork at office$127k - $249k
...scale solutions that have the ability to impact our customer’s most crucial workloads. Role OverviewWe are seeking a talented Site Reliability Engineer (SRE) with a strong infrastructure background. This role requires engineers to have a customer-first mindset to ensure...WebsiteLocal areaRemote workWorldwideFlexible hours- ...Network Reliability Engineer Hybrid At Cloudflare, we are on a mission to help build a better Internet. Today the company runs one... ...Skills, Knowledge, and Experience ~3 years of relevant Network/Site Reliability Engineering experience ~ BA/BS in Computer...WebsiteLocal area
- ...operational performance and availability of critical business platforms and cloud services. With a strong technical background in site reliability engineering, the ideal applicant will have excellent communication skills and a focus on continuous improvement through...Website
- ...motivated software engineer to join our Production Platform Organization. You will build the infrastructure to collect, store, and make reliability data accessible for monitoring needs, working with Product Managers and SREs to measure service quality for enterprise customers....Website
$74.1k - $148.3k
...forecasting, software performance analysis, and system tuning. As a Site Reliability Engineer, you will solve interesting technical challenges by... .... You will usually get called in during major incidents as an SME, when the source of a problem is unclear. You will have the...WebsiteTemporary workImmediate startFlexible hours$161.5k - $190k
Job Title Director, Operations SME - Data Centers & Critical Env Job Description Summary We are seeking a strategic and... ..., ensuring best‑in‑class operational execution, site start‑up support, reliability, and continuous improvement across the data center portfolio...WebsiteContract workLocal areaFlexible hours- Selby Jennings is seeking a Senior Site Reliability Engineer to scale and support critical workflow orchestration and automation platforms across the organization. The role sits in Platform Engineering, delivering highly available and resilient infrastructure for business...Website
- ...experiences. You will design scalable data pipelines, integrate LLM capabilities, and collaborate across teams to deliver reliable, production-grade AI systems on-site in Austin. You will mentor teammates and drive best practices while focusing on reliability, observability, and...Website
- City of Austin seeks a Compliance Analyst Senior to oversee Reliability Requirements for electric operations and energy markets. You will... ...Operations, a strong knowledge of reliability standards, and the ability to travel across sites as needed. #J-18808-Ljbffr austintexasWebsite
- ...hardware operations, network infrastructure, and third-party data center vendors to ensure operability and reliability. The role emphasizes driving cost efficiency, maintaining safety and environmental standards, and guiding site-level programs. #J-18808-Ljbffr GoogleWebsite
- ..., Inc is seeking a Senior Electrical Maintenance Specialist to support multi-site industrial facilities in the Austin/Round Rock region. The role focuses on electrical maintenance, reliability, and high voltage system troubleshooting across 13.8kV equipment and substations...Website
- ...engage directly with the customer to optimize torque tools and fastening equipment, driving operational improvements and reliability. Reporting to the Site Manager, you will perform maintenance, calibration, troubleshooting, and network-related support on tooling, ensuring...Website
- ...advance your career. THE ROLEAMD is seeking a Principal Cluster Reliability Architect to define and drive the reliability strategy for next... ....Operational Readiness & Day-2 OperationsPartner with Site Reliability Engineering (SRE) and Platform Operations teams to...Website
- ...DC engineering teams, hardware operations, network infrastructure, and vendors to ensure operability, maintainability, and reliability across sites. The role drives cost efficiency, process improvements, and incident response, while upholding safety and environmental...Website
- ...that matter to millions of clients, and to grow your career in one of the most exciting areas of technology today.As a Senior AI Site Reliability Engineer you will support reliability efforts for cutting-edge GenAI applications that enhance the client experience and...WebsiteFull time
$169.5k - $233k
...Protection Engineer / Subject Matter Expert (SME) to support our rapidly growing Data... ...fire protection designs that meet stringent reliability, operational, and code requirements. You... ...and oversee hydraulic calculations for site fire water distribution networks, fire pumps...WebsiteFull timeFor contractorsWork at officeLocal areaRemote work- ...efficiency to support 100% uptime. You will lead daily engineering and facility operations on-site in Austin, manage PM programs, and coordinate vendors while focusing on reliability and safe operations. The role requires a bachelor’s in electrical or mechanical...Website
- Recognized as the No. 1 site trusted by real estate professionals, Realtor.com has been at the forefront of online real estate for... ...confidence through expert guidance.We are seeking a Senior Site Reliability Engineer to join our newly formed Operations Excellence organization...WebsiteWork at officeLocal area
- ...10s to 100s of GWs. Fly the plane while it's being built. Sites come online in pieces, and you keep the live ones running flawlessly... ...you set become the standard. Role Scope Own reliability for named customer workloads: their clusters, their SLAs, their...Website
- ...network engineers, systems architects, and game studio developers. This is an ownership role: driving technical direction, influencing reliability from architecture review through production operation, and closing the gap between what engineering ships and what players...Website
- ...SRE to join our Platform Engineering team as the operations owner of our observability platforms. You’ll be responsible for the reliability, scalability, and continued evolution of the tools that give our engineering organization visibility into everything they build and...WebsiteFull time
- ...table and foster a dynamic work environment where new ideas thrive. Are you ready to join our team and make an impact?As a Senior Site Reliability Engineer at TeamViewer, you’ll be a key player in ensuring the reliability, scalability, and performance of our Azure-based...WebsiteTemporary workCasual workWorldwide
$116k - $159.5k
...join our winning team that is focused on Quality Engineering and Reliability Engineering. We are working with engineers and scientists... ...equipment failures during stress tests in the lab, or at customer sites. Our engineering judgment is in demand every day as we are striving...WebsiteFull timeWorldwide- ...opening for SharePoint Online Migration Specialist / Techno-Functional SME rec 807555 This position is 6 months, with the option of... ...migrating data to Sharepoint specifically from a Canvas LMS site Highly desired Prior experience migrating data to Sharepoint...WebsiteLocal areaRemote work
$109.65k - $182.76k
...and encrypt data to make the connected world more secure.Austin, TX - Hybrid (3 days a week)Position SummaryWe are seeking a Site Reliability Engineer to ensure the high level of service and operation excellence for the development of the innovative and ambitious Telecommunication...WebsiteFull timeLocal area3 days per week$144k - $209k
...and product de-risk at an early stage of development.Lead system reliability efforts by working with other organizations to define... ...plan.Develop mission profiles for chasis, rack from integration sites to field (data centers) that help predict field reliability.Implement...WebsiteContract workWorldwide- ...ReliabilityServe as a primary escalation point for production support across each of the developer tooling ecosystemsOwn the operational reliability of developer tooling ecosystems, including Python toolchains (uv, Poetry, etc.), .NET SDKs, and IDE configurations and...WebsiteFull timeLocal area
- ...fully intend for the selected candidate for this role to work on site in the specified location(s).Retail Web Technologies builds high... ...millions of Schwab clients. In this role, you’ll contribute to reliable, secure, and scalable web experiences by designing, building, testing...WebsiteFull timeWork at office
- ...Job Description Job Description Senior Site Reliability Engineer - Developer Productivity & Tooling Location: Austin, TX Area (Remote-First) Requirement: Candidates must be within commuting distance of Austin. About the Role A leading asset management...WebsiteLocal areaRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to K8 Site Reliability SME. Be the first to apply!
- website content developer Austin, TX
- site leader Austin, TX
- on-site clinical research associate (traveling/remote) Austin, TX
- on site coordinator Austin, TX
- official site Austin, TX
- site recruiter Austin, TX
- historic site Austin, TX
- IT site lead Austin, TX
- junior website developer Austin, TX
- site safety Austin, TX


