K8 Site Reliability SME
Bitdeer (NASDAQ: BTDR)
Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure. Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to support the AI revolution. Bitdeer handles complex processes involved in computing such as equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial intelligence. Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia. To learn more, visit Position Overview You run the control plane where AIOps meets tenants — where topology-aware scheduling, self-healing, and agent-driven remediation actually execute. NeoCloud is building an AI-operated GPU cloud. Kubernetes is where all of that lands on real customer workloads: the topology-aware scheduler places jobs on the right NVLink domain, the operator drains and reschedules around predicted faults, and the tenant boundary is enforced against a Bare-Metal-as-a-Service backend. In this role you design, deploy, and operate that control plane — and you make sure the AIOps substrate can reach in and remediate without a human on the pager. What You'll Own Production Kubernetes clusters optimized for GPU workloads at scale (100–10,000 GPUs). Nvidia GPU operator, device plugin, MIG configuration, and GPU time-slicing policies. Topology-aware scheduling: GPU locality, NVLink domain awareness, network rail affinity. Custom Resource Definitions (CRDs) for GPU workload lifecycle management. AI framework integrations: Slurm on K8S, Ray on K8S, Kubeflow. Multi-tenant isolation: namespaces, network policies, resource quotas, RBAC, pod security standards. Bare-Metal as a Service (BMaaS): automated provisioning, tenant onboarding, lifecycle, reclamation. Terraform providers and modules for infrastructure-as-code across GPU clusters. SLIs/SLOs for cluster availability, job completion rates, and provisioning latency. Incident management: runbook automation, escalation, post-incident reviews.Monitoring stack: Prometheus, Grafana, Alertmanager, PagerDuty. GPU node failure handling: automated detection, drain/cordon/taint, workload rescheduling. Feed the AIOps substrate The remediation-actuator and workflow engine land here — you make the control plane safe for automated action. Your CRDs are the schema the platform's predictors and remediators write against. Every human intervention you do this quarter becomes an autonomous workflow next quarter. What Success Looks Like In Year 1 Automated drain/reschedule around predicted GPU faults, at scale, without customer impact. BMaaS live for external tenants with self-service onboarding. Cluster availability and job-completion SLOs published and met. Job Requirement 5+ years in Kubernetes operations, with at least 2 years managing GPU workloads on K8S Deep understanding of Nvidia GPU operator, device plugin, and GPU scheduling in K8S Experience with topology-aware scheduling and GPU-specific resource management Hands-on experience building multi-tenant K8S platforms with strong isolation guaranteesExperience with bare-metal server provisioning and lifecycle automation (Ironic, MAAS, or custom) Proficiency in Terraform, Helm, and GitOps workflows (ArgoCD/Flux) Strong SRE background: SLI/SLO frameworks, incident management, capacity planning Experience with Prometheus, Grafana, and alerting at scale Strong programming skills in Go or Python for operator/CRD development AIOps aptitude — you think of the K8S control plane as an execution surface for automated remediation, not just a scheduler. You've either wired an autoscaler/remediator loop into K8S or you can design one. Runbook-as-code mindset — every SRE playbook you write should be executable by the platform. Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, color, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union. #J-18808-Ljbffr Bitdeer (NASDAQ: BTDR)
- ...Required U.S. Citizenship / No clearance needed / 100% remote within the US Staff Site Reliability Engineer / Cloud SME Location: 100% remote in the continental US Type: Long-term contract (3+ years) Role Summary As the Staff SRE/Cloud SME, you will be...WebsiteLong term contractRemote work
$98.58k - $138.02k
...office locations: Austin, TX; Irvine, CA; or Akron, OH.TheSite Reliability Engineer IIwillbe responsible forsupporting, enhancing, and maintaining... .... Qualified candidates willdemonstrategrowingexpertisein site reliability practices, with skills in incident response, system...WebsiteFull timeWork at office$161.5k - $190k
...Job Title Director, Operations SME – Data Centers & Critical Env Job Description Summary We are seeking a strategic and... ..., ensuring best‑in‑class operational execution, site start‑up support, reliability, and continuous improvement across the data center portfolio...WebsiteContract workLocal areaFlexible hours- Electrical SME - Data Center OperationsJLL empowers you to shape a brighter way.Our people... ...data center facilities. You'll ensure the reliable operation of high-voltage electrical... ...United States without sponsorship.Location:On-site - Austin, TXIf this job description resonates...WebsiteDaily paidFlexible hours
- ...maintain a shared robotics software stack for HERO deployments across sites. You will collaborate with researchers and engineers to... ...robust, deployable systems, develop ROS 2 packages, and ensure reliability, testing, and deployment readiness. #J-18808-Ljbffr Phase2 TechnologyWebsite
- ...and deploy a common robotics software infrastructure for HERO deployments across multiple sites. You will own software stack architecture, integrate components, and ensure reliability in long-running robot operations. You will work with researchers and the Center Robotics...Website
- ...Installer in Austin, TX. This customer service role emphasizes reliability and determination as you load and unload gear, set drape, furniture... ...will participate in daily activities in the warehouse and on site, support a crew leader, and provide outstanding service to...Website
- ...channel. We work with enterprise retailers and move fast. Our infra has to match. The role We\'re looking for a Senior SRE to own the reliability, scalability, and operational posture of Satsuma\'s multi-cloud infrastructure. You\'ll be the person who keeps things running,...Website
- ...Technician 3 to support product, package, and process qualification activities within Reliability Engineering. You will execute reliability stress tests on semiconductor devices in an on-site lab, document results, and help with qualification readiness and schedules. You...Website
- ...Austin-based role seeks a Product Development Specialist to advance reliability and qualification activities for electromechanical systems. You... ...to ensure manufacturability and durable performance. This on-site position requires experience in product development,...Website
- ...plant operations, including chillers, boilers, and building automation. You will guide a team of operators and engineers, ensuring reliable, safe performance and client-ready communication daily. The role emphasizes hands-on leadership, preventive maintenance, and...Website
- ...Production Platform Organization. You will build infrastructure to collect, store, and make reliability data accessible for monitoring needs, working with Product Managers and Site Reliability Engineers on service quality for enterprise customers. You will define...Website
- ...and performance. The role will set strategy, governance, and standards for the fleet, ensuring safe, reliable, high-quality, and cost-effective execution across sites and functions. The VP will sponsor enterprise-wide programs, lead organizational transformation, and align...Website
- Charles Schwab seeks an engineer to join our Site Reliability/Platform Reliability team in Austin. You will help support cloud and login platforms, drive automation, and contribute to production operations and incident response. You’ll work with platform teams to improve...Website
- ...efficiency to support 100% uptime. You will lead daily engineering and facility operations on-site in Austin, manage PM programs, and coordinate vendors while focusing on reliability and safe operations. The role requires a bachelor’s in electrical or mechanical...Website
- ...local travel, supporting commercial landscape maintenance across sites. You will work with Crew Leaders and Account Managers, operate... ...prior landscaping experience is required, but a valid license and reliable transportation are essential. #J-18808-Ljbffr Clean ScapesWebsiteHourly payLocal area
- KNAPP is seeking a Reliability Technician in Lancaster, TX to maintain and improve the performance of automated systems. You will define standards... ..., and strong problem-solving skills. You will collaborate with site teams, document findings, and support customer service...Website
- ...configurations Key Responsibilities: Build and operate scalable and reliable infrastructure. Collaborate with development teams to improve... ...insurance Vision insurance 401(k) Get notified about new Site Reliability Engineer jobs in Austin, Texas Metropolitan Area ....WebsiteFull timeRemote work
$60k - $135k
...and task automation • Identify application reliability and availability improvements and build... ...Serve as technical subject matter expert (SME) for cross-functional engineering Teams;... ...Git, Jira, Confluence Mandatory Skills: Site Reliability Engineering (SRE). Experience...WebsiteMinimum wageLocal areaRelocation- ...Role: Site Reliability Engineer Location: Southlake / Austin, TX - Onsite 4 days weekly Duration: 12 Months Job Summary We are seeking a motivated Site Reliability Engineer (Contractor) with 3 to 5 years of experience in automation, cloud infrastructure, and production...WebsiteFor contractors
- ...ownership for production‑critical applications at our Austin and Warren centers. You will ensure availability and performance using Site Reliability Engineering practices, across Linux, Windows, Kubernetes, cloud, Oracle, SQL Server, and PostgreSQL environments. You will...Website
- RESPONSIBILITIES: Kforce has a client seeking a remote Site Reliability Engineer to join their team. We are seeking a highly motivated Site Reliability Engineer (SRE) to help build, scale, and maintain cloud infrastructure, CI/CD pipelines, and deployment automation...WebsiteHourly payContract workWork experience placementRemote work
- ...operations to deliver the Brivo Security Suite—a revolutionary platform combining AI, access control, and video intelligence. As a Site Reliability Engineer at Brivo, you will bridge the gap between development and operations. You will ensure our global platform remains...WebsiteWorldwide
- ...who build great products and contribute to our growth, we’re looking to add an Equipment Reliability Engineer located in Austin, TX LUSA.Reporting to Manager, as part of the site engineering team, the Equipment Reliability Engineer is responsible for Equipment Reliability...WebsiteFull timeVisa sponsorshipFlexible hours
- ...Job Description Job Description Sr. Software Engineer - Site Reliability About ShipperHQ: ShipperHQ is a trusted leader in the e-commerce shipping space, with over 15 years of experience helping merchants deliver better checkout experiences. Founded in 2009, we...WebsiteFull timeWork at office
- ...Recognized as the No. 1 site trusted by real estate professionals, Realtor.com has been at the forefront of online real estate for... ...confidence through expert guidance.We are seeking a Senior Site Reliability Engineer to join our newly formed Operations Excellence organization...WebsiteWork at officeLocal area
- ...projects with a specific or niche skill set. Since our inception, Reliable Software has been offering IT consulting services to the clients... ...Term: Contract Interview Process: Phone then Skype / On-Site Remote Option: No Required: Tasks & Duties...WebsiteFull timeContract workLocal areaImmediate startRemote work
- ...results and best in class outcomesVisionary in future focused problem-solvingExceptional in execution and impactThe RoleAs a Senior Site Reliability Engineer at Proofpoint you will develop a deep understanding of the various servicesand applications that come together to...WebsiteFull timeFlexible hours
$90k - $130k
...Site Reliability Engineer Austin, TX $90,000 - $130,000 a year Profession: Engineering Job Type: Contract Full Time Location: Austin, Texas Schedule: Full-Time Pay Range: Competitive pay, based on experience and qualifications. Make a Difference as...WebsiteFull timeContract workFlexible hours- ...table and foster a dynamic work environment where new ideas thrive. Are you ready to join our team and make an impact?As a Senior Site Reliability Engineer at TeamViewer, you’ll be a key player in ensuring the reliability, scalability, and performance of our Azure-based...WebsiteTemporary workCasual workWorldwide
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to K8 Site Reliability SME. Be the first to apply!
- IT site lead Austin, TX
- site safety Austin, TX
- website content developer Austin, TX
- site leader Austin, TX
- on-site clinical research associate (traveling/remote) Austin, TX
- junior website developer Austin, TX
- historic site Austin, TX
- on site coordinator Austin, TX
- site recruiter Austin, TX
- construction site safety Austin, TX




