Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer

Doist

Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure. Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to support the AI revolution. Bitdeer handles complex processes involved in computing such as equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial intelligence. Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia. To learn more, visit Position Overview You run the control plane where AIOps meets tenants — where topology-aware scheduling, self-healing, and agent-driven remediation actually execute. NeoCloud is building an AI-operated GPU cloud. Kubernetes is where all of that lands on real customer workloads: the topology-aware scheduler places jobs on the right NVLink domain, the operator drains and reschedules around predicted faults, and the tenant boundary is enforced against a Bare-Metal-as-a-Service backend. In this role you design, deploy, and operate that control plane — and you make sure the AIOps substrate can reach in and remediate without a human on the pager. What you'll own Production Kubernetes clusters optimized for GPU workloads at scale (100–10,000 GPUs). Nvidia GPU operator, device plugin, MIG configuration, and GPU time-slicing policies. Topology-aware scheduling: GPU locality, NVLink domain awareness, network rail affinity. Custom Resource Definitions (CRDs) for GPU workload lifecycle management. AI framework integrations: Slurm on K8S, Ray on K8S, Kubeflow. Multi-tenant isolation: namespaces, network policies, resource quotas, RBAC, pod security standards. Bare-Metal as a Service (BMaaS): automated provisioning, tenant onboarding, lifecycle, reclamation. Terraform providers and modules for infrastructure-as-code across GPU clusters. SLIs/SLOs for cluster availability, job completion rates, and provisioning latency. Incident management: runbook automation, escalation, post-incident reviews. Monitoring stack: Prometheus, Grafana, Alertmanager, PagerDuty. GPU node failure handling: automated detection, drain/cordon/taint, workload rescheduling. Feed the AIOps substrate The remediation-actuator and workflow engine land here — you make the control plane safe for automated action. Your CRDs are the schema the platform's predictors and remediators write against. Every human intervention you do this quarter becomes an autonomous workflow next quarter. What success looks like in year 1 Automated drain/reschedule around predicted GPU faults, at scale, without customer impact. BMaaS live for external tenants with self-service onboarding. Cluster availability and job-completion SLOs published and met. Job Requirement 5+ years in Kubernetes operations, with at least 2 years managing GPU workloads on K8S Deep understanding of Nvidia GPU operator, device plugin, and GPU scheduling in K8S Experience with topology-aware scheduling and GPU-specific resource management Hands-on experience building multi-tenant K8S platforms with strong isolation guarantees Experience with bare-metal server provisioning and lifecycle automation (Ironic, MAAS, or custom) Proficiency in Terraform, Helm, and GitOps workflows (ArgoCD/Flux) Strong SRE background: SLI/SLO frameworks, incident management, capacity planning Experience with Prometheus, Grafana, and alerting at scale Strong programming skills in Go or Python for operator/CRD development AIOps aptitude — you think of the K8S control plane as an execution surface for automated remediation, not just a scheduler. You've either wired an autoscaler/remediator loop into K8S or you can design one. Runbook-as-code mindset — every SRE playbook you write should be executable by the platform. Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, color, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union. #J-18808-Ljbffr

Vacancy posted 7 hours ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in Austin, TX vacancy
  • Recognized as the No. 1 site trusted by real estate professionals, Realtor.com has been at the forefront of online real estate...  ...confidence through expert guidance.We are seeking a Senior Site Reliability Engineer to join our newly formed Operations Excellence organization,... 
    Suggested
    Work at office
    Local area

    Realtor.com

    Austin, TX
    18 hours ago
  •  ...Description:About the Role: We are looking for a Senior SRE to join our Platform Engineering team as the operations owner of our observability platforms. You’ll be responsible for the reliability, scalability, and continued evolution of the tools that give our engineering... 
    Suggested
    Full time

    Dimensional Fund Advisors

    Austin, TX
    4 days ago
  • $98.58k - $138.02k

     ...Northern California / Silicon Valley Region / Denver, COProduct Engineering - DevOps /Full Time /HybridRestaurant365 is a SaaS company...  ...office locations: Austin, TX; Irvine, CA; or Akron, OH. The Site Reliability Engineer II will be responsible for supporting, enhancing,... 
    Suggested
    Full time
    Work at office

    Restaurant 365

    Austin, TX
    3 days ago
  • $109.65k - $182.76k

     ...encrypt data to make the connected world more secure.Austin, TX - Hybrid (3 days a week)Position SummaryWe are seeking a Site Reliability Engineer to ensure the high level of service and operation excellence for the development of the innovative and ambitious Telecommunication... 
    Suggested
    Full time
    Local area
    3 days per week

    Thales Group

    Austin, TX
    1 day ago
  •  ...and foster a dynamic work environment where new ideas thrive. Are you ready to join our team and make an impact?As a Senior Site Reliability Engineer at TeamViewer, you’ll be a key player in ensuring the reliability, scalability, and performance of our Azure-based SaaS... 
    Suggested
    Temporary work
    Casual work
    Worldwide

    TeamViewer

    Austin, TX
    4 days ago
  •  ...importance of in-office collaboration and fully intend for the selected candidate for this role to work on site in the specified location(s).As a Senior Reliability Engineer, you will help shape the reliability, scalability, and operational excellence of mission-critical... 
    Full time
    Work at office

    The Charles Schwab Corporation

    Austin, TX
    1 day ago
  •  ...across multiple clouds and regions while partnering with network engineers, systems architects, and game studio developers. This is an ownership role: driving technical direction, influencing reliability from architecture review through production operation, and closing... 

    2K Games

    Austin, TX
    1 day ago
  • $152k - $241.5k

     ...infrastructure platforms for automated host lifecycle management, fleet reliability/auto-healing, E2E observability or data-driven operations (...  ...languages such as Python, Go, Perl, or Ruby.Mentored other engineers and influenced technical direction through design reviews,... 
    Full time

    Nvidia

    Austin, TX
    4 days ago
  •  ...Dimensional leverages the rapidly evolving state of the art to engineer scalable, innovative, and research driven solutions to improve...  ...each of the developer tooling ecosystemsOwn the operational reliability of developer tooling ecosystems, including Python toolchains (... 
    Full time
    Local area

    Dimensional Fund Advisors

    Austin, TX
    18 hours ago
  • $127k - $249k

    The TeamPlatform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions...  ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper).... 
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    Austin, TX
    1 day ago
  • $152k - $195k

     ...Senior Site Reliability Engineer Austin, TX (Hybrid) SecurityScorecard is the global leader in cybersecurity ratings, with over 12 million companies continuously rated, operating in 64 countries. Founded in 2013 by security and risk experts Dr. Alex Yampolskiy and... 

    SecurityScorecard

    Austin, TX
    18 hours ago
  • $75.7k - $136.3k

     ...solve complex challenges? Do you have a passion for automation and building systems that scale? Join our highly skilled Site Reliability Engineering team! Our team designs, develops, and manages applications and infrastructure that support Akamai Cloud's products and... 
    Work experience placement
    Work at office

    Akamai

    Austin, TX
    2 days ago
  • $127k - $249k

     ...MongoDB, Inc. is seeking an experienced Senior or Staff Engineer for their SRE, InfraSec team, responsible for guiding the security of cloud-based infrastructure. The role involves hands-on technical work and mentorship of a small team while collaborating with engineering... 
    Remote work
    Flexible hours

    RTL2 Fernsehen GmbH & Co. KG

    Austin, TX
    4 days ago
  •  ...commercialization, and mass production to change the world for the better. JOB SUMMARY We are seeking an experienced Site Reliability Engineer to own and maintain the deployment of our cloud-based infrastructure to customer sites. In this role, you will work... 
    Full time
    Local area

    Synthesia

    Austin, TX
    3 days ago
  •  ...METRIX IT SOLUTIONS INC is seeking a senior database engineer to design, deploy, and manage multi-region CockroachDB clusters in production. The role focuses on high availability, data consistency, and scalable capacity planning for global deployments. You will monitor... 

    METRIX IT SOLUTIONS INC

    Austin, TX
    4 days ago
  •  ...the layer where data becomes decisions, and decisions make the advantage. About the Role Gallatin is looking for a Site Reliability Engineer to keep our production systems running with the reliability our national security customers require. You'll work at the... 
    Full time
    Local area

    Gallatin

    Austin, TX
    5 days ago
  • Bitdeer Technologies Group is seeking an L1 NOC/Data Center Operations technician to monitor and respond to incidents in NeoCloud's US GPU data centers during 8AM–8PM PST shifts. You will follow SOPs, escalate complex cases, and provide ground truth data to the platform...
    Shift work
    Night shift

    Doist

    Austin, TX
    6 hours ago
  •  ...Job Description Job Description Sr. Software Engineer - Site Reliability About ShipperHQ: ShipperHQ is a trusted leader in the e-commerce shipping space, with over 15 years of experience helping merchants deliver better checkout experiences. Founded in 2009, we... 
    Full time
    Work at office

    ShipperHQ

    Austin, TX
    a month ago
  • $74.1k - $148.3k

     ...systems. Facilitate service capacity planning and demand forecasting, software performance analysis, and system tuning. As a Site Reliability Engineer, you will solve interesting technical challenges by defining, designing, deploying, and solving key Oracle Cloud services,... 
    Temporary work
    Immediate start
    Flexible hours

    Oracle

    Austin, TX
    3 days ago
  •  ...Job Description Job Description Senior Site Reliability Engineer - Developer Productivity & Tooling Location: Austin, TX Area (Remote-First) Requirement: Candidates must be within commuting distance of Austin. About the Role A leading asset management... 
    Local area
    Remote work

    Selby Jennings

    Austin, TX
    14 days ago
  • $168k - $200k

     ...is passionate about creating transformative change in healthcare. What We’re Looking For We’re looking for a Senior Site Reliability Engineer to join our Data & ML Platform team. You’ll be at the forefront of building and operating a resilient, observable, and... 

    Datavant

    Austin, TX
    3 days ago
  •  ...operational performance and availability of critical business platforms and cloud services. With a strong technical background in site reliability engineering, the ideal applicant will have excellent communication skills and a focus on continuous improvement through automation.... 

    Take-Two Interactive

    Austin, TX
    4 days ago
  • $127k - $249k

     ...A leading technology company is seeking an experienced Senior or Staff Engineer for their SRE, InfraSec team in Austin. This role focuses on leading the design and implementation of security solutions for cloud platforms while mentoring a team. Candidates should have... 

    MongoDB

    Austin, TX
    4 days ago
  • $110.7k - $171.8k

     ...components Participation in on-call rotation as a platform reliability escalation point Incident response, post-incident reviews,...  ..., and internal control requirements. Collaborate with engineering teams across the organization to influence platform adoption,... 
    Work experience placement
    Work at office
    Local area

    Visa

    Austin, TX
    3 days ago
  • $112.11k - $190.66k

     ...world more secure. This is for a hybrid role in Austin, TX. Position Summary We are looking for an experienced SRE - Site Reliability Engineer to work with our North American Team. Your responsibility will be to help design and build tools and infrastructure that... 
    Full time
    Work at office
    Local area
    Monday to Friday

    Thales

    Austin, TX
    5 days ago
  • $127k - $249k

     ...Central time zones. We are looking for an experienced Senior Engineer for our SRE, Atlas team to support, maintain and grow the Atlas...  ...crucial workloads. Role OverviewWe are seeking a talented Site Reliability Engineer (SRE) with a strong infrastructure background. This... 
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    Austin, TX
    2 days ago
  • $167.18k - $203.61k

     ...remotely part of the weekTravel %NoWork ShiftJob DescriptionCox Automotive Corporate Services, LLCLEAD SITE RELIABILITY ENGINEERJob Description: Lead Site Reliability Engineer positions offered by Cox Automotive Corporate Services, LLC (Austin, Texas). Lead the... 
    Full time
    Work at office
    Remote work
    Flexible hours

    Cox Enterprises

    Austin, TX
    2 days ago
  • $210.6k - $305.1k

     ...Minimum Qualifications:  You have led a distributed team of 5+ engineers, can demonstrate strong technical vision for your team, and ensure...  ..., and basic life insurance. Please see the Cisco careers site to discover more benefits and perks. Employees may be eligible... 
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    Austin, TX
    1 day ago
  • $127k - $249k

    We are looking for an experienced Senior or Staff Engineer for our SRE, InfraSec team, to guide the security of our cloud-based infrastructure. As a Staff SRE, you will be very hands-on technically while also mentoring a small team of SREs.The InfraSec team collaborates... 
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    Austin, TX
    4 days ago
  •  ...Artificial Intelligence at Schwab. We are an integrated product, engineering, strategy and risk team, all based in San Francisco. We help...  ...the most exciting areas of technology today.As a Senior AI Site Reliability Engineer you will support reliability efforts for cutting-... 
    Full time

    The Charles Schwab Corporation

    Austin, TX
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!