Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer

Bitdeer Technologies Group

Software Engineer On The SRE / Monitoring Platform TeamBitdeer is building an AI-operated GPU cloud — a global fleet of self-built and OEM-rented data centers running the world's most valuable compute, run by a platform that observes, protects, and operates the fleet. The SRE Platform team builds the monitoring and automation substrate that every other squad — storage, network, GPU, K8S, and L1 operators — depends on. Their signals become the system you help build.As an entry-level Software Engineer on the SRE / Monitoring Platform team, you contribute to the NeoCloud SRE platform — the multi-region system that observes, protects, and operates a GPU rental fleet across self-built and OEM-rented data centers. You join a bounded context led by a senior engineer, take well-scoped components from design to production code that ships through GitOps + the CICD release pipeline, follows the Plugin Framework conventions, meets declared SLOs, and stays drift-free.This is a build + learn role. You write code, write tests, and operate what you build under the guidance of a senior engineer. You participate in on-call as a shadow before taking primary. Within 12 months, you should be delivering components independently within your assigned area and growing toward owning a sub-context.Key ResponsibilitiesCollection + Storage — help build collection-agent, metrics-store / logs-store / traces-store / profiles-store, enrichment-service, collection-monitor. Write ingestion, query, and storage-path code.Alert + Correlation + SLO — contribute to alert-engine-framework, alert-correlation, slo-framework; implement and tune default alert rules.Topology + Cluster-Health — contribute to topology-service, cluster-health-rollup, OSS-SRE-tool collection plugins for K8s / Slurm / Ray / Volcano / Kueue / KubeRay.Remediation + Workflow + Jobs — help build remediation-actuator, orchestration / workflow components, inspection probes, job-scheduler.Observability instrumentation — instrument services with metrics, logs, and traces via OpenTelemetry; build dashboards; write runbooks an on-call can follow.Test discipline — write unit / integration / contract tests for everything you ship; participate in chaos and soak tests led by senior engineers.Why this is a great first roleGreenfield with a well-defined vision. The Plugin Framework, GitOps pipeline, and SLO framework are decided; you build components inside them with a clear blueprint — not from a blank page.You learn the full observability stack at production scale — ingest, query, storage — by building it, not just using it.Mentorship-heavy. You work directly with senior and principal engineers who own the architecture; their expertise becomes your growth path.Job Requirements0-2 years of software engineering experience (new graduates with strong projects or internships welcome).Solid fundamentals in one programming language — Go (preferred), Python, Java, or Rust. You can write clean, tested, readable code and explain your design choices.CS fundamentals — data structures, algorithms, concurrency, basic networking (TCP / and operating-system concepts (processes, threads, I/O). You can reason about correctness and performance.Distributed systems basics — you understand the ideas behind idempotency, retries, back-pressure, caching, and eventual consistency, even if you haven't operated them at scale yet. Eagerness to go deep.Monitoring / observability exposure — some hands-on with Prometheus, Grafana, Loki, or similar; can write a basic PromQL query and instrument a service. Eagerness to learn the ingest, query, and storage path of a real observability stack.Familiarity with Linux and the shell; comfort reading system logs and using standard debugging tools.Kubernetes basics — understand Pods, Services, Deployments; have run something on K8s (a project, lab, or internship).Git + CI basics — branching, pull requests, and have used a CI pipeline (GitHub Actions, GitLab CI, or similar).Test discipline — you write unit and integration tests as a habit, not an afterthought.Communication — clear written and verbal English; can write a good PR description and ask good questions.Curiosity and a learning mindset — the most important qualifier. You're excited to learn GPU / AI infrastructure, AIOps, distributed systems, and observability at production scale.Nice-to-HavesInternship or project in monitoring / observability, telemetry pipelines, or platform / SRE tooling.Exposure to GPU / AI-infra — DCGM, InfiniBand / RoCE, Kubernetes GPU Operator, Slurm / Ray. Interest counts more than depth.Exposure to AIOps / ML-adjacent tooling (anomaly detection, alert correlation).Contributions to open-source observability or cloud-native projects.Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, color, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union.

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in Austin, TX vacancy
  •  ...and foster a dynamic work environment where new ideas thrive. Are you ready to join our team and make an impact?As a Senior Site Reliability Engineer at TeamViewer, you’ll be a key player in ensuring the reliability, scalability, and performance of our Azure-based SaaS... 
    Suggested
    Temporary work
    Casual work
    Worldwide

    TeamViewer

    Austin, TX
    2 days ago
  • $98.58k - $138.02k

     ...Northern California / Silicon Valley Region / Denver, COProduct Engineering - DevOps /Full Time /HybridRestaurant365 is a SaaS company...  ...office locations: Austin, TX; Irvine, CA; or Akron, OH. The Site Reliability Engineer II will be responsible for supporting, enhancing,... 
    Suggested
    Full time
    Work at office

    Restaurant 365

    Austin, TX
    5 days ago
  •  ...importance of in-office collaboration and fully intend for the selected candidate for this role to work on site in the specified location(s).As a Senior Reliability Engineer, you will help shape the reliability, scalability, and operational excellence of mission-critical... 
    Suggested
    Full time
    Work at office

    The Charles Schwab Corporation

    Austin, TX
    3 days ago
  • $167.7k - $245.2k

     ...very effective.We’re looking for talented engineers with a software or operations background...  ...development teams to ensure the reliability, performance and security of our infrastructure...  ...insurance. Please see the Cisco careers site to discover more benefits and perks.... 
    Suggested
    Full time
    Temporary work
    Work at office
    Local area
    Flexible hours
    1 day per week

    CISCO Systems

    Austin, TX
    5 days ago
  •  ...Dimensional leverages the rapidly evolving state of the art to engineer scalable, innovative, and research driven solutions to improve...  ...each of the developer tooling ecosystemsOwn the operational reliability of developer tooling ecosystems, including Python toolchains (... 
    Suggested
    Full time
    Local area

    Dimensional Fund Advisors

    Austin, TX
    2 days ago
  • $109.65k - $182.76k

     ...encrypt data to make the connected world more secure.Austin, TX - Hybrid (3 days a week)Position SummaryWe are seeking a Site Reliability Engineer to ensure the high level of service and operation excellence for the development of the innovative and ambitious Telecommunication... 
    Full time
    Local area
    3 days per week

    Thales Group

    Austin, TX
    4 days ago
  • Recognized as the No. 1 site trusted by real estate professionals, Realtor.com has been at the forefront of online real estate...  ...confidence through expert guidance.We are seeking a Senior Site Reliability Engineer to join our newly formed Operations Excellence organization,... 
    Work at office
    Local area

    Realtor.com

    Austin, TX
    2 days ago
  • $127k - $249k

    The TeamPlatform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions...  ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper).... 
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    Austin, TX
    3 days ago
  • $127k - $249k

     ...MongoDB, Inc. is seeking an experienced Senior or Staff Engineer for their SRE, InfraSec team, responsible for guiding the security of cloud-based infrastructure. The role involves hands-on technical work and mentorship of a small team while collaborating with engineering... 
    Remote work
    Flexible hours

    RTL2 Fernsehen GmbH & Co. KG

    Austin, TX
    1 day ago
  •  ...Key Responsibilities: Build and operate scalable and reliable infrastructure. Collaborate with development teams to improve...  ...Vision insurance 401(k) Get notified about new Site Reliability Engineer jobs in Austin, Texas Metropolitan Area . Site Reliability... 
    Full time
    Remote work

    Altimetrik

    Austin, TX
    5 days ago
  • A leading company is seeking a Site Reliability Engineer to join their Platform Infrastructure team. This remote role involves building reliable infrastructure, collaborating with development teams, and ensuring robust integrations with third-party services. Ideal candidates... 
    Remote work

    Altimetrik

    Austin, TX
    5 days ago
  •  ...Cloudflare is seeking a highly motivated software engineer to join our Production Platform Organization. You will build the infrastructure to collect, store, and make reliability data accessible for monitoring needs, working with Product Managers and SREs to measure service... 

    WebHosting.coop

    Austin, TX
    1 day ago
  • $127k - $249k

     ...Central time zones. We are looking for an experienced Senior Engineer for our SRE, Atlas team to support, maintain and grow the Atlas...  ...crucial workloads. Role OverviewWe are seeking a talented Site Reliability Engineer (SRE) with a strong infrastructure background. This... 
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    Austin, TX
    4 days ago
  • $192.4k - $275.8k

     ...CloudOps— the team that keeps Splunk Cloud running for some of the world's most demanding enterprise customers, blending Site Reliability Engineering, Systems Engineering, and Service Engineering disciplines at a scale very few teams ever get to operate at. When the... 
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    Austin, TX
    1 day ago
  • $167.18k - $203.61k

     ...remotely part of the weekTravel %NoWork ShiftJob DescriptionCox Automotive Corporate Services, LLCLEAD SITE RELIABILITY ENGINEERJob Description: Lead Site Reliability Engineer positions offered by Cox Automotive Corporate Services, LLC (Austin, Texas). Lead the... 
    Full time
    Work at office
    Remote work
    Flexible hours

    Cox Enterprises

    Austin, TX
    4 days ago
  •  ...IDR is seeking a Site Reliability Engineer to join one of our top clients for an opportunity in Austin, TX.  This role focuses on supporting and administering enterprise platform environments with an emphasis on improving reliability, governance, and automation. The... 
    Temporary work

    IDR

    Austin, TX
    7 days ago
  •  ...Required U.S. Citizenship / No clearance needed / 100% remote within the US  Staff Site Reliability Engineer / Cloud SME Location: 100% remote in the continental US  Type: Long-term contract (3+ years) Role Summary As the Staff SRE/Cloud SME, you will be... 
    Long term contract
    Remote work

    ASCENDING

    Austin, TX
    17 days ago
  •  ...assisted developers or autonomous agents is reliable, secure, and maintainable.Integrating...  ...descriptionAs a member of one of our engineering teams, you'll be a key player in making...  ...team of Engineers (Cloud Engineers and Site Reliability Engineers), providing guidance... 
    Full time
    Relocation
    Flexible hours

    SonarSource

    Austin, TX
    1 day ago
  •  ...Artificial Intelligence at Schwab. We are an integrated product, engineering, strategy and risk team, all based in San Francisco. We help...  ...the most exciting areas of technology today.As a Senior AI Site Reliability Engineer you will support reliability efforts for cutting-... 
    Full time

    The Charles Schwab Corporation

    Austin, TX
    1 day ago
  • $127k - $249k

    We are looking for an experienced Senior or Staff Engineer for our SRE, InfraSec team, to guide the security of our cloud-based infrastructure. As a Staff SRE, you will be very hands-on technically while also mentoring a small team of SREs.The InfraSec team collaborates... 
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    Austin, TX
    1 day ago
  • $100.1k - $180.2k

     ...world's top brands, offering comprehensive engineering, supply chain, and manufacturing...  ...industries and a vast network of over 100 sites worldwide, Jabil combines global reach with...  ...the globe.Jabil is seeking a Lead Site Reliability Infrastructure and Security Engineer with... 
    Full time
    Temporary work
    Work at office
    Local area
    Remote work
    Worldwide

    Jabil Circuit

    Austin, TX
    1 day ago
  • $140k - $215k

     ...operate at the intersection of our Core Platform and Embedded Reliability charters: building the foundational libraries, services, and...  ...product group depends on, while embedding directly with product engineering teams and their leadership to drive reliability outcomes at... 
    Full time
    Work experience placement
    Work at office
    Local area
    2 days per week
    3 days per week

    CrowdStrike

    Austin, TX
    1 day ago
  •  ...to join IBM in a full‑time role between December 2027 and August 2028 upon successful completion of their degree. As a Site Reliability Engineer, you will work in an agile, collaborative environment to build, deploy, configure, and maintain systems for the IBM client... 
    Full time
    Contract work
    Part time
    Fixed term contract
    Internship
    Worldwide
    Flexible hours
    Shift work

    IBM

    Austin, TX
    2 days ago
  • Job Description:About the Role:We are looking for a Senior SRE to join our Platform Engineering team where you’ll own the reliability, scalability, and operational excellence of our workflow orchestration platforms - primarily Apache Airflow and Broadcom Automic/UC4. This... 
    Full time

    Dimensional Fund Advisors

    Austin, TX
    5 days ago
  •  ...Principal Site Reliability Engineer About ShipperHQ: ShipperHQ is a trusted leader in the e-commerce shipping space, with over 15 years of experience helping merchants deliver better checkout experiences. Founded in 2009, we power shipping logic and checkout optimization... 
    Full time
    Work at office

    ShipperHQ

    Austin, TX
    more than 2 months ago
  •  ...performance analysis, and system tuning. May telecommute. (385.37335)Employer will accept a Master’s degree in Computer Science, Engineering, or related technical field and 6 years of experience in the job offered or in a Network or Software Developer-related occupation... 
    Temporary work
    Remote work
    Flexible hours

    Oracle Corporation

    Austin, TX
    1 day ago
  •  ...role and responsibilities IBM is seeking a motivated and detail-oriented IT Administrator intern with an interest in Site Reliability Engineering (SRE) and/or Networking Reliabillity Engineering (NRE) to join our team. These roles offers hands-on experience in enterprise... 
    Full time
    Contract work
    Part time
    Fixed term contract
    Internship
    Shift work

    IBM

    Austin, TX
    12 hours ago
  • $132.23k - $176.31k

     ...future of AI‑ready connectivity, join us today. The Role We are seeking a highly skilled and proactive Senior Lead Site Reliability Engineer (SRE) to join our team, focusing on production support and performance optimization across our portal ecosystem. This role... 
    Full time
    Temporary work
    Remote work

    Lumen

    Austin, TX
    2 days ago
  •  ...in-office collaboration and fully intend for the selected candidate for this role to work on site in the specified location(s).We are seeking a Kafka Site Reliability Engineer to help build, operate, and continuously improve Schwab's enterprise streaming platform ecosystem... 
    Full time
    Work at office

    The Charles Schwab Corporation

    Austin, TX
    1 day ago
  •  ...Job Description Job Description Sr. Software Engineer - Site Reliability About ShipperHQ: ShipperHQ is a trusted leader in the e-commerce shipping space, with over 15 years of experience helping merchants deliver better checkout experiences. Founded in 2009, we... 
    Full time
    Work at office

    ShipperHQ

    Austin, TX
    9 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!