Site Reliability Engineer
Bitdeer Technologies Group
Software Engineer On The SRE / Monitoring Platform TeamBitdeer is building an AI-operated GPU cloud — a global fleet of self-built and OEM-rented data centers running the world's most valuable compute, run by a platform that observes, protects, and operates the fleet. The SRE Platform team builds the monitoring and automation substrate that every other squad — storage, network, GPU, K8S, and L1 operators — depends on. Their signals become the system you help build.As an entry-level Software Engineer on the SRE / Monitoring Platform team, you contribute to the NeoCloud SRE platform — the multi-region system that observes, protects, and operates a GPU rental fleet across self-built and OEM-rented data centers. You join a bounded context led by a senior engineer, take well-scoped components from design to production code that ships through GitOps + the CICD release pipeline, follows the Plugin Framework conventions, meets declared SLOs, and stays drift-free.This is a build + learn role. You write code, write tests, and operate what you build under the guidance of a senior engineer. You participate in on-call as a shadow before taking primary. Within 12 months, you should be delivering components independently within your assigned area and growing toward owning a sub-context.Key ResponsibilitiesCollection + Storage — help build collection-agent, metrics-store / logs-store / traces-store / profiles-store, enrichment-service, collection-monitor. Write ingestion, query, and storage-path code.Alert + Correlation + SLO — contribute to alert-engine-framework, alert-correlation, slo-framework; implement and tune default alert rules.Topology + Cluster-Health — contribute to topology-service, cluster-health-rollup, OSS-SRE-tool collection plugins for K8s / Slurm / Ray / Volcano / Kueue / KubeRay.Remediation + Workflow + Jobs — help build remediation-actuator, orchestration / workflow components, inspection probes, job-scheduler.Observability instrumentation — instrument services with metrics, logs, and traces via OpenTelemetry; build dashboards; write runbooks an on-call can follow.Test discipline — write unit / integration / contract tests for everything you ship; participate in chaos and soak tests led by senior engineers.Why this is a great first roleGreenfield with a well-defined vision. The Plugin Framework, GitOps pipeline, and SLO framework are decided; you build components inside them with a clear blueprint — not from a blank page.You learn the full observability stack at production scale — ingest, query, storage — by building it, not just using it.Mentorship-heavy. You work directly with senior and principal engineers who own the architecture; their expertise becomes your growth path.Job Requirements0-2 years of software engineering experience (new graduates with strong projects or internships welcome).Solid fundamentals in one programming language — Go (preferred), Python, Java, or Rust. You can write clean, tested, readable code and explain your design choices.CS fundamentals — data structures, algorithms, concurrency, basic networking (TCP / and operating-system concepts (processes, threads, I/O). You can reason about correctness and performance.Distributed systems basics — you understand the ideas behind idempotency, retries, back-pressure, caching, and eventual consistency, even if you haven't operated them at scale yet. Eagerness to go deep.Monitoring / observability exposure — some hands-on with Prometheus, Grafana, Loki, or similar; can write a basic PromQL query and instrument a service. Eagerness to learn the ingest, query, and storage path of a real observability stack.Familiarity with Linux and the shell; comfort reading system logs and using standard debugging tools.Kubernetes basics — understand Pods, Services, Deployments; have run something on K8s (a project, lab, or internship).Git + CI basics — branching, pull requests, and have used a CI pipeline (GitHub Actions, GitLab CI, or similar).Test discipline — you write unit and integration tests as a habit, not an afterthought.Communication — clear written and verbal English; can write a good PR description and ask good questions.Curiosity and a learning mindset — the most important qualifier. You're excited to learn GPU / AI infrastructure, AIOps, distributed systems, and observability at production scale.Nice-to-HavesInternship or project in monitoring / observability, telemetry pipelines, or platform / SRE tooling.Exposure to GPU / AI-infra — DCGM, InfiniBand / RoCE, Kubernetes GPU Operator, Slurm / Ray. Interest counts more than depth.Exposure to AIOps / ML-adjacent tooling (anomaly detection, alert correlation).Contributions to open-source observability or cloud-native projects.Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, color, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union.
- ...and foster a dynamic work environment where new ideas thrive. Are you ready to join our team and make an impact?As a Senior Site Reliability Engineer at TeamViewer, you’ll be a key player in ensuring the reliability, scalability, and performance of our Azure-based SaaS...SuggestedTemporary workCasual workWorldwide
$98.58k - $138.02k
...Northern California / Silicon Valley Region / Denver, COProduct Engineering - DevOps /Full Time /HybridRestaurant365 is a SaaS company... ...office locations: Austin, TX; Irvine, CA; or Akron, OH. The Site Reliability Engineer II will be responsible for supporting, enhancing,...SuggestedFull timeWork at office- ...importance of in-office collaboration and fully intend for the selected candidate for this role to work on site in the specified location(s).As a Senior Reliability Engineer, you will help shape the reliability, scalability, and operational excellence of mission-critical...SuggestedFull timeWork at office
$167.7k - $245.2k
...very effective.We’re looking for talented engineers with a software or operations background... ...development teams to ensure the reliability, performance and security of our infrastructure... ...insurance. Please see the Cisco careers site to discover more benefits and perks....SuggestedFull timeTemporary workWork at officeLocal areaFlexible hours1 day per week- ...Dimensional leverages the rapidly evolving state of the art to engineer scalable, innovative, and research driven solutions to improve... ...each of the developer tooling ecosystemsOwn the operational reliability of developer tooling ecosystems, including Python toolchains (...SuggestedFull timeLocal area
$109.65k - $182.76k
...encrypt data to make the connected world more secure.Austin, TX - Hybrid (3 days a week)Position SummaryWe are seeking a Site Reliability Engineer to ensure the high level of service and operation excellence for the development of the innovative and ambitious Telecommunication...Full timeLocal area3 days per week- Recognized as the No. 1 site trusted by real estate professionals, Realtor.com has been at the forefront of online real estate... ...confidence through expert guidance.We are seeking a Senior Site Reliability Engineer to join our newly formed Operations Excellence organization,...Work at officeLocal area
$127k - $249k
The TeamPlatform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions... ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper)....Work at officeLocal areaRemote workWorldwideFlexible hours$127k - $249k
...MongoDB, Inc. is seeking an experienced Senior or Staff Engineer for their SRE, InfraSec team, responsible for guiding the security of cloud-based infrastructure. The role involves hands-on technical work and mentorship of a small team while collaborating with engineering...Remote workFlexible hours- ...Key Responsibilities: Build and operate scalable and reliable infrastructure. Collaborate with development teams to improve... ...Vision insurance 401(k) Get notified about new Site Reliability Engineer jobs in Austin, Texas Metropolitan Area . Site Reliability...Full timeRemote work
- A leading company is seeking a Site Reliability Engineer to join their Platform Infrastructure team. This remote role involves building reliable infrastructure, collaborating with development teams, and ensuring robust integrations with third-party services. Ideal candidates...Remote work
- ...Cloudflare is seeking a highly motivated software engineer to join our Production Platform Organization. You will build the infrastructure to collect, store, and make reliability data accessible for monitoring needs, working with Product Managers and SREs to measure service...
$127k - $249k
...Central time zones. We are looking for an experienced Senior Engineer for our SRE, Atlas team to support, maintain and grow the Atlas... ...crucial workloads. Role OverviewWe are seeking a talented Site Reliability Engineer (SRE) with a strong infrastructure background. This...Local areaRemote workWorldwideFlexible hours$192.4k - $275.8k
...CloudOps— the team that keeps Splunk Cloud running for some of the world's most demanding enterprise customers, blending Site Reliability Engineering, Systems Engineering, and Service Engineering disciplines at a scale very few teams ever get to operate at. When the...Full timeTemporary workLocal areaFlexible hours$167.18k - $203.61k
...remotely part of the weekTravel %NoWork ShiftJob DescriptionCox Automotive Corporate Services, LLCLEAD SITE RELIABILITY ENGINEERJob Description: Lead Site Reliability Engineer positions offered by Cox Automotive Corporate Services, LLC (Austin, Texas). Lead the...Full timeWork at officeRemote workFlexible hours- ...IDR is seeking a Site Reliability Engineer to join one of our top clients for an opportunity in Austin, TX. Â This role focuses on supporting and administering enterprise platform environments with an emphasis on improving reliability, governance, and automation. The...Temporary work
- ...Required U.S. Citizenship / No clearance needed / 100% remote within the US Staff Site Reliability Engineer / Cloud SME Location: 100% remote in the continental US Type: Long-term contract (3+ years) Role Summary As the Staff SRE/Cloud SME, you will be...Long term contractRemote work
- ...assisted developers or autonomous agents is reliable, secure, and maintainable.Integrating... ...descriptionAs a member of one of our engineering teams, you'll be a key player in making... ...team of Engineers (Cloud Engineers and Site Reliability Engineers), providing guidance...Full timeRelocationFlexible hours
- ...Artificial Intelligence at Schwab. We are an integrated product, engineering, strategy and risk team, all based in San Francisco. We help... ...the most exciting areas of technology today.As a Senior AI Site Reliability Engineer you will support reliability efforts for cutting-...Full time
$127k - $249k
We are looking for an experienced Senior or Staff Engineer for our SRE, InfraSec team, to guide the security of our cloud-based infrastructure. As a Staff SRE, you will be very hands-on technically while also mentoring a small team of SREs.The InfraSec team collaborates...Local areaRemote workWorldwideFlexible hours$100.1k - $180.2k
...world's top brands, offering comprehensive engineering, supply chain, and manufacturing... ...industries and a vast network of over 100 sites worldwide, Jabil combines global reach with... ...the globe.Jabil is seeking a Lead Site Reliability Infrastructure and Security Engineer with...Full timeTemporary workWork at officeLocal areaRemote workWorldwide$140k - $215k
...operate at the intersection of our Core Platform and Embedded Reliability charters: building the foundational libraries, services, and... ...product group depends on, while embedding directly with product engineering teams and their leadership to drive reliability outcomes at...Full timeWork experience placementWork at officeLocal area2 days per week3 days per week- ...to join IBM in a full‑time role between December 2027 and August 2028 upon successful completion of their degree. As a Site Reliability Engineer, you will work in an agile, collaborative environment to build, deploy, configure, and maintain systems for the IBM client...Full timeContract workPart timeFixed term contractInternshipWorldwideFlexible hoursShift work
- Job Description:About the Role:We are looking for a Senior SRE to join our Platform Engineering team where you’ll own the reliability, scalability, and operational excellence of our workflow orchestration platforms - primarily Apache Airflow and Broadcom Automic/UC4. This...Full time
- ...Principal Site Reliability Engineer About ShipperHQ: ShipperHQ is a trusted leader in the e-commerce shipping space, with over 15 years of experience helping merchants deliver better checkout experiences. Founded in 2009, we power shipping logic and checkout optimization...Full timeWork at office
- ...performance analysis, and system tuning. May telecommute. (385.37335)Employer will accept a Master’s degree in Computer Science, Engineering, or related technical field and 6 years of experience in the job offered or in a Network or Software Developer-related occupation...Temporary workRemote workFlexible hours
- ...role and responsibilities IBM is seeking a motivated and detail-oriented IT Administrator intern with an interest in Site Reliability Engineering (SRE) and/or Networking Reliabillity Engineering (NRE) to join our team. These roles offers hands-on experience in enterprise...Full timeContract workPart timeFixed term contractInternshipShift work
$132.23k - $176.31k
...future of AI‑ready connectivity, join us today. The Role We are seeking a highly skilled and proactive Senior Lead Site Reliability Engineer (SRE) to join our team, focusing on production support and performance optimization across our portal ecosystem. This role...Full timeTemporary workRemote work- ...in-office collaboration and fully intend for the selected candidate for this role to work on site in the specified location(s).We are seeking a Kafka Site Reliability Engineer to help build, operate, and continuously improve Schwab's enterprise streaming platform ecosystem...Full timeWork at office
- ...Job Description Job Description Sr. Software Engineer - Site Reliability About ShipperHQ: ShipperHQ is a trusted leader in the e-commerce shipping space, with over 15 years of experience helping merchants deliver better checkout experiences. Founded in 2009, we...Full timeWork at office
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
- site reliability engineer Austin, TX
- site reliability engineer sre Austin, TX
- site reliability engineer remote Austin, TX
- website coordinator Austin, TX
- on-site clinical research associate (traveling/remote) Austin, TX
- site safety Austin, TX
- junior website developer Austin, TX
- construction site safety Austin, TX
- IT site lead Austin, TX
- website content developer Austin, TX



