SRE Platform Software Engineer
Bitdeer Technologies Group
Role Description
Bitdeer is building an AI-operated GPU cloud — a global fleet of self-built and OEM-rented data centers running the world's most valuable compute, run by a platform that observes, protects, and operates the fleet. The SRE Platform team builds the monitoring and automation substrate that every other squad — storage, network, GPU, K8S, and L1 operators — depends on. Their signals become the system you help build.
As an entry-level Software Engineer on the SRE / Monitoring Platform team, you contribute to the NeoCloud SRE platform — the multi-region system that observes, protects, and operates a GPU rental fleet across self-built and OEM-rented data centers. You join a bounded context led by a senior engineer, take well-scoped components from design to production code that ships through GitOps + the CICD release pipeline, follows the Plugin Framework conventions, meets declared SLOs, and stays drift-free.
This is a build + learn role. You write code, write tests, and operate what you build under the guidance of a senior engineer. You participate in on-call as a shadow before taking primary. Within 12 months, you should be delivering components independently within your assigned area and growing toward owning a sub-context.
Key Responsibilities
- Collection + Storage — help build collection-agent, metrics-store / logs-store / traces-store / profiles-store, enrichment-service, collection-monitor. Write ingestion, query, and storage-path code.
- Alert + Correlation + SLO — contribute to alert-engine-framework, alert-correlation, slo-framework; implement and tune default alert rules.
- Topology + Cluster-Health — contribute to topology-service, cluster-health-rollup, OSS-SRE-tool collection plugins for K8s / Slurm / Ray / Volcano / Kueue / KubeRay.
- Remediation + Workflow + Jobs — help build remediation-actuator, orchestration / workflow components, inspection probes, job-scheduler.
- Observability instrumentation — instrument services with metrics, logs, and traces via OpenTelemetry; build dashboards; write runbooks an on-call can follow.
- Test discipline — write unit / integration / contract tests for everything you ship; participate in chaos and soak tests led by senior engineers.
Qualifications
- 0-2 years of software engineering experience (new graduates with strong projects or internships welcome).
- Solid fundamentals in one programming language — Go (preferred), Python, Java, or Rust. You can write clean, tested, readable code and explain your design choices.
- CS fundamentals — data structures, algorithms, concurrency, basic networking (TCP / and operating-system concepts (processes, threads, I/O). You can reason about correctness and performance.
- Distributed systems basics — you understand the ideas behind idempotency, retries, back-pressure, caching, and eventual consistency, even if you haven't operated them at scale yet. Eagerness to go deep.
- Monitoring / observability exposure — some hands-on with Prometheus, Grafana, Loki, or similar; can write a basic PromQL query and instrument a service. Eagerness to learn the ingest, query, and storage path of a real observability stack.
- Familiarity with Linux and the shell; comfort reading system logs and using standard debugging tools.
- Kubernetes basics — understand Pods, Services, Deployments; have run something on K8s (a project, lab, or internship).
- Git + CI basics — branching, pull requests, and have used a CI pipeline (GitHub Actions, GitLab CI, or similar).
- Test discipline — you write unit and integration tests as a habit, not an afterthought.
- Communication — clear written and verbal English; can write a good PR description and ask good questions.
- Curiosity and a learning mindset — the most important qualifier. You're excited to learn GPU / AI infrastructure, AIOps, distributed systems, and observability at production scale.
Nice-to-Haves
- Internship or project in monitoring / observability, telemetry pipelines, or platform / SRE tooling.
- Exposure to GPU / AI-infra — DCGM, InfiniBand / RoCE, Kubernetes GPU Operator, Slurm / Ray. Interest counts more than depth.
- Exposure to AIOps / ML-adjacent tooling (anomaly detection, alert correlation).
- Contributions to open-source observability or cloud-native projects.
Company Description
Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure. Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to support the AI revolution. Bitdeer handles complex processes involved in computing such as equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial intelligence. Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia.
- ...Senior DevOps/Platform Engineer (SRE) 100 % Remote (W/M/D) Easybill is the leading provider of cloud-based invoicing software and has over 18 years of successful market presence. Our platform supports companies in efficiently managing their financial processes....SoftwareRemote workFlexible hoursNight shift
$83.52k - $125.28k
...organization, apply now.We are currently seeking a Lead ML Platform Engineer (SRE / FTE / Onsite) to join our team in Charlotte, North Carolina... ...secrets management, encryption, audit logging, and secure software delivery.Experience implementing observability, monitoring,...SoftwareFull timeTemporary workWork at officeRemote workFlexible hours- ...visionary and hands-on Cloud SRE Architect to lead the design,... ...next-generation public cloud platform. This role will oversee the end... ...infrastructure, platform engineering, and AI systems, capable of bridging... ...~7+ years of production software engineering experience, including...SoftwareFull timeContract work
- ...running the world's most valuable compute, run by a platform that observes, protects, and operates the fleet. The SRE Platform team builds the monitoring and... ...system you help build. As an entry-level Software Engineer on the SRE / Monitoring Platform team, you contribute...SoftwareFull timeContract workTemporary workInternshipLocal area
- ...working parameters for hardware/software compatibility. Defines system... ...in systems design. Performs engineering studies and analyses, analyze... ...and dynamic AWS Cloud Platform Engineer to provide hands-on... ...Engineering role Experience with SRE principles and transformation...SoftwareImmediate startRemote work
$1,000 per month
...mobile-first financial wellness platform designed to help individuals... ...the AI infrastructure our engineers use every day. When this team... ...of gaps.As a Senior DevOps / SRE Engineer on this team, you'll... ...'ve written and shipped real software, not only infrastructure code...SoftwareTemporary workWork at officeImmediate startRemote workFlexible hours- ...seeking experienced Developer & Infrastructure Experts to evaluate AI-powered workflows across software development, cloud infrastructure, DevOps, SRE, and platform engineering. You will test AI-generated commands, configurations, and workflows and provide practical...SoftwareFor contractorsRemote work
- ...the future, today.We are looking for an experienced Software Engineer with a strong background in Platform Engineering to build and scale our modern infrastructure... ...reliability.Qualifications:6+ years in DevOps, SRE, or Platform Engineering plus a Bachelor's degree in...SoftwareWork at officeWork from homeFlexible hours3 days per week
- ...Location: Remote, United States As a System Engineer, Sr. on our team, your main role is to... .... Use your hardware and software experience to help strengthen the systems... ...application owners to understand their platform designs and how they operate across different...SoftwareContract workWork at officeRemote work
$106.5k - $177.5k
...Software Engineer- Site Reliability Engineering (SRE) DC, MD, VA, CA The Site Reliability Engineering discipline at Noctua Technology, LLC is a strategic force driving digital transformation. We treat operations as a software engineering challenge, focusing on the...SoftwareRemote work$190.9k - $334.1k
...DescriptionIt all started when engineer Fred Luddy wrote code that... ...reinvention. Our ServiceNow AI platform brings together any AI, any... ...Staff Reliability Engineer - SRE & AIOps to drive infrastructure... ...or industry.12+ years in software engineering or infrastructure...SoftwareWork at officeImmediate startRemote workFlexible hours$149.4k - $202k
...Senior Software Engineer- Site Reliability Engineering (SRE) DC, MD, VA, CA The Site Reliability Engineering discipline at Noctua Technology, LLC is... ...of SLIs and SLOs across multiple services or entire platforms, ensuring alignment with business goals. Design...SoftwareRemote work$185k - $200k
Back to All JobsSenior Site Reliability Engineer (SRE) Dayton, OH (Remote) full time Top... ...Engineer (SRE) to support AFRL's Google Cloud Platform environment. This engineer will help... ..., cloud/platform engineering, DevOps, software engineering, or a related discipline....SoftwareFull timeRemote work$207k - $300k
...as system design consulting, developing software platforms and frameworks, capacity planning and... ...roadmap, and reliability strategy for Home SRE.Minimum qualifications:Bachelor’s... ...Master's degree in Computer Science or Engineering, or a related field.Experience architecting...SoftwareWorldwide$1,000 per month
...driven support. Requirements Demonstrated DevOps/SRE depth and a genuine backend software engineering background with shipped production software.... ...tooling, MCP, model gateways, or internal AI developer platforms. Preferred prior SOC 2 or PCI-DSS experience,...SoftwareFull timeTemporary workWork at officeImmediate startRemote workFlexible hours$207k - $300k
Lead a team of Software/Systems Engineers on projects for users and be directly responsible for uptime.... ...Language Model.Site Reliability Engineering (SRE) combines software and systems... ...building the next generation of Google platforms, we make Google's product portfolio possible...Software- ...high‑scale, high‑reliability platforms that support card manufacturing... .... Bringing together engineering, product, delivery, and commercialization... ...be doing As Principal SRE - define meaningful SLIs,... ...years of experience in SRE, Software engineering and architecture...SoftwareWork at office
$100k - $150k
...Technologies is a technology consulting and software development company delivering cloud,... ...potential. Job Title: DevOps & SRE Engineer Location: 100% Remote (U.S.) Position... ...problems, and continually pushing the platform toward higher reliability with lower...SoftwareFull timeH1bLocal areaImmediate startRemote work$250k - $300k
We're looking for a hands-on engineering leader to build and own the Release... ...Reliability Engineering (SRE) functions for Infinia. This... ...role: you'll define how our software is built, secured, released,... ...of a market-leading storage platform — this is your role.What You'...SoftwareRemote work- ...manage a portfolio of data management platforms and mobile offerings in support of our... ...Healthcare Informatics platform. As a Software Engineer, you must possess world class technical... ...Knowledge of Site Reliability Engineering (SRE) concepts and practices is a plus....SoftwareFull timeTemporary workWork experience placementLive inRemote workFlexible hours
- ...based vulnerability management platform built for today’s dynamic IT... ...Responsible for developing software, tools, and scripts to automate... ...Collaboration with cloud engineers in understanding new cloud technologies... ...~2+ years of related SRE experience ~ Apply core software...SoftwareFull timeWork experience placementRemote work
- ...SRE Ropes assessment required As there is internal competition I ask that you only... ...focused SRE with strong cloud automation and software engineering skills who can leverage AI/LLMs to automate operations and improve platform reliability at scale. Key...SoftwareImmediate startRemote work
$175k - $215k
...Sr. Manager, Site Reliability Engineer provides strategic leadership across multiple SRE teams and their managers, ensuring... ...operational excellence across commerce platforms. Develops and executes... ...understanding of how SRE integrates with software development, security, and...Software$160k - $185k
...OverviewThe Sr. Manager, Site Reliability Engineering (SRE) leads the strategy, execution, and... ...responsible for ensuring that both digital platforms and in-club systems run reliably,... ...This role operates at the intersection of software engineering and infrastructure, driving...SoftwareWork at officeLocal areaRemote workWork from home- ...This Role As a Senior Application Support Engineer, you will help power DTCC's global... ...of Institutional Trade Processing (ITP) platforms that support cross-border equity and debt... ...Leveraging Site Reliability Engineering (SRE) principles, you will support a portfolio...Remote workFlexible hours
- ...of Feeld , GT is looking for a Senior Site Reliability Engineer (SRE) to join a fast-growing consumer mobile product in the online... .... The ideal profile is someone who started in backend/software engineering and has moved into SRE or reliability-focused work...SoftwareFull timeRemote work
- ...real estate closing and title insurance software. A division of Fidelity National... ...seeking a well-rounded Site Reliability Engineer (SRE) to join our Cloud Operations Team in... ...will develop in-depth knowledge of the platforms that run our solutions. This position will...SoftwareHourly payWork at officeRemote work
- ...partner is looking for an IT Automation & Platform Engineer based in Australia. This is a senior... ...combines business systems engineering, software development, automation, security, and... ...is desirable. ~ Experience applying SRE concepts, including SLIs and SLOs, to...SoftwareFull timeRemote work
$78k - $185k
...remaining days of the week.ABOUT THE TEAMThe Platform Engineering team is part of Parametric IT's... ...system reliability in all phases of the Software Development Lifecycle (SDLC). As a team... ...Platform Engineering (i.e., DevOps, SysOps, SRE) experience with a proven track record...SoftwareTemporary workWork at officeLocal areaRemote workWorldwide3 days per week$87.95k - $162.88k
...Senior AI DevOps Engineer (AI Ops / Platform Engineering)NTT DATA strives to hire exceptional, innovative... ...the next generation of AI-powered software delivery and cloud operations. In this... ...Platform Engineering, DevOps, Security, SRE, and AI teams to design secure, scalable...SoftwareTemporary workWork at officeRemote workFlexible hours3 days per week
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to SRE Platform Software Engineer. Be the first to apply!
- site reliability engineer remote Remote
- site reliability engineer Remote
- site reliability engineer sre Remote
- senior platform engineer Remote
- platform developer Remote
- platform engineer Remote
- client platform engineer Remote
- platform engineering manager Remote
- data platform engineer Remote
- cybersecurity software engineer Remote


