SRE Monitoring Platform Software Engineer (Entry Level) [Remote]
Bitdeer Technologies Group
- Remote job
Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure.
Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to support the AI revolution. Bitdeer handles complex processes involved in computing such as equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial intelligence.
Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia.
To learn more, visit (Position Overview
Bitdeer is building an AI-operated GPU cloud — a global fleet of self-built and OEM-rented data centers running the world's most valuable compute, run by a platform that observes, protects, and operates the fleet. The SRE Platform team builds the monitoring and automation substrate that every other squad — storage, network, GPU, K8S, and L1 operators — depends on. Their signals become the system you help build.
As an entry-level Software Engineer on the SRE / Monitoring Platform team, you contribute to the NeoCloud SRE platform — the multi-region system that observes, protects, and operates a GPU rental fleet across self-built and OEM-rented data centers. You join a bounded context led by a senior engineer, take well-scoped components from design to production code that ships through GitOps + the CICD release pipeline, follows the Plugin Framework conventions, meets declared SLOs, and stays drift-free.
This is a build + learn role. You write code, write tests, and operate what you build under the guidance of a senior engineer. You participate in on-call as a shadow before taking primary. Within 12 months, you should be delivering components independently within your assigned area and growing toward owning a sub-context.
Key Responsibilities
Where you'll contribute (guided by a senior engineer)
- Collection + Storage — help build collection-agent, metrics-store / logs-store / traces-store / profiles-store, enrichment-service, collection-monitor. Write ingestion, query, and storage-path code.
- Alert + Correlation + SLO — contribute to alert-engine-framework, alert-correlation, slo-framework; implement and tune default alert rules.
- Topology + Cluster-Health — contribute to topology-service, cluster-health-rollup, OSS-SRE-tool collection plugins for K8s / Slurm / Ray / Volcano / Kueue / KubeRay.
- Remediation + Workflow + Jobs — help build remediation-actuator, orchestration / workflow components, inspection probes, job-scheduler.
- Observability instrumentation — instrument services with metrics, logs, and traces via OpenTelemetry; build dashboards; write runbooks an on-call can follow.
- Test discipline — write unit / integration / contract tests for everything you ship; participate in chaos and soak tests led by senior engineers.
Why this is a great first role
- Greenfield with a well-defined vision. The Plugin Framework, GitOps pipeline, and SLO framework are decided; you build components inside them with a clear blueprint — not from a blank page.
- You learn the full observability stack at production scale — ingest, query, storage — by building it, not just using it.
- Mentorship-heavy. You work directly with senior and principal engineers who own the architecture; their expertise becomes your growth path.
Job Requirements
- 0-2 years of software engineering experience (new graduates with strong projects or internships welcome).
- Solid fundamentals in one programming language — Go (preferred), Python, Java, or Rust. You can write clean, tested, readable code and explain your design choices.
- CS fundamentals — data structures, algorithms, concurrency, basic networking (TCP / and operating-system concepts (processes, threads, I/O). You can reason about correctness and performance.
- Distributed systems basics — you understand the ideas behind idempotency, retries, back-pressure, caching, and eventual consistency, even if you haven't operated them at scale yet. Eagerness to go deep.
- Monitoring / observability exposure — some hands-on with Prometheus, Grafana, Loki, or similar; can write a basic PromQL query and instrument a service. Eagerness to learn the ingest, query, and storage path of a real observability stack.
- Familiarity with Linux and the shell; comfort reading system logs and using standard debugging tools.
- Kubernetes basics — understand Pods, Services, Deployments; have run something on K8s (a project, lab, or internship).
- Git + CI basics — branching, pull requests, and have used a CI pipeline (GitHub Actions, GitLab CI, or similar).
- Test discipline — you write unit and integration tests as a habit, not an afterthought.
- Communication — clear written and verbal English; can write a good PR description and ask good questions.
- Curiosity and a learning mindset — the most important qualifier. You're excited to learn GPU / AI infrastructure, AIOps, distributed systems, and observability at production scale.
Nice-to-Haves
- Internship or project in monitoring / observability, telemetry pipelines, or platform / SRE tooling.
- Exposure to GPU / AI-infra — DCGM, InfiniBand / RoCE, Kubernetes GPU Operator, Slurm / Ray. Interest counts more than depth.
- Exposure to AIOps / ML-adjacent tooling (anomaly detection, alert correlation).
- Contributions to open-source observability or cloud-native projects.
--------------------------------------------------------------------
Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, color, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union.
$145k
...Platform Engineer at Akuna Capital Akuna Capital is... ...and security, and monitor our hybrid-cloud environments... ...Engineering, SRE, DevOps, or... ...pipelines and modern software delivery practices... ...Please note that level/title will be... ...take a look at our Entry Level and Intern positions...SoftwareEntry levelInternshipWork at officeRemote work- ...currently seeking a Site Reliability / Platform Engineer - Remote position with a Leading... ...are seeking a highly hands-on SRE / Platform Engineer for a hybrid software development, Site Reliability... ...across an environment where monitoring and operational data have historically...SoftwareHourly payPermanent employmentContract workTemporary workRemote work
$130k - $180k
...seeking a Cloud Site Reliability Engineer (SRE)to work in our Arlington, VA... ...itself. That means writing software and automation that lets... ...systems across our federal cloud platform (AWS GovCloud, IL5 zero-... ...& Incidents Build monitoring, logging, alerting, and tracing...SoftwareFor contractorsWork at officeRemote workShift work- Platform Engineer (CI, CD, CT, Ruby, AWS, Azure, Scripting, Monitoring) in VA or NYC AWS, Azure, CD, CI, CT, Docker, Perl, Python, REST API, SQL Location: Virginia... ...initiative and enjoy working with engineers to make the software development process as painless as possible ·...SoftwarePermanent employmentFull timeRemote work
- ...parameters for hardware/software compatibility.... ...design. Performs engineering studies and... ...functions. Education Level: Bachelor's Degree... ...dynamic AWS Cloud Platform Engineer to provide... ...role Experience with SRE principles and... ...a plus, Platform Monitoring, Observability, &...SoftwareImmediate startRemote work
- ...Description Environmental Monitoring Technician – Water &... ...monitoring, civil engineering, instrumentation or... ...technology. We are open to entry-level candidates and will... ...WATERWAI® monitoring platform**. Identify... ...equipment, WATERWAI software, wireless communications...SoftwareEntry levelFor contractorsRemote work
$106.5k - $177.5k
...Software Engineer- Site Reliability Engineering (SRE) DC, MD, VA, CA The Site Reliability Engineering discipline... ...using Infrastructure as Code (IaC), monitor it through advanced observability... ...defining and monitoring Service Level Objectives (SLOs), and implementing...SoftwareRemote work- ...Senior Application Support Engineer, you will help power... ...Trade Processing (ITP) platforms that support cross-... ...Reliability Engineering (SRE) principles, you will support... ...schedules and service-level commitments.Change,... ...operational risk.Enhance monitoring, alerting,...Remote workFlexible hours
- ...States As a System Engineer, Sr. on our team, your... ...Team (SET) to generate monitoring/observability... ...skills in enterprise-level triage and incident resolution... ...Use your hardware and software experience to help strengthen... ...to understand their platform designs and how they...SoftwareContract workWork at officeRemote work
- ...DescriptionThe AI Inference Engineer plays a critical... ..., and monitoring system performance... ...orchestration. Ensure software solutions are optimized... ...against service-level agreements (SLAs).... ..., and cloud platforms such as AWS, GCP,... ...Background in MLOps or SRE roles focused on...SoftwareFull timeLocal areaImmediate start
- ...Our client seeks an Associate Security Engineer, Cyber Monitoring for a role on the Network Monitoring... ...a motivated, detail‑oriented entry‑level Cyber Security Engineer to join a Cyber... ...Information Technology, Computer Science, Software Engineering, Computer Engineering)...SoftwareEntry levelHourly payLocal areaRemote workEarly shift
$3,200 - $4,800 per month
...need an English-fluent Mid Level Cloud / DevOps (SRE) Engineer , based in Latin America... ...Implement and maintain monitoring, logging, alerting, and... .... Collaborate with software engineers, Tech Leads, and... ...container orchestration platforms. Experience with Terraform...SoftwareFull timeContract workRemote work- ...Remote50% Remote The Lead Monitoring and Observability Engineer (IC3) serves as a senior technical... ...who guides monitoring-platform evolution, improves... ...Cloud, AppDev, Security, and SRE teams to establish monitoring... ...monitoring, or service-level reporting. • Experience leading...Remote work
$116.4k - $194k
...Fraud and Financial Crime platforms. Leads production support, release engineering, observability,... ...resiliency, streamline software delivery, eliminate manual... ...validation, and post-deployment monitoring.Partner with... ...initiatives.Understanding of SRE principles, reliability...SoftwareFull timeWork experience placementWork from home- ...reliability, observability, and cloud platform engineering excellence across a large,... ...personally raise the bar on SRE and observability practices,... ...with a dedicated SRE/monitoring team and a centralized... ...progressive, hands-on experience in software engineering with an...SoftwareFull timeRemote work2 days per week
$207k - $300k
...consulting, developing software platforms and frameworks,... ...live by measuring and monitoring availability, latency... ...Computer Science or Engineering.Experience mentoring... ...Reliability Engineering (SRE) combines software and... ...meet stringent Service Level Objectives. You will...Software- ...Senior Kubernetes-focused SRE SRE Ropes assessment required... ...strong cloud automation and software engineering skills who can leverage AI/... ...operations and improve platform reliability at scale. Key... ...Develop observability and monitoring capabilities using tools such...SoftwareImmediate startRemote work
- ...portfolio of data management platforms and mobile offerings in... ...Informatics platform. As a Software Engineer, you must possess world class... ...Site Reliability Engineering (SRE) concepts and practices is a... ...cloud is a plus. Experience monitoring infrastructure with monitoring...SoftwareFull timeTemporary workWork experience placementLive inRemote workFlexible hours
$175k - $215k
...Site Reliability Engineer provides strategic... ...leadership across multiple SRE teams and their... ...across commerce platforms. Develops and... ...advanced telemetry and monitoring practices,... ...SRE integrates with software development, security... ...familiarity with FAANG-level engineering...Software$100k - $180k
...Site Reliability Engineer (SRE) - Remote Bright Vision... ...technology consulting and software development company... ...continually pushing the platform toward higher... ...continually refine service-level objectives (SLOs), service... ...comprehensive monitoring, logging, and tracing...SoftwareFull timeH1bLocal areaRemote workVisa sponsorship- ...Vantor is seeking a Senior Platform Engineer, AI Infrastructure to... ...the intersection of software engineering, cloud... ...systems that build, deploy, monitor, and operate mission-... ...to DevOps, SRE, developer productivity... ...practices, including service-level objectives, monitoring...SoftwarePermanent employmentFull timeRemote work
$83.6k - $196k
...You are: A cloud and platform engineer who builds the... ...technologies. Implement monitoring, logging, alerting, observability... ...engineering, SRE, or comparable work, or... ...automation, and modern software delivery practices. Experience... ...role, skill set, and level of experience....SoftwareWork at officeLocal areaRemote work$78k - $185k
...week.ABOUT THE TEAMThe Platform Engineering team is part of... ...in all phases of the Software Development Lifecycle... ...KubernetesProvide observability via monitoring tools such as Datadog... ....e., DevOps, SysOps, SRE) experience with a... ...Type:Full timeJob Level:ProfessionalPosted...SoftwareTemporary workWork at officeLocal areaRemote workWorldwide3 days per week- ...robust suite of fintech software enables us to support... ...is seeking a Platform Engineer to join our Cloud Engineering... ...and Cloud Storage Monitor, troubleshoot, and resolve... ...of DevOps and SRE principles including CI... ...#engineering #mid-level #full-time #APEX Please...SoftwareFull timeWork experience placementWork at officeWork from home3 days per week
$175k - $215k
...Site Reliability Engineer At Disney, we're... ...leadership for multiple SRE teams, fostering a... ...telemetry and monitoring practices,... ...Expertise in cloud platforms (AWS, GCP, Azure),... ...SRE integrates with software development, security... ...familiarity with FAANG-level engineering...SoftwareLocal area$40 per hour
A technology solutions provider is seeking a remote Junior SRE/DevOps Engineer. The ideal candidate should have foundational knowledge of Site Reliability Engineering (SRE) and Kubernetes. Responsibilities include gaining experience in a DevOps-driven environment. Applicants...Long term contractInternshipRemote work$180.5k - $236.91k
.... We're hiring a Senior Software Engineer, Cloud Infrastructure / SRE to join our Engineering... ...a full stack technology platform and a relentless focus on... ...development of Service-Level Objectives (SLOs) for systems... ...: Proficiency with monitoring using tools like Prometheus...SoftwareFull timeWork at officeRemote work- ...technology consulting and software development company... ...Job Title: Kafka Platform Engineer Location: 100% Remote... ...Build comprehensive monitoring, alerting, logging, and... ...data engineers, DevOps, SRE, and enterprise... ...experience . ~ Expert-level knowledge of Kafka internals...SoftwareFull timeH1bLocal areaRemote workVisa sponsorship
- ...not just building software—we’re reimagining... ...specialty-specific cloud platform that places... ...a Manager, Cloud Engineering at Modernizing Medicine... ...industry-standard SRE practices.Foster a... ...observability and monitoring solutions using... ...support junior and mid-level team members,...SoftwareFull timeFixed term contractWork at officeRemote workFlexible hours
$160k - $210k
...category-leading enterprise software that unleashes that... ...team, we build the platforms and systems the entire... ...This is a software engineering role. You will not be... ..., engineer, and build SRE platform systems and capabilities... ...in livesite monitoring rotations, handle escalations...SoftwareWork at officeImmediate startRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to SRE Monitoring Platform Software Engineer (Entry Level) [Remote]. Be the first to apply!




