SRE Monitoring Platform Software Engineer (Entry Level) [Remote]
Bitdeer Technologies Group
- Remote job
Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure.
Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to support the AI revolution. Bitdeer handles complex processes involved in computing such as equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial intelligence.
Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia.
To learn more, visit (Position Overview
Bitdeer is building an AI-operated GPU cloud — a global fleet of self-built and OEM-rented data centers running the world's most valuable compute, run by a platform that observes, protects, and operates the fleet. The SRE Platform team builds the monitoring and automation substrate that every other squad — storage, network, GPU, K8S, and L1 operators — depends on. Their signals become the system you help build.
As an entry-level Software Engineer on the SRE / Monitoring Platform team, you contribute to the NeoCloud SRE platform — the multi-region system that observes, protects, and operates a GPU rental fleet across self-built and OEM-rented data centers. You join a bounded context led by a senior engineer, take well-scoped components from design to production code that ships through GitOps + the CICD release pipeline, follows the Plugin Framework conventions, meets declared SLOs, and stays drift-free.
This is a build + learn role. You write code, write tests, and operate what you build under the guidance of a senior engineer. You participate in on-call as a shadow before taking primary. Within 12 months, you should be delivering components independently within your assigned area and growing toward owning a sub-context.
Key Responsibilities
Where you'll contribute (guided by a senior engineer)
- Collection + Storage — help build collection-agent, metrics-store / logs-store / traces-store / profiles-store, enrichment-service, collection-monitor. Write ingestion, query, and storage-path code.
- Alert + Correlation + SLO — contribute to alert-engine-framework, alert-correlation, slo-framework; implement and tune default alert rules.
- Topology + Cluster-Health — contribute to topology-service, cluster-health-rollup, OSS-SRE-tool collection plugins for K8s / Slurm / Ray / Volcano / Kueue / KubeRay.
- Remediation + Workflow + Jobs — help build remediation-actuator, orchestration / workflow components, inspection probes, job-scheduler.
- Observability instrumentation — instrument services with metrics, logs, and traces via OpenTelemetry; build dashboards; write runbooks an on-call can follow.
- Test discipline — write unit / integration / contract tests for everything you ship; participate in chaos and soak tests led by senior engineers.
Why this is a great first role
- Greenfield with a well-defined vision. The Plugin Framework, GitOps pipeline, and SLO framework are decided; you build components inside them with a clear blueprint — not from a blank page.
- You learn the full observability stack at production scale — ingest, query, storage — by building it, not just using it.
- Mentorship-heavy. You work directly with senior and principal engineers who own the architecture; their expertise becomes your growth path.
Job Requirements
- 0-2 years of software engineering experience (new graduates with strong projects or internships welcome).
- Solid fundamentals in one programming language — Go (preferred), Python, Java, or Rust. You can write clean, tested, readable code and explain your design choices.
- CS fundamentals — data structures, algorithms, concurrency, basic networking (TCP / and operating-system concepts (processes, threads, I/O). You can reason about correctness and performance.
- Distributed systems basics — you understand the ideas behind idempotency, retries, back-pressure, caching, and eventual consistency, even if you haven't operated them at scale yet. Eagerness to go deep.
- Monitoring / observability exposure — some hands-on with Prometheus, Grafana, Loki, or similar; can write a basic PromQL query and instrument a service. Eagerness to learn the ingest, query, and storage path of a real observability stack.
- Familiarity with Linux and the shell; comfort reading system logs and using standard debugging tools.
- Kubernetes basics — understand Pods, Services, Deployments; have run something on K8s (a project, lab, or internship).
- Git + CI basics — branching, pull requests, and have used a CI pipeline (GitHub Actions, GitLab CI, or similar).
- Test discipline — you write unit and integration tests as a habit, not an afterthought.
- Communication — clear written and verbal English; can write a good PR description and ask good questions.
- Curiosity and a learning mindset — the most important qualifier. You're excited to learn GPU / AI infrastructure, AIOps, distributed systems, and observability at production scale.
Nice-to-Haves
- Internship or project in monitoring / observability, telemetry pipelines, or platform / SRE tooling.
- Exposure to GPU / AI-infra — DCGM, InfiniBand / RoCE, Kubernetes GPU Operator, Slurm / Ray. Interest counts more than depth.
- Exposure to AIOps / ML-adjacent tooling (anomaly detection, alert correlation).
- Contributions to open-source observability or cloud-native projects.
--------------------------------------------------------------------
Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, color, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union.
$83.52k - $125.28k
...currently seeking a Lead ML Platform Engineer (SRE / FTE / Onsite) to join our... ...to build, validate, deploy, monitor, and operate predictive models... ...dashboards, alerts, service-level indicators, service-level... ..., audit logging, and secure software delivery.Experience implementing...SoftwareFull timeTemporary workWork at officeRemote workFlexible hours$172.54k - $231k
...resilience. (15%) Proactively monitor system health,... ...enhance the ELK/EFK Logging platform across on-premise and... ...solutions based on SRE principles. (15%)... ...other members of the engineering organization. (5%) What... ...the job offered or as a Software Architect, Software Engineer...SoftwareFull timeWork at officeWork from home3 days per week- Platform Engineer (CI, CD, CT, Ruby, AWS, Azure, Scripting, Monitoring) in VA or NYC AWS, Azure, CD, CI, CT, Docker, Perl, Python, REST API, SQL Location: Virginia... ...initiative and enjoy working with engineers to make the software development process as painless as possible ·...SoftwarePermanent employmentFull timeRemote work
- ...parameters for hardware/software compatibility.... ...design. Performs engineering studies and... ...functions. Education Level: Bachelor's Degree... ...dynamic AWS Cloud Platform Engineer to provide... ...role Experience with SRE principles and... ...a plus, Platform Monitoring, Observability, &...SoftwareImmediate startRemote work
$1,000 per month
...financial wellness platform designed to help... ...infrastructure our engineers use every day. When... ...a Senior DevOps / SRE Engineer on this... ...and shipped real software, not only infrastructure... ...observability/monitoring (Datadog, OpenTelemetry... ...PTOYour actual level and base salary...SoftwareTemporary workWork at officeImmediate startRemote workFlexible hours- ...Our client seeks an Associate Security Engineer, Cyber Monitoring for a role on the Network Monitoring... ...a motivated, detail‑oriented entry‑level Cyber Security Engineer to join a Cyber... ...Information Technology, Computer Science, Software Engineering, Computer Engineering)...SoftwareEntry levelHourly payLocal areaRemote workEarly shift
- ...We are looking for an experienced Software Engineer with a strong background in Platform Engineering to build and scale... ...platform resilience by improving monitoring, logging, and observability using... ...Qualifications:6+ years in DevOps, SRE, or Platform Engineering plus a Bachelor...SoftwareWork at officeWork from homeFlexible hours3 days per week
- ...DescriptionThe AI Inference Engineer plays a critical... ..., and monitoring system performance... ...orchestration. Ensure software solutions are optimized... ...against service-level agreements (SLAs).... ..., and cloud platforms such as AWS, GCP,... ...Background in MLOps or SRE roles focused on...SoftwareFull timeLocal areaImmediate start
$149.4k - $202k
...Senior Software Engineer- Site Reliability Engineering (SRE) DC, MD, VA, CA The Site Reliability... ...Infrastructure as Code (IaC), monitor it through advanced... ...and monitoring Service Level Objectives (SLOs), and implementing... ...services or entire platforms, ensuring alignment with...SoftwareRemote work- OB SUMMARYThe Systems Engineer - Site Reliability Engineering (SRE) is responsible for the reliability... ...such as Service Level Objectives, Error... ...Management, Observability & Monitoring, Blameless Postmortems,... ...(including hardware and software) to enter data and/ or process...SoftwareFull timeFor contractorsWork at officeRemote workFlexible hoursShift work
$212.5k - $270k
...applied to a range of vehicle platforms and product use cases. The... ...problems related to Waymo Fleet Monitoring and building reusable... ...with Product, UX, and other engineers to design and develop internal... ...training and education, and skill level. Your recruiter can share more...SoftwareFull timeRemote work$207k - $300k
...design consulting, developing software platforms and frameworks, capacity... ...they are live by measuring and monitoring availability, latency and... ...reliability strategy for Home SRE.Minimum qualifications:Bachelor... ...in Computer Science or Engineering, or a related field.Experience...SoftwareWorldwide- ...Senior Site Reliability Engineer to help us mature and... ...our multi-cloud SaaS platform. Most of our footprint... ...treats infrastructure like software. You'll have... ...run by hand Build monitoring, alerting, and observability... ...~6+ years in SRE, DevOps, or infrastructure...SoftwareRemote workFlexible hours
- ...Senior Site Reliability Engineer with direct experience... ...AI or machine learning platforms in large-scale production... ...infrastructure-only SRE role then this is the position... ...applications Monitoring the performance, availability... ...logs, traces, service-level indicators, and...Permanent employmentRemote work
- ...Lead SRE Position in India This position... ...Site Reliability Engineering function... ...across cloud-native platforms. The position combines... ...including Service Level Indicators, Service... ...Establish and monitor key metrics covering... ...experience across software engineering, DevOps...SoftwareRemote work
- ...vulnerability management platform built for today’s... ...defined service level objectives... ...Responsible for developing software, tools, and... ..., management, and monitoring of production systems... ...with cloud engineers in understanding new... ...years of related SRE experience ~ Apply...SoftwareFull timeWork experience placementRemote work
- ...looking for a Site Reliability Engineer (SRE) to join a fast-growing... ...model, strong observability, monitoring and reliable incident response... ...who started in backend/software engineering and has moved into... ...for key services and product-level metrics. Improve monitoring...SoftwareFull timeRemote work
- ...writing, shipping, and running software effortless and efficient. We work closely with product-engineering to identify friction in the... ...Support, and guide engineers on SRE related topics Partner... ...~ Experience with monitoring / alerting (primarily with Prometheus...SoftwareLocal areaRemote work
$85 - $95 per hour
...Enterprise Site Reliability Engineering (SRE) with our consumer... ...to a full-time AVP-level position based on... ...reliability a core product and platform capability.... ...requirements throughout the software development lifecycle... ..., orchestration, monitoring, logging, distributed...SoftwareHourly payPermanent employmentFull timeContract workImmediate startRemote work$40 per hour
A technology solutions provider is seeking a remote Junior SRE/DevOps Engineer. The ideal candidate should have foundational knowledge of Site Reliability Engineering (SRE) and Kubernetes. Responsibilities include gaining experience in a DevOps-driven environment. Applicants...Long term contractInternshipRemote work- ...portfolio of data management platforms and mobile offerings in... ...Informatics platform. As a Software Engineer, you must possess world class... ...Site Reliability Engineering (SRE) concepts and practices is a... ...cloud is a plus. Experience monitoring infrastructure with monitoring...SoftwareFull timeTemporary workWork experience placementLive inRemote workFlexible hours
$100k - $180k
...Site Reliability Engineer (SRE) - Remote Bright Vision... ...technology consulting and software development company... ...continually pushing the platform toward higher... ...continually refine service-level objectives (SLOs), service... ...comprehensive monitoring, logging, and tracing...SoftwareFull timeH1bLocal areaImmediate startRemote workVisa sponsorship$180.5k - $236.91k
.... We're hiring a Senior Software Engineer, Cloud Infrastructure / SRE to join our Engineering... ...a full stack technology platform and a relentless focus on... ...development of Service-Level Objectives (SLOs) for systems... ...: Proficiency with monitoring using tools like Prometheus...SoftwareRemote jobFull timeWork at office- ...SRE Engineer – Very Critical SFO CA (100% Remote) End Client: Airbnb Exp : 6 to 8 years only, don’t get profiles beyond 8+ years... ...Centres Operation preferably with volume above 5K TPS Monitoring & alerting setup experience with AWS, New Relic, DataDog, PagerDuty...Remote work
$307k - $427k
...velocity.Ensure that SRE principles (SLOs,... ...into the shipped software so external... ...incident response, monitoring, and debugging paradigms... ...who act as Level 3 support or manage... ...experienced software engineers of large-scale projects... ...of Google platforms, we make Google's...Software- ...Platform Engineering Manager Sia Partners is looking for a... ...pivotal in bridging high-level product vision with... ...support AI, data, and software workloads Key... ...Platform / DevOps / SRE engineering teams, ensuring... ...management, cost monitoring, and cost optimization...SoftwareRemote work
$78k - $185k
...week.ABOUT THE TEAMThe Platform Engineering team is part of... ...in all phases of the Software Development Lifecycle... ...KubernetesProvide observability via monitoring tools such as Datadog... ....e., DevOps, SysOps, SRE) experience with a... ...Type:Full timeJob Level:ProfessionalPosted...SoftwareTemporary workWork at officeLocal areaRemote workWorldwide3 days per week$105k - $158k
Platform Engineer - Healthcare Cloud PlatformSpecialist I - DevOps EngineeringWho We... ...· Site Reliability Engineering (SRE) · Establish and monitor Service Level Indicators (SLIs), Service Level... ...observability. · Automate patching, software deployment, inventory management...SoftwareFull timeTemporary workPart timeWork at officeLocal areaRemote workFlexible hours$139.2k - $235.2k
...intelligent orchestration platform for DevSecOps.... ..., more secure software faster. The same... ...a Senior Platform Engineer on the Orbit team,... ...and architecture-level features across GitLab... ...the deployment, monitoring, and operations of... ...intelligence, and SRE teams. What you...SoftwareFull timeRemote workFlexible hours- ...Gen. About the Role: Engine by Gen is a leader in... ...premier embedded finance platform for enterprise... ...years of experience in software engineering, platform engineering, or DevOps/SRE roles. ~ Strong systems... ...Kubernetes, Terraform, CI/CD, monitoring). ~ Experience...SoftwareFull timeFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to SRE Monitoring Platform Software Engineer (Entry Level) [Remote]. Be the first to apply!



