Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

SRE Monitoring Platform Software Engineer (Entry Level)

Full-time

Bitdeer Technologies Group

Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure.

Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to support the AI revolution. Bitdeer handles complex processes involved in computing such as equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial intelligence.

Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia.

To learn more, visit (

Position Overview  

Bitdeer is building an AI-operated GPU cloud — a global fleet of self-built and OEM-rented data centers running the world's most valuable compute, run by a platform that observes, protects, and operates the fleet. The SRE Platform team builds the monitoring and automation substrate that every other squad — storage, network, GPU, K8S, and L1 operators — depends on. Their signals become the system you help build.

As an entry-level Software Engineer on the SRE / Monitoring Platform team, you contribute to the NeoCloud SRE platform — the multi-region system that observes, protects, and operates a GPU rental fleet across self-built and OEM-rented data centers. You join a bounded context led by a senior engineer, take well-scoped components from design to production code that ships through GitOps + the CICD release pipeline, follows the Plugin Framework conventions, meets declared SLOs, and stays drift-free.

This is a build + learn role. You write code, write tests, and operate what you build under the guidance of a senior engineer. You participate in on-call as a shadow before taking primary. Within 12 months, you should be delivering components independently within your assigned area and growing toward owning a sub-context.

Key Responsibilities

Where you'll contribute (guided by a senior engineer)

  • Collection + Storage — help build collection-agent, metrics-store / logs-store / traces-store / profiles-store, enrichment-service, collection-monitor. Write ingestion, query, and storage-path code.
  • Alert + Correlation + SLO — contribute to alert-engine-framework, alert-correlation, slo-framework; implement and tune default alert rules.
  • Topology + Cluster-Health — contribute to topology-service, cluster-health-rollup, OSS-SRE-tool collection plugins for K8s / Slurm / Ray / Volcano / Kueue / KubeRay.
  • Remediation + Workflow + Jobs — help build remediation-actuator, orchestration / workflow components, inspection probes, job-scheduler.
  • Observability instrumentation — instrument services with metrics, logs, and traces via OpenTelemetry; build dashboards; write runbooks an on-call can follow.
  • Test discipline — write unit / integration / contract tests for everything you ship; participate in chaos and soak tests led by senior engineers.

Why this is a great first role

  • Greenfield with a well-defined vision. The Plugin Framework, GitOps pipeline, and SLO framework are decided; you build components inside them with a clear blueprint — not from a blank page.
  • You learn the full observability stack at production scale — ingest, query, storage — by building it, not just using it.
  • Mentorship-heavy. You work directly with senior and principal engineers who own the architecture; their expertise becomes your growth path.

Job Requirements

  • 0-2 years of software engineering experience (new graduates with strong projects or internships welcome).
  • Solid fundamentals in one programming language — Go (preferred), Python, Java, or Rust. You can write clean, tested, readable code and explain your design choices.
  • CS fundamentals — data structures, algorithms, concurrency, basic networking (TCP / and operating-system concepts (processes, threads, I/O). You can reason about correctness and performance.
  • Distributed systems basics — you understand the ideas behind idempotency, retries, back-pressure, caching, and eventual consistency, even if you haven't operated them at scale yet. Eagerness to go deep.
  • Monitoring / observability exposure — some hands-on with Prometheus, Grafana, Loki, or similar; can write a basic PromQL query and instrument a service. Eagerness to learn the ingest, query, and storage path of a real observability stack.
  • Familiarity with Linux and the shell; comfort reading system logs and using standard debugging tools.
  • Kubernetes basics — understand Pods, Services, Deployments; have run something on K8s (a project, lab, or internship).
  • Git + CI basics — branching, pull requests, and have used a CI pipeline (GitHub Actions, GitLab CI, or similar).
  • Test discipline — you write unit and integration tests as a habit, not an afterthought.
  • Communication — clear written and verbal English; can write a good PR description and ask good questions.
  • Curiosity and a learning mindset — the most important qualifier. You're excited to learn GPU / AI infrastructure, AIOps, distributed systems, and observability at production scale.

Nice-to-Haves

  • Internship or project in monitoring / observability, telemetry pipelines, or platform / SRE tooling.
  • Exposure to GPU / AI-infra — DCGM, InfiniBand / RoCE, Kubernetes GPU Operator, Slurm / Ray. Interest counts more than depth.
  • Exposure to AIOps / ML-adjacent tooling (anomaly detection, alert correlation).
  • Contributions to open-source observability or cloud-native projects.

--------------------------------------------------------------------

Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, color, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union.

Vacancy posted 12 days ago
Similar jobs that could be interesting for youBased on the SRE Monitoring Platform Software Engineer (Entry Level) in Remote vacancy
  • $130k - $180k

     ...seeking a Cloud Site Reliability Engineer (SRE)to work in our Arlington, VA...  ...itself. That means writing software and automation that lets...  ...systems across our federal cloud platform (AWS GovCloud, IL5 zero-...  ...& Incidents Build monitoring, logging, alerting, and tracing... 
    Software
    For contractors
    Work at office
    Remote work
    Shift work

    ECS Federal

    Arlington, VA
    13 hours ago
  • $106.5k - $177.5k

     ...The Site Reliability Engineering discipline at Noctua Technology...  ...treat operations as a software engineering challenge,...  ...as Code (IaC), monitor it through advanced observability...  ...Reliability Engineer (SRE) to join our dynamic...  ...and monitoring Service Level Objectives (SLOs), and... 
    Software
    Remote work

    Noctua Technology

    United States
    2 days ago
  • $1,000 per month

     ...financial wellness platform designed to help...  ...infrastructure our engineers use every day. When...  ...a Senior DevOps / SRE Engineer on this...  ...and shipped real software, not only infrastructure...  ...observability/monitoring (Datadog, OpenTelemetry...  ...PTOYour actual level and base salary... 
    Software
    Temporary work
    Work at office
    Immediate start
    Remote work
    Flexible hours

    Creditly Corp

    Philadelphia, PA
    2 days ago
  •  ...Senior Application Support Engineer, you will help power...  ...Trade Processing (ITP) platforms that support cross-...  ...Reliability Engineering (SRE) principles, you will support...  ...schedules and service-level commitments.Change,...  ...operational risk.Enhance monitoring, alerting,... 
    Suggested
    Remote work
    Flexible hours

    DTCC- The Depository Trust & Clearing Corporation

    Boston, MA
    2 hours ago
  •  ...parameters for hardware/software compatibility....  ...design. Performs engineering studies and...  ...functions. Education Level: Bachelor's Degree...  ...dynamic AWS Cloud Platform Engineer to provide...  ...role Experience with SRE principles and...  ...a plus, Platform Monitoring, Observability, &... 
    Software
    Immediate start
    Remote work

    Mindlance

    Reston, VA
    3 days ago
  • $207k - $300k

     ...design consulting, developing software platforms and frameworks, capacity...  ...they are live by measuring and monitoring availability, latency and...  ...reliability strategy for Home SRE.Minimum qualifications:Bachelor...  ...in Computer Science or Engineering, or a related field.Experience... 
    Software
    Worldwide

    Google

    San Francisco, CA
    2 days ago
  • $207k - $300k

     ...consulting, developing software platforms and frameworks,...  ...live by measuring and monitoring availability, latency...  ...Computer Science or Engineering.Experience mentoring...  ...Reliability Engineering (SRE) combines software and...  ...meet stringent Service Level Objectives. You will... 
    Software

    Google

    New York, NY
    1 day ago
  • $160k - $210k

     ...category-leading enterprise software that unleashes that...  ...team, we build the platforms and systems the entire...  ...This is a software engineering role. You will not be...  ..., engineer, and build SRE platform systems and capabilities...  ...in livesite monitoring rotations, handle escalations... 
    Software
    Work at office
    Immediate start
    Remote work

    UiPath

    Denver, CO
    2 days ago
  •  ...category-leading enterprise software that unleashes that...  ...team, we build the platforms and systems the entire...  ...This is a software engineering role. You will not be...  ..., engineer, and build SRE platform systems and capabilities...  ...in livesite monitoring rotations, handle escalations... 
    Software
    Work at office
    Immediate start
    Remote work

    Socket

    Denver, CO
    2 days ago
  • $180.5k - $236.91k

     .... We're hiring a Senior Software Engineer, Cloud Infrastructure / SRE to join our Engineering...  ...a full stack technology platform and a relentless focus on...  ...development of Service-Level Objectives (SLOs) for systems...  ...: Proficiency with monitoring using tools like Prometheus... 
    Software
    Full time
    Work at office
    Remote work

    Oscar Health Insurance

    Boston, MA
    13 hours ago
  •  ...TVRoku is the #1 TV streaming platform in the U.S., Canada, and...  ...talented and experienced Senior Software Engineer, MLOps/DevOps, to join the...  ...strong background in DevOps/SRE practices, cloud infrastructure...  ...evaluation, deployment, and monitoring — on top of a modern, cloud-... 
    Software
    Work at office
    Local area
    Remote work
    Monday to Thursday
    Flexible hours

    Roku

    Austin, TX
    1 day ago
  •  ...Technologies develops software for customers in the...  ...looking for an experienced SRE Engineer for a long-term...  ...once and reuse it across platforms. It helps confirm age,...  ...with CI/CD, monitoring, logging and incident...  ...English (Intermediate level and higher).   Working... 
    Software
    Full time
    Remote work
    Flexible hours

    Nitka Technologies Inc

    United States
    3 days ago
  •  ...POSITION Platform Engineer / SRE-DevOps Engineer (Redshift) REQUIRED SKILLS Strong hands-on experience administering...  ...concurrency scaling, and data distribution. Experience with monitoring, alerting, logging, and production incident management.... 

    Akaasa Technologies

    Remote
    2 days ago
  • $40 per hour

    A technology solutions provider is seeking a remote Junior SRE/DevOps Engineer. The ideal candidate should have foundational knowledge of Site Reliability Engineering (SRE) and Kubernetes. Responsibilities include gaining experience in a DevOps-driven environment. Applicants... 
    Long term contract
    Internship
    Remote work

    BayOne Solutions

    United States
    13 hours ago
  •  ...Senior Site Reliability Engineer to help us mature and...  ...our multi-cloud SaaS platform. Most of our footprint...  ...treats infrastructure like software. You'll have...  ...run by hand Build monitoring, alerting, and observability...  ...~​​6+ years in SRE, DevOps, or infrastructure... 
    Software
    Remote work
    Flexible hours

    Dental Intelligence

    United States
    2 days ago
  •  ...DescriptionThe AI Inference Engineer plays a critical...  ..., and monitoring system performance...  ...orchestration. Ensure software solutions are optimized...  ...against service-level agreements (SLAs)....  ..., and cloud platforms such as AWS, GCP,...  ...Background in MLOps or SRE roles focused on... 
    Software
    Full time
    Local area
    Immediate start

    F5 Networks

    San Jose, CA
    2 days ago
  •  ...DevOps Engineer Key Responsibilities Technical Skills • Strong hands-on experience...  ...Azure DevOps, etc.) • Strong experience in monitoring & observability tools (CloudWatch,...  ...Linux system administration DevOps / SRE Expertise • Strong understanding of DevOps... 
    Remote work

    Omni Inclusive

    United States
    5 days ago
  •  ...portfolio of data management platforms and mobile offerings in...  ...Informatics platform. As a Software Engineer, you must possess world class...  ...Site Reliability Engineering (SRE) concepts and practices is a...  ...cloud is a plus. Experience monitoring infrastructure with monitoring... 
    Software
    Full time
    Temporary work
    Work experience placement
    Live in
    Remote work
    Flexible hours

    ResMed

    San Diego, CA
    1 day ago
  •  ...writing, shipping, and running software effortless and efficient. We work closely with product-engineering to identify friction in the...  ...Support, and guide engineers on SRE related topics Partner...  ...~ Experience with monitoring / alerting (primarily with Prometheus... 
    Software
    Local area
    Remote work

    Kraken

    United States
    3 days ago
  • $185k - $200k

     ...JobsSenior Site Reliability Engineer (SRE) Dayton, OH (Remote)...  ...AFRL's Google Cloud Platform environment. This...  .... This is a senior-level position where deep,...  ...repository management, monitoring, and platform operations...  ...engineering, DevOps, software engineering, or a related... 
    Software
    Full time
    Remote work

    Metronome LLC

    Dayton, OH
    3 days ago
  •  ...Senior Kubernetes-focused SRE SRE Ropes assessment required...  ...strong cloud automation and software engineering skills who can leverage AI/...  ...operations and improve platform reliability at scale. Key...  ...Develop observability and monitoring capabilities using tools such... 
    Software
    Immediate start
    Remote work

    RIT Solutions

    Atlanta, GA
    5 days ago
  • $100k - $130k

     ...We are looking for a Platform Integration Engineer to join our engineering...  ...with DevOps/SRE to define deployment...  ...Qualifications ~3+ years of software engineering...  ...management and data quality monitoring ~ Working experience...  ...-Remote Position Level Associate Country... 
    Software
    Work experience placement
    Local area
    Remote work
    Flexible hours

    Huron

    Chicago, IL
    18 hours ago
  •  ...Job title : Platform Engineer / SRE-DevOps Engineer Amazon Redshift Location: Remote Duration: 3+Months Key Responsibilities...  ..., database, and data platform deployments. Monitor and optimize Redshift performance, storage, workloads, and... 
    Remote work

    Conch Technologies Inc

    New York, NY
    2 days ago
  •  ...closing and title insurance software. A division of Fidelity...  ...rounded Site Reliability Engineer (SRE) to join our Cloud Operations...  ...in-depth knowledge of the platforms that run our solutions....  ...to refine our service level indicators monitoring capabilities with the goal... 
    Software
    Hourly pay
    Work at office
    Remote work

    SoftPro

    Raleigh, NC
    1 day ago
  • $307k - $427k

     ...velocity.Ensure that SRE principles (SLOs,...  ...into the shipped software so external...  ...incident response, monitoring, and debugging paradigms...  ...who act as Level 3 support or manage...  ...experienced software engineers of large-scale projects...  ...of Google platforms, we make Google's... 
    Software

    Google

    Sunnyvale, TX
    3 days ago
  • $167k - $196.5k

     ...the data collaboration platform of choice for the...  ...platforms.The Global SRE team is responsible for...  ...Senior Site Reliability Engineer who is excited about establishing...  ...is important (Software Engineer, Site...  ...& Product Reliability monitoring and alerting Maintain... 
    Software
    Full time
    Work at office
    Remote work
    Work from home
    Flexible hours
    Night shift

    LiveRamp

    San Francisco, CA
    1 day ago
  •  ...SRE / DevOps Engineer Canada / Remote 6+ Months Contract Position Requirements Collaborate closely with Development teams to improve...  ...tools like Sonar. Manage application integration and monitoring using DataDog, SumoLogic, and alert systems like MS Teams.... 
    Contract work
    Remote work

    Veracity

    United States
    5 days ago
  •  ...Site Reliability Engineer (SRE) Location: Remote Shift Timings: 5:30 PM to 3:00 AM IST...  ...strong background in both log and metrics monitoring stacks, specifically ELK (Elasticsearch...  ..., and performance of our customer's platforms and services, bridging the gap between... 
    Remote work
    Shift work

    InOrg Global

    United States
    3 days ago
  • $165k - $225k

     ...Sr. Site Reliability Engineer (SRE) Chicago, IL or Remote Moonlite...  ..., network engineers, and platform engineering team, you'll architect...  ...SLIs, SLOs, and monitoring to meet enterprise reliability...  ...engineers, network engineers, and software developers. Preferred... 
    Software
    Remote work
    Flexible hours

    Moonlite AI

    Chicago, IL
    5 days ago
  • $100k - $180k

     ...Site Reliability Engineer (SRE) - Remote Bright Vision...  ...technology consulting and software development company...  ...continually pushing the platform toward higher...  ...continually refine service-level objectives (SLOs), service...  ...comprehensive monitoring, logging, and tracing... 
    Software
    Full time
    H1b
    Local area
    Immediate start
    Remote work
    Visa sponsorship

    Bright Vision Technologies

    Cranberry, PA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to SRE Monitoring Platform Software Engineer (Entry Level). Be the first to apply!