Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

SRE Monitoring Platform Software Engineer (Entry Level) [Remote]

Full-time

Bitdeer Technologies Group

Remote
  • Remote job

Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure.

Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to support the AI revolution. Bitdeer handles complex processes involved in computing such as equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial intelligence.

Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia.

To learn more, visit  (

Position Overview  

Bitdeer is building an AI-operated GPU cloud — a global fleet of self-built and OEM-rented data centers running the world's most valuable compute, run by a platform that observes, protects, and operates the fleet. The SRE Platform team builds the monitoring and automation substrate that every other squad — storage, network, GPU, K8S, and L1 operators — depends on. Their signals become the system you help build.

As an entry-level Software Engineer on the SRE / Monitoring Platform team, you contribute to the NeoCloud SRE platform — the multi-region system that observes, protects, and operates a GPU rental fleet across self-built and OEM-rented data centers. You join a bounded context led by a senior engineer, take well-scoped components from design to production code that ships through GitOps + the CICD release pipeline, follows the Plugin Framework conventions, meets declared SLOs, and stays drift-free.

This is a build + learn role. You write code, write tests, and operate what you build under the guidance of a senior engineer. You participate in on-call as a shadow before taking primary. Within 12 months, you should be delivering components independently within your assigned area and growing toward owning a sub-context.

Key Responsibilities

Where you'll contribute (guided by a senior engineer)

  • Collection + Storage — help build collection-agent, metrics-store / logs-store / traces-store / profiles-store, enrichment-service, collection-monitor. Write ingestion, query, and storage-path code.
  • Alert + Correlation + SLO — contribute to alert-engine-framework, alert-correlation, slo-framework; implement and tune default alert rules.
  • Topology + Cluster-Health — contribute to topology-service, cluster-health-rollup, OSS-SRE-tool collection plugins for K8s / Slurm / Ray / Volcano / Kueue / KubeRay.
  • Remediation + Workflow + Jobs — help build remediation-actuator, orchestration / workflow components, inspection probes, job-scheduler.
  • Observability instrumentation — instrument services with metrics, logs, and traces via OpenTelemetry; build dashboards; write runbooks an on-call can follow.
  • Test discipline — write unit / integration / contract tests for everything you ship; participate in chaos and soak tests led by senior engineers.

Why this is a great first role

  • Greenfield with a well-defined vision. The Plugin Framework, GitOps pipeline, and SLO framework are decided; you build components inside them with a clear blueprint — not from a blank page.
  • You learn the full observability stack at production scale — ingest, query, storage — by building it, not just using it.
  • Mentorship-heavy. You work directly with senior and principal engineers who own the architecture; their expertise becomes your growth path.

Job Requirements

  • 0-2 years of software engineering experience (new graduates with strong projects or internships welcome).
  • Solid fundamentals in one programming language — Go (preferred), Python, Java, or Rust. You can write clean, tested, readable code and explain your design choices.
  • CS fundamentals — data structures, algorithms, concurrency, basic networking (TCP / and operating-system concepts (processes, threads, I/O). You can reason about correctness and performance.
  • Distributed systems basics — you understand the ideas behind idempotency, retries, back-pressure, caching, and eventual consistency, even if you haven't operated them at scale yet. Eagerness to go deep.
  • Monitoring / observability exposure — some hands-on with Prometheus, Grafana, Loki, or similar; can write a basic PromQL query and instrument a service. Eagerness to learn the ingest, query, and storage path of a real observability stack.
  • Familiarity with Linux and the shell; comfort reading system logs and using standard debugging tools.
  • Kubernetes basics — understand Pods, Services, Deployments; have run something on K8s (a project, lab, or internship).
  • Git + CI basics — branching, pull requests, and have used a CI pipeline (GitHub Actions, GitLab CI, or similar).
  • Test discipline — you write unit and integration tests as a habit, not an afterthought.
  • Communication — clear written and verbal English; can write a good PR description and ask good questions.
  • Curiosity and a learning mindset — the most important qualifier. You're excited to learn GPU / AI infrastructure, AIOps, distributed systems, and observability at production scale.

Nice-to-Haves

  • Internship or project in monitoring / observability, telemetry pipelines, or platform / SRE tooling.
  • Exposure to GPU / AI-infra — DCGM, InfiniBand / RoCE, Kubernetes GPU Operator, Slurm / Ray. Interest counts more than depth.
  • Exposure to AIOps / ML-adjacent tooling (anomaly detection, alert correlation).
  • Contributions to open-source observability or cloud-native projects.

--------------------------------------------------------------------

Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, color, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union.

Vacancy posted 16 days ago
Similar jobs that could be interesting for youBased on the SRE Monitoring Platform Software Engineer (Entry Level) [Remote] in Remote vacancy
  • $83.52k - $125.28k

     ...currently seeking a Lead ML Platform Engineer (SRE / FTE / Onsite) to join our...  ...to build, validate, deploy, monitor, and operate predictive models...  ...dashboards, alerts, service-level indicators, service-level...  ..., audit logging, and secure software delivery.Experience implementing... 
    Software
    Full time
    Temporary work
    Work at office
    Remote work
    Flexible hours

    NTT DATA

    Charlotte, NC
    1 day ago
  • $172.54k - $231k

     ...resilience. (15%) Proactively monitor system health,...  ...enhance the ELK/EFK Logging platform across on-premise and...  ...solutions based on SRE principles. (15%)...  ...other members of the engineering organization. (5%) What...  ...the job offered or as a Software Architect, Software Engineer... 
    Software
    Full time
    Work at office
    Work from home
    3 days per week

    Verizon

    Basking Ridge, NJ
    1 day ago
  • Platform Engineer (CI, CD, CT, Ruby, AWS, Azure, Scripting, Monitoring) in VA or NYC AWS, Azure, CD, CI, CT, Docker, Perl, Python, REST API, SQL Location: Virginia...  ...initiative and enjoy working with engineers to make the software development process as painless as possible ·... 
    Software
    Permanent employment
    Full time
    Remote work

    DBA Web Technologies

    New York, NY
    a month ago
  •  ...parameters for hardware/software compatibility....  ...design. Performs engineering studies and...  ...functions. Education Level: Bachelor's Degree...  ...dynamic AWS Cloud Platform Engineer to provide...  ...role Experience with SRE principles and...  ...a plus, Platform Monitoring, Observability, &... 
    Software
    Immediate start
    Remote work

    Mindlance

    Reston, VA
    1 day ago
  • $1,000 per month

     ...financial wellness platform designed to help...  ...infrastructure our engineers use every day. When...  ...a Senior DevOps / SRE Engineer on this...  ...and shipped real software, not only infrastructure...  ...observability/monitoring (Datadog, OpenTelemetry...  ...PTOYour actual level and base salary... 
    Software
    Temporary work
    Work at office
    Immediate start
    Remote work
    Flexible hours

    Creditly Corp

    Philadelphia, PA
    5 days ago
  •  ...Our client seeks an Associate Security Engineer, Cyber Monitoring for a role on the Network Monitoring...  ...a motivated, detail‑oriented entry‑level Cyber Security Engineer to join a Cyber...  ...Information Technology, Computer Science, Software Engineering, Computer Engineering)... 
    Software
    Entry level
    Hourly pay
    Local area
    Remote work
    Early shift

    Eliassen Group

    Aiken, SC
    22 days ago
  •  ...We are looking for an experienced Software Engineer with a strong background in Platform Engineering to build and scale...  ...platform resilience by improving monitoring, logging, and observability using...  ...Qualifications:6+ years in DevOps, SRE, or Platform Engineering plus a Bachelor... 
    Software
    Work at office
    Work from home
    Flexible hours
    3 days per week

    Zions Bancorporation

    Midvale, UT
    2 days ago
  •  ...DescriptionThe AI Inference Engineer plays a critical...  ..., and monitoring system performance...  ...orchestration. Ensure software solutions are optimized...  ...against service-level agreements (SLAs)....  ..., and cloud platforms such as AWS, GCP,...  ...Background in MLOps or SRE roles focused on... 
    Software
    Full time
    Local area
    Immediate start

    F5 Networks

    San Jose, CA
    5 days ago
  • $149.4k - $202k

     ...Senior Software Engineer- Site Reliability Engineering (SRE) DC, MD, VA, CA The Site Reliability...  ...Infrastructure as Code (IaC), monitor it through advanced...  ...and monitoring Service Level Objectives (SLOs), and implementing...  ...services or entire platforms, ensuring alignment with... 
    Software
    Remote work

    Noctua Technology

    Virginia, MN
    5 days ago
  • OB SUMMARYThe Systems Engineer - Site Reliability Engineering (SRE) is responsible for the reliability...  ...such as Service Level Objectives, Error...  ...Management, Observability & Monitoring, Blameless Postmortems,...  ...(including hardware and software) to enter data and/ or process... 
    Software
    Full time
    For contractors
    Work at office
    Remote work
    Flexible hours
    Shift work

    Marriott International

    Bethesda, MD
    3 days ago
  • $212.5k - $270k

     ...applied to a range of vehicle platforms and product use cases. The...  ...problems related to Waymo Fleet Monitoring and building reusable...  ...with Product, UX, and other engineers to design and develop internal...  ...training and education, and skill level. Your recruiter can share more... 
    Software
    Full time
    Remote work

    Waymo

    San Francisco, CA
    1 day ago
  • $207k - $300k

     ...design consulting, developing software platforms and frameworks, capacity...  ...they are live by measuring and monitoring availability, latency and...  ...reliability strategy for Home SRE.Minimum qualifications:Bachelor...  ...in Computer Science or Engineering, or a related field.Experience... 
    Software
    Worldwide

    Google

    San Francisco, CA
    5 days ago
  •  ...Senior Site Reliability Engineer to help us mature and...  ...our multi-cloud SaaS platform. Most of our footprint...  ...treats infrastructure like software. You'll have...  ...run by hand Build monitoring, alerting, and observability...  ...~​​6+ years in SRE, DevOps, or infrastructure... 
    Software
    Remote work
    Flexible hours

    Dental Intelligence

    United States
    4 days ago
  •  ...Senior Site Reliability Engineer with direct experience...  ...AI or machine learning platforms in large-scale production...  ...infrastructure-only SRE role then this is the position...  ...applications Monitoring the performance, availability...  ...logs, traces, service-level indicators, and... 
    Permanent employment
    Remote work

    HTC Global Services Inc

    Seattle, WA
    5 days ago
  •  ...Lead SRE Position in India This position...  ...Site Reliability Engineering function...  ...across cloud-native platforms. The position combines...  ...including Service Level Indicators, Service...  ...Establish and monitor key metrics covering...  ...experience across software engineering, DevOps... 
    Software
    Remote work

    Jobgether

    United States
    5 hours ago
  •  ...vulnerability management platform built for today’s...  ...defined service level objectives...  ...Responsible for developing software, tools, and...  ..., management, and monitoring of production systems...  ...with cloud engineers in understanding new...  ...years of related SRE experience ~ Apply... 
    Software
    Full time
    Work experience placement
    Remote work

    Tenable

    Remote
    a month ago
  •  ...looking for a Site Reliability Engineer (SRE) to join a fast-growing...  ...model, strong observability, monitoring and reliable incident response...  ...who started in backend/software engineering and has moved into...  ...for key services and product-level metrics. Improve monitoring... 
    Software
    Full time
    Remote work

    GT

    Remote
    17 days ago
  •  ...writing, shipping, and running software effortless and efficient. We work closely with product-engineering to identify friction in the...  ...Support, and guide engineers on SRE related topics Partner...  ...~ Experience with monitoring / alerting (primarily with Prometheus... 
    Software
    Local area
    Remote work

    Kraken

    United States
    1 day ago
  • $85 - $95 per hour

     ...Enterprise Site Reliability Engineering (SRE) with our consumer...  ...to a full-time AVP-level position based on...  ...reliability a core product and platform capability....  ...requirements throughout the software development lifecycle...  ..., orchestration, monitoring, logging, distributed... 
    Software
    Hourly pay
    Permanent employment
    Full time
    Contract work
    Immediate start
    Remote work

    Genesis10

    Pittsburgh, PA
    5 days ago
  • $40 per hour

    A technology solutions provider is seeking a remote Junior SRE/DevOps Engineer. The ideal candidate should have foundational knowledge of Site Reliability Engineering (SRE) and Kubernetes. Responsibilities include gaining experience in a DevOps-driven environment. Applicants... 
    Long term contract
    Internship
    Remote work

    BayOne Solutions

    United States
    2 days ago
  •  ...portfolio of data management platforms and mobile offerings in...  ...Informatics platform. As a Software Engineer, you must possess world class...  ...Site Reliability Engineering (SRE) concepts and practices is a...  ...cloud is a plus. Experience monitoring infrastructure with monitoring... 
    Software
    Full time
    Temporary work
    Work experience placement
    Live in
    Remote work
    Flexible hours

    ResMed

    San Diego, CA
    4 days ago
  • $100k - $180k

     ...Site Reliability Engineer (SRE) - Remote Bright Vision...  ...technology consulting and software development company...  ...continually pushing the platform toward higher...  ...continually refine service-level objectives (SLOs), service...  ...comprehensive monitoring, logging, and tracing... 
    Software
    Full time
    H1b
    Local area
    Immediate start
    Remote work
    Visa sponsorship

    Bright Vision Technologies

    United States
    1 day ago
  • $180.5k - $236.91k

     .... We're hiring a Senior Software Engineer, Cloud Infrastructure / SRE to join our Engineering...  ...a full stack technology platform and a relentless focus on...  ...development of Service-Level Objectives (SLOs) for systems...  ...: Proficiency with monitoring using tools like Prometheus... 
    Software
    Remote job
    Full time
    Work at office

    Oscar Health

    Boston, MA
    1 day ago
  •  ...SRE Engineer – Very Critical SFO CA (100% Remote) End Client: Airbnb Exp : 6 to 8 years only, don’t get profiles beyond 8+ years...  ...Centres Operation preferably with volume above 5K TPS Monitoring & alerting setup experience with AWS, New Relic, DataDog, PagerDuty... 
    Remote work

    My3Tech Inc

    United States
    3 days ago
  • $307k - $427k

     ...velocity.Ensure that SRE principles (SLOs,...  ...into the shipped software so external...  ...incident response, monitoring, and debugging paradigms...  ...who act as Level 3 support or manage...  ...experienced software engineers of large-scale projects...  ...of Google platforms, we make Google's... 
    Software

    Google

    Sunnyvale, TX
    17 hours ago
  •  ...Platform Engineering Manager Sia Partners is looking for a...  ...pivotal in bridging high-level product vision with...  ...support AI, data, and software workloads Key...  ...Platform / DevOps / SRE engineering teams, ensuring...  ...management, cost monitoring, and cost optimization... 
    Software
    Remote work

    SIA

    United States
    5 days ago
  • $78k - $185k

     ...week.ABOUT THE TEAMThe Platform Engineering team is part of...  ...in all phases of the Software Development Lifecycle...  ...KubernetesProvide observability via monitoring tools such as Datadog...  ....e., DevOps, SysOps, SRE) experience with a...  ...Type:Full timeJob Level:ProfessionalPosted... 
    Software
    Temporary work
    Work at office
    Local area
    Remote work
    Worldwide
    3 days per week

    Morgan Stanley

    Seattle, WA
    1 day ago
  • $105k - $158k

    Platform Engineer - Healthcare Cloud PlatformSpecialist I - DevOps EngineeringWho We...  ...· Site Reliability Engineering (SRE) · Establish and monitor Service Level Indicators (SLIs), Service Level...  ...observability. · Automate patching, software deployment, inventory management... 
    Software
    Full time
    Temporary work
    Part time
    Work at office
    Local area
    Remote work
    Flexible hours

    UST Global

    Chicago, IL
    4 days ago
  • $139.2k - $235.2k

     ...intelligent orchestration platform for DevSecOps....  ..., more secure software faster. The same...  ...a Senior Platform Engineer on the Orbit team,...  ...and architecture-level features across GitLab...  ...the deployment, monitoring, and operations of...  ...intelligence, and SRE teams. What you... 
    Software
    Full time
    Remote work
    Flexible hours

    GitLab

    Remote
    3 days ago
  •  ...Gen. About the Role: Engine by Gen is a leader in...  ...premier embedded finance platform for enterprise...  ...years of experience in software engineering, platform engineering, or DevOps/SRE roles. ~ Strong systems...  ...Kubernetes, Terraform, CI/CD, monitoring). ~ Experience... 
    Software
    Full time
    Flexible hours

    Gen Digital

    New York, NY
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to SRE Monitoring Platform Software Engineer (Entry Level) [Remote]. Be the first to apply!