Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

SRE Monitoring Platform Software Engineer (Entry Level)

Full-time

Bitdeer

Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure.

Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to support the AI revolution. Bitdeer handles complex processes involved in computing such as equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial intelligence.

Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia.

To learn more, visit [ Position Overview

Bitdeer is building an AI-operated GPU cloud — a global fleet of self-built and OEM-rented data centers running the world's most valuable compute, run by a platform that observes, protects, and operates the fleet. The SRE Platform team builds the monitoring and automation substrate that every other squad — storage, network, GPU, K8S, and L1 operators — depends on. Their signals become the system you help build.

As an entry-level Software Engineer on the SRE / Monitoring Platform team, you contribute to the NeoCloud SRE platform — the multi-region system that observes, protects, and operates a GPU rental fleet across self-built and OEM-rented data centers. You join a bounded context led by a senior engineer, take well-scoped components from design to production code that ships through GitOps + the CICD release pipeline, follows the Plugin Framework conventions, meets declared SLOs, and stays drift-free.

This is a build + learn role. You write code, write tests, and operate what you build under the guidance of a senior engineer. You participate in on-call as a shadow before taking primary. Within 12 months, you should be delivering components independently within your assigned area and growing toward owning a sub-context.

Key Responsibilities

Where you'll contribute (guided by a senior engineer)

  • Collection + Storage — help build collection-agent, metrics-store / logs-store / traces-store / profiles-store, enrichment-service, collection-monitor. Write ingestion, query, and storage-path code.

  • Alert + Correlation + SLO — contribute to alert-engine-framework, alert-correlation, slo-framework; implement and tune default alert rules.

  • Topology + Cluster-Health — contribute to topology-service, cluster-health-rollup, OSS-SRE-tool collection plugins for K8s / Slurm / Ray / Volcano / Kueue / KubeRay.

  • Remediation + Workflow + Jobs — help build remediation-actuator, orchestration / workflow components, inspection probes, job-scheduler.

  • Observability instrumentation — instrument services with metrics, logs, and traces via OpenTelemetry; build dashboards; write runbooks an on-call can follow.

  • Test discipline — write unit / integration / contract tests for everything you ship; participate in chaos and soak tests led by senior engineers.

Why this is a great first role
  • Greenfield with a well-defined vision. The Plugin Framework, GitOps pipeline, and SLO framework are decided; you build components inside them with a clear blueprint — not from a blank page.

  • You learn the full observability stack at production scale — ingest, query, storage — by building it, not just using it.

  • Mentorship-heavy. You work directly with senior and principal engineers who own the architecture; their expertise becomes your growth path.

Job Requirements
  • 0-2 years of software engineering experience (new graduates with strong projects or internships welcome).

  • Solid fundamentals in one programming language — Go (preferred), Python, Java, or Rust. You can write clean, tested, readable code and explain your design choices.

  • CS fundamentals — data structures, algorithms, concurrency, basic networking (TCP / and operating-system concepts (processes, threads, I/O). You can reason about correctness and performance.

  • Distributed systems basics — you understand the ideas behind idempotency, retries, back-pressure, caching, and eventual consistency, even if you haven't operated them at scale yet. Eagerness to go deep.

  • Monitoring / observability exposure — some hands-on with Prometheus, Grafana, Loki, or similar; can write a basic PromQL query and instrument a service. Eagerness to learn the ingest, query, and storage path of a real observability stack.

  • Familiarity with Linux and the shell; comfort reading system logs and using standard debugging tools.

  • Kubernetes basics — understand Pods, Services, Deployments; have run something on K8s (a project, lab, or internship).

  • Git + CI basics — branching, pull requests, and have used a CI pipeline (GitHub Actions, GitLab CI, or similar).

  • Test discipline — you write unit and integration tests as a habit, not an afterthought.

  • Communication — clear written and verbal English; can write a good PR description and ask good questions.

  • Curiosity and a learning mindset — the most important qualifier. You're excited to learn GPU / AI infrastructure, AIOps, distributed systems, and observability at production scale.

Nice-to-Haves
  • Internship or project in monitoring / observability, telemetry pipelines, or platform / SRE tooling.

  • Exposure to GPU / AI-infra — DCGM, InfiniBand / RoCE, Kubernetes GPU Operator, Slurm / Ray. Interest counts more than depth.

  • Exposure to AIOps / ML-adjacent tooling (anomaly detection, alert correlation).

  • Contributions to open-source observability or cloud-native projects.

-------------------------------------------------------------------- Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, color, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union.
Vacancy posted 12 days ago
Similar jobs that could be interesting for youBased on the SRE Monitoring Platform Software Engineer (Entry Level) in San Jose, CA vacancy
  • $101k - $161k

     ...intelligence, and software-defined networking...  ...awards, such as Best Engineering Team, Best Company...  ...(CVaaS) global SRE team. SREs at Arista...  ...GCP (Google Cloud Platform) and GKE (Google...  ...microservices stack, monitoring infrastructure,...  ...EngineeringExperience level: Mid-Senior... 
    Software

    Arista Networks

    Santa Clara, CA
    5 days ago
  •  ...DescriptionThe AI Inference Engineer plays a critical...  ..., and monitoring system performance...  ...orchestration. Ensure software solutions are optimized...  ...against service-level agreements (SLAs)....  ..., and cloud platforms such as AWS, GCP,...  ...Background in MLOps or SRE roles focused on... 
    Software
    Full time
    Local area
    Immediate start

    F5 Networks

    San Jose, CA
    3 days ago
  • $176.1k - $308.2k

     ...DescriptionIt all started when engineer Fred Luddy wrote...  ...Our ServiceNow AI platform brings together any...  ...AI Engineer - SRE/DevOps you will: Build software solutions to solve...  ..., observability, monitoring, high availability,...  ...qualifications, skill level, competencies, and... 
    Software
    Work experience placement
    Work at office
    Immediate start
    Remote work
    Flexible hours

    ServiceNow

    Santa Clara, CA
    15 hours ago
  •  ...related to our SaaS platform, including data...  ...inconsistencies, software bugs, and integration...  ...Reliability & Monitoring Maintain,...  ...reliability. Optimize SRE workflows with AI...  ...functionally with Engineering to define and...  .../C++ or other low-level systems languages.... 
    Software
    Work at office
    Local area
    3 days per week

    ThoughtSpot

    Mountain View, CA
    5 days ago
  • $155k - $230k

     ...unified data security platform addresses...  ...Infrastructure & Platform Engineer to help architect,...  ..., CI/CD, and software development. You will...  ..., testing, monitoring, and operational workflows...  ...engineering, SRE, DevOps, or related...  ...who would require entry into the H-1B lottery... 
    Software
    Temporary work
    H1b
    Worldwide

    Fortanix

    Santa Clara, CA
    6 days ago
  • $96.8k - $306.4k

    Drives cross-group platform initiatives (e.g.,...  ...partnership with SRE and security.Only...  ...remains posted.Career Level - IC5Key ResponsibilitiesPlatform Software Development:Set...  ...lifecycle; coaches engineers across teams or...  ...software error logging, monitoring, and observability... 
    Software
    Temporary work
    Flexible hours
    Shift work

    Oracle Corporation

    Santa Clara, CA
    4 days ago
  • $148.75k - $361k

     ...TVRoku is the #1 TV streaming platform in the U.S., Canada, and...  ...talented and experienced Senior Software Engineer, MLOps/DevOps, to join the...  ...strong background in DevOps/SRE practices, cloud infrastructure...  ...evaluation, deployment, and monitoring — on top of a modern, cloud-... 
    Software
    Work at office
    Local area
    Remote work
    Monday to Thursday
    Flexible hours

    Roku

    San Jose, CA
    2 days ago
  •  ...are seeking an experienced IT SRE Team Lead to build and run...  ...right candidate will bring a software engineering mindset to IT operations,...  ...and SaaS integrations with monitoring, alerting, and on-call workflows...  ...experience with identity platforms (Okta, Entra), endpoint management... 
    Software

    Cerebras Systems

    Sunnyvale, CA
    2 days ago
  •  ...NVIDIA Corporation in Santa Clara, CA is seeking an Engineering Manager to lead a team of SRE, platform, and software engineers responsible for the AI Platform Runtime and related production services. You will set the team's strategy, roadmaps, and drive highly available... 
    Software

    Jobleads-US

    Santa Clara, CA
    6 days ago
  •  ...Senior/Lead SRE Platform Services Engineer Technical Leader (Remote) Experience : 12 - 15 years of experience...  ...or platform programs. ~ Strong software engineering skills in Python, Go,...  ...ownership. Mentor senior and mid-level engineers; raise the quality of designs... 
    Software
    Remote job
    Permanent employment
    Full time

    Jobleads-US

    San Jose, CA
    3 days ago
  •  ...and detail-oriented Site Reliability Engineer II (SRE II) to join our 24/7 Operations team in...  ...provide round-the-clock, eyes-on-glass monitoring, proactive incident response,...  ...primary responsibility is maintaining platform availability, hardware reliability, and... 
    Full time
    Local area
    Immediate start
    Shift work
    Night shift
    Afternoon shift
    Weekday work

    F5 Networks

    San Jose, CA
    15 hours ago
  •  ...Position- SRE Engineer Duration-Contract Location- San Jose, C JD...  ...infrastructure deployments, validations and monitoring to improve operational tasks • Experience...  ...CICD pipelines (preferred) • Cloud platform knowledge (specifically AWS) is required... 
    Contract work
    Immediate start

    Syntricate Technologies

    San Jose, CA
    16 hours ago
  •  ...Synopsys is the leader in engineering solutions from silicon...  ...You are a strong platform engineer with a passion...  ...comfortable working across software, systems, and...  ...integrations that enhance monitoring, alerting, incident management...  ...with infrastructure, SRE, platform engineering,... 
    Software

    Synopsys

    Sunnyvale, CA
    2 days ago
  • $176.1k - $308.2k

     ...data.Veza's Access Graph platform maps an organization's...  ...privileged access monitoring, non-human identity security...  ..., and AI agents. For engineers joining Veza today,...  ...a passionate Staff Software Engineer to join Veza'...  ...independently driving workstream-level design reviewsSet the... 
    Software
    Work at office
    Remote work
    Flexible hours

    ServiceNow

    Santa Clara, CA
    5 days ago
  • $80 per hour

     .../hr Summary As Staff Software Engineer in the team, you will...  ...and performance of our platforms. Responsibilities...  ...resolving bottlenecks via monitoring, logging, and metrics...  ...to mentor and up-level team members. Proven...  ...Contract Seniority level Entry level #J-18808-Ljbffr... 
    Software
    Entry level
    Contract work

    Aditi Consulting

    San Jose, CA
    3 days ago
  • $187k - $270.7k

    ## AIOPs Observability/SRE LeadApply: San Jose, California, United...  ...the needs of the business, engineering organizations, laboratories,...  ...engineers who drive platform stability, observability, automation...  ...secure, and high-performance monitoring architectures that support both... 
    Full time
    Work at office
    Flexible hours

    Jobleads-US

    San Jose, CA
    3 days ago
  • $132.3k - $152.3k

     ...looking for a Staff Engineer DevOps to join our Platform Engineering & Operations...  ...management, account‑level security posture,...  ...on top of existing monitoring stacks. Help support...  ...available, enterprise‑grade software and platform...  ...Reliability Engineering (SRE), with hands‑on... 
    Software
    Full time
    Work at office
    Local area

    Lendistry

    Santa Clara, CA
    5 days ago
  • $179k - $219k

     ...are looking for a Staff Software Engineer to join the FortiSOC(SIEM) Cloud Platform team in Santa Clara, US...  ...frameworks. Enhance monitoring, observability,...  ...Partner with product, QA, SRE, and security teams to...  ...market, job type, and job level. Exact salary offers will... 
    Software
    Full time
    Local area
    Worldwide

    Jobleads-US

    Santa Clara, CA
    4 days ago
  • $140k - $160k

     ...re looking for a Senior Software Engineer to join the FortiSOC (SIEM) Cloud Platform team in Santa Clara, CA...  ...provisions, upgrades, monitors and backs up their deployment...  ...Work with product, QA, SRE and security teams to...  ..., job type, and job level. Exact salary offers... 
    Software
    Full time
    Local area

    Jobleads-US

    Santa Clara, CA
    5 days ago
  • $170k - $225k

     ...United StatesProducts - Engineering /Fulltime /HybridOver...  ...our Cloud Networking platform. In this role, you will...  ...Management, DevOps/SRE, and Customer Support...  ...experience).10+ years of software quality engineering experience...  ...network management/monitoring platforms or IoT/edge... 
    Software
    Full time
    Shift work

    Extreme Networks, Inc.

    San Jose, CA
    2 days ago
  • $178k - $321k

     ...capability: a multi-agent platform (Hive Mind), agentic...  ...This is a two-person engineering team: you deploy,...  ...grade here means service levels sized for an internal...  ...cycle workflows. Working software wins arguments;...  ...(CAAT) and continuous-monitoring data foundation: analytics... 
    Software

    OKX

    San Jose, CA
    4 days ago
  • $208k - $333.5k

    Site Reliability Engineering (SRE) at NVIDIA is an engineering field focused...  ...availability. It combines software and systems engineering...  ...and operating resilient AI platform capabilities at enterprise scale...  ...operational practices using service-level indicators, service-level... 
    Software
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  •  ...Success & Support, RevOps, and Platform Engineering. We are embedding AI into...  ...and reliable through monitoring, alerting, and data integrity...  ...What You'll Bring 8+ years of software engineering experience,...  ...enterprise scale; Architect-level certification LLM security:... 
    Software
    Full time
    Local area

    F5 Networks

    San Jose, CA
    1 day ago
  • $92.5k - $209.5k

     ...capabilities. We are seeking a Software Developer 3 to help build and operate AI and cloud platform capabilities supporting a...  ...you will work with experienced engineers, product management, and customer...  ...Reliability & Operations Add monitoring, logging, and telemetry to... 
    Software
    Temporary work
    Flexible hours

    Oracle

    Santa Clara, CA
    3 days ago
  •  ...RoleAs a Senior or Staff Software Engineer in Lambda’s Cloud...  ...cloud. Our teams own platform capabilities across compute...  ..., security, and SRE teams to resolve dependencies...  ...mentorship; at Staff level, lead cross-team...  ...testing, staged rollout, monitoring, incident response, and... 
    Software
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    1 day ago
  •  ...high-performance SRE function to support...  ...by the Wafer-Scale Engine (WSE). This team...  ...and mentor them as platform engineers.You will...  ...delivering and running software reliably and at...  ....Mentor mid-level SREs, support critical...  ...accuracy SLOs, or drift monitoring.Prior work on... 
    Software
    Entry level
    Shift work

    Cerebras Systems

    Sunnyvale, CA
    2 days ago
  • $65 - $85 per hour

     ...bring a Site Reliability Engineer (Contract) to the team...  ...with NVIDIA Software groups such as Graphics...  ...processor hardware. The SRE will build the next generation...  ...Fleet monitoring and recovery of assets...  ...collaboration with the platform engineering team. Experience... 
    Software
    Full time
    Contract work
    Worldwide

    Sustainable Talent

    Santa Clara, CA
    2 days ago
  •  ...internal tooling and platform capabilities to improve...  ...partnering closely with engineering, data platform,...  ...engineering, data platform, SRE, security and business...  ...(provisioning, CI/CD, monitoring, cost optimization,...  ...engineers (experience as a software engineer or technical... 
    Software

    CyberCoders

    San Jose, CA
    3 days ago
  • $280k - $380k

     ...is the #1 TV streaming platform in the U.S., Canada,...  ...About the TeamOur DevOps/SRE team runs an active-...  ...reliability and automation, engineering systems that perform...  ...Engineering) Senior Software Engineer to join our...  ...bottlenecks through detailed monitoring and profiling and... 
    Software
    Work at office
    Local area
    Remote work
    Monday to Thursday
    Flexible hours

    Roku

    San Jose, CA
    2 days ago
  •  ...United StatesJob Title: Senior Kubernetes Platform Engineer / ArchitectLocation: Milpitas,...  ...development teams, security teams, and SRE teams to deliver a robust cloud-native...  ...for development teams.Automate platform monitoring, compliance, and operational processes.... 

    Apptad

    Milpitas, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to SRE Monitoring Platform Software Engineer (Entry Level). Be the first to apply!