Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

SRE Monitoring Platform Software Engineer (Early Career / Temporary)

Full-time

Bitdeer Technologies Group

Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure.

Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to support the AI revolution. Bitdeer handles complex processes involved in computing such as equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial intelligence.

Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia.

To learn more, visit (

Position Overview  

Bitdeer is building an AI-operated GPU cloud — a global fleet of self-built and OEM-rented data centers running the world's most valuable compute, run by a platform that observes, protects, and operates the fleet. The SRE Platform team builds the monitoring and automation substrate that every other squad — storage, network, GPU, K8S, and L1 operators — depends on. Their signals become the system you help build.

As an entry-level Software Engineer on the SRE / Monitoring Platform team, you contribute to the NeoCloud SRE platform — the multi-region system that observes, protects, and operates a GPU rental fleet across self-built and OEM-rented data centers. You join a bounded context led by a senior engineer, take well-scoped components from design to production code that ships through GitOps + the CICD release pipeline, follows the Plugin Framework conventions, meets declared SLOs, and stays drift-free.

This is a build + learn role. You write code, write tests, and operate what you build under the guidance of a senior engineer. You participate in on-call as a shadow before taking primary. Within 12 months, you should be delivering components independently within your assigned area and growing toward owning a sub-context.

Key Responsibilities

Where you'll contribute (guided by a senior engineer)

  • Collection + Storage — help build collection-agent, metrics-store / logs-store / traces-store / profiles-store, enrichment-service, collection-monitor. Write ingestion, query, and storage-path code.
  • Alert + Correlation + SLO — contribute to alert-engine-framework, alert-correlation, slo-framework; implement and tune default alert rules.
  • Topology + Cluster-Health — contribute to topology-service, cluster-health-rollup, OSS-SRE-tool collection plugins for K8s / Slurm / Ray / Volcano / Kueue / KubeRay.
  • Remediation + Workflow + Jobs — help build remediation-actuator, orchestration / workflow components, inspection probes, job-scheduler.
  • Observability instrumentation — instrument services with metrics, logs, and traces via OpenTelemetry; build dashboards; write runbooks an on-call can follow.
  • Test discipline — write unit / integration / contract tests for everything you ship; participate in chaos and soak tests led by senior engineers.

Why this is a great first role

  • Greenfield with a well-defined vision. The Plugin Framework, GitOps pipeline, and SLO framework are decided; you build components inside them with a clear blueprint — not from a blank page.
  • You learn the full observability stack at production scale — ingest, query, storage — by building it, not just using it.
  • Mentorship-heavy. You work directly with senior and principal engineers who own the architecture; their expertise becomes your growth path.

Job Requirements

  • 0-2 years of software engineering experience (new graduates with strong projects or internships welcome).
  • Solid fundamentals in one programming language — Go (preferred), Python, Java, or Rust. You can write clean, tested, readable code and explain your design choices.
  • CS fundamentals — data structures, algorithms, concurrency, basic networking (TCP / and operating-system concepts (processes, threads, I/O). You can reason about correctness and performance.
  • Distributed systems basics — you understand the ideas behind idempotency, retries, back-pressure, caching, and eventual consistency, even if you haven't operated them at scale yet. Eagerness to go deep.
  • Monitoring / observability exposure — some hands-on with Prometheus, Grafana, Loki, or similar; can write a basic PromQL query and instrument a service. Eagerness to learn the ingest, query, and storage path of a real observability stack.
  • Familiarity with Linux and the shell; comfort reading system logs and using standard debugging tools.
  • Kubernetes basics — understand Pods, Services, Deployments; have run something on K8s (a project, lab, or internship).
  • Git + CI basics — branching, pull requests, and have used a CI pipeline (GitHub Actions, GitLab CI, or similar).
  • Test discipline — you write unit and integration tests as a habit, not an afterthought.
  • Communication — clear written and verbal English; can write a good PR description and ask good questions.
  • Curiosity and a learning mindset — the most important qualifier. You're excited to learn GPU / AI infrastructure, AIOps, distributed systems, and observability at production scale.

Nice-to-Haves

  • Internship or project in monitoring / observability, telemetry pipelines, or platform / SRE tooling.
  • Exposure to GPU / AI-infra — DCGM, InfiniBand / RoCE, Kubernetes GPU Operator, Slurm / Ray. Interest counts more than depth.
  • Exposure to AIOps / ML-adjacent tooling (anomaly detection, alert correlation).
  • Contributions to open-source observability or cloud-native projects.

--------------------------------------------------------------------

Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, color, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union.

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the SRE Monitoring Platform Software Engineer (Early Career / Temporary) in San Jose, CA vacancy
  • $70 - $80 per hour

     ...seeking a highly skilled Software Development Engineer to design, develop, and scale...  ..., data pipelines, and platform extensions across Gainsight...  ...governance, testing, and monitoring.• Collaboration: Partner with...  ...available for this temporary role may include the following... 
    Temporary work
    Software
    Contract work

    TEKsystems

    San Jose, CA
    1 day ago
  • $101k - $161k

     ...artificial intelligence, and software-defined networking to...  ...awards, such as Best Engineering Team, Best Company for...  ...(CVaaS) global SRE team. SREs at Arista combine...  ...GCP (Google Cloud Platform) and GKE (Google Kubernetes...  ...microservices stack, monitoring infrastructure, and... 
    Software

    Arista Networks

    Santa Clara, CA
    1 day ago
  • $96.8k - $306.4k

    Drives cross-group platform initiatives (e.g.,...  ...partnership with SRE and security.Only...  ...job remains posted.Career Level - IC5Key ResponsibilitiesPlatform...  ...Software Development:Set...  ...lifecycle; coaches engineers across teams or...  ...software error logging, monitoring, and observability... 
    Temporary work
    Software
    Flexible hours
    Shift work

    Oracle Corporation

    Santa Clara, CA
    4 days ago
  • $190.9k - $334.1k

     ...DescriptionIt all started when engineer Fred Luddy wrote code...  .... Our ServiceNow AI platform brings together any AI...  ...Engineer - SRE & AIOps to drive infrastructure...  ...stack, including monitoring platforms, incident management...  ...industry.12+ years in software engineering or... 
    Software
    Work at office
    Immediate start
    Remote work
    Flexible hours

    ServiceNow

    Santa Clara, CA
    2 days ago
  •  ...responsibilities of a Technical Support Engineer within a SaaS (Software as a Service) environment...  ...Reliability Engineering (SRE).The ideal candidate has a...  ...AI Security Public SaaS platform, operating AI inference...  ...:Proactive Monitoring & Uptime AssuranceMonitor... 
    Software
    Full time
    Local area

    F5 Networks

    San Jose, CA
    3 days ago
  • $65k - $85k

     ...DevOps and AI Security Engineer to help support,...  ...master’s graduate, or an early-career professional with hands...  ...deployment pipelines, monitor security posture, document...  ...customer-facing web platforms. Deployment Support...  ..., and repeatable software delivery. Documentation... 
    Temporary work
    Software
    Entry level
    Permanent employment
    Full time
    Internship
    Remote work

    Renesas Electronics

    San Jose, CA
    8 days ago
  • $98.9k - $228.7k

     ...evolving our observability platform, working across Kubernetes,...  ...This is a hands-on, on-call engineering role: you will own the systems...  ...or improves observability or monitoring platforms at production...  ...your skills and advance your career in a collaborative, growth-focused... 
    Permanent employment
    Full time
    Work at office
    Remote work

    Zoom

    San Jose, CA
    4 days ago
  • $100k - $170k

     ...are seeking a  DevOps Engineer who is eager to have...  ...testing, and release platforms, taking us from square...  ...scalable, highly-available software systems in AWS...  ...Implement comprehensive monitoring, alerting, and observability...  ...of experience in SRE, DevOps, or Platform Engineering... 
    Software
    Full time
    Work at office
    Immediate start
    Visa sponsorship
    Night shift

    E-Space

    Saratoga, CA
    4 days ago
  •  ...high-performance SRE function to support...  ...by the Wafer-Scale Engine (WSE). This team...  ...partner with our early-career SRE sub-team, who...  ...and mentor them as platform engineers.You will...  ...delivering and running software reliably and at...  ...accuracy SLOs, or drift monitoring.Prior work on... 
    Software
    Entry level
    Shift work

    Cerebras Systems

    Sunnyvale, CA
    2 days ago
  • $148.75k - $361k

     ...TVRoku is the #1 TV streaming platform in the U.S., Canada, and...  ...talented and experienced Senior Software Engineer, MLOps/DevOps, to join the...  ...strong background in DevOps/SRE practices, cloud infrastructure...  ...evaluation, deployment, and monitoring — on top of a modern, cloud-... 
    Software
    Work at office
    Local area
    Remote work
    Monday to Thursday
    Flexible hours

    Roku

    San Jose, CA
    3 days ago
  • $98.9k - $228.7k

     ...evolving our observability platform, working across Kubernetes,...  ...This is a hands-on, on-call engineering role: you will own the systems...  ...or improves observability or monitoring platforms at production...  ...your skills and advance your career in a collaborative, growth-focused... 
    Permanent employment
    Full time
    Work at office
    Remote work

    Zoom

    San Jose, CA
    2 days ago
  •  ...are seeking an experienced IT SRE Team Lead to build and run...  ...right candidate will bring a software engineering mindset to IT operations,...  ...and SaaS integrations with monitoring, alerting, and on-call workflows...  ...experience with identity platforms (Okta, Entra), endpoint management... 
    Software

    Cerebras Systems

    Sunnyvale, CA
    3 days ago
  •  ...bounded contexts of the NeoCloud SRE platform — the multi-region substrate...  ...enrichment-service, collection-monitor. Alert, Correlation & SLO: alert-engine-framework, alert-correlation, slo...  ...-platform. Qualifications Software Engineering Experience: 7+ years... 
    Software
    Full time
    Contract work
    Local area

    Bitdeer

    San Jose, CA
    22 days ago
  • $75k - $95k

     ...collaboration tools. Monitor system health, performance...  ...hardware, software, network, and connectivity...  ...VMware virtualization platforms. Hands‑on experience with...  ...supporting construction, engineering, architecture, or real...  ...Experience deploying temporary or field office networks... 
    Temporary work
    Software
    Full time
    Work at office

    SWENSON

    San Jose, CA
    5 days ago
  • $153k - $204k

     ...pioneers, CoreWeave delivers a platform of technology, tools, and...  ...talented and experienced Senior Software Engineer to join our Network Datapath...  ...services. As a Datapath Monitoring/Observability Engineer, you...  ...is needed, please contact: careers@coreweave.com. Export Control... 
    Temporary work
    Software
    Permanent employment
    Full time
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    Sunnyvale, CA
    4 days ago
  •  ...preparing faculty appointments, planning and monitoring faculty workload and maintaining...  ...Resource Analyst with faculty recruitment, temporary faculty evaluations and RTP as needed Independently...  ..., Skills & Abilities Knowledge of software applications: Docusign, word processing,... 
    Temporary work
    Software
    Work experience placement
    Work at office

    SupportFinity™

    San Jose, CA
    5 days ago
  • $106k - $159k

     ...transforms your career—and the lives of...  ...searching for a Lead AI Engineer to architect,...  ...Coordinate with platform, security, data,...  ...Python and modern software engineering practices...  ...and Security· AI Monitoring and Observability...  ....Full-time temporary employees receive... 
    Temporary work
    Software
    Full time
    Part time
    Work at office
    Local area
    Flexible hours

    UST Global

    Santa Clara, CA
    1 day ago
  • $248k - $391k

     ...high-performance engineering organization that...  ...production-grade software systems at global...  ...automated, AI-driven platforms that scale with NVIDIA...  ...calibration, and career development that...  ...: Infrastructure, SRE, DevOps, or Production...  ...~ Experience with monitoring and observability... 
    Software
    Full time

    NVIDIA

    Santa Clara, CA
    4 days ago
  •  ...Position- SRE Engineer Duration-Contract Location- San Jose, C JD...  ...infrastructure deployments, validations and monitoring to improve operational tasks • Experience...  ...CICD pipelines (preferred) • Cloud platform knowledge (specifically AWS) is required... 
    Contract work
    Immediate start

    Syntricate Technologies

    San Jose, CA
    2 days ago
  • $176.6k - $265k

     ...The AI Inference Engineer plays a critical...  ...infrastructure, and monitoring system performance...  .... ~ Ensure software solutions are optimized...  ..., and cloud platforms such as AWS , GCP...  ...Background in MLOps or SRE roles focused on...  ...~ Advancing your career in a fast-paced,... 
    Software
    Full time
    Local area
    Immediate start

    F5

    San Jose, CA
    5 days ago
  • $178k - $321k

     ...function has an early but working AI-native...  ...: a multi-agent platform (Hive Mind),...  ...is a two-person engineering team: you deploy,...  ...workflows. Working software wins arguments; migrate...  ...) and continuous-monitoring data foundation:...  ...or external careers site.Notice:All official... 
    Software

    OKX

    San Jose, CA
    10 hours ago
  • $170k - $225k

     ...United StatesProducts - Engineering /Fulltime /HybridOver...  ...our Cloud Networking platform. In this role, you will...  ...Management, DevOps/SRE, and Customer Support...  ...experience).10+ years of software quality engineering experience...  ...network management/monitoring platforms or IoT/edge... 
    Software
    Full time
    Remote work
    Flexible hours
    Shift work

    Extreme Networks, Inc.

    San Jose, CA
    2 days ago
  • $150k - $200k

     ...everything possible.The Backend & AI Platform Engineer is responsible for creating the backend...  ...position reports to the Senior Manager, Software Engineering - Automation and is part...  ...that coordinate planning, execution, monitoring, and recovery across heterogeneous AI... 
    Software
    Full time
    Part time

    Danaher Corporation

    San Jose, CA
    3 days ago
  •  ...internal tooling and platform capabilities to improve...  ...partnering closely with engineering, data platform,...  ...engineering, data platform, SRE, security and business...  ...(provisioning, CI/CD, monitoring, cost optimization,...  ...engineers (experience as a software engineer or technical... 
    Software

    CyberCoders

    San Jose, CA
    3 days ago
  • $280k - $380k

     ...is the #1 TV streaming platform in the U.S., Canada,...  ...About the TeamOur DevOps/SRE team runs an active-...  ...reliability and automation, engineering systems that perform...  ...Engineering) Senior Software Engineer to join our...  ...bottlenecks through detailed monitoring and profiling and... 
    Software
    Work at office
    Local area
    Remote work
    Monday to Thursday
    Flexible hours

    Roku

    San Jose, CA
    3 days ago
  • $157.2k - $254.1k

     ...meaningful work of your career alongside people who...  ...comprehensive AI security platform. Organizations are...  ...Machine Learning Inference Engineer, you will serve as a...  ...practices in MLOps, software engineering, and...  ...model deployment, robust monitoring, and operational excellence... 
    Software
    Full time
    Work at office

    Palo Alto Networks

    Santa Clara, CA
    1 day ago
  •  ...t mean fewer humans — it means humans focus on judgment calls the platform can't yet make, and every judgment call trains the platform to do it next time. In this L1 role you cover front-line monitoring and incident response for NeoCloud's US GPU DCs during the 8AM–8PM... 
    Full time
    Local area
    Shift work
    Night shift

    Bitdeer Technologies Group

    San Jose, CA
    17 days ago
  • $70k - $90k

     ...network intelligence platform. As a Client Program...  ...resolution while monitoring program performance through...  ...well-suited for an early-career analyst who enjoys...  ...Provide first-line software support to internal and...  ...degree in business, engineering, analytics, operations... 
    Temporary work
    Software
    Entry level
    Full time
    Internship
    Summer holiday
    Work at office
    Remote work
    Worldwide
    Flexible hours
    Shift work

    Picarro

    Santa Clara, CA
    3 days ago
  • $144k - $230k

     ...are seeking a passionate AI Tools Engineer to join the Site Reliability Engineering (SRE) Data Team. We encourage...  ...Frameworks, recommending the best platforms, toolsets, and architectural approaches...  ...pipelines.Experience applying monitoring and visualization tools, such as... 
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    3 days ago
  •  ...most meaningful work of your career alongside people who are...  ...visionary Senior Principal Engineer/Architect to serve as the technical...  ...authority for our global SRE and Platform Engineering initiatives...  ...a deep empathy for internal software developers as your primary customers... 
    Software
    Full time
    Work at office
    Visa sponsorship
    Work visa
    Flexible hours

    Palo Alto Networks, Inc.

    Santa Clara, CA
    10 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to SRE Monitoring Platform Software Engineer (Early Career / Temporary). Be the first to apply!