Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Sr. SRE Platform Software Engineer

Full-time

Bitdeer

About Bitdeer Technologies Group

Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure.

Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to support the AI revolution. Bitdeer handles complex processes involved in computing such as equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial intelligence.

Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia.

To learn more, visit [ Position Overview

Build and operate one or more bounded contexts of the NeoCloud SRE platform — the multi-region substrate that observes, protects, and operates a GPU rental fleet across self-built and OEM-rented data centers. You take an architect-approved design and turn it into production code that ships through GitOps + the CICD release pipeline, ride the Plugin Framework conventions, meet declared SLOs, and stay drift-free.

This is the build + run role. You don't only ship code; you ship a service that other squads, cloud-service teams, and tenants depend on. You take the on-call pager for what you build.

Key Responsibilities _ You will own 1-2 of these:_
  • Collection & Storage: collection-agent, customer-sdk-gateway, metrics-store, logs-store, traces-store, profiles-store, analytics-lake, enrichment-service, collection-monitor.
  • Alert, Correlation & SLO: alert-engine-framework, alert-correlation, slo-framework, default M-series alert rules.
  • Topology, Cluster-Health & Cluster Platform Services: topology-service, cluster-health-rollup, OSS-SRE-tool collection plugins for K8s, Slurm, Ray, Volcano, Kueue, and KubeRay.
  • Fault-Prediction: prediction-engine-framework and built-in predictors (GPU, Link, Disk, XPA, Straggler, SDC, Stranded GPU).
  • Remediation, Workflow, Inspection & Jobs: remediation-actuator, orchestration-substrate (workflow engine), inspection-orchestrator, job-scheduler, NCCL-baseline inspection probe.
  • Hardware Lifecycle & DC Ops: hardware-lifecycle, dc-operations, boot-provisioning, rolling-upgrade, bare-metal-bmc-service, auto-discovery, ZTP D0–D5 pipeline, IPMI bare-metal management.
  • Identity, Secrets, Tenant-Config & CMDB: iam-service, secrets-service, tenant-sre-config, cmdb-cache, schema registry.
  • Customer-Bridge, Ticketing & SRE Platform Portal: customer-bridge, customer-ticketing, sre-operation-system, Customer Console BFF, SRE Console BFF.
  • Backup, DR & Meta-Monitor: backup-orchestrator, meta-monitor, external-watcher integration (Datadog or equivalent).
  • CI/CD, GitOps, Plugin Framework & SRE Image Registry: cicd-pipeline, gitops-sync, plugin-registry, sre-image-registry.
  • Self-Improving Agent: agent-control-plane, agent-discovery, agent-codegen, agent-sandbox, per-Region LLM gateway.
  • Global SRE Management: maintenance-window-orchestrator, change-management, capacity-planner, cost-optimizer, gpu-efficiency-dashboard, network-stability-dashboard, patching-orchestrator, artifact-management, compat-matrix-service, security-platform.
Qualifications
  • Software Engineering Experience: 7+ years of production software engineering experience, including 2 or more years operating what you built (real on-call experience, not just shipping code).
  • Programming Languages: Production-depth mastery of at least one systems-grade language—Go (preferred), Rust, or Java. Proficiency in Python for tooling and SDK work.
  • Distributed Systems Fundamentals: Strong grasp of at-least-once vs. exactly-once trade-offs, idempotency, back-pressure, leader election, consistent hashing, gossip, and fan-out. Ability to evaluate CRDT vs. Raft vs. Paxos and select the right tool for the job.
  • Multi-Region Observability Stack: Experience at production scale with Prometheus, VictoriaMetrics, Mimir, Thanos, Loki, Elasticsearch, Tempo, Jaeger, or OpenTelemetry. Must have built or substantively contributed to the ingest, query, or storage paths of these systems.
  • GitOps & CI/CD: Hands-on experience with Argo, Flux, Helm, Kustomize, Cosign signing, signed-bundle promotion, and blast-radius-aware rollouts.
  • Kubernetes Operator Pattern: Proven experience writing a controller or CRD handling real production traffic, with a deep understanding of watch-cache mechanics, leader election, and reconcile loops.
  • mTLS & Secrets Management: Experience executing end-to-end mTLS bootstrap with certificate rotation. Hands-on experience with HashiCorp Vault or cloud KMS (AWS KMS / GCP KMS).
  • SQL & Time-Series Data: Ability to read a Prometheus query plan, build a recording-rule strategy, and write SQL that joins per-tenant telemetry against analytics-lake tables.
  • Testing Discipline: Rigorous approach to unit, integration, contract, chaos, and soak testing. Experience writing and maintaining your own comprehensive tests.
  • Technical Writing Fluency: Ability to author clear design docs that align with existing platform architecture, create runbooks optimized for 3 AM on-call responses, and write intent-driven PR descriptions.

**Preferred Qualifications (GPU / AI-Infra Context)

** Experience in at least one of the following areas is a strong plus:
  • NVIDIA Internals: Deep understanding of DCGM and NVIDIA driver internals, including XID semantics and MIG / vGPU partitioning.
  • Networking & Fabrics: Experience with InfiniBand or RoCE fabrics, including subnet managers, partitioning, optical health, and NCCL collective tracing.
  • HPC Storage: Experience managing Lustre, NetApp, Pure, DDN, VAST, or NVMe-oF under multi-tenant loads.
  • Hardware Management: Hands-on experience with BMC, IPMI, and Redfish at OEM scale (Supermicro, Dell, HPE, Lenovo).
  • Cluster Platform Internals: Familiarity with Kubernetes GPU Operator, Slurm controller, or Ray GCS.
  • BS/MS in Computer Science or similar
  • Hyperscale or NeoCloud experience
-------------------------------------------------------------------- Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, color, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union.
Vacancy posted 12 days ago
Similar jobs that could be interesting for youBased on the Sr. SRE Platform Software Engineer in San Jose, CA vacancy
  •  ...looking for a visionary Senior Principal Engineer/Architect to serve as the technical authority for our global SRE and Platform Engineering initiatives across the US and India...  ..., coupled with a deep empathy for internal software developers as your primary customers. The... 
    Senior
    Software
    Full time
    Work at office
    Visa sponsorship
    Work visa
    Flexible hours

    Palo Alto Networks, Inc.

    Santa Clara, CA
    18 hours ago
  • $101k - $161k

     ...artificial intelligence, and software-defined networking to provide...  ...prestigious awards, such as Best Engineering Team, Best Company for...  ...-as-a-Service (CVaaS) global SRE team. SREs at Arista combine...  ...Familiarity with GCP (Google Cloud Platform) and GKE (Google Kubernetes... 
    Senior
    Software

    Arista Networks

    Santa Clara, CA
    1 day ago
  •  ...Together, we advance your career. THE ROLEWe are hiring AI / ML Platform Engineers to build the platform layer that makes AI-for-engineering...  ...comfortable working across ML, infrastructure, and hardware/software tooling, and you can partner with research teams without... 
    Senior
    Software

    AMD

    Santa Clara, CA
    4 days ago
  •  ...DevOps EngineerAs a DevOps engineer, the candidate will be responsible for the development...  ...automated deployments (cloud and on-premise), software-defined networking, resilient storage,...  ...experience. 3+ years of experience in SRE/DevOps. Excellent Python, bash, and scripting... 
    Senior
    Software

    Professional Recruiters

    Santa Clara, CA
    4 days ago
  • $148.75k - $361k

     ...watches TVRoku is the #1 TV streaming platform in the U.S., Canada, and Mexico, and we...  ...seeking a talented and experienced Senior Software Engineer, MLOps/DevOps, to join the Advertising...  ...has a strong background in DevOps/SRE practices, cloud infrastructure management... 
    Senior
    Software
    Work at office
    Local area
    Remote work
    Monday to Thursday
    Flexible hours

    Roku

    San Jose, CA
    3 days ago
  • $200k - $322k

     ...lasting impact on the world.Ready to build the platforms that make AI at scale possible? NVIDIA is seeking a Senior Staff Platform Engineer to architect, build, and scale...  ...distributing models, datasets, containers, software artifacts, or other large objects across globally... 
    Senior
    Software
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  •  ...solutions spanning all phases of the Software Development Lifecycle (SDLC)...  ...-level issues Guides junior engineers Operates with little day-to-...  ...PagerDuty. Collaborate with SRE and cloud teams to design and...  ...solutions. Solid backend or platform engineering experience in... 
    Senior
    Software
    Full time
    Work experience placement
    Immediate start

    Paypal

    San Jose, CA
    18 hours ago
  • $176k - $276k

     ...to join us!NVIDIA invites applications for a Senior DevOps Platform Engineer skilled in Platform and Release Engineering to join the Metropolis...  ...Jenkins, GitHub/GitLab Actions and Runners for Metropolis software products.Develop and manage Kubernetes-based platform... 
    Senior
    Software
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  •  ...and efficiency, and real-time platform management. Our technology...  ...technology takes more than great engineering—it takes a team of...  ...DescriptionWe are seeking an experienced Sr. Staff Firmware Engineer with...  ...hardware, firmware, and software integrity. Key Responsibilities... 
    Senior
    Software

    Axiado Corporation

    San Jose, CA
    18 hours ago
  • $200k - $322k

     ...developer to build the next generation AI platforms and products that improve business efficiency and productivity. This engineer is expected to be familiar with concepts of...  ...architecture, development, and scaling of our software systems. This role will give an opportunity... 
    Senior
    Software
    Full time
    Immediate start

    Nvidia

    Santa Clara, CA
    3 days ago
  • $170k - $277k

     ...great outcomes.Job SummaryAs a Principal Software Engineer within the Engineering team, you will...  ...Collaborate proactively with Product Management, SRE, and Quality Engineering to deliver...  ...with container orchestration platforms, specifically Docker and Kubernetes.Comprehensive... 
    Senior
    Software
    Full time
    Work at office

    Palo Alto Networks

    Santa Clara, CA
    4 days ago
  • $125k - $185k

     ...we all win with precision. Job Description Your Career The Cortex Vulnerability Intelligence Platform team is looking for a Principal Software Engineer (full stack) to join our team. This team is responsible for vulnerability intelligence collection, processing... 
    Senior
    Software
    Full time
    Casual work
    Work at office

    Palo Alto Networks

    San Jose, CA
    18 hours ago
  • $184k - $287.5k

     ...”Are you willing to challenge yourself and build phenomenal software alongside some of the smartest people in the world? Join us...  ...of technological advancement. We’re hiring a Senior Backend/Platform Engineer to build and maintain the core infrastructure behind NVIDIA... 
    Senior
    Software
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    18 hours ago
  • $170k - $277k

     ...Summary We are seeking a Senior Principal Software Engineer who is first and foremost a software...  ...passion for building innovative tools, platforms, and infrastructure that enable...  ...empower engineering, platform engineering, SRE, and R&D teams, we'd love to hear from... 
    Senior
    Software
    Full time
    Work at office
    Visa sponsorship
    Work visa

    Palo Alto Networks, Inc.

    Santa Clara, CA
    18 hours ago
  •  ...Palo Alto Networks, Inc. is seeking a Senior Principal Software Engineer to lead development of tools, platforms, and infrastructure enabling engineering excellence at scale. This role combines deep software engineering expertise with strategic thinking and cross-team... 
    Senior
    Software

    Jobleads-US

    Santa Clara, CA
    18 hours ago
  • $152k - $241.5k

     ...most thoughtful people in the world.We’re looking for a highly motivated, creative Factory System Software and Diagnostics Integration engineer to join the Datacenter Platform Software team. You will play a crucial role in coordinating factory projects, guiding cross-... 
    Senior
    Software
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $176k - $276k

     ...integrate, and operate the Kubernetes-based platform and shared services used to provision,...  ..., and service enablement. We build software and automation to standardize how network...  ...environments.We are looking for a hands-on senior engineer to own the lifecycle and automation of... 
    Senior
    Software
    Full time
    Remote work
    Weekend work

    Nvidia

    Santa Clara, CA
    1 day ago
  • $96.8k - $306.4k

    Drives cross-group platform initiatives (e.g., identity, config, API...  ...in partnership with SRE and security.Only Oracle brings...  ...Key ResponsibilitiesPlatform Software Development:Set group-wide guardrails...  ...lifecycle; coaches engineers across teams or units to drive... 
    Software
    Temporary work
    Flexible hours
    Shift work

    Oracle Corporation

    Santa Clara, CA
    3 days ago
  • $178k - $321k

     ...working AI-native capability: a multi-agent platform (Hive Mind), agentic workflows, data...  ...just implement it. This is a two-person engineering team: you deploy, debug, and hotfix your...  ...daily and board-cycle workflows. Working software wins arguments; migrate by strangling,... 
    Senior
    Software

    OKX

    San Jose, CA
    18 hours ago
  • $193.3k - $261.5k

     ...learning clusters. Our team builds virtual platforms — full-system C++ and SystemC models of these custom SoCs — that let software teams start development months before silicon...  ...silicon. We're looking for a software engineer to build and own the models and infrastructure... 
    Senior
    Software
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    3 days ago
  • $280k - $380k

     ...TVRoku is the #1 TV streaming platform in the U.S., Canada, and...  ...disciplines.About the TeamOur DevOps/SRE team runs an active-active,...  ...reliability and automation, engineering systems that perform under...  ...Reliability Engineering) Senior Software Engineer to join our dynamic... 
    Senior
    Software
    Work at office
    Local area
    Remote work
    Monday to Thursday
    Flexible hours

    Roku

    San Jose, CA
    3 days ago
  •  ...Jose, California, United StatesProducts - Engineering /Fulltime /HybridOver 50,000 customers...  ...to join the Extreme team.JOB DESCRIPTION Sr. Principal Engineer/ Principal Engineer - Cloud Platform Reports To: Director of Software Systems Engineering Position Summary We are... 
    Senior
    Software
    Full time
    Remote work
    Flexible hours

    Extreme Networks, Inc.

    San Jose, CA
    18 hours ago
  • $176k - $276k

     ...of artificial intelligence.Join our team of innovative engineers who develop and maintain software facilitating GPU communication, driving groundbreaking...  .... We're looking for highly motivated EngOps and Platform Engineers to boost execution efficiency while managing... 
    Senior
    Software
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $200k - $322k

     ...make a lasting impact on the world.This Senior Staff Client Platform Engineer role is a high-impact technical leadership position dedicated...  ..., security, and developer productivity through advanced software engineering practices.What you’ll be doing:Architect and build... 
    Senior
    Software
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $245k - $350k

     ...data lake to power our cloud-native Zero Trust Exchange platform. This innovation protects our customers from cyberattacks...  ...three days a week in San Jose, reporting to the Director,Software Development Engineering ZIdentity team. You will be responsible for design and development... 
    Senior
    Software
    Full time
    Work at office
    Local area
    3 days per week

    Zscaler

    San Jose, CA
    1 day ago
  •  ...Overview Senior Data Platform Engineer - Direct-Hire/FTE - Remote (US) This is a hands-on Senior Data Platform Engineer role that will require...  ...of Data Lakes, Data Warehouses is MUST to have Software development, coding expertise in the Data engineering field... 
    Senior
    Software
    Full time
    Work experience placement
    Local area
    Remote work
    Flexible hours

    INSPYR Solutions

    San Jose, CA
    4 days ago
  •  ...developer to build the next generation AI platforms and products that improve business...  ...architecture, development, and scaling of software systems in a collaborative, agile environment...  ..., with mentorship responsibilities for engineers and a strong emphasis on #J-18808-... 
    Senior
    Software

    NVIDIA

    Santa Clara, CA
    1 day ago
  • Apple is seeking a Senior Software Engineer for the Data Platform to design, develop, and deploy scalable data processing systems at an immense scale in a cloud environment. You will work across Go, Java, and Scala to build highly reliable services, leveraging IaC with... 
    Senior
    Software

    Apple Inc.

    Cupertino, CA
    2 days ago
  • Qualcomm Technologies, Inc. is seeking an experienced Senior Engineer to join the DDR Customer Engineering Software team. You will enable DDR memory subsystems, bring up platforms across firmware, bootloaders, and OS environments, and work with customers to resolve complex... 
    Senior
    Software

    Nutanix

    Santa Clara, CA
    18 hours ago
  •  ...Greetings From Rootshell Inc, One of our client is looking for Sr Software engineer in Santa Clara, CA Job Title : Sr Software engineer Location : Remote ( Preferably Bay Area) Duration: Long Term Must Have skills : Automatic Testing... 
    Senior
    Software
    Work experience placement
    Remote work

    Rootshell Enterprise Technologies

    Santa Clara, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Sr. SRE Platform Software Engineer. Be the first to apply!