Sr. SRE Platform Software Engineer
Bitdeer Technologies Group
About Bitdeer Technologies Group
Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure.
Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to support the AI revolution. Bitdeer handles complex processes involved in computing such as equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial intelligence.
Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia.
To learn more, visit (Position Overview
Build and operate one or more bounded contexts of the NeoCloud SRE platform — the multi-region substrate that observes, protects, and operates a GPU rental fleet across self-built and OEM-rented data centers. You take an architect-approved design and turn it into production code that ships through GitOps + the CICD release pipeline, ride the Plugin Framework conventions, meet declared SLOs, and stay drift-free.
This is the build + run role. You don't only ship code; you ship a service that other squads, cloud-service teams, and tenants depend on. You take the on-call pager for what you build.
Key Responsibilities
You will own 1-2 of these:
- Collection & Storage: collection-agent, customer-sdk-gateway, metrics-store, logs-store, traces-store, profiles-store, analytics-lake, enrichment-service, collection-monitor.
- Alert, Correlation & SLO: alert-engine-framework, alert-correlation, slo-framework, default M-series alert rules.
- Topology, Cluster-Health & Cluster Platform Services: topology-service, cluster-health-rollup, OSS-SRE-tool collection plugins for K8s, Slurm, Ray, Volcano, Kueue, and KubeRay.
- Fault-Prediction: prediction-engine-framework and built-in predictors (GPU, Link, Disk, XPA, Straggler, SDC, Stranded GPU).
- Remediation, Workflow, Inspection & Jobs: remediation-actuator, orchestration-substrate (workflow engine), inspection-orchestrator, job-scheduler, NCCL-baseline inspection probe.
- Hardware Lifecycle & DC Ops: hardware-lifecycle, dc-operations, boot-provisioning, rolling-upgrade, bare-metal-bmc-service, auto-discovery, ZTP D0–D5 pipeline, IPMI bare-metal management.
- Identity, Secrets, Tenant-Config & CMDB: iam-service, secrets-service, tenant-sre-config, cmdb-cache, schema registry.
- Customer-Bridge, Ticketing & SRE Platform Portal: customer-bridge, customer-ticketing, sre-operation-system, Customer Console BFF, SRE Console BFF.
- Backup, DR & Meta-Monitor: backup-orchestrator, meta-monitor, external-watcher integration (Datadog or equivalent).
- CI/CD, GitOps, Plugin Framework & SRE Image Registry: cicd-pipeline, gitops-sync, plugin-registry, sre-image-registry.
- Self-Improving Agent: agent-control-plane, agent-discovery, agent-codegen, agent-sandbox, per-Region LLM gateway.
- Global SRE Management: maintenance-window-orchestrator, change-management, capacity-planner, cost-optimizer, gpu-efficiency-dashboard, network-stability-dashboard, patching-orchestrator, artifact-management, compat-matrix-service, security-platform.
Qualifications
- Software Engineering Experience: 7+ years of production software engineering experience, including 2 or more years operating what you built (real on-call experience, not just shipping code).
- Programming Languages: Production-depth mastery of at least one systems-grade language—Go (preferred), Rust, or Java. Proficiency in Python for tooling and SDK work.
- Distributed Systems Fundamentals: Strong grasp of at-least-once vs. exactly-once trade-offs, idempotency, back-pressure, leader election, consistent hashing, gossip, and fan-out. Ability to evaluate CRDT vs. Raft vs. Paxos and select the right tool for the job.
- Multi-Region Observability Stack: Experience at production scale with Prometheus, VictoriaMetrics, Mimir, Thanos, Loki, Elasticsearch, Tempo, Jaeger, or OpenTelemetry. Must have built or substantively contributed to the ingest, query, or storage paths of these systems.
- GitOps & CI/CD: Hands-on experience with Argo, Flux, Helm, Kustomize, Cosign signing, signed-bundle promotion, and blast-radius-aware rollouts.
- Kubernetes Operator Pattern: Proven experience writing a controller or CRD handling real production traffic, with a deep understanding of watch-cache mechanics, leader election, and reconcile loops.
- mTLS & Secrets Management: Experience executing end-to-end mTLS bootstrap with certificate rotation. Hands-on experience with HashiCorp Vault or cloud KMS (AWS KMS / GCP KMS).
- SQL & Time-Series Data: Ability to read a Prometheus query plan, build a recording-rule strategy, and write SQL that joins per-tenant telemetry against analytics-lake tables.
- Testing Discipline: Rigorous approach to unit, integration, contract, chaos, and soak testing. Experience writing and maintaining your own comprehensive tests.
- Technical Writing Fluency: Ability to author clear design docs that align with existing platform architecture, create runbooks optimized for 3 AM on-call responses, and write intent-driven PR descriptions.
Preferred Qualifications (GPU / AI-Infra Context)
Experience in at least one of the following areas is a strong plus:
- NVIDIA Internals: Deep understanding of DCGM and NVIDIA driver internals, including XID semantics and MIG / vGPU partitioning.
- Networking & Fabrics: Experience with InfiniBand or RoCE fabrics, including subnet managers, partitioning, optical health, and NCCL collective tracing.
- HPC Storage: Experience managing Lustre, NetApp, Pure, DDN, VAST, or NVMe-oF under multi-tenant loads.
- Hardware Management: Hands-on experience with BMC, IPMI, and Redfish at OEM scale (Supermicro, Dell, HPE, Lenovo).
- Cluster Platform Internals: Familiarity with Kubernetes GPU Operator, Slurm controller, or Ray GCS.
- BS/MS in Computer Science or similar
- Hyperscale or NeoCloud experience
--------------------------------------------------------------------
Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, color, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union.
$101k - $161k
...artificial intelligence, and software-defined networking to provide... ...prestigious awards, such as Best Engineering Team, Best Company for... ...-as-a-Service (CVaaS) global SRE team. SREs at Arista combine... ...Familiarity with GCP (Google Cloud Platform) and GKE (Google Kubernetes...SeniorSoftware$167.7k - $245.2k
...and control.As a Senior Site Reliability Engineer (SRE), you will build, operate, and... ...Splunk Agent Observability's deployment platform and production infrastructure. You will... ...Infrastructure as Code tools• Collaborate with software engineers and customers to design...SeniorSoftwareFull timeTemporary workLocal areaFlexible hours2 days per week- ...Together, we advance your career. THE ROLEWe are hiring AI / ML Platform Engineers to build the platform layer that makes AI-for-engineering... ...comfortable working across ML, infrastructure, and hardware/software tooling, and you can partner with research teams without...SeniorSoftware
$148.75k - $361k
...watches TVRoku is the #1 TV streaming platform in the U.S., Canada, and Mexico, and we... ...seeking a talented and experienced Senior Software Engineer, MLOps/DevOps, to join the Advertising... ...has a strong background in DevOps/SRE practices, cloud infrastructure management...SeniorSoftwareWork at officeLocal areaRemote workMonday to ThursdayFlexible hours$100k - $170k
...We are seeking a DevOps Engineer who is eager to have an immediate... ..., testing, and release platforms, taking us from square one to... ...highly-scalable, highly-available software systems in AWS Architect and... ...~5+ years of experience in SRE, DevOps, or Platform Engineering...SoftwareFull timeWork at officeImmediate startVisa sponsorshipNight shift- ...solutions spanning all phases of the Software Development Lifecycle (SDLC)... ...-level issues Guides junior engineers Operates with little day-to-... ...PagerDuty. Collaborate with SRE and cloud teams to design and... ...solutions. Solid backend or platform engineering experience in...SeniorSoftwareFull timeWork experience placementImmediate start
$170k - $277k
...great outcomes.Job SummaryAs a Principal Software Engineer within the Engineering team, you will... ...Collaborate proactively with Product Management, SRE, and Quality Engineering to deliver... ...with container orchestration platforms, specifically Docker and Kubernetes.Comprehensive...SeniorSoftwareFull timeWork at office- ...and efficiency, and real-time platform management. Our technology... ...technology takes more than great engineering—it takes a team of... ...DescriptionWe are seeking an experienced Sr. Staff Firmware Engineer with... ...hardware, firmware, and software integrity. Key Responsibilities...SeniorSoftware
$144k - $175k
...members.We are looking for a highly motivated and versatile Senior Platform Engineer to join our platform engineering team. This unique role is... ...Platform Engineering, DevOps, Site Reliability Engineering (SRE), or a related discipline.Deep expertise in cloud platforms (AWS...SeniorLocal area$125k - $185k
...we all win with precision. Job Description Your Career The Cortex Vulnerability Intelligence Platform team is looking for a Principal Software Engineer (full stack) to join our team. This team is responsible for vulnerability intelligence collection, processing...SeniorSoftwareFull timeCasual workWork at office$184k - $287.5k
...”Are you willing to challenge yourself and build phenomenal software alongside some of the smartest people in the world? Join us... ...of technological advancement. We’re hiring a Senior Backend/Platform Engineer to build and maintain the core infrastructure behind NVIDIA...SeniorSoftwareFull timeRemote work$207.4k - $259.2k
...highly experienced and passionate Sr. Staff Site Reliability Engineer (SRE) to join our growing team. In this... ...chat service, potentially leveraging platforms like OpenRouter or similar alternatives... ...reliability is built into the software development lifecycle from inception...SeniorSoftwarePermanent employmentLocal area$152k - $241.5k
...most thoughtful people in the world.We’re looking for a highly motivated, creative Factory System Software and Diagnostics Integration engineer to join the Datacenter Platform Software team. You will play a crucial role in coordinating factory projects, guiding cross-...SeniorSoftwareFull time$176k - $276k
...integrate, and operate the Kubernetes-based platform and shared services used to provision,... ..., and service enablement. We build software and automation to standardize how network... ...environments.We are looking for a hands-on senior engineer to own the lifecycle and automation of...SeniorSoftwareFull timeRemote workWeekend work$178k - $321k
...working AI-native capability: a multi-agent platform (Hive Mind), agentic workflows, data... ...just implement it. This is a two-person engineering team: you deploy, debug, and hotfix your... ...daily and board-cycle workflows. Working software wins arguments; migrate by strangling,...SeniorSoftware- ...Jose, California, United StatesProducts - Engineering /Fulltime /HybridOver 50,000 customers... ...to join the Extreme team.JOB DESCRIPTION Sr. Principal Engineer/ Principal Engineer - Cloud Platform Reports To: Director of Software Systems Engineering Position Summary We are...SeniorSoftwareFull timeRemote workFlexible hours
$176k - $276k
...of artificial intelligence.Join our team of innovative engineers who develop and maintain software facilitating GPU communication, driving groundbreaking... .... We're looking for highly motivated EngOps and Platform Engineers to boost execution efficiency while managing...SeniorSoftwareFull time$200k - $322k
...make a lasting impact on the world.This Senior Staff Client Platform Engineer role is a high-impact technical leadership position dedicated... ..., security, and developer productivity through advanced software engineering practices.What you’ll be doing:Architect and build...SeniorSoftwareFull time$193.3k - $261.5k
...learning clusters. Our team builds virtual platforms — full-system C++ and SystemC models of these custom SoCs — that let software teams start development months before silicon... ...silicon. We're looking for a software engineer to build and own the models and infrastructure...SeniorSoftwareInternshipLocal areaFlexible hours$96.8k - $306.4k
Drives cross-group platform initiatives (e.g., identity, config, API... ...in partnership with SRE and security.Only Oracle brings... ...Key ResponsibilitiesPlatform Software Development:Set group-wide guardrails... ...lifecycle; coaches engineers across teams or units to drive...SoftwareTemporary workFlexible hoursShift work$280k - $380k
...TVRoku is the #1 TV streaming platform in the U.S., Canada, and... ...disciplines.About the TeamOur DevOps/SRE team runs an active-active,... ...reliability and automation, engineering systems that perform under... ...Reliability Engineering) Senior Software Engineer to join our dynamic...SeniorSoftwareWork at officeLocal areaRemote workMonday to ThursdayFlexible hours$245k - $350k
...data lake to power our cloud-native Zero Trust Exchange platform. This innovation protects our customers from cyberattacks... ...three days a week in San Jose, reporting to the Director,Software Development Engineering ZIdentity team. You will be responsible for design and development...SeniorSoftwareFull timeWork at officeLocal area3 days per week$173.5k - $331.05k
...mission to transform the future of creativity by building SDKs and platform libraries that power data‑driven insights and AI‑enabled experiences across Creative Cloud. We are seeking a software engineer with strong development and computer science fundamentals and experience...SeniorSoftwareFull timeTemporary workLocal areaWorldwide- ...Job Description Job Description Java SRE Engineer Onsite San Francisco Bay Area Infrastructure Engineer (2 Positions) We are looking for an experienced Java SRE / Platform Engineer to support large-scale cloud migrations and production systems on AWS and...
- ...immediate opening with my client. If you are looking for a new project, please send me a copy of your updated resumes Title: Sr. SRE / DevOps Engineer Location: Sunnyvale, CA (Only Local candidate) Client Interview In-Person Job Summary For this role, we are...SeniorLocal areaImmediate start
- ...Greetings From Rootshell Inc, One of our client is looking for Sr Software engineer in Santa Clara, CA Job Title : Sr Software engineer Location : Remote ( Preferably Bay Area) Duration: Long Term Must Have skills : Automatic Testing...SeniorSoftwareWork experience placementRemote work
- ...networking technology. Our distributed services platform expands AMD’s data center product portfolio with a high‑performance DPU and software stack deployed at scale across major... ...a high‑impact Design Verification Engineer with strong technical depth, ownership, and...SeniorSoftware
$92.5k - $209.5k
...Job Description Owns moderately complex components within platform services or SDKs; leads team-level improvements to integration frameworks... .... Responsibilities Key Responsibilities Platform Software Development: Own a bounded platform component (service...SeniorSoftwareTemporary workImmediate startFlexible hoursShift work$184k - $287.5k
...can make a lasting impact on the world.We are now looking for a senior software engineer to join our Hardware Infrastructure team! Our team is responsible for delivering highly available platform services and performant web applications which integrate with various data...SeniorSoftwareFull time- ...DevOps Support Engineer Client is currently seeking multiple DevOps Support Engineers to join our team in Santa Clara, CA. As a member... ...like AWS, VMware, OpenStack is a must. Knowledge of software application frameworks. Expert knowledge in cloud infrastructure...SeniorSoftware
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Sr. SRE Platform Software Engineer. Be the first to apply!
- site reliability engineer San Jose, CA
- platform developer San Jose, CA
- senior platform engineer San Jose, CA
- platform engineer San Jose, CA
- agile software developer San Jose, CA
- software developer internship no experience San Jose, CA
- intermediate software engineer San Jose, CA
- software engineer staff San Jose, CA
- experienced software developer San Jose, CA
- work from home software developer San Jose, CA



