SRE L1 Support/Cloud Platform Ops Engineer
Bitdeer Technologies Group
About Bitdeer Technologies Group
Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure.
Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to support the AI revolution. Bitdeer handles complex processes involved in computing such as equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial intelligence.
Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia.
To learn more, visit (Position Overview
You are the first human in the loop — the escalation target when the AIOps system needs a decision, and the source of ground truth that turns novel incidents into new automations.
NeoCloud is building an AI-operated GPU cloud. That doesn't mean fewer humans — it means humans focus on judgment calls the platform can't yet make, and every judgment call trains the platform to do it next time. In this L1 role you cover front-line monitoring and incident response for NeoCloud's US GPU DCs during the 8AM–8PM PST shift. You execute SOPs, escalate the hard cases, and feed the AIOps substrate the ground truth it needs to learn from novel incidents.
What you'll own
- Monitor GPU cluster health, network status, storage systems, and environmental sensors via centralized dashboards.
- Respond to alerts and execute runbooks for common incidents: GPU errors, link flaps, node failures, storage alerts.
- Perform hardware triage: identify failed GPUs, NICs, PSUs, disks, and cables from monitoring data and physical inspection.
- Execute standard remediation: GPU reset, node drain/reboot, link re-seat, BMC recovery.
- Collect diagnostic data for L2/SME escalation: logs, DCGM output, network diagnostics, hardware health reports.
- Manage incident tickets from creation through resolution or escalation (ServiceNow/Jira).
- Perform physical DC tasks: cable installation, hardware swap-outs, rack and stack, labeling (on-site roles).
- Execute structured shift handoffs at 8AM and 8PM PST with the APAC operations team.
- Maintain and update operational runbooks based on recurring issues.
- Assist with hardware deployment, firmware updates, and inventory management under SME guidance.
Feed the AIOps substrate
- Every novel incident you resolve is data the platform team needs — you tag it, describe it, and hand it back so it becomes an automation.
- Every runbook you touch should get closer to being executable by the platform, not by you.
- Your handoff notes are structured signal, not free-form email.
Why this role is different from a NOC job
- You are not the last line of defense — the platform is. You are the training signal.
- Growth path is real: strong L1s here move into SME roles, or into the platform team as automation authors.
Job Requirement:
- 2+ years in NOC, data center operations, or IT support role
- Basic Linux system administration (command line, log analysis, service management)
- Familiarity with monitoring tools (Prometheus, Grafana, Nagios, or equivalent)
- Experience with ticketing systems (ServiceNow, Jira Service Management)
- Ability to perform physical data center tasks: rack and stack, cabling, hardware replacement
- Strong communication skills for shift handoffs, incident documentation, and escalation
- Ability to work 8AM-8PM PST shift schedule (12-hour shifts with rotation)
- Curiosity about automation — you don't just execute the runbook, you notice when it's the third time this month and ask what should change.
- Comfort with structured data — you understand that how you file a ticket matters, because it may train a model that decides how the next one is filed.
--------------------------------------------------------------------
Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, color, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union.
- ...computational infrastructure to support the AI revolution.... ...also offers advanced cloud capabilities to... ...contexts of the NeoCloud SRE platform — the multi-region substrate... ...& SLO: alert-engine-framework, alert-correlation... ...Hardware Lifecycle & DC Ops: hardware-lifecycle, dc...CloudFull timeContract workLocal area
$98.9k - $228.7k
...reliably at scale. We collaborate across engineering, product, and operations to solve... ...principlesUtilize monitoring tools, observability platforms, and metrics collection systems to... ...rotations and demonstrate experience supporting production systems in high-availability...CloudFull timeWork at officeRemote work$120k - $180k
...most advanced AI-native platform. We work on large scale... ...starts with you.Sr. SRE & DevOps EngineerAbout... ...Role:At CrowdStrike, our engineering organization depends on... ...infrastructure spanning multiple cloud providers and regions,... ...and collaborate - Support engineering teams with...CloudFull timeWork experience placementWork at officeLocal area$101k - $161k
...leader in data-driven, client-to-cloud networking for large data... ...prestigious awards, such as Best Engineering Team, Best Company for... ...-as-a-Service (CVaaS) global SRE team. SREs at Arista combine... ...Familiarity with GCP (Google Cloud Platform) and GKE (Google Kubernetes...Cloud$152k - $190k
...security data lake to power our cloud-native Zero Trust Exchange platform. This innovation protects... ...for a Staff Software Engineer (Service Platform &... ...capabilities and AI-driven SRE practices for a global fleet... ...are built to last and support a high-growth, global organization...CloudFull timeTemporary workWork at officeLocal area$103k - $143k
...managing Windows Server environments and supporting Windows-based applications. We... ...principles. We expect experience with cloud platforms such as AWS or Azure and virtualization... ...DevOps More: We are hiring a Platform Engineer IV to support mission-critical, remote...CloudFull timeRemote work$163.5k - $212.4k
...experienced kernel or hypervisor engineer who wants to work hands-on... ...of NIO’s in-vehicle compute platform. You will join the core... ...trusted execution (TrustZone, OP-TEE). In collaboration with AI and cloud teams, you’ll also support emerging LLM-based...CloudFull timeTemporary workFlexible hours$100k - $170k
...are seeking a DevOps Engineer who is eager to have an... ..., testing, and release platforms, taking us from square... ...using Terraform for AWS cloud resources Develop and... ...scaling strategies to support mission-critical operations... ...years of experience in SRE, DevOps, or Platform...CloudFull timeWork at officeImmediate startVisa sponsorshipNight shift$200k - $322k
...in A.I, video games, cloud and enterprise computing... ...autonomous driving platforms of some of the world’s... ...Management with solid engineering background who will be... ...deployment.Experience supporting or leading data campaign... ...domains.Exposure to ML Ops, data operations,...CloudFull time- Lambda, The Superintelligence Cloud, is a leader in AI cloud... ...home day is currently Tuesday.Engineering at Lambda is responsible for... ...networking, and RBAC across the platform.Lead incident response, root-... ...Platform, Infrastructure, or SRE roles, including running Kubernetes...CloudWork at officeLocal areaWork from homeFlexible hours
$100k
...software models, compilers, platforms, networking, and... ...backend or infrastructure engineer with experience building... ...frameworks.Collaborate with SRE, infrastructure, and deployment teams to support large-scale on-prem and... ...differs from cloud-native environments at scale...CloudPermanent employment$167.7k - $245.2k
...and control.As a Senior Site Reliability Engineer (SRE), you will build, operate, and... ...Splunk Agent Observability's deployment platform and production infrastructure. You will own the operational backbone supporting both cloud and air-gapped customer deployments, develop...CloudFull timeTemporary workLocal areaFlexible hours2 days per week- ...Job Description Job Description Java SRE Engineer Onsite San Francisco Bay Area Infrastructure Engineer (2... ...Positions) We are looking for an experienced Java SRE / Platform Engineer to support large-scale cloud migrations and production systems on AWS and...Cloud
$96.8k - $306.4k
Drives cross-group platform initiatives (e.g., identity... ...implementations in partnership with SRE and security.Only... ...the way in AI and cloud solutions that impact... ...competitive benefits that support our people with... ...development lifecycle; coaches engineers across teams or units...CloudTemporary workFlexible hoursShift work- ...delivering an AI-powered platform that governs and secures... ...complex, distributed, cloud-native systems. As a Staff Platform Engineer, you will play a critical... ...technical guidance and support Participate in on-... ...of experience as a Staff SRE with a strong focus on building...Cloud
$110k - $185k
...industry leader in data-driven, client-to-cloud networking for large data center, campus... ...prestigious awards, such as Best Engineering Team, Best Company for Diversity, Compensation... ...engineers, hardware engineers, field support engineers and software test engineers to...CloudFull time- Selector is building an operational intelligence platform for digital infrastructure. We are hiring a software engineering associate to learn our analytics solution,... ..., CS degree, Python basics, Docker/Kubernetes, cloud familiarity (AWS/GCP/Azure), and strong communication...Cloud
$140k - $215k
...world’s most advanced AI-native platform. We work on large scale... ...CrowdStrike, Site Reliability Engineering (SRE) is at the forefront of... ...reliability and scalability of our cloud-native security platform. In... ...empowered to succeed. We support veterans and individuals with...CloudFull timeWork experience placementWork at officeLocal area- ...DevOps Support Engineer Client is currently seeking multiple DevOps Support Engineers to join our team in Santa Clara, CA. As a member... ...troubleshooting the various issues that our clients face with our Client’s Cloud Center product suite. You will also be a key point of...Cloud
- ...DevSecOpsTechnical understanding of a range of enterprise IT and cloud-based architectures and technologies (AWS, Azure, Databricks... ...practice product and cloud security across AWS or Azure or GCP by supporting the implementation of the firm's security standards in...CloudApprenticeshipShift work
$176k - $276k
...invites applications for a Senior DevOps Platform Engineer skilled in Platform and Release... ...Kubernetes-based platform infrastructure supporting AI/ML workloads on NVIDIA Data Center GPUs... ...scale.Background in BareMetal and hybrid cloud (AWS, GCP, Azure) environment...CloudFull time$123k - $191k
...industry leader in data-driven, client-to-cloud networking for large data center, campus... ...prestigious awards, such as Best Engineering Team, Best Company for Diversity, Compensation... ...CPU vendors to integrate their kernel support into EOS. You will also help to bring-up...Cloud- ...partners closely with customer engineering and cross-functional Arm... ...some of Arm's most important cloud and AI infrastructure Partners... ...architecture, silicon, software, platform, and operations organizations... ...long-term strategic growth. Support critical customer concerns...Cloud
$240k - $260k
...security, delivering an AI-powered platform that governs and secures... ...You define the standards ML engineers and scientists build on, and... ...database: operate Pgvector (Cloud SQL) for POC and Qdrant on GKE... ...artificial intelligence (AI) tools to support parts of the hiring process,...Cloud- ...supercomputing, high-performance computing, cloud, and AI. Whether you’re designing next-... ...of AMD, Inc., is hiring SMTS IT Engineer to Design strategies for enterprise databases... ...and Data Security for various big Data Platforms. Enterprise Data Warehousing scalability...Cloud
$230k - $250k
...networking, giving engineers and AI agents the... ...a groundbreaking platform that transforms how... ...across every major cloud and vendor... ...keep the lights on" SRE role. As our first... ...to HaveExperience supporting enterprise or federal... ...Role Is NotA pure ops or NOC role — you...CloudNight shift- ...infrastructure to support the AI... ...offers advanced cloud capabilities to... ...hands-on Cloud SRE Architect to lead... ...generation public cloud platform. This role will... ...infrastructure, platform engineering and AI systems,... ...layered model (L1–L6), tier... ...infrastructure — NVIDIA GPU ops (DCGM, MIG, vGPU...CloudFull timeContract workLocal areaShift work
$262.7k - $355.4k
...partners closely with customer engineering and cross-functional Arm... ...some of Arm’s most important cloud and AI infrastructure Partners... ...architecture, silicon, software, platform, and operations organizations... ...long-term strategic growth.Support critical customer concerns and...CloudWork at officeLocal area$215.18k
...engagement across various industries. This group supports Intel's mission to deliver innovative... ...win opportunities and provides co-engineering to drive product to production that results... ...Intel's product portfolio, including cloud, AI, connectivity, and edge solutions.Develop...CloudFull timeLocal areaImmediate startShift work$184k - $287.5k
...ll be immersed in a diverse, supportive environment where everyone is... ...a part of NVIDIA GeForce NOW cloud team that allows users to play... ...at high resolutions. As an engineer on the GeForce NOW team, you... ...of the next generation cloud platform. You will be at the forefront...CloudFull time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to SRE L1 Support/Cloud Platform Ops Engineer. Be the first to apply!
- IT developer San Jose, CA
- technical support engineer San Jose, CA
- remote support engineer San Jose, CA
- software technical support engineer San Jose, CA
- support engineer San Jose, CA
- IT engineer San Jose, CA
- remote network administrator / IT support engineer San Jose, CA
- operations support system engineer San Jose, CA
- lab support engineer San Jose, CA
- customer support engineer San Jose, CA



