SRE L1 Support/Cloud Platform Ops Engineer
Bitdeer Technologies Group
About Bitdeer Technologies Group
Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure.
Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to support the AI revolution. Bitdeer handles complex processes involved in computing such as equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial intelligence.
Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia.
To learn more, visit (Position Overview
You are the first human in the loop — the escalation target when the AIOps system needs a decision, and the source of ground truth that turns novel incidents into new automations.
NeoCloud is building an AI-operated GPU cloud. That doesn't mean fewer humans — it means humans focus on judgment calls the platform can't yet make, and every judgment call trains the platform to do it next time. In this L1 role you cover front-line monitoring and incident response for NeoCloud's US GPU DCs during the 8AM–8PM PST shift. You execute SOPs, escalate the hard cases, and feed the AIOps substrate the ground truth it needs to learn from novel incidents.
What you'll own
- Monitor GPU cluster health, network status, storage systems, and environmental sensors via centralized dashboards.
- Respond to alerts and execute runbooks for common incidents: GPU errors, link flaps, node failures, storage alerts.
- Perform hardware triage: identify failed GPUs, NICs, PSUs, disks, and cables from monitoring data and physical inspection.
- Execute standard remediation: GPU reset, node drain/reboot, link re-seat, BMC recovery.
- Collect diagnostic data for L2/SME escalation: logs, DCGM output, network diagnostics, hardware health reports.
- Manage incident tickets from creation through resolution or escalation (ServiceNow/Jira).
- Perform physical DC tasks: cable installation, hardware swap-outs, rack and stack, labeling (on-site roles).
- Execute structured shift handoffs at 8AM and 8PM PST with the APAC operations team.
- Maintain and update operational runbooks based on recurring issues.
- Assist with hardware deployment, firmware updates, and inventory management under SME guidance.
Feed the AIOps substrate
- Every novel incident you resolve is data the platform team needs — you tag it, describe it, and hand it back so it becomes an automation.
- Every runbook you touch should get closer to being executable by the platform, not by you.
- Your handoff notes are structured signal, not free-form email.
Why this role is different from a NOC job
- You are not the last line of defense — the platform is. You are the training signal.
- Growth path is real: strong L1s here move into SME roles, or into the platform team as automation authors.
Job Requirement:
- 2+ years in NOC, data center operations, or IT support role
- Basic Linux system administration (command line, log analysis, service management)
- Familiarity with monitoring tools (Prometheus, Grafana, Nagios, or equivalent)
- Experience with ticketing systems (ServiceNow, Jira Service Management)
- Ability to perform physical data center tasks: rack and stack, cabling, hardware replacement
- Strong communication skills for shift handoffs, incident documentation, and escalation
- Ability to work 8AM-8PM PST shift schedule (12-hour shifts with rotation)
- Curiosity about automation — you don't just execute the runbook, you notice when it's the third time this month and ask what should change.
- Comfort with structured data — you understand that how you file a ticket matters, because it may train a model that decides how the next one is filed.
--------------------------------------------------------------------
Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, color, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union.
- ...Senior/Lead SRE Platform Services Engineer Technical Leader (Remote) Experience : 12 - 15 years of experience... ..., distributed systems, or cloud operations. ~ Demonstrated technical... ...strategy for Platform Services work supporting FedRAMP High and IL5. Own key architectural...CloudRemote jobPermanent employmentFull time
- ...infrastructure to support the AI revolution.... ...also offers advanced cloud capabilities to customers... ...compute, run by a platform that observes,... ...operates the fleet. The SRE Platform team... ...network, GPU, K8S, and L1 operators —... ...entry-level Software Engineer on the SRE / Monitoring...CloudFull timeContract workInternshipLocal area
- ...NVIDIA Corporation in Santa Clara, CA is seeking an Engineering Manager to lead a team of SRE, platform, and software engineers responsible for the AI Platform... ...You will recruit and develop engineers, partner with Cloud, Security, Networking, and AI/ML teams, improve...Cloud
$208k - $333.5k
Site Reliability Engineering (SRE) at NVIDIA is an engineering field focused... ..., Kubernetes, public cloud, observability, capacity management... ...and operating resilient AI platform capabilities at enterprise... ...resilient distributed systems that support enterprise AI agent products...CloudFull time- ...Corporation in Santa Clara, CA, is seeking an Engineering Manager for Site Reliability Engineering... ...lead a team responsible for NVIDIA's AI Platform Runtime and related production services.... ...scale. You will partner with Cloud, Platform, Security, and AI/ML organizations...Cloud
$152k - $190k
...Zscaler Zero Trust Exchange️ platform protects thousands of... ...’s largest in-line cloud security platform.We believe... ...for a Staff Software Engineer (Service Platform &... ...capabilities and AI-driven SRE practices for a global... ...are built to last and support a high-growth, global...CloudFull timeTemporary workWork at officeLocal area$101k - $161k
...leader in data-driven, client-to-cloud networking for large data... ...prestigious awards, such as Best Engineering Team, Best Company for... ...-as-a-Service (CVaaS) global SRE team. SREs at Arista combine... ...Familiarity with GCP (Google Cloud Platform) and GKE (Google Kubernetes...Cloud$130.7k - $261.3k
...scientists.THE OPPORTUNITYThis Senior Staff ML Ops Engineer position can work out of our Santa Clara... ...operational excellence of enterprise AI platforms. This role bridges advanced AI... ...OnArchitect a highly available, secure, scalable cloud/on-prem hybrid ML infrastructure that...CloudShift work$163.5k - $212.4k
...experienced kernel or hypervisor engineer who wants to work hands-on... ...of NIO’s in-vehicle compute platform. You will join the core... ...trusted execution (TrustZone, OP-TEE). In collaboration with AI and cloud teams, you’ll also support emerging LLM-based...CloudFull timeTemporary workFlexible hours- ...career.THE ROLE:We are seeking a hands-on Platform Engineer to build and operate the... ...with GitOps operating models.Experience supporting AI/ML infrastructure, model serving, GPU... ...experience in software engineering, DevOps, SRE, cloud infrastructure, or platform engineering...Cloud
$160k - $320k
...our Nutanix Kubernetes Platform (NKP) team in San Jose.... ...scalable, enterprise-grade cloud-native solutions at the... ...combine strong backend engineering skills with deep... ...based on Kubernetes. It supports AI/ML workloads, GPU infrastructure... ...product management, SRE, support, and other...CloudWork at officeLocal areaRemote workRelocation package3 days per week$140k - $155k
...networking solutions for Data Center, Cloud Computing, Enterprise IT,... ..., passionate, and committed engineers, technologists, and business... ...available in the enterprise platform.Strengthen Preventive... ...Protection & Data Governance to support Microsoft 365 platform protection...CloudWork at officeWorldwide- ...StatesJob Title: Senior Kubernetes Platform Engineer / ArchitectLocation: Milpitas... ...design, build, automate, and support enterprise-grade container... ..., container orchestration, cloud-native technologies, platform... ...teams, security teams, and SRE teams to deliver a robust cloud...Cloud
- ...purpose-built AI inference silicon and the supporting infrastructure. This role builds and leads the Site Reliability Engineering function from the ground up, owning the infrastructure... ...rely on across on-prem, colocation, and cloud environments. You will hire and grow the...Cloud
- ...ServiceNow is seeking a Staff Voice AI Engineer – SRE/DevOps in Santa Clara. You will design and deliver cloud-native solutions for Voice AI, ensuring high availability... ..., and integrate LLMs into real-time voice platforms. You will mentor teammates, enforce best practices...Cloud
$152k - $241.5k
NVIDIA DGX Cloud builds and operates large-scale GPU infrastructure... ...We are looking for Software Engineers with SRE or Production Engineering... ...to production service supporting an IaaS production environment... ...with hardware, networking, platform, data center operations, and...CloudPermanent employmentFull time- Lambda, The Superintelligence Cloud, is a leader in AI cloud... ...home day is currently Tuesday.Engineering at Lambda is responsible for... ...upgrades, and scaling.Similar to SRE-level coding abilities (small... ...networking, and RBAC across the platform.Lead incident response, root-...CloudWork at officeLocal areaWork from homeFlexible hours
- ...Palo Alto Networks is seeking a Senior DevOps Engineer to design and operate cloud platforms across GCP, AWS, and global data centers, expanding AI-driven automation for SRE/DevOps. You will build intelligent systems that predict incidents, automate root cause analysis...Cloud
- ...Role 1 - Core Platform Engineer (L1, breadth-first) The engineering first line of defense:... ...GitLab CI, AWX, etc.) for bare-metal or cloud infrastructure. TCP/IP and network... ...exposure. Sourcing note: This is a deep SRE profile, not a pure generalist. The...CloudNight shift
- ...SRE Engineer St Louis, MO (Onsite from day 1) Client Required Skills: • Bachelor'... ...• Experience with performance tuning of cloud-native applications • Experience working... ...improve software quality • Hands on CI/CD platforms (Jenkins, Bamboo, Concourse, etc)...Cloud
- ...Position- SRE Engineer Duration-Contract Location- San Jose, C JD... ...back-ups, DR planning • Creating and supporting automation scripts (shell/ansible/python... ...building CICD pipelines (preferred) • Cloud platform knowledge (specifically AWS) is required...CloudContract workImmediate start
$96.8k - $306.4k
Drives cross-group platform initiatives (e.g., identity... ...implementations in partnership with SRE and security.Only... ...the way in AI and cloud solutions that impact... ...competitive benefits that support our people with... ...development lifecycle; coaches engineers across teams or units...CloudTemporary workFlexible hoursShift work- ...detail-oriented Site Reliability Engineer II (SRE II) to join our 24/7... ..., and continuous compliance support across our hybrid environment... ...responsibility is maintaining platform availability, hardware reliability... ...-premises data centers and cloud regions.• On-Call...CloudFull timeLocal areaImmediate startShift workNight shiftAfternoon shiftWeekday work
- ...DescriptionThe AI Inference Engineer plays a critical role in the... ...including Docker, Kubernetes, and cloud platforms such as AWS, GCP, and Azure.... ...). Background in MLOps or SRE roles focused on high-performance... ...and hardware solutions to support real-time AI applications. Advancing...CloudFull timeLocal areaImmediate start
- ...Job Title: Senior Kubernetes Platform Engineer / Architect Location:... ...design, build, automate, and support enterprise-grade container platforms... ..., container orchestration, cloud-native technologies, platform... ...teams, security teams, and SRE teams to deliver a robust...Cloud
$155k - $230k
...spreads across various clouds and devices, traditional... ...unified data security platform addresses vulnerabilities... ...& Platform Engineer to help architect, build... ...software engineering, SRE, DevOps, or related fields... ...sponsorship. We are able to support H-1B transfers for...CloudTemporary workH1bWorldwide- ...Job Description Job Description Java SRE Engineer Onsite San Francisco Bay Area Infrastructure Engineer (2... ...Positions) We are looking for an experienced Java SRE / Platform Engineer to support large-scale cloud migrations and production systems on AWS and...Cloud
$272k - $431.25k
...100% hands-on Storage Services Software engineer to join the block storage group. You will... ...of what is possible today and define the platform of tomorrow.At NVIDIA, we work, think and... ...and reuse existing solutions.Knowledge of cloud computing concepts, including virtualization...CloudFull time$65 - $85 per hour
...a Site Reliability Engineer (Contract) to the team... ...experience. The Cloud team is an... ...Driverless Cars to support their infrastructure... ...processor hardware. The SRE will build the next... ...call and rotational L1 support for... ...collaboration with the platform engineering team....CloudFull timeContract workWorldwide$176.1k - $308.2k
...ServiceNow in Santa Clara, CA, is hiring a Staff Voice AI Engineer – SRE/DevOps to build cloud-native Voice AI solutions, ensure high availability,... ...delivery, mentor peers, and integrate LLMs into voice platforms and real-time systems. Base pay ranges from $176,100 to...Cloud
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to SRE L1 Support/Cloud Platform Ops Engineer. Be the first to apply!
- operations support system engineer San Jose, CA
- technical support engineer San Jose, CA
- remote support engineer San Jose, CA
- IT developer San Jose, CA
- tech support engineer San Jose, CA
- support engineer San Jose, CA
- lab support engineer San Jose, CA
- customer support engineer San Jose, CA
- software technical support engineer San Jose, CA
- junior application support engineer San Jose, CA




