Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

SRE L1 Support/Cloud Platform Ops Engineer

Full-time

Bitdeer Technologies Group

About Bitdeer Technologies Group

Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure.

Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to support the AI revolution. Bitdeer handles complex processes involved in computing such as equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial intelligence.

Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia.

To learn more, visit (

Position Overview

You are the first human in the loop — the escalation target when the AIOps system needs a decision, and the source of ground truth that turns novel incidents into new automations.

NeoCloud is building an AI-operated GPU cloud. That doesn't mean fewer humans — it means humans focus on judgment calls the platform can't yet make, and every judgment call trains the platform to do it next time. In this L1 role you cover front-line monitoring and incident response for NeoCloud's US GPU DCs during the 8AM–8PM PST shift. You execute SOPs, escalate the hard cases, and feed the AIOps substrate the ground truth it needs to learn from novel incidents.

What you'll own

  • Monitor GPU cluster health, network status, storage systems, and environmental sensors via centralized dashboards.
  • Respond to alerts and execute runbooks for common incidents: GPU errors, link flaps, node failures, storage alerts.
  • Perform hardware triage: identify failed GPUs, NICs, PSUs, disks, and cables from monitoring data and physical inspection.
  • Execute standard remediation: GPU reset, node drain/reboot, link re-seat, BMC recovery.
  • Collect diagnostic data for L2/SME escalation: logs, DCGM output, network diagnostics, hardware health reports.
  • Manage incident tickets from creation through resolution or escalation (ServiceNow/Jira).
  • Perform physical DC tasks: cable installation, hardware swap-outs, rack and stack, labeling (on-site roles).
  • Execute structured shift handoffs at 8AM and 8PM PST with the APAC operations team.
  • Maintain and update operational runbooks based on recurring issues.
  • Assist with hardware deployment, firmware updates, and inventory management under SME guidance.

Feed the AIOps substrate

  • Every novel incident you resolve is data the platform team needs — you tag it, describe it, and hand it back so it becomes an automation.
  • Every runbook you touch should get closer to being executable by the platform, not by you.
  • Your handoff notes are structured signal, not free-form email.

Why this role is different from a NOC job

  • You are not the last line of defense — the platform is. You are the training signal.
  • Growth path is real: strong L1s here move into SME roles, or into the platform team as automation authors.

Job Requirement:

  • 2+ years in NOC, data center operations, or IT support role
  • Basic Linux system administration (command line, log analysis, service management)
  • Familiarity with monitoring tools (Prometheus, Grafana, Nagios, or equivalent)
  • Experience with ticketing systems (ServiceNow, Jira Service Management)
  • Ability to perform physical data center tasks: rack and stack, cabling, hardware replacement
  • Strong communication skills for shift handoffs, incident documentation, and escalation
  • Ability to work 8AM-8PM PST shift schedule (12-hour shifts with rotation)
  • Curiosity about automation — you don't just execute the runbook, you notice when it's the third time this month and ask what should change.
  • Comfort with structured data — you understand that how you file a ticket matters, because it may train a model that decides how the next one is filed.

--------------------------------------------------------------------

Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, color, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union.

Vacancy posted a month ago
Similar jobs that could be interesting for youBased on the SRE L1 Support/Cloud Platform Ops Engineer in San Jose, CA vacancy
  •  ...infrastructure to support the AI revolution....  ...also offers advanced cloud capabilities to customers...  ...compute, run by a platform that observes,...  ...operates the fleet. The SRE Platform team...  ...network, GPU, K8S, and L1 operators —...  ...entry-level Software Engineer on the SRE / Monitoring... 
    Cloud
    Remote job
    Full time
    Contract work
    Internship
    Local area

    Bitdeer Technologies Group

    San Jose, CA
    17 days ago
  • $108k - $172.5k

    The NVIDIA DGX Cloud organization is looking for passionate software support engineers to partner closely with our internal customers to support them on our internal platforms. This partnership requires you to gain a deep understanding of the customer needs, how their... 
    Cloud
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    4 days ago
  • $101k - $161k

     ...leader in data-driven, client-to-cloud networking for large data...  ...prestigious awards, such as Best Engineering Team, Best Company for...  ...-as-a-Service (CVaaS) global SRE team. SREs at Arista combine...  ...Familiarity with GCP (Google Cloud Platform) and GKE (Google Kubernetes... 
    Cloud

    Arista Networks

    Santa Clara, CA
    3 days ago
  • $152k - $190k

     ...security data lake to power our cloud-native Zero Trust Exchange platform. This innovation protects...  ...for a Staff Software Engineer (Service Platform &...  ...capabilities and AI-driven SRE practices for a global fleet...  ...are built to last and support a high-growth, global organization... 
    Cloud
    Full time
    Temporary work
    Work at office
    Local area

    Zscaler

    San Jose, CA
    1 day ago
  • $130.7k - $261.3k

     ...scientists.THE OPPORTUNITYThis Senior Staff ML Ops Engineer position can work out of our Santa Clara...  ...operational excellence of enterprise AI platforms. This role bridges advanced AI...  ...OnArchitect a highly available, secure, scalable cloud/on-prem hybrid ML infrastructure that... 
    Cloud
    Shift work

    Abbott

    Santa Clara, CA
    4 days ago
  • $163.5k - $212.4k

     ...experienced kernel or hypervisor engineer who wants to work hands-on...  ...of NIO’s in-vehicle compute platform. You will join the core...  ...trusted execution (TrustZone, OP-TEE). In collaboration with AI and cloud teams, you’ll also support emerging LLM-based... 
    Cloud
    Full time
    Temporary work
    Flexible hours

    NIO USA, INC

    San Jose, CA
    21 days ago
  •  ...career.THE ROLE:We are seeking a hands-on Platform Engineer to build and operate the...  ...with GitOps operating models.Experience supporting AI/ML infrastructure, model serving, GPU...  ...experience in software engineering, DevOps, SRE, cloud infrastructure, or platform engineering... 
    Cloud

    AMD

    San Jose, CA
    2 days ago
  • $140k - $155k

     ...networking solutions for Data Center, Cloud Computing, Enterprise IT,...  ..., passionate, and committed engineers, technologists, and business...  ...available in the enterprise platform.Strengthen Preventive...  ...Protection & Data Governance to support Microsoft 365 platform protection... 
    Cloud
    Work at office
    Worldwide

    Super Micro Computer

    San Jose, CA
    1 day ago
  • $200k - $322k

     ...in A.I, video games, cloud and enterprise computing...  ...autonomous driving platforms of some of the world’s...  ...Management with solid engineering background who will be...  ...deployment.Experience supporting or leading data campaign...  ...domains.Exposure to ML Ops, data operations,... 
    Cloud
    Full time

    Nvidia

    Santa Clara, CA
    5 hours ago
  • Lambda, The Superintelligence Cloud, is a leader in AI cloud...  ...home day is currently Tuesday.Engineering at Lambda is responsible for...  ...networking, and RBAC across the platform.Lead incident response, root-...  ...Platform, Infrastructure, or SRE roles, including running Kubernetes... 
    Cloud
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    3 days ago
  • $100k - $170k

     ...are seeking a  DevOps Engineer who is eager to have an...  ..., testing, and release platforms, taking us from square...  ...using Terraform for AWS cloud resources Develop and...  ...scaling strategies to support mission-critical operations...  ...years of experience in SRE, DevOps, or Platform... 
    Cloud
    Full time
    Work at office
    Immediate start
    Visa sponsorship
    Night shift

    E-Space

    Saratoga, CA
    21 days ago
  • $196k - $310.5k

     ...looking for a Senior Security Engineer passionate about platform and device security! This...  ...infrastructure across cloud, on-premises, and managed device...  ...of the charter is building SRE agents. These are...  ...protecting the platforms that support some of the most advanced computing... 
    Cloud
    Full time
    Remote work
    Worldwide

    Nvidia

    Santa Clara, CA
    1 day ago
  • $155k - $230k

     ...spreads across various clouds and devices, traditional...  ...unified data security platform addresses vulnerabilities...  ...& Platform Engineer to help architect, build...  ...software engineering, SRE, DevOps, or related fields...  ...sponsorship. We are able to support H-1B transfers for... 
    Cloud
    Temporary work
    H1b
    Worldwide

    Fortanix

    Santa Clara, CA
    14 days ago
  • $96.8k - $306.4k

    Drives cross-group platform initiatives (e.g., identity...  ...implementations in partnership with SRE and security.Only...  ...the way in AI and cloud solutions that impact...  ...competitive benefits that support our people with...  ...development lifecycle; coaches engineers across teams or units... 
    Cloud
    Temporary work
    Flexible hours
    Shift work

    Oracle Corporation

    Santa Clara, CA
    6 days ago
  • $207k - $300k

     ...team of Software/Systems Engineers on projects for users...  ...Reliability Engineering (SRE) combines software and...  ...ensures that Google Cloud's services—both our internally...  ...that provides the support and mentorship needed to...  ...-generation of Google platforms, we make Google's product... 
    Cloud

    Google

    San Jose, CA
    1 day ago
  •  ...DescriptionThe AI Inference Engineer plays a critical role in the...  ...including Docker, Kubernetes, and cloud platforms such as AWS, GCP, and Azure....  ...). Background in MLOps or SRE roles focused on high-performance...  ...and hardware solutions to support real-time AI applications. Advancing... 
    Cloud
    Full time
    Local area
    Immediate start

    F5 Networks

    San Jose, CA
    1 day ago
  •  ...Job Description Job Description Java SRE Engineer Onsite San Francisco Bay Area Infrastructure Engineer (2...  ...Positions) We are looking for an experienced Java SRE / Platform Engineer to support large-scale cloud migrations and production systems on AWS and... 
    Cloud

    Eitacies Inc

    Santa Clara, CA
    23 days ago
  • $207k - $300k

     ...team of Site Reliability Engineerings (SREs) and security...  ...technical mentorship while supporting the team's expansion...  ..., and enterprise Cloud customers.Extensive background...  ...Engineering (SRE) combines software and...  ...generation of Google platforms, we make Google's product... 
    Cloud
    Shift work

    Google

    Sunnyvale, CA
    4 days ago
  • $272k - $431.25k

     ...100% hands-on Storage Services Software engineer to join the block storage group. You will...  ...of what is possible today and define the platform of tomorrow.At NVIDIA, we work, think and...  ...and reuse existing solutions.Knowledge of cloud computing concepts, including virtualization... 
    Cloud
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $110k - $185k

     ...industry leader in data-driven, client-to-cloud networking for large data center, campus...  ...prestigious awards, such as Best Engineering Team, Best Company for Diversity, Compensation...  ...engineers, hardware engineers, field support engineers and software test engineers to... 
    Cloud
    Full time

    Arista Networks

    Santa Clara, CA
    1 day ago
  • $90k - $110k

     ...solutions for Data Center, Cloud Computing, Enterprise...  ..., and committed engineers, technologists, and business...  ...Job Summary Join us in supporting our Global Service network...  ...escalated from the L1 team. Apply troubleshooting...  ..., and enterprise GPU platforms. Hands-on experience... 
    Cloud
    Work experience placement
    Worldwide
    Night shift
    Weekend work
    Afternoon shift

    Super Micro Computer Spain, S.L.

    San Jose, CA
    2 days ago
  •  ...supercomputing, high-performance computing, cloud, and AI. Whether you’re designing next-...  ...CPU, GPU, AI, and adaptive computing platforms. This technology provides software-executable...  ...a Virtual Platform Functional Modeling Engineer, you will play a key role in defining... 
    Cloud

    AMD

    San Jose, CA
    1 day ago
  • $130k - $170k

     ...industry leader in data-driven, client-to-cloud networking for large data center, campus...  ...prestigious awards, such as Best Engineering Team, Best Company for Diversity, Compensation...  ..., where you’ll collaborate and be supported by like minded sales professionals. This... 
    Cloud
    Full time

    Arista Networks, Inc.

    Santa Clara, CA
    1 day ago
  • $115k - $135k

     ...and networking solutions for Data Center, Cloud Computing, Enterprise IT, Hadoop/ Big...  ...seek talented, passionate, and committed engineers, technologists, and business leaders to...  ...Service Quality & Customer Experience Analyst supports global Service Operations by evaluating... 
    Cloud
    Worldwide

    Super Micro Computer

    San Jose, CA
    1 day ago
  • $140k - $215k

     ...world’s most advanced AI-native platform. We work on large scale...  ...CrowdStrike, Site Reliability Engineering (SRE) is at the forefront of...  ...reliability and scalability of our cloud-native security platform. In...  ...empowered to succeed. We support veterans and individuals with... 
    Cloud
    Full time
    Work experience placement
    Work at office
    Local area

    CrowdStrike

    Sunnyvale, CA
    5 days ago
  • $200k - $322k

     ...ll be immersed in a diverse, supportive environment where everyone is...  ...the world.Ready to build the platforms that make AI at scale possible...  ...a Senior Staff Platform Engineer to architect, build, and scale...  ...role spans distributed systems, cloud, networking, content delivery... 
    Cloud
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  •  ...DevSecOpsTechnical understanding of a range of enterprise IT and cloud-based architectures and technologies (AWS, Azure, Databricks...  ...practice product and cloud security across AWS or Azure or GCP by supporting the implementation of the firm's security standards in... 
    Cloud
    Apprenticeship
    Shift work

    McKinsey & Company

    San Jose, CA
    5 days ago
  • $130k - $170k

     ...industry leader in data-driven, client-to-cloud networking for large data center, campus...  ...prestigious awards, such as Best Engineering Team, Best Company for Diversity, Compensation...  ..., where you’ll collaborate and be supported by like minded sales professionals. This... 
    Cloud

    Arista Networks

    Santa Clara, CA
    3 days ago
  •  ...Job Title 12+ years in platform engineering, SRE, or DevOps. Experience with HPC clusters (Slurm, PBS, Grid Engine). Cloud infrastructure expertise (GCP/AWS preferred). Proficiency with Terraform, Ansible, Prometheus, Grafana, ELK. Strong Linux administration... 
    Cloud

    Saxon Global

    Mountain View, CA
    3 days ago
  • $230k - $250k

     ...networking, giving engineers and AI agents the...  ...a groundbreaking platform that transforms how...  ...across every major cloud and vendor...  ...keep the lights on" SRE role. As our first...  ...to HaveExperience supporting enterprise or federal...  ...Role Is NotA pure ops or NOC role — you... 
    Cloud
    Night shift

    Forward Networks

    Santa Clara, CA
    5 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to SRE L1 Support/Cloud Platform Ops Engineer. Be the first to apply!