Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

SRE L1 Support/Cloud Platform Ops Engineer

Bitdeer

About Bitdeer Technologies Group Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure. Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to support the AI revolution. Bitdeer handles complex processes involved in computing such as equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial intelligence. Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia. To learn more, visit Position Overview You are the first human in the loop — the escalation target when the AIOps system needs a decision, and the source of ground truth that turns novel incidents into new automations. NeoCloud is building an AI-operated GPU cloud. That doesn't mean fewer humans — it means humans focus on judgment calls the platform can't yet make, and every judgment call trains the platform to do it next time. In this L1 role you cover front-line monitoring and incident response for NeoCloud's US GPU DCs during the 8AM–8PM PST shift. You execute SOPs, elevate the hard cases, and feed the AIOps substrate the ground truth it needs to learn from novel incidents. What you'll own Monitor GPU cluster health, network status, storage systems, and environmental sensors via centralized dashboards. Respond to alerts and execute runbooks for common incidents: GPU errors, link flaps, node failures, storage alerts. Perform hardware triage: identify failed GPUs, NICs, PSUs, disks, and cables from monitoring data and physical inspection. Execute standard remediation: GPU reset, node drain/reboot, link re-seat, BMC recovery. Collect diagnostic data for L2/SME escalation: logs, DCGM output, network diagnostics, hardware health reports. Manage incident tickets from creation through resolution or escalation (ServiceNow/Jira). Perform physical DC tasks: cable installation, hardware swap-outs, rack and stack, labeling (on-site roles). Execute structured shift handoffs at 8AM and 8PM PST with the APAC operations team. Maintain and update operational runbooks based on recurring issues. Assist with hardware deployment, firmware updates, and inventory management under SME guidance. Feed the AIOps substrate Every novel incident you resolve is data the platform team needs — you tag it, describe it, and hand it back so it becomes an automation. Every runbook you touch should get closer to being executable by the platform, not by you. Your handoff notes are structured signal, not free-form email. Why this role is different from a NOC job You are not the last line of defense — the platform is. You are the training signal. Growth path is real: strong L1s here move into SME roles, or into the platform team as automation authors. Job Requirement: 2+ years in NOC, data center operations, or IT support role Basic Linux system administration (command line, log analysis, service management) Familiarity with monitoring tools (Prometheus, Grafana, Nagios, or equivalent) Experience with ticketing systems (ServiceNow, Jira Service Management) #J-18808-Ljbffr Bitdeer

Vacancy posted 19 hours ago
Similar jobs that could be interesting for youBased on the SRE L1 Support/Cloud Platform Ops Engineer in San Jose, CA vacancy
  • Bitdeer Technologies Group is seeking an L1 NOC/US Data Center operator to support NeoCloud's GPU DCs during 8AM-8PM PST shifts. You will monitor GPU clusters, networks, and storage, respond to alerts, and execute runbooks for common incidents across shore-to-APAC handoffs... 
    Cloud
    Shift work
    Night shift

    Bitdeer (NASDAQ: BTDR)

    San Jose, CA
    3 days ago
  • $94k - $130k

    8-12+ years of experience in SRE, platform engineering, infrastructure, distributed systems, or cloud operations. Demonstrated technical leadership across multi-team...  ...technical strategy for Platform Services work supporting FedRAMP High and IL5. Own key architectural decisions... 
    Cloud
    Permanent employment

    Tata Consultancy Services

    San Jose, CA
    3 days ago
  • Bitdeer Technologies Group seeks an L1 NOC/Operations associate to monitor GPU data center health and respond to alerts. You will...  ...NeoCloud US GPU DCs, logging findings to build automations for the platform and aiding inventory and firmware tasks as directed. #J-18808-... 
    Cloud
    Shift work
    Night shift

    Bitdeer

    San Jose, CA
    3 days ago
  •  ...Inc. is seeking a Hybrid Technical Support Engineer with a focus on Site Reliability Engineering...  ...scale an AI Security Public SaaS platform. You will monitor and troubleshoot...  ...to work across Web technologies, cloud architectures, and SRE practices in a dynamic, #J-18808-... 
    Cloud

    F5 Networks, Inc.

    San Jose, CA
    3 days ago
  •  ...responsibilities of a Technical Support Engineer within a SaaS (Software as a...  ...Reliability Engineering (SRE).The ideal candidate has a strong...  ...an AI Security Public SaaS platform, operating AI inference workloads...  ...of SaaS environments and cloud-based architectures (preferably... 
    Cloud
    Full time
    Local area

    F5 Networks

    San Jose, CA
    1 day ago
  •  ...infrastructure to support the AI revolution....  ...also offers advanced cloud capabilities to customers...  ...compute, run by a platform that observes,...  ...operates the fleet. The SRE Platform team...  ...network, GPU, K8S, and L1 operators —...  ...entry-level Software Engineer on the SRE / Monitoring... 
    Cloud
    Remote job
    Full time
    Contract work
    Temporary work
    Internship
    Local area

    Bitdeer Technologies Group

    San Jose, CA
    a month ago
  • $98.9k - $228.7k

     ...and evolving our observability platform, working across Kubernetes, Terraform, and cloud infrastructure to build the systems...  .... This is a hands-on, on-call engineering role: you will own the systems...  ...a production contextBackground supporting 24/7 or mission-critical... 
    Cloud
    Permanent employment
    Full time
    Work at office
    Remote work

    Zoom

    San Jose, CA
    1 day ago
  • $120k - $180k

     ...most advanced AI-native platform. We work on large scale...  ...starts with you.Sr. SRE & DevOps EngineerAbout...  ...Role:At CrowdStrike, our engineering organization depends on...  ...infrastructure spanning multiple cloud providers and regions,...  ...and collaborate - Support engineering teams with... 
    Cloud
    Full time
    Work experience placement
    Work at office
    Local area

    CrowdStrike

    Sunnyvale, CA
    1 day ago
  • $170k - $277k

     ...when it’s needed. This model supports real-time problem-solving,...  ...visionary Senior Principal Engineer/Architect to serve as the technical...  ...authority for our global SRE and Platform Engineering initiatives...  ...pipelines scale seamlessly across cloud-native environments.... 
    Cloud
    Full time
    Work at office
    Visa sponsorship
    Work visa
    Flexible hours

    Palo Alto Networks, Inc.

    Santa Clara, CA
    1 day ago
  • $101k - $161k

     ...leader in data-driven, client-to-cloud networking for large data...  ...prestigious awards, such as Best Engineering Team, Best Company for...  ...-as-a-Service (CVaaS) global SRE team. SREs at Arista combine...  ...Familiarity with GCP (Google Cloud Platform) and GKE (Google Kubernetes... 
    Cloud

    Arista Networks

    Santa Clara, CA
    4 hours ago
  • $152k - $190k

     ...security data lake to power our cloud-native Zero Trust Exchange platform. This innovation protects...  ...for a Staff Software Engineer (Service Platform &...  ...capabilities and AI-driven SRE practices for a global fleet...  ...are built to last and support a high-growth, global organization... 
    Cloud
    Full time
    Temporary work
    Work at office
    Local area

    Zscaler

    San Jose, CA
    3 days ago
  • $163.5k - $212.4k

     ...experienced kernel or hypervisor engineer who wants to work hands-on...  ...of NIO’s in-vehicle compute platform. You will join the core...  ...trusted execution (TrustZone, OP-TEE). In collaboration with AI and cloud teams, you’ll also support emerging LLM-based... 
    Cloud
    Full time
    Temporary work
    Flexible hours

    NIO USA, INC

    San Jose, CA
    3 days ago
  • $100k - $170k

     ...We are seeking a DevOps Engineer who is eager to have an...  ..., testing, and release platforms, taking us from square...  ...using Terraform for AWS cloud resourcesDevelop and...  ...scaling strategies to support mission-critical operationsAutomate...  ...years of experience in SRE, DevOps, or Platform... 
    Cloud
    Full time
    Work at office
    Immediate start
    Visa sponsorship
    Night shift

    eSpace

    Saratoga, CA
    1 day ago
  • $200k - $322k

     ...in A.I, video games, cloud and enterprise computing...  ...autonomous driving platforms of some of the world’s...  ...Management with solid engineering background who will be...  ...deployment.Experience supporting or leading data campaign...  ...domains.Exposure to ML Ops, data operations,... 
    Cloud
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • Lambda, The Superintelligence Cloud, is a leader in AI cloud...  ...home day is currently Tuesday.Engineering at Lambda is responsible for...  ...networking, and RBAC across the platform.Lead incident response, root-...  ...Platform, Infrastructure, or SRE roles, including running Kubernetes... 
    Cloud
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    4 hours ago
  • $100k

     ...software models, compilers, platforms, networking, and...  ...backend or infrastructure engineer with experience building...  ...frameworks.Collaborate with SRE, infrastructure, and deployment teams to support large-scale on-prem and...  ...differs from cloud-native environments at scale... 
    Cloud
    Permanent employment

    Tenstorrent

    Santa Clara, CA
    4 hours ago
  •  ...looking to hire a Senior Infrastructure, SRE & AI Platforms Manager to help set the long-term...  ...managing large, globally distributed engineering teams. Description In this role, you...  ...technical talent across global sites. Cloud & Distributed Compute Expertise: Demonstrated... 
    Cloud

    Socket.dev

    Cupertino, CA
    3 days ago
  • $153k - $230k

     ...storage pioneer to data platform, closing fiscal 2026...  ...ROLE As a Technical Support Engineer for Portworx® by...  ...public, private, and hybrid cloud environments. You will...  ...technical support work/SRE Container & Cloud...  ...this, contact us at TA-Ops@purestorage.com if you’... 
    Cloud
    Full time
    Work at office
    Flexible hours

    Everpure

    Santa Clara, CA
    4 days ago
  • $190.9k - $334.1k

     ...DescriptionIt all started when engineer Fred Luddy wrote code...  .... Our ServiceNow AI platform brings together any AI,...  ...Reliability Engineer - SRE & AIOps to drive...  ...elimination across our hybrid cloud and data center...  ...operational practices that support high-velocity application... 
    Cloud
    Work at office
    Immediate start
    Remote work
    Flexible hours

    ServiceNow

    Santa Clara, CA
    1 day ago
  • $96.8k - $306.4k

    Drives cross-group platform initiatives (e.g., identity...  ...implementations in partnership with SRE and security.Only...  ...the way in AI and cloud solutions that impact...  ...competitive benefits that support our people with...  ...development lifecycle; coaches engineers across teams or units... 
    Cloud
    Temporary work
    Flexible hours
    Shift work

    Oracle Corporation

    Santa Clara, CA
    3 days ago
  • $88.4k - $143k

    A leading cybersecurity firm in Santa Clara is seeking a Technical Support Engineer for its Cortex Cloud Compute team. This role involves providing technical support to customers, managing complex issues, and collaborating with multi-functional teams. Candidates should... 
    Cloud

    Palo Alto Networks

    Santa Clara, CA
    2 days ago
  • $207k - $300k

     ...team of Site Reliability Engineerings (SREs) and security...  ...technical mentorship while supporting the team's expansion...  ..., and enterprise Cloud customers.Extensive background...  ...Engineering (SRE) combines software and...  ...generation of Google platforms, we make Google's product... 
    Cloud
    Shift work

    Google

    Sunnyvale, CA
    4 hours ago
  •  ...Job Description Job Description Java SRE Engineer Onsite San Francisco Bay Area Infrastructure Engineer (2...  ...Positions) We are looking for an experienced Java SRE / Platform Engineer to support large-scale cloud migrations and production systems on AWS and... 
    Cloud

    Eitacies Inc

    Santa Clara, CA
    5 days ago
  • $272k - $431.25k

     ...100% hands-on Storage Services Software engineer to join the block storage group. You will...  ...of what is possible today and define the platform of tomorrow.At NVIDIA, we work, think and...  ...and reuse existing solutions.Knowledge of cloud computing concepts, including virtualization... 
    Cloud
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  •  ...delivering an AI-powered platform that governs and secures...  ...complex, distributed, cloud-native systems. As a Staff Platform Engineer, you will play a critical...  ...technical guidance and support   Participate in on-...  ...of experience as a Staff SRE with a strong focus on building... 
    Cloud

    Saviynt

    Milpitas, CA
    5 days ago
  • $215.2k - $245.6k

    Lead AI Engineer (Gen AI Platform Services: Agentic AI, Agent Guardrails, Agent Evaluation, Agent Memory...  ...Design, develop, test, deploy, and support AI software components including foundation...  ...and responsible AI solutions on cloud platforms (e.g. AWS, Google Cloud, Azure... 
    Cloud
    Full time
    Part time
    Local area

    Capital One

    San Jose, CA
    3 days ago
  •  ...supercomputing, high-performance computing, cloud, and AI. Whether you’re designing next-...  ...CPU, GPU, AI, and adaptive computing platforms. This technology provides software-executable...  ...a Virtual Platform Functional Modeling Engineer, you will play a key role in defining... 
    Cloud

    AMD

    San Jose, CA
    1 day ago
  • $140k - $215k

     ...world’s most advanced AI-native platform. We work on large scale...  ...CrowdStrike, Site Reliability Engineering (SRE) is at the forefront of...  ...reliability and scalability of our cloud-native security platform. In...  ...empowered to succeed. We support veterans and individuals with... 
    Cloud
    Full time
    Work experience placement
    Work at office
    Local area

    CrowdStrike

    Sunnyvale, CA
    2 days ago
  • Senior Platform Engineer (Embedded Linux) Senior Platform Engineer (Embedded Linux) at Matic Robots Company Overview Each year, 2.5 trillion...  ...data processing performed by the robot itself, not in the cloud. Our Approach Before the iPhone, consumers adopted several... 
    Cloud
    Immediate start
    Work from home

    Matic Robots

    Mountain View, CA
    3 days ago
  •  ...Client is currently seeking multiple DevOps Support Engineers to join our team in Santa Clara, CA. As a member of the customer success team...  ...troubleshooting the various issues that our clients face with our Client’s Cloud Center product suite. You will also be a key point of... 
    Cloud

    BayInfotech

    Santa Clara, CA
    5 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to SRE L1 Support/Cloud Platform Ops Engineer. Be the first to apply!