Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

GPU Cloud Ops SRE (L1) - Incident Response & Automation

Bitdeer

Bitdeer Technologies Group seeks an L1 NOC/Operations associate to monitor GPU data center health and respond to alerts. You will execute runbooks, triage hardware, and escalate complex incidents to SME teams as needed. You will work within 8AM–8PM PST shifts for NeoCloud US GPU DCs, logging findings to build automations for the platform and aiding inventory and firmware tasks as directed. #J-18808-Ljbffr Bitdeer

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the GPU Cloud Ops SRE (L1) - Incident Response & Automation in San Jose, CA vacancy
  •  ...also offers advanced cloud capabilities to customers...  ...that turns novel incidents into new automations. NeoCloud is building an AI-operated GPU cloud. That doesn't mean...  ...it next time. In this L1 role you cover front-...  ...monitoring and incident response for NeoCloud's US GPU... 
    Cloud
    Shift work
    Night shift

    Bitdeer

    San Jose, CA
    21 hours ago
  • Bitdeer Technologies Group is seeking an L1 NOC/US Data Center operator to support NeoCloud's GPU DCs during 8AM-8PM PST shifts. You will monitor GPU clusters...  ...to alerts, and execute runbooks for common incidents across shore-to-APAC handoffs. Your role includes... 
    Cloud
    Shift work
    Night shift

    Bitdeer (NASDAQ: BTDR)

    San Jose, CA
    3 days ago
  • $190.9k - $334.1k

     ...Luddy wrote code that automated a tedious task for his...  ...Reliability Engineer - SRE & AIOps to drive infrastructure...  ...across our hybrid cloud and data center...  ...intervention, accelerate incident remediation, and enable...  ...enable data-driven incident response.Establish SLO frameworks... 
    Cloud
    Work at office
    Immediate start
    Remote work
    Flexible hours

    ServiceNow

    Santa Clara, CA
    1 day ago
  • $120k

     ...managed services, enterprise AI and automation frameworks, cloud and hybrid infrastructure support,...  ...unmistakably people‑driven.  Key Responsibilities: ~ Serve as the first line of...  ...Monitor and respond to facility-related incidents, including: ~ High humidity ~... 
    Cloud
    Permanent employment
    San Jose, CA
    more than 2 months ago
  • $208.8k

    Responsibilities Team Intro TikTok video system is a world-leading...  ...ownership models, incident response protocols,...  ...for observability and automation initiatives that enhance...  ...comprehensive knowledge of cloud-native architectures...  ...design/development or SRE experience in video... 
    Cloud
    Temporary work
    Local area
    Shift work

    TikTok USDS Joint Venture

    San Jose, CA
    2 days ago
  •  ...also offers advanced cloud capabilities to...  ...building an AI-operated GPU cloud — a global...  ...the fleet. The SRE Platform team...  ...the monitoring and automation substrate that every...  ...network, GPU, K8S, and L1 operators —...  ...context. Key Responsibilities Where you'll... 
    Cloud
    Remote job
    Full time
    Contract work
    Temporary work
    Internship
    Local area

    Bitdeer Technologies Group

    San Jose, CA
    a month ago
  •  ...environment, and collaborate with engineering to automate and improve reliability. The role emphasizes automation, incident resolution, and customer-centric support,...  ...to work across Web technologies, cloud architectures, and SRE practices in a dynamic, #J-18808-Ljbffr... 
    Cloud

    F5 Networks, Inc.

    San Jose, CA
    3 days ago
  •  ...Bitdeer also offers advanced cloud capabilities to...  ...building an AI-operated GPU cloud where East-West bandwidth...  .... Convert every incident into a labeled example...  ...for telemetry-driven ops — you've either built dashboards...  ...today should become automation next quarter. -----... 
    Cloud
    Full time
    Local area

    Bitdeer

    San Jose, CA
    17 days ago
  •  ...The Superintelligence Cloud, is a leader in AI cloud...  .... One person, one GPU.If you'd like to build...  ...quality, because everything automated downstream depends on...  ...top-talent engineers responsible for the deployment and...  ...participation in our Incident Management and Review... 
    Cloud
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    2 days ago
  •  ...operations. Bitdeer also offers advanced cloud capabilities to customers with high...  ...our AI cloud to the world — and drive the automation that turns EVPN/BGP ops from tickets into policy. NeoCloud is building an AI-operated GPU cloud spanning 4 US DCs, APAC sites, and... 
    Cloud
    Full time
    Local area

    Bitdeer Technologies Group

    San Jose, CA
    18 days ago
  • $103k - $154k

     ...team. In this role, you will monitor security alerts, conduct incident response, and collaborate with senior analysts to enhance security...  ...least 5 years of experience in Information Security, emphasizing cloud initiatives. The position offers a hybrid work arrangement... 
    Cloud

    Arista Networks, Inc.

    Santa Clara, CA
    1 day ago
  • McAfee is seeking a Senior SOC Analyst with extensive experience in incident response and security automation, located in San Jose, CA. This role involves analyzing security incidents, mitigating threats, and leading cross-functional initiatives to maintain enterprise security... 
    Cloud

    McAfee

    San Jose, CA
    2 days ago
  • $101k - $161k

     ...data-driven, client-to-cloud networking for large...  ...Service (CVaaS) global SRE team. SREs at Arista combine...  ...at scale. We are responsible for our global...  ...believe in building highly automated and self-sustaining environments...  ...sustainable incident response and blameless... 
    Cloud

    Arista Networks

    Santa Clara, CA
    7 hours ago
  •  ...across any environment—on-premises, in the cloud, at the edge, or in sovereign data...  ...driven workloads, Mirantis delivers the automation, GPU orchestration, and policy-driven control...  ..., demo and benchmark environments, RFP response libraries. Set and hold a technical... 
    Cloud
    Full time

    Mirantis

    San Jose, CA
    10 days ago
  •  ...PositionThe Opportunity: The Global Security Monitoring and Incident Response (MIR) team at Roche strives to keep our networks and users safe...  ...signs of abuse or compromise of on-premise as well as cloud resources. All team members share a set of core responsibilities... 
    Cloud
    Full time
    Weekend work

    Roche

    San Jose, CA
    2 days ago
  •  ...times faster than GPU-based hyperscale cloud inference...  ...high-performance SRE function to support...  ...and operational automation. This role starts...  ...reliability from an ops-only burden to a...  ...support critical incident escalations, and...  ...observability, incident response, and SLO-based... 
    Cloud
    Shift work

    Cerebras Systems

    Sunnyvale, CA
    7 hours ago
  •  ...Superintelligence Cloud, is a leader in AI...  ...superintelligence. One person, one GPU.If you'd like to...  ...seeking a Senior Incident Manager to lead critical incident response across our AI data...  ...implementation of automation and observability...  ...frameworks (ITIL, SRE, or equivalent)... 
    Cloud
    Work at office
    Local area
    Flexible hours

    Lambda Labs

    San Jose, CA
    3 days ago
  • $256k - $414k

     ...the global leader in cloud gaming, dedicated to making...  ...networking for GPU-based cloud infrastructure...  ...frameworks to automate provisioning, scaling,...  ...hardware vendors, and SRE groups to influence technology...  ...tolerance and lead incident response and root cause analysis... 
    Cloud
    Full time
    Local area

    Nvidia

    Santa Clara, CA
    7 hours ago
  •  ...Work with your global SRE team to optimize operations...  ...in our use of cloud resources and our developer...  ...managing SLOsWrite code and automation to reduce operational...  ...analysis meetings for incidents to learn and drive improvement...  ...global IRC (incident response coordination) for all... 
    Cloud
    Flexible hours

    Sumo Logic

    San Jose, CA
    7 hours ago
  •  ...10 times faster than GPU-based hyperscale cloud inference services. This...  ...a high-performance SRE function to support...  ...understand their pain points, automate their toil, and...  ...reliability from an ops-only burden to a...  ...SREs, support critical incident escalations, and use... 
    Cloud
    Shift work

    Cerebras Systems

    Sunnyvale, CA
    1 day ago
  • Lambda is building a GPU-accelerated Kubernetes-based AI cloud orchestration platform. The Senior Software Engineer will shape architecture, reliability, and automation for Kubernetes-based infrastructure powering AI workloads at scale. This role requires four days in the... 
    Cloud
    Work at office
    Work from home

    Front Door Defense

    San Jose, CA
    5 days ago
  • $168.2k - $310.1k

     ...skilled and experienced Staff Cyber Incident Responder. This senior role is pivotal in our incident response efforts, providing skilled...  ..., tools, and methodologies to automate and improve our incident...  ...proactive and iterative hunts through cloud and enterprise networks,... 
    Cloud
    Full time
    Temporary work
    Local area
    Worldwide

    Adobe Systems

    San Jose, CA
    3 days ago
  • $148.75k - $361k

     ...strong background in DevOps/SRE practices, cloud infrastructure management,...  ...across AWS and GCP, including GPU/TPU-based training and...  ...on-call rotation, leading incident response and root-cause analysis for...  ...ML infrastructure through automation, resilience engineering, disaster... 
    Cloud
    Work at office
    Local area
    Remote work
    Monday to Thursday
    Flexible hours

    Roku

    San Jose, CA
    2 days ago
  • $120k - $180k

     ...contribute to a culture of responsible AI adoption,...  ...cybersecurity starts with you.Sr. SRE & DevOps EngineerAbout...  ...spanning multiple cloud providers and regions,...  ...— building automation, hardening security, establishing...  ...on-call and incidents - Participate in on-call... 
    Cloud
    Full time
    Work experience placement
    Work at office
    Local area

    CrowdStrike

    Sunnyvale, CA
    1 day ago
  •  ...site-builder / network automation platform. This role...  ...Infrastructure as Code & Cloud ServicesDesign and...  ...Observability, Reliability & Incident ReadinessDesign and...  ..., and lifecycle responsibilities.Work closely with software...  ...Platform Engineering, SRE, or Infrastructure Engineering... 
    Cloud
    Local area

    Omni Inclusive

    Santa Clara, CA
    2 days ago
  • $140k - $215k

     ...contribute to a culture of responsible AI adoption,...  ...Reliability Engineering (SRE) is at the forefront of...  ...and scalability of our cloud-native security platform...  ...with AI-driven automation, moving from reactive...  ...operationsLead major incident response and facilitate... 
    Cloud
    Full time
    Work experience placement
    Work at office
    Local area

    CrowdStrike

    Sunnyvale, CA
    2 days ago
  •  ...responding to security events from multiple sources, and driving incident response from detection through containment and recovery. The...  ...in security monitoring, experience with SIEM and CrowdStrike, and strong cloud security knowledge. #J-18808-Ljbffr Omnissa, LLC
    Cloud

    Omnissa, LLC

    Mountain View, CA
    4 days ago
  • $152k - $248k

     ...LinkedIn's DC Hardware Automation & Tooling...  ...Hardware Capacity, Cloud, Network, Systems,...  ...and enterprise DC-Ops/Engineering. The infrastructure...  ...build, applying SRE disciplines to...  ...maintainability.Responsibilities:Design and build...  ...call rotation, lead incident response, and... 
    Cloud
    For contractors
    Work at office
    Flexible hours

    LinkedIn

    Mountain View, CA
    1 day ago
  •  ...Company, Inc. is looking for a cybersecurity incidents expert to monitor and respond to events across enterprise and multi-cloud environments. You will conduct triage,...  ...developing detection use cases and improving SOC automation. You will also perform threat hunting and... 
    Cloud

    McKinsey & Company, Inc.

    San Jose, CA
    2 days ago
  • $120k - $155k

     ...Senior IT Automation Engineer Note- This role focuses on IT infrastructure...  ...hybrid IT environment. Responsibilities Design and implement...  ...automation for hybrid cloud and on-prem environments...  ...schedule Respond to critical incidents and support after-hours deployments... 
    Cloud

    Omni Vision Inc

    Santa Clara, CA
    5 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to GPU Cloud Ops SRE (L1) - Incident Response & Automation. Be the first to apply!