GPU Cloud Ops SRE (L1) - Incident Response & Automation
Bitdeer
Bitdeer Technologies Group seeks an L1 NOC/Operations associate to monitor GPU data center health and respond to alerts. You will execute runbooks, triage hardware, and escalate complex incidents to SME teams as needed. You will work within 8AM–8PM PST shifts for NeoCloud US GPU DCs, logging findings to build automations for the platform and aiding inventory and firmware tasks as directed. #J-18808-Ljbffr Bitdeer
Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the GPU Cloud Ops SRE (L1) - Incident Response & Automation in San Jose, CA vacancy
- ...also offers advanced cloud capabilities to customers... ...that turns novel incidents into new automations. NeoCloud is building an AI-operated GPU cloud. That doesn't mean... ...it next time. In this L1 role you cover front-... ...monitoring and incident response for NeoCloud's US GPU...CloudShift workNight shift
- Bitdeer Technologies Group is seeking an L1 NOC/US Data Center operator to support NeoCloud's GPU DCs during 8AM-8PM PST shifts. You will monitor GPU clusters... ...to alerts, and execute runbooks for common incidents across shore-to-APAC handoffs. Your role includes...CloudShift workNight shift
$190.9k - $334.1k
...Luddy wrote code that automated a tedious task for his... ...Reliability Engineer - SRE & AIOps to drive infrastructure... ...across our hybrid cloud and data center... ...intervention, accelerate incident remediation, and enable... ...enable data-driven incident response.Establish SLO frameworks...CloudWork at officeImmediate startRemote workFlexible hours$120k
...managed services, enterprise AI and automation frameworks, cloud and hybrid infrastructure support,... ...unmistakably people‑driven. Key Responsibilities: ~ Serve as the first line of... ...Monitor and respond to facility-related incidents, including: ~ High humidity ~...CloudPermanent employment$208.8k
Responsibilities Team Intro TikTok video system is a world-leading... ...ownership models, incident response protocols,... ...for observability and automation initiatives that enhance... ...comprehensive knowledge of cloud-native architectures... ...design/development or SRE experience in video...CloudTemporary workLocal areaShift work- ...also offers advanced cloud capabilities to... ...building an AI-operated GPU cloud — a global... ...the fleet. The SRE Platform team... ...the monitoring and automation substrate that every... ...network, GPU, K8S, and L1 operators —... ...context. Key Responsibilities Where you'll...CloudRemote jobFull timeContract workTemporary workInternshipLocal area
- ...environment, and collaborate with engineering to automate and improve reliability. The role emphasizes automation, incident resolution, and customer-centric support,... ...to work across Web technologies, cloud architectures, and SRE practices in a dynamic, #J-18808-Ljbffr...Cloud
- ...Bitdeer also offers advanced cloud capabilities to... ...building an AI-operated GPU cloud where East-West bandwidth... .... Convert every incident into a labeled example... ...for telemetry-driven ops — you've either built dashboards... ...today should become automation next quarter. -----...CloudFull timeLocal area
- ...The Superintelligence Cloud, is a leader in AI cloud... .... One person, one GPU.If you'd like to build... ...quality, because everything automated downstream depends on... ...top-talent engineers responsible for the deployment and... ...participation in our Incident Management and Review...CloudWork at officeLocal areaWork from homeFlexible hours
- ...operations. Bitdeer also offers advanced cloud capabilities to customers with high... ...our AI cloud to the world — and drive the automation that turns EVPN/BGP ops from tickets into policy. NeoCloud is building an AI-operated GPU cloud spanning 4 US DCs, APAC sites, and...CloudFull timeLocal area
$103k - $154k
...team. In this role, you will monitor security alerts, conduct incident response, and collaborate with senior analysts to enhance security... ...least 5 years of experience in Information Security, emphasizing cloud initiatives. The position offers a hybrid work arrangement...Cloud- McAfee is seeking a Senior SOC Analyst with extensive experience in incident response and security automation, located in San Jose, CA. This role involves analyzing security incidents, mitigating threats, and leading cross-functional initiatives to maintain enterprise security...Cloud
$101k - $161k
...data-driven, client-to-cloud networking for large... ...Service (CVaaS) global SRE team. SREs at Arista combine... ...at scale. We are responsible for our global... ...believe in building highly automated and self-sustaining environments... ...sustainable incident response and blameless...Cloud- ...across any environment—on-premises, in the cloud, at the edge, or in sovereign data... ...driven workloads, Mirantis delivers the automation, GPU orchestration, and policy-driven control... ..., demo and benchmark environments, RFP response libraries. Set and hold a technical...CloudFull time
- ...PositionThe Opportunity: The Global Security Monitoring and Incident Response (MIR) team at Roche strives to keep our networks and users safe... ...signs of abuse or compromise of on-premise as well as cloud resources. All team members share a set of core responsibilities...CloudFull timeWeekend work
- ...times faster than GPU-based hyperscale cloud inference... ...high-performance SRE function to support... ...and operational automation. This role starts... ...reliability from an ops-only burden to a... ...support critical incident escalations, and... ...observability, incident response, and SLO-based...CloudShift work
- ...Superintelligence Cloud, is a leader in AI... ...superintelligence. One person, one GPU.If you'd like to... ...seeking a Senior Incident Manager to lead critical incident response across our AI data... ...implementation of automation and observability... ...frameworks (ITIL, SRE, or equivalent)...CloudWork at officeLocal areaFlexible hours
$256k - $414k
...the global leader in cloud gaming, dedicated to making... ...networking for GPU-based cloud infrastructure... ...frameworks to automate provisioning, scaling,... ...hardware vendors, and SRE groups to influence technology... ...tolerance and lead incident response and root cause analysis...CloudFull timeLocal area- ...Work with your global SRE team to optimize operations... ...in our use of cloud resources and our developer... ...managing SLOsWrite code and automation to reduce operational... ...analysis meetings for incidents to learn and drive improvement... ...global IRC (incident response coordination) for all...CloudFlexible hours
- ...10 times faster than GPU-based hyperscale cloud inference services. This... ...a high-performance SRE function to support... ...understand their pain points, automate their toil, and... ...reliability from an ops-only burden to a... ...SREs, support critical incident escalations, and use...CloudShift work
- Lambda is building a GPU-accelerated Kubernetes-based AI cloud orchestration platform. The Senior Software Engineer will shape architecture, reliability, and automation for Kubernetes-based infrastructure powering AI workloads at scale. This role requires four days in the...CloudWork at officeWork from home
$168.2k - $310.1k
...skilled and experienced Staff Cyber Incident Responder. This senior role is pivotal in our incident response efforts, providing skilled... ..., tools, and methodologies to automate and improve our incident... ...proactive and iterative hunts through cloud and enterprise networks,...CloudFull timeTemporary workLocal areaWorldwide$148.75k - $361k
...strong background in DevOps/SRE practices, cloud infrastructure management,... ...across AWS and GCP, including GPU/TPU-based training and... ...on-call rotation, leading incident response and root-cause analysis for... ...ML infrastructure through automation, resilience engineering, disaster...CloudWork at officeLocal areaRemote workMonday to ThursdayFlexible hours$120k - $180k
...contribute to a culture of responsible AI adoption,... ...cybersecurity starts with you.Sr. SRE & DevOps EngineerAbout... ...spanning multiple cloud providers and regions,... ...— building automation, hardening security, establishing... ...on-call and incidents - Participate in on-call...CloudFull timeWork experience placementWork at officeLocal area- ...site-builder / network automation platform. This role... ...Infrastructure as Code & Cloud ServicesDesign and... ...Observability, Reliability & Incident ReadinessDesign and... ..., and lifecycle responsibilities.Work closely with software... ...Platform Engineering, SRE, or Infrastructure Engineering...CloudLocal area
$140k - $215k
...contribute to a culture of responsible AI adoption,... ...Reliability Engineering (SRE) is at the forefront of... ...and scalability of our cloud-native security platform... ...with AI-driven automation, moving from reactive... ...operationsLead major incident response and facilitate...CloudFull timeWork experience placementWork at officeLocal area- ...responding to security events from multiple sources, and driving incident response from detection through containment and recovery. The... ...in security monitoring, experience with SIEM and CrowdStrike, and strong cloud security knowledge. #J-18808-Ljbffr Omnissa, LLCCloud
$152k - $248k
...LinkedIn's DC Hardware Automation & Tooling... ...Hardware Capacity, Cloud, Network, Systems,... ...and enterprise DC-Ops/Engineering. The infrastructure... ...build, applying SRE disciplines to... ...maintainability.Responsibilities:Design and build... ...call rotation, lead incident response, and...CloudFor contractorsWork at officeFlexible hours- ...Company, Inc. is looking for a cybersecurity incidents expert to monitor and respond to events across enterprise and multi-cloud environments. You will conduct triage,... ...developing detection use cases and improving SOC automation. You will also perform threat hunting and...Cloud
$120k - $155k
...Senior IT Automation Engineer Note- This role focuses on IT infrastructure... ...hybrid IT environment. Responsibilities Design and implement... ...automation for hybrid cloud and on-prem environments... ...schedule Respond to critical incidents and support after-hours deployments...Cloud
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to GPU Cloud Ops SRE (L1) - Incident Response & Automation. Be the first to apply!
Related searches
- cloud administrator San Jose, CA
- aws cloud San Jose, CA
- vp cloud San Jose, CA
- junior cloud administrator San Jose, CA
- cloud engineer azure San Jose, CA
- senior cloud service delivery manager San Jose, CA
- oracle cloud technical San Jose, CA
- cloud security San Jose, CA
- cloud San Jose, CA
- cloud admin San Jose, CA


