Site Reliability Engineer
Bitdeer
About Bitdeer Technologies Group Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure. Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to support the AI revolution. Bitdeer handles complex processes involved in computing such as equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial intelligence. Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia. To learn more, visit Position Overview You are the first human in the loop — the escalation target when the AIOps system needs a decision, and the source of ground truth that turns novel incidents into new automations. NeoCloud is building an AI-operated GPU cloud. That doesn't mean fewer humans — it means humans focus on judgment calls the platform can't yet make, and every judgment call trains the platform to do it next time. In this L1 role you cover front-line monitoring and incident response for NeoCloud's US GPU DCs during the 8AM–8PM PST shift. You execute SOPs, elevate the hard cases, and feed the AIOps substrate the ground truth it needs to learn from novel incidents. What you'll own Monitor GPU cluster health, network status, storage systems, and environmental sensors via centralized dashboards. Respond to alerts and execute runbooks for common incidents: GPU errors, link flaps, node failures, storage alerts. Perform hardware triage: identify failed GPUs, NICs, PSUs, disks, and cables from monitoring data and physical inspection. Execute standard remediation: GPU reset, node drain/reboot, link re-seat, BMC recovery. Collect diagnostic data for L2/SME escalation: logs, DCGM output, network diagnostics, hardware health reports. Manage incident tickets from creation through resolution or escalation (ServiceNow/Jira). Perform physical DC tasks: cable installation, hardware swap-outs, rack and stack, labeling (on-site roles). Execute structured shift handoffs at 8AM and 8PM PST with the APAC operations team. Maintain and update operational runbooks based on recurring issues. Assist with hardware deployment, firmware updates, and inventory management under SME guidance. Feed the AIOps substrate Every novel incident you resolve is data the platform team needs — you tag it, describe it, and hand it back so it becomes an automation. Every runbook you touch should get closer to being executable by the platform, not by you. Your handoff notes are structured signal, not free-form email. Why this role is different from a NOC job You are not the last line of defense — the platform is. You are the training signal. Growth path is real: strong L1s here move into SME roles, or into the platform team as automation authors. Job Requirement: 2+ years in NOC, data center operations, or IT support role Basic Linux system administration (command line, log analysis, service management) Familiarity with monitoring tools (Prometheus, Grafana, Nagios, or equivalent) Experience with ticketing systems (ServiceNow, Jira Service Management) #J-18808-Ljbffr
- ...and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Chief Data & Analytics Office (CDAO) AI/ML & Data Platforms team,, you will solve complex...SuggestedWork at office
$140k - $150k
WORK OPTION: Remote_________________The NBA is hiring a Senior Site Reliability Engineer (SRE) - Messaging & Collaboration to ensure the availability, performance, and reliability of enterprise messaging and collaboration platforms, including Microsoft Exchange Online (...SuggestedFull timeTemporary workLocal areaRemote workWeekend work$90k - $120k
As a Performance II-Epic, your role is to provide reliability engineering services through observability and performance engineering techniques.... ...passion for optimizing operational efficiency. You will use Site Reliability Engineering practices to deliver a seamless user...SuggestedFull timePart timeWork experience placementRemote workFlexible hours- ...Overview We are seeking an experienced Site Reliability Engineer (SRE) – Microsoft Hyper-V & Private Cloud to operate highly available private cloud and Virtual Desktop Infrastructure (VDI) platforms based on Microsoft Hyper-V. This role combines deep Hyper-V expertise...Suggested
$175k - $185k
...together. Come join our team as we develop new ways to improve the lives of working Americans. About the role: As the Senior Site Reliability Engineer, you will lead Branch’s effort to achieve greater reliability, performance, scalability, capacity and observability of our...SuggestedDaily paidRemote workHome officeFlexible hours- ...We are seeking an experienced Site Reliability Engineer (SRE) – Microsoft Hyper-V & Private Cloud to operate highly available private cloud and Virtual Desktop Infrastructure (VDI) platforms based on Microsoft Hyper-V. This role combines deep Hyper-V expertise with modern...Local area
- ...Goldman Sachs is seeking a Vice President in Compliance Engineering SRE for Dallas. The role combines software and systems engineering... ...monitoring, and collaborate with cross-functional teams to deliver reliable, compliant platforms for regulatory risk management. #J-18808-...
- ...itD is seeking a Site Reliability Engineer to develop and enhance automation solutions that improve the reliability, scalability, and operational efficiency of large-scale cloud infrastructure. The ideal candidate will bring hands-on experience in site reliability engineering...Work experience placementRemote work
- ...is responsible for partnering with leaders across engineering and technology to define objective reliability goals for services. Key responsibilities include composing... .... Position Summary: The Senior Azure Site Reliability Engineer acts as an advanced senior individual...Work at officeShift workDay shift
- ...is a recognized, award-winning leader in supply chain AI and a FedRAMP® authorized provider to the federal government. Site Reliability Engineer Location: U.S. (Hybrid) This role requires U.S. citizenship and eligibility for a U.S. security clearance. Role Summary...Work at officeWork from homeFlexible hours
$155k - $175k
...Next! Summary We are seeking a highly skilled and experienced Site Reliability Manager to join our team to ensure the reliability,... ...performance of our systems and services. You will lead a team of engineers focusing on three core pillars: Application Reliability, DevSecOps...Work experience placementH1bWork at officeLocal area- ...exceptional professionals for this role. JOB DESCRIPTION Elevate your engineering prowess to unprecedented levels by joining a team of... ...and position yourself among the top echelon in site reliability. As an Associate Site Reliability Engineer at JPMorgan Chase...Worldwide
- ...Site Reliability Engineer Location: Schaumburg, IL or Secaucus, NJ (Hybrid) Mandatory Skills: Python/R and ML libraries (scikit-learn, TensorFlow, PyTorch), Data analysis and visualization (Pandas, NumPy, Power BI/Tableau), SQL and database management Key Responsibilities...Local area
- ...Site Reliability Engineer As a Site Reliability Engineer, your role is to provide reliability engineering services through observability and performance engineering techniques. Using monitoring and performance tools to deliver detailed feedback to product owners and...Work experience placement
- ...Site Reliability Engineer III There's nothing more exciting than being at the center of a rapidly growing field in technology and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems. As a Site Reliability Engineer...
- ...globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability. As a Lead Site Reliability Engineer at JPMorgan Chase within the Cloud Foundational Services team, you hold a leadership role in your team, demonstrate...
- ...contributing to revolutionary projects. You've discovered the perfect environment to have a major impact. As a Principal Site Reliability Engineer at JPMorgan Chase within the Corporate Technology Team, you draw upon your advanced knowledge to identify new opportunities...
$130k - $180k
...alongside some of the most experienced and innovative leaders and engineers in the field. Where we work Headquartered in Amsterdam and... ...an in-house AI R&D team. The role Nebius is looking for a Site Reliability Engineer in Hardware Infrastructure team. You’re welcome to...Temporary workWork at officeImmediate startRemote workFlexible hours- ...Partner with software developers, platform engineers, and IT staff to improve system design,... ...requirements, service quality, reliability, security, and compliance needs. Drive continuous... ...Required: 8+ years of experience in Site Reliability Engineering, DevOps, Platform...Work at officeRemote work
- ...exceptional professionals for this role. JOB DESCRIPTION Elevate your engineering prowess to unprecedented levels by joining a team of... ...and position yourself among the top echelon in site reliability. As a Sr Lead Site Reliability Engineer at JPMorgan Chase within...
- ...are looking for people just like you. Join our team and help us develop game-changing, high-quality solutions. As a Lead Site Reliability Engineer at JPMorganChase within the Corporate sector, Enterprise Technology team, you are an integral part of a team that develops...Work at office
$113.1k - $232.3k
Position Summary Lead Applied AI Site Reliability Engineer II Role Overview: As a Lead Applied AI Site Reliability Engineer II, you will actively engage in your engineering craft, taking a hands-on approach to the reliability, performance, and operational integrity...Work at officeLocal areaVisa sponsorshipFlexible hours3 days per week- Elevate your engineering prowess to unprecedented levels by joining a team of exceptionally gifted professionals and position yourself among the top echelon in site reliability.As a Senior Lead Site Reliability Engineer at JPMorgan Chase within the Chief Data & Analytics...Work at office
- ...we serve.The Information Technology group delivers secure, reliable technology solutions that enable DTCC to be the trusted... ...scalability, and performance of enterprise platforms.As a Principal Site Reliability Engineer (SRE), you will drive operational excellence across...Remote workFlexible hours
- Compliance Engineering, Site Reliability Engineering, Vice President, Dallas location_on Dallas, TX, United States We are Compliance Engineering, a global team of more than 300 engineers and scientists who work on the most complex, mission-critical problems. We build and...Full timeTemporary workWork at office
- Jack Henry & Associates, Inc. is seeking a Senior Site Reliability Engineer to drive modernization across a large-scale hybrid cloud footprint, with emphasis on re-architecting on-prem workloads to Google Cloud Platform. The role involves implementing SRE practices, IaC...
$60 - $65 per hour
...SRE Engineer (W2) Jersey City, NJ (Onsite) 6 Months Contract to Hire Job Description: Proficient in application development skills for more than one technology as well as multiple design techniques. Working proficiency in development toolset to design, develop, test,...Full timeContract workWork experience placement- ...storage tanks, water metering, energy metering, gas monitoring, and asset management. Our founders are hardcore telecommunications engineers with combined 200 + years of experience in designing, optimizing and performance engineering; for several mid - large wireless...Work experience placement
- ...storage tanks, water metering, energy metering, gas monitoring, and asset management. Our founders are hardcore telecommunications engineers with combined 200 + years of experience in designing, optimizing and performance engineering; for several mid - large wireless...
$189.59k - $228k
Pershing LLC seeks Senior Vice President, Release Train Engineer in Jersey City, NJ, to perform Agile Release Train (ART) or product. Facilitate development events and processes and assist the teams in delivering value. Communicate with stakeholders, escalate impediments...Temporary workWork at officeRemote workWorldwideFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
- on-site clinical research associate (traveling/remote) Brooklyn, NY
- website coordinator Brooklyn, NY
- junior website developer Brooklyn, NY
- site leader Brooklyn, NY
- historic site Brooklyn, NY
- website content developer Brooklyn, NY
- construction site safety Brooklyn, NY
- official site Brooklyn, NY
- site services specialist Brooklyn, NY
- on site coordinator Brooklyn, NY


