Site Reliability Engineer
GMI Cloud
About GMI
GMI Cloud is a fast-growing, AI-native infrastructure company delivering high-performance GPU compute, inference services, and infrastructure for AI agents.
Following 8x ARR growth, GMI Cloud continues to scale rapidly across the U.S. and APAC. As a Reference Platform NVIDIA Cloud Partner (NCP) and a validated leading NCP across both markets, we power production AI for leading AI-native companies including Fireworks AI, Cartesia, Reflection, and OpenRouter.
From large-scale compute to optimized inference and agentic workloads, GMI Cloud gives AI teams the infrastructure they need to build, deploy, and scale on one unified cloud.
One cloud for compute, inference, and agents.
Role Overview
We are seeking a skilled Site Reliability Engineer to join the GMI Global Infrastructure team. This role is hands-on and critical to ensuring the stability, efficiency, and reliability of the large-scale high performance AI/ML clusters in our data center. The ideal candidate will bring expertise in system-level troubleshooting, AI cluster maintenance, and operational excellence to ensure maximum performance for our infrastructure. Experience with large-scale infrastructure automation is considered a strong plus.
Responsibilities
- Design, implement and maintain scalable AI/ML infrastructure solutions.
- Proactively monitor GPU cluster health, performance and troubleshoot issues across compute, accelerator, and storage systems.
- Automate deployment, configuration and management of infrastructure resources.
- Manage GPU node lifecycle workflows, including provisioning, scaling, maintenance, decommissioning and upgrades of GPU nodes.
- Implement CI/CD pipelines for infrastructure deployment and orchestration.
- Ensure security, compliance and best practices across infrastructure.
- Manage incident response related to Infrastructure resources (GPU, CPU, Storage, Network).
- Handle customer provisioning requests for GPU resources, including onboarding, configuration and troubleshooting; resolve customer service requests related to GPU infrastructure, ensuring high customer satisfaction.
- Stay current with emerging GPU hardware and software technologies, integrating improvements as appropriate.
- Regional/international travel to GMI data center locations.
Qualifications
- Bachelor’s degree in Computer Science or related field.
- Over 3+ years of experience in data center operations, infrastructure, or systems engineering.
- Proven experience in site reliability engineering and infrastructure automation (e.g. Ansible, Terraform)
- Familiarity with containers orchestration platform (e.g. Kubernetes, Nvidia GPU operator, Nvidia Network operator, CNI, CSI) and job scheduling systems (e.g. Slurm).
- Familiarity with Linux system administration and scripting (Python, Bash).
- Familiarity with logging and monitoring tools such as Prometheus, Grafana, Loki.
- Good knowledge of GPU architecture, Nvidia CUDA, NCCL, or related AI/ML frameworks - added advantage.
- Strong troubleshooting skills and ability to analyze system logs and performance metrics.
- Excellent communication and teamwork abilities.
Meeting every qualification is not required—if you’re excited about this role, we’d love to hear from you. We believe diverse perspectives and experiences strengthen our team.
$107.9k - $195.05k
Site Reliability EngineerLocation: Hickam Air Force Base, HawaiiClearance: TS/SCILeidos has an opening for a highly qualified TS/SCI cleared Site Reliability Engineer at Hickam Air Force Base, Hawaii for the Decision Advantage Business Area in Defense sector. This is an...SuggestedFull time$141.8k - $195k
...their best work, grow fast, and bring their full selves to the herd.Why You'll Love This RoleCribl Inc is seeking a Senior Site Reliability Engineer to join our mission where you will unlock the value of all observability data, as we expand our team in the U.S. Cribl...SuggestedRemote work$168k - $200k
...is passionate about creating transformative change in healthcare. What We're Looking For We're looking for a Senior Site Reliability Engineer to join our Data & ML Platform team. You'll be at the forefront of building and operating a resilient, observable, and...Suggested$121.4k - $218.6k
...solve complex challenges? Do you have a passion for automation and building systems that scale? Join our highly skilled Site Reliability Engineering team! Our team designs, develops, and manages applications and infrastructure that support Akamai Cloud's products and...SuggestedWork experience placementWork at office$95k - $171k
.... Opportunities exist to focus on GPU infrastructure, Kubernetes, and ensuring reliability for AI workloads within Akamai's serverless inference platform. As an Site Reliability Engineer II, you will be responsible for: Building and maintaining dashboards, alerts...SuggestedPermanent employmentWork experience placementWork at officeRemote workWork from homeWorldwideFlexible hours$75.7k - $136.3k
...solve complex challenges? Do you have a passion for automation and building systems that scale? Join our highly skilled Site Reliability Engineering team! Our team designs, develops, and manages applications and infrastructure that support Akamai Cloud's products and...Work experience placementWork at office$81.1k - $187k
...Job Description We are looking for a Site Reliability Engineer 3 to support mission-critical cloud services and production operations. The role focuses on improving service reliability, reducing operational risk, automating repetitive tasks, and driving faster detection...Temporary workImmediate startFlexible hoursShift work$169.3k - $304.7k
...in building and maintaining fast, efficient, scalable, and reliable routing software and infrastructure that is responsible... ...growth and stability of our global platform. As a Principal Site Reliability Engineer - Network, you will be responsible for: Architecting,...Work experience placementWork at office$114.6k - $234.6k
...creates durable fixes and preventive controls. Designs performance, reliability, and fault-tolerance improvements for drivers, services, and... ...development lifecycle; provides guidance and coaching to engineers to drive improvements. Utilizes advanced knowledge to develop...Temporary workFlexible hoursShift work$61.9k - $141k
Application Systems EngineerThe Opportunity: As an application engineer and administrator on our team, you’ll be integral to architecting... ...our total benefits by visiting the Resource page on our Careers site and reviewing Our Employee Benefits page.Salary at Booz Allen is...Full timeContract workPart timeWork at officeLocal areaRemote work$73.45k - $132.78k
...career is good business.Leidos is seeking a Lead Distribution Engineer in Oahu, HIwho is passionate about electric utility design engineering... ...Team is the go-to for utilities and mobile operators who need reliable power and telecommunication expertise. We've worked with over 5...Full timeWork at officeLocal areaRelocationRelocation packageFlexible hoursNight shift- ...health, transportation, and citizen services. We’re a nonprofit engineering, applied research, and advanced technology organization working... ...clearance.This position requires a minimum of 3 days a week on-site.Preferred Qualifications:Bachelor’s Degree and 10 years of...InternshipLocal area3 days per week
$83.59k - $150k
...Hemisphere. Join our team on STINGRAI and support mission-critical operations through advanced intelligence, surveillance, cyber, engineering, and all-domain capabilities that strengthen regional security, enable informed decision-making, and help defend our nation today...Full timeWork experience placementWork at officeLocal areaWorldwide- ...#LI-HybridJob RequirementsBachelor's degree or its equivalent in Information Technology, Computer Science, or related computer or engineering field and three (3) years of systems analysis and programming experience. May substitute a higher level of degree in Information...Work experience placement
- ...#LI-HybridJob RequirementsBachelor's degree or its equivalent in Information Technology, Computer Science, or related computer or engineering field and five years of systems analysis and programming experience. May substitute a higher level of degree in Information...Temporary workWork experience placement
$125k - $200k
...locate missing children, and more.The RoleForward Deployed Software Engineers (FDSEs) understand our customers’ greatest pain points and... ...Strategists. We also work externally with our customers, often on site, to understand and solve their problems.Trust: We trust each...Full timeWork experience placementWork at officeRemote workWork from homeRelocation package- ...Platform Engineer NeuBird AI is scaling rapidly and we need a platform engineer who can build the internal tools and infrastructure... ...across the engineering organization. You'll partner with SRE on reliability requirements, work with security on compliance needs, and...Flexible hours
$250.6k - $362.6k
...comprehensive security outcomes, as a Principal Engineer. The team delivers secure, scalable... ...networking, with a strong emphasis on reliability, interoperability, and long-term... ...insurance. Please see the Cisco careers site to discover more benefits and perks. Employees...Full timeTemporary workLocal areaRemote workFlexible hours$220k - $270k
...systems long after they leave the lab. As a Senior DevOps Engineer based at Schofield Barracks, you will own what you deploy. That... ...native applications in DDIL environments This role is hybrid/on site at customer sites at Schofield Barracks in Oߵahu, Hawaiߵi. You...Full time$120k - $140k
...Forward Deployment Engineer Fractal Analytics is a strategic AI partner to Fortune 500 companies with a vision to power every human... ...Ensure solutions meet enterprise requirements for security, reliability, maintainability, quality, and production readiness. Troubleshoot...Hourly payFull timeLocal areaRemote work- ...Remote United States Remote United States Full time JR0037944 Job Title: Junior Solutions Engineer About Trellix Trellix is a global company redefining the future of cybersecurity. The company’s comprehensive, open, and native cybersecurity platform...Full timeRemote workFlexible hours
$113.4k - $170.2k
...project solutions to life.Your OpportunityThe role is to work as a team member on various projects under the guidance of a Resident Engineer (RE) or Construction Manager (CM). The role will work on challenging and diverse tasks in a consulting and construction management...Full timeContract workTemporary workPart timeFor contractorsCasual workLocal areaFlexible hours- ...creativity, problem-solving, and impact-driven innovation. Role Description This is a full-time remote role for a Software Engineer. The Software Engineer will be responsible for designing, implementing, testing, and maintaining software applications, with a...Full timeRemote work
$166k - $253k
...Software Engineer, Connected Warfare (Active Clearance) Anduril Industries is a defense technology company with a mission to transform... ...code and design reviews—while championing scalability, reliability, testability, portability, and maintainability. Drive software...Full timeWork experience placementImmediate start$186.07k - $218.9k
...surges.”learn more about working at Coinbase. Senior Software Engineer (EAA) The EAA Compliance CXAE team, part of Coinbase's... ...capabilities that benefit multiple teams across the organization. Own reliability for Tier-1 compliance systems by anticipating potential issues...Local area$100k - $120k
...partnerships, while our focus on creativity and innovative solutions empowers our customer communities to thrive. The Senior Software Engineer on the New Ventures team designs, creates, maintains, audits, and improves software applications by performing coding, debugging,...Full timeTemporary workLocal areaRemote work- ...Software Engineer Anywhere Type: Contract-to-Hire Category: Engineer Industry: Financial Services Workplace Type:... ...reviews. Identify code metrics and perform system risk and reliability analysis. Apply object-oriented design and analysis. Develop...Hourly payContract workLocal areaRemote work
$125k - $150k
...across cloud and on-premises environments while working with technologies from NVIDIA, Microsoft, and Google. This role is ideal for engineers who love solving complex problems, building production software, and learning new technologies in a fast-moving AI landscape....Full timeRemote workWorldwideFlexible hours$223.2k - $279k
...important decisions. On the Public Sector Engineering team, you'll embed directly with U.S.... ...deploy, and maintain solutions at customer sites, from prototype to stable production.... ...Us: At Scale, our mission is to develop reliable AI systems for the world's most important...Full timeRelocation$77.1k - $123.3k
...helps millions of learners improve their lives and achieve their dreams through education. What you'll do here: As a Software Engineer, you will deliver a world-class experience for learners and instructors on our Cengage Learning Platforms (CLP). Working on a...Full timeLive inLocal areaWorldwide
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
- official site Honolulu, HI
- site services specialist Honolulu, HI
- construction site safety Honolulu, HI
- IT site lead Honolulu, HI
- site leader Honolulu, HI
- site safety Honolulu, HI
- historic site Honolulu, HI
- junior website developer Honolulu, HI
- website coordinator Honolulu, HI
- on-site clinical research associate (traveling/remote) Honolulu, HI



