Site Reliability Engineer
GMI Cloud
About GMI
GMI Cloud is a fast-growing, AI-native infrastructure company delivering high-performance GPU compute, inference services, and infrastructure for AI agents.
Following 8x ARR growth, GMI Cloud continues to scale rapidly across the U.S. and APAC. As a Reference Platform NVIDIA Cloud Partner (NCP) and a validated leading NCP across both markets, we power production AI for leading AI-native companies including Fireworks AI, Cartesia, Reflection, and OpenRouter.
From large-scale compute to optimized inference and agentic workloads, GMI Cloud gives AI teams the infrastructure they need to build, deploy, and scale on one unified cloud.
One cloud for compute, inference, and agents.
Role Overview
We are seeking a skilled Site Reliability Engineer to join the GMI Global Infrastructure team. This role is hands-on and critical to ensuring the stability, efficiency, and reliability of the large-scale high performance AI/ML clusters in our data center. The ideal candidate will bring expertise in system-level troubleshooting, AI cluster maintenance, and operational excellence to ensure maximum performance for our infrastructure. Experience with large-scale infrastructure automation is considered a strong plus.
Responsibilities
- Design, implement and maintain scalable AI/ML infrastructure solutions.
- Proactively monitor GPU cluster health, performance and troubleshoot issues across compute, accelerator, and storage systems.
- Automate deployment, configuration and management of infrastructure resources.
- Manage GPU node lifecycle workflows, including provisioning, scaling, maintenance, decommissioning and upgrades of GPU nodes.
- Implement CI/CD pipelines for infrastructure deployment and orchestration.
- Ensure security, compliance and best practices across infrastructure.
- Manage incident response related to Infrastructure resources (GPU, CPU, Storage, Network).
- Handle customer provisioning requests for GPU resources, including onboarding, configuration and troubleshooting; resolve customer service requests related to GPU infrastructure, ensuring high customer satisfaction.
- Stay current with emerging GPU hardware and software technologies, integrating improvements as appropriate.
- Regional/international travel to GMI data center locations.
Qualifications
- Bachelor’s degree in Computer Science or related field.
- Over 3+ years of experience in data center operations, infrastructure, or systems engineering.
- Proven experience in site reliability engineering and infrastructure automation (e.g. Ansible, Terraform)
- Familiarity with containers orchestration platform (e.g. Kubernetes, Nvidia GPU operator, Nvidia Network operator, CNI, CSI) and job scheduling systems (e.g. Slurm).
- Familiarity with Linux system administration and scripting (Python, Bash).
- Familiarity with logging and monitoring tools such as Prometheus, Grafana, Loki.
- Good knowledge of GPU architecture, Nvidia CUDA, NCCL, or related AI/ML frameworks - added advantage.
- Strong troubleshooting skills and ability to analyze system logs and performance metrics.
- Excellent communication and teamwork abilities.
Meeting every qualification is not required—if you’re excited about this role, we’d love to hear from you. We believe diverse perspectives and experiences strengthen our team.
$62k - $141k
Site Reliability EngineerThe Opportunity: Engineering to make a system more resilient and efficient frees up time and money to build more capabilities. Whether you come from a background in network engineering, systems administration, or software development, if you have...SuggestedFull timeContract workPart timeLocal areaRemote work$130k - $200k
...Site Reliability Engineer (Aurora, CO; Herndon, VA) Summary Position Title: Site Reliability Engineer Position ID: TA247 Location(s): On-site; Aurora, CO; Herndon, VA Application Deadline: September 30, 2026 Security Clearance Requirement: An active TS/SCI...SuggestedTemporary workLocal area$135k - $155k
...while also making it easy for buyers at Fortune 1000 companies to tap into global manufacturing capacity.Xometry is seeking a Site Reliability Engineer II to join our Site Reliability Engineering (SRE) Organization. In this role as an individual contributor, you will guide...SuggestedFlexible hours$114.4k - $125.4k
...excellence of the Salesforce GovCloud! Are you passionate about ensuring the reliability and performance of mission-critical cloud services? Salesforce is seeking a talented Site Reliability Engineer to join our dynamic team, supporting our GovCloud environment. As a key...SuggestedFull timeLocal areaShift workNight shift$98.58k - $138.02k
...This role requires a hybrid work schedule based out of one of our office locations: Austin, TX; Irvine, CA; or Akron, OH. Site Reliability Engineer II will be responsible for supporting, enhancing, and maintaining Restaurant365’s cloud infrastructure and applications....SuggestedWork at office$110k - $125k
...The Site Reliability Engineer maintains and improves the availability, performance, resilience, and operational recoverability of enterprise identity, credential, and access-management services. This role supports cloud identity and directory platforms, including Microsoft...Contract workWork at office- ...profitable developer-tooling company whose product is used by engineering teams at thousands of software companies for application... ...well-resourced group of nine. As Senior SRE you will lead reliability initiatives across the platform — from defining and driving SLOs...
- ...Protecting others requires a team that works together with trust and cares deeply about carrying out our mission. The Site Reliability Engineering team at Todyl exists to make our platform reliable, secure, and easy for engineering teams to ship to. We do that by...Full timeTemporary workLocal areaFlexible hoursShift work
$95k - $134k
...helped build. For more information, visit Job Application Deadline: 10/31/2026 The Opportunity DAT is looking for a Site Reliability Engineer to join our SRE platform team. This position will work hybrid in Denver, CO Candidate profile DAT is seeking an...Temporary workFor contractorsWork experience placementWork at officeLocal areaImmediate startFlexible hours- ...Site Reliability Engineer - Greenwood Village, CO (Hybrid) About the Role: Join a forward-thinking engineering team as a Site Reliability Engineer, specializing in enterprise-scale experimentation and configuration management platforms. In this operations-focused...Contract work
- ...focusing on private cloud systems supporting 5G wireless systems. This position will focus on platform monitoring, logging, and reliability aspects supporting the Mobile Core team. A critical goal is to gather metrics of the platform during stress and load events to ensure...
$95k - $171k
.... Opportunities exist to focus on GPU infrastructure, Kubernetes, and ensuring reliability for AI workloads within Akamai's serverless inference platform. As an Site Reliability Engineer II, you will be responsible for: Building and maintaining dashboards, alerts...Permanent employmentWork experience placementWork at officeRemote workWork from homeWorldwideFlexible hours$141.8k - $195k
...their best work, grow fast, and bring their full selves to the herd.Why You'll Love This RoleCribl Inc is seeking a Senior Site Reliability Engineer to join our mission where you will unlock the value of all observability data, as we expand our team in the U.S. Cribl...Remote work$168k - $200k
...is passionate about creating transformative change in healthcare. What We're Looking For We're looking for a Senior Site Reliability Engineer to join our Data & ML Platform team. You'll be at the forefront of building and operating a resilient, observable, and...- ...We are seeking an experienced Site Reliability Engineer (SRE) to join the Applied AI and Data Science program. This role focuses on deploying, monitoring, and optimizing cloud-based applications and infrastructure to ensure high availability and performance. The...
$81.1k - $187k
...Job Description We are looking for a Site Reliability Engineer 3 to support mission-critical cloud services and production operations. The role focuses on improving service reliability, reducing operational risk, automating repetitive tasks, and driving faster detection...Temporary workImmediate startFlexible hoursShift work- ...At Todyl, our Application Platform Engineering team is dedicated to building infrastructure... ...work will not only directly impact the reliability and security of our platform but also empower... ...to grow our team and is hiring two Site Reliability Engineers (SRE I and SRE II)...Local area
$121.4k - $218.6k
...solve complex challenges? Do you have a passion for automation and building systems that scale? Join our highly skilled Site Reliability Engineering team! Our team designs, develops, and manages applications and infrastructure that support Akamai Cloud's products and...Work experience placementWork at office$140k - $165k
...Summary of Job ResponsibilitiesThe Senior Site Reliability Engineer (SRE) is Colorado PERA’s technical owner for production reliability, observability, and platform automation across a hybrid AWS, Azure and on-premises environment. This position leads the design, implementation...Full timeWork at officeWork from homeAfternoon shift3 days per week$127k - $249k
Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions... ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper). As...Work at officeLocal areaRemote workWorldwideFlexible hours$192.4k - $275.8k
...CloudOps— the team that keeps Splunk Cloud running for some of the world's most demanding enterprise customers, blending Site Reliability Engineering, Systems Engineering, and Service Engineering disciplines at a scale very few teams ever get to operate at. When the...Full timeTemporary workLocal areaFlexible hours$110k - $155k
...global, with headquarters in Denver, Colorado, and offices across the U.S., Canada, and India. We are seeking a Senior Site Reliability Engineer to own the reliability, scalability, performance, and operational integrity of critical production services. This role is...Contract workWork at officeWork from homeFlexible hours$160k - $180k
...global, with headquarters in Denver, Colorado, and offices across the U.S., Canada, and India. We are seeking a Principal Site Reliability Engineer to define the strategic vision and own the enterprise‑wide reliability, scalability, and performance of our critical...Contract workTemporary workWork at officeWork from homeFlexible hours$175k - $220k
...with headquarters in Denver, Colorado, and offices across the U.S., Canada, and India. Role Summary The Director, Site Reliability Engineering (SRE) will lead reliability, performance, and observability initiatives for a portfolio of Vertafore products. This role...Contract workTemporary workWork at officeWork from homeFlexible hours$169.3k - $304.7k
...in building and maintaining fast, efficient, scalable, and reliable routing software and infrastructure that is responsible... ...growth and stability of our global platform. As a Principal Site Reliability Engineer - Network, you will be responsible for: Architecting,...Work experience placementWork at office$130k - $200k
Summary Position Title: Site Reliability Engineer Position ID: TA247 Location(s): On-site; Aurora, CO; Herndon, VA Application Deadline: August 31, 2026 Security Clearance Requirement: TS/SCI Security Clearance with Polygraph Job Description Trusted Space...Full timeTemporary workLocal area$119k - $170k
...impact at the company pioneering security transformation in the AI era? Join us at Zscaler.RoleWe are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler...Full timeWork at officeLocal areaRemote workShift work3 days per week$69.4k - $158k
...specifications makes you an integral part of delivering a customer-focused engineering solution.As a Software and Systems Engineer on our team, you... ...total benefits by visiting the Resource page on our Careers site and reviewing Our Employee Benefits page.Salary at Booz Allen...Full timeContract workPart timeWork at officeLocal areaRemote work$94.35k - $136.85k
Systems Software Engineer (Associate or Experienced)Company:The Boeing CompanyThe Boeing Company has an exciting opportunity for a Systems Software Engineer to join the FishTools program supporting the Space Mission Systems Software Team in Chantilly, VA, Mesa, AZ, or...Permanent employmentFull timeWork experience placementCurrently hiringImmediate startVisa sponsorshipWork visaRelocation packageFlexible hoursShift work$146k - $234k
ResponsibilitiesPeraton is seeking a Software Test Engineer in Herndon, VA or Aurora, CO location to support Intelligence Community and DoD customers as part of a talented, high-performing team. As part of this team you will work with emerging service and distributed computing...Contract workShift work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
- official site Aurora, CO
- site services specialist Aurora, CO
- construction site safety Aurora, CO
- IT site lead Aurora, CO
- site leader Aurora, CO
- site safety Aurora, CO
- historic site Aurora, CO
- junior website developer Aurora, CO
- website coordinator Aurora, CO
- on-site clinical research associate (traveling/remote) Aurora, CO



