Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer

GMI Cloud

About GMI

GMI Cloud is a fast-growing, AI-native infrastructure company delivering high-performance GPU compute, inference services, and infrastructure for AI agents.

Following 8x ARR growth, GMI Cloud continues to scale rapidly across the U.S. and APAC. As a Reference Platform NVIDIA Cloud Partner (NCP) and a validated leading NCP across both markets, we power production AI for leading AI-native companies including Fireworks AI, Cartesia, Reflection, and OpenRouter.

From large-scale compute to optimized inference and agentic workloads, GMI Cloud gives AI teams the infrastructure they need to build, deploy, and scale on one unified cloud.

One cloud for compute, inference, and agents.

Role Overview

We are seeking a skilled Site Reliability Engineer to join the GMI Global Infrastructure team. This role is hands-on and critical to ensuring the stability, efficiency, and reliability of the large-scale high performance AI/ML clusters in our data center. The ideal candidate will bring expertise in system-level troubleshooting, AI cluster maintenance, and operational excellence to ensure maximum performance for our infrastructure. Experience with large-scale infrastructure automation is considered a strong plus.

Responsibilities

  • Design, implement and maintain scalable AI/ML infrastructure solutions.
  • Proactively monitor GPU cluster health, performance and troubleshoot issues across compute, accelerator, and storage systems.
  • Automate deployment, configuration and management of infrastructure resources.
  • Manage GPU node lifecycle workflows, including provisioning, scaling, maintenance, decommissioning and upgrades of GPU nodes.
  • Implement CI/CD pipelines for infrastructure deployment and orchestration.
  • Ensure security, compliance and best practices across infrastructure.
  • Manage incident response related to Infrastructure resources (GPU, CPU, Storage, Network).
  • Handle customer provisioning requests for GPU resources, including onboarding, configuration and troubleshooting; resolve customer service requests related to GPU infrastructure, ensuring high customer satisfaction.
  • Stay current with emerging GPU hardware and software technologies, integrating improvements as appropriate.
  • Regional/international travel to GMI data center locations.

Qualifications

  • Bachelor’s degree in Computer Science or related field.
  • Over 3+ years of experience in data center operations, infrastructure, or systems engineering.
  • Proven experience in site reliability engineering and infrastructure automation (e.g. Ansible, Terraform)
  • Familiarity with containers orchestration platform (e.g. Kubernetes, Nvidia GPU operator, Nvidia Network operator, CNI, CSI) and job scheduling systems (e.g. Slurm).
  • Familiarity with Linux system administration and scripting (Python, Bash).
  • Familiarity with logging and monitoring tools such as Prometheus, Grafana, Loki.
  • Good knowledge of GPU architecture, Nvidia CUDA, NCCL, or related AI/ML frameworks - added advantage.
  • Strong troubleshooting skills and ability to analyze system logs and performance metrics.
  • Excellent communication and teamwork abilities.

Meeting every qualification is not required—if you’re excited about this role, we’d love to hear from you. We believe diverse perspectives and experiences strengthen our team.

Vacancy posted 4 hours ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in Aurora, CO vacancy
  • $62k - $141k

    Site Reliability EngineerThe Opportunity: Engineering to make a system more resilient and efficient frees up time and money to build more capabilities. Whether you come from a background in network engineering, systems administration, or software development, if you have... 
    Suggested
    Full time
    Contract work
    Part time
    Local area
    Remote work

    Booz Allen Hamilton

    Aurora, CO
    15 hours ago
  • $130k - $200k

     ...Site Reliability Engineer (Aurora, CO; Herndon, VA) Summary Position Title: Site Reliability Engineer Position ID: TA247 Location(s): On-site; Aurora, CO; Herndon, VA Application Deadline: September 30, 2026 Security Clearance Requirement: An active TS/SCI... 
    Suggested
    Temporary work
    Local area

    Trusted Space, LLC

    Aurora, CO
    22 hours ago
  • $135k - $155k

     ...while also making it easy for buyers at Fortune 1000 companies to tap into global manufacturing capacity.Xometry is seeking a Site Reliability Engineer II to join our Site Reliability Engineering (SRE) Organization. In this role as an individual contributor, you will guide... 
    Suggested
    Flexible hours

    Thomas

    Denver, CO
    2 days ago
  • $114.4k - $125.4k

     ...excellence of the Salesforce GovCloud! Are you passionate about ensuring the reliability and performance of mission-critical cloud services? Salesforce is seeking a talented Site Reliability Engineer to join our dynamic team, supporting our GovCloud environment. As a key... 
    Suggested
    Full time
    Local area
    Shift work
    Night shift

    Salesforce

    Denver, CO
    15 hours ago
  • $98.58k - $138.02k

     ...This role requires a hybrid work schedule based out of one of our office locations: Austin, TX; Irvine, CA; or Akron, OH. Site Reliability Engineer II will be responsible for supporting, enhancing, and maintaining Restaurant365’s cloud infrastructure and applications.... 
    Suggested
    Work at office

    Restaurant365

    Denver, CO
    2 days ago
  • $110k - $125k

     ...The Site Reliability Engineer maintains and improves the availability, performance, resilience, and operational recoverability of enterprise identity, credential, and access-management services. This role supports cloud identity and directory platforms, including Microsoft... 
    Contract work
    Work at office

    ASM Research, An Accenture Federal Services Company

    Denver, CO
    2 days ago
  •  ...profitable developer-tooling company whose product is used by engineering teams at thousands of software companies for application...  ...well-resourced group of nine. As Senior SRE you will lead reliability initiatives across the platform — from defining and driving SLOs... 

    Kovoro

    Denver, CO
    22 hours ago
  •  ...Protecting others requires a team that works together with trust and cares deeply about carrying out our mission. The Site Reliability Engineering team at Todyl exists to make our platform reliable, secure, and easy for engineering teams to ship to. We do that by... 
    Full time
    Temporary work
    Local area
    Flexible hours
    Shift work

    Todyl

    Denver, CO
    22 hours ago
  • $95k - $134k

     ...helped build. For more information, visit Job Application Deadline: 10/31/2026 The Opportunity DAT is looking for a Site Reliability Engineer to join our SRE platform team. This position will work hybrid in Denver, CO Candidate profile DAT is seeking an... 
    Temporary work
    For contractors
    Work experience placement
    Work at office
    Local area
    Immediate start
    Flexible hours

    Worky Ltd

    Denver, CO
    2 days ago
  •  ...Site Reliability Engineer - Greenwood Village, CO (Hybrid) About the Role: Join a forward-thinking engineering team as a Site Reliability Engineer, specializing in enterprise-scale experimentation and configuration management platforms. In this operations-focused... 
    Contract work

    Indotronix International Corporation

    Greenwood Village, CO
    2 days ago
  •  ...focusing on private cloud systems supporting 5G wireless systems. This position will focus on platform monitoring, logging, and reliability aspects supporting the Mobile Core team. A critical goal is to gather metrics of the platform during stress and load events to ensure... 

    Software Technology Inc

    Denver, CO
    1 day ago
  • $95k - $171k

     .... Opportunities exist to focus on GPU infrastructure, Kubernetes, and ensuring reliability for AI workloads within Akamai's serverless inference platform. As an Site Reliability Engineer II, you will be responsible for: Building and maintaining dashboards, alerts... 
    Permanent employment
    Work experience placement
    Work at office
    Remote work
    Work from home
    Worldwide
    Flexible hours

    Akamai

    Denver, CO
    22 hours ago
  • $141.8k - $195k

     ...their best work, grow fast, and bring their full selves to the herd.Why You'll Love This RoleCribl Inc is seeking a Senior Site Reliability Engineer to join our mission where you will unlock the value of all observability data, as we expand our team in the U.S. Cribl... 
    Remote work

    Cribl

    Denver, CO
    22 hours ago
  • $168k - $200k

     ...is passionate about creating transformative change in healthcare. What We're Looking For We're looking for a Senior Site Reliability Engineer to join our Data & ML Platform team. You'll be at the forefront of building and operating a resilient, observable, and... 

    Datavant

    Denver, CO
    1 day ago
  •  ...We are seeking an experienced Site Reliability Engineer (SRE) to join the Applied AI and Data Science program. This role focuses on deploying, monitoring, and optimizing cloud-based applications and infrastructure to ensure high availability and performance. The... 

    Compunnel

    Greenwood Village, CO
    2 days ago
  • $81.1k - $187k

     ...Job Description We are looking for a Site Reliability Engineer 3 to support mission-critical cloud services and production operations. The role focuses on improving service reliability, reducing operational risk, automating repetitive tasks, and driving faster detection... 
    Temporary work
    Immediate start
    Flexible hours
    Shift work

    Oracle

    Denver, CO
    1 day ago
  •  ...At Todyl, our Application Platform Engineering team is dedicated to building infrastructure...  ...work will not only directly impact the reliability and security of our platform but also empower...  ...to grow our team and is hiring two Site Reliability Engineers (SRE I and SRE II)... 
    Local area

    Todyl

    Denver, CO
    2 days ago
  • $121.4k - $218.6k

     ...solve complex challenges? Do you have a passion for automation and building systems that scale? Join our highly skilled Site Reliability Engineering team! Our team designs, develops, and manages applications and infrastructure that support Akamai Cloud's products and... 
    Work experience placement
    Work at office

    Akamai

    Denver, CO
    22 hours ago
  • $140k - $165k

     ...Summary of Job ResponsibilitiesThe Senior Site Reliability Engineer (SRE) is Colorado PERA’s technical owner for production reliability, observability, and platform automation across a hybrid AWS, Azure and on-premises environment. This position leads the design, implementation... 
    Full time
    Work at office
    Work from home
    Afternoon shift
    3 days per week

    Colorado-Public-Employees

    Denver, CO
    22 hours ago
  • $127k - $249k

    Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions...  ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper). As... 
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    Denver, CO
    4 days ago
  • $192.4k - $275.8k

     ...CloudOps— the team that keeps Splunk Cloud running for some of the world's most demanding enterprise customers, blending Site Reliability Engineering, Systems Engineering, and Service Engineering disciplines at a scale very few teams ever get to operate at. When the... 
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    Denver, CO
    3 days ago
  • $110k - $155k

     ...global, with headquarters in Denver, Colorado, and offices across the U.S., Canada, and India. We are seeking a Senior Site Reliability Engineer to own the reliability, scalability, performance, and operational integrity of critical production services. This role is... 
    Contract work
    Work at office
    Work from home
    Flexible hours

    Vertafore

    Denver, CO
    a month ago
  • $160k - $180k

     ...global, with headquarters in Denver, Colorado, and offices across the U.S., Canada, and India. We are seeking a Principal Site Reliability Engineer to define the strategic vision and own the enterprise‑wide reliability, scalability, and performance of our critical... 
    Contract work
    Temporary work
    Work at office
    Work from home
    Flexible hours

    Vertafore

    Denver, CO
    2 days ago
  • $175k - $220k

     ...with headquarters in Denver, Colorado, and offices across the U.S., Canada, and India. Role Summary The Director, Site Reliability Engineering (SRE) will lead reliability, performance, and observability initiatives for a portfolio of Vertafore products. This role... 
    Contract work
    Temporary work
    Work at office
    Work from home
    Flexible hours

    Vertafore

    Denver, CO
    2 days ago
  • $169.3k - $304.7k

     ...in building and maintaining fast, efficient, scalable, and reliable routing software and infrastructure that is responsible...  ...growth and stability of our global platform. As a Principal Site Reliability Engineer - Network, you will be responsible for: Architecting,... 
    Work experience placement
    Work at office

    Akamai

    Denver, CO
    2 days ago
  • $130k - $200k

    Summary Position Title: Site Reliability Engineer Position ID: TA247 Location(s): On-site; Aurora, CO; Herndon, VA Application Deadline: August 31, 2026 Security Clearance Requirement: TS/SCI Security Clearance with Polygraph Job Description Trusted Space... 
    Full time
    Temporary work
    Local area

    Trusted Space, LLC

    Aurora, CO
    22 hours ago
  • $119k - $170k

     ...impact at the company pioneering security transformation in the AI era? Join us at Zscaler.RoleWe are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler... 
    Full time
    Work at office
    Local area
    Remote work
    Shift work
    3 days per week

    Zscaler

    Denver, CO
    2 days ago
  • $69.4k - $158k

     ...specifications makes you an integral part of delivering a customer-focused engineering solution.As a Software and Systems Engineer on our team, you...  ...total benefits by visiting the Resource page on our Careers site and reviewing Our Employee Benefits page.Salary at Booz Allen... 
    Full time
    Contract work
    Part time
    Work at office
    Local area
    Remote work

    Booz Allen Hamilton

    Aurora, CO
    22 hours ago
  • $94.35k - $136.85k

    Systems Software Engineer (Associate or Experienced)Company:The Boeing CompanyThe Boeing Company has an exciting opportunity for a Systems Software Engineer to join the FishTools program supporting the Space Mission Systems Software Team in Chantilly, VA, Mesa, AZ, or... 
    Permanent employment
    Full time
    Work experience placement
    Currently hiring
    Immediate start
    Visa sponsorship
    Work visa
    Relocation package
    Flexible hours
    Shift work

    Boeing

    Aurora, CO
    22 hours ago
  • $146k - $234k

    ResponsibilitiesPeraton is seeking a Software Test Engineer in Herndon, VA or Aurora, CO location to support Intelligence Community and DoD customers as part of a talented, high-performing team. As part of this team you will work with emerging service and distributed computing... 
    Contract work
    Shift work

    Peraton Corporation

    Aurora, CO
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!