Site Reliability Engineer
GMI Cloud
About GMI
GMI Cloud is a fast-growing, AI-native infrastructure company delivering high-performance GPU compute, inference services, and infrastructure for AI agents.
Following 8x ARR growth, GMI Cloud continues to scale rapidly across the U.S. and APAC. As a Reference Platform NVIDIA Cloud Partner (NCP) and a validated leading NCP across both markets, we power production AI for leading AI-native companies including Fireworks AI, Cartesia, Reflection, and OpenRouter.
From large-scale compute to optimized inference and agentic workloads, GMI Cloud gives AI teams the infrastructure they need to build, deploy, and scale on one unified cloud.
One cloud for compute, inference, and agents.
Role Overview
We are seeking a skilled Site Reliability Engineer to join the GMI Global Infrastructure team. This role is hands-on and critical to ensuring the stability, efficiency, and reliability of the large-scale high performance AI/ML clusters in our data center. The ideal candidate will bring expertise in system-level troubleshooting, AI cluster maintenance, and operational excellence to ensure maximum performance for our infrastructure. Experience with large-scale infrastructure automation is considered a strong plus.
Responsibilities
- Design, implement and maintain scalable AI/ML infrastructure solutions.
- Proactively monitor GPU cluster health, performance and troubleshoot issues across compute, accelerator, and storage systems.
- Automate deployment, configuration and management of infrastructure resources.
- Manage GPU node lifecycle workflows, including provisioning, scaling, maintenance, decommissioning and upgrades of GPU nodes.
- Implement CI/CD pipelines for infrastructure deployment and orchestration.
- Ensure security, compliance and best practices across infrastructure.
- Manage incident response related to Infrastructure resources (GPU, CPU, Storage, Network).
- Handle customer provisioning requests for GPU resources, including onboarding, configuration and troubleshooting; resolve customer service requests related to GPU infrastructure, ensuring high customer satisfaction.
- Stay current with emerging GPU hardware and software technologies, integrating improvements as appropriate.
- Regional/international travel to GMI data center locations.
Qualifications
- Bachelor’s degree in Computer Science or related field.
- Over 3+ years of experience in data center operations, infrastructure, or systems engineering.
- Proven experience in site reliability engineering and infrastructure automation (e.g. Ansible, Terraform)
- Familiarity with containers orchestration platform (e.g. Kubernetes, Nvidia GPU operator, Nvidia Network operator, CNI, CSI) and job scheduling systems (e.g. Slurm).
- Familiarity with Linux system administration and scripting (Python, Bash).
- Familiarity with logging and monitoring tools such as Prometheus, Grafana, Loki.
- Good knowledge of GPU architecture, Nvidia CUDA, NCCL, or related AI/ML frameworks - added advantage.
- Strong troubleshooting skills and ability to analyze system logs and performance metrics.
- Excellent communication and teamwork abilities.
Meeting every qualification is not required—if you’re excited about this role, we’d love to hear from you. We believe diverse perspectives and experiences strengthen our team.
- Senior Site Reliability Engineer (SRE)Salt Lake City, UTAre you passionate about building highly reliable, scalable cloud platforms that power mission-critical applications? We're partnering with an innovative technology company that's investing heavily in platform reliability...SuggestedWork at officeRemote work1 day per week
- ...P osition Name: Site Reliability Engineer Location: Salt Lake,USA Salary: ***/Annum Detailed Job Description. Well versed in Application Monitoring tools (Splunk, Extrahop, App Dynamics, Prometheus Grafana) Good understanding of JVM and Database metrics...Suggested
$121.4k - $218.6k
...solve complex challenges? Do you have a passion for automation and building systems that scale? Join our highly skilled Site Reliability Engineering team! Our team designs, develops, and manages applications and infrastructure that support Akamai Cloud's products and...SuggestedWork experience placementWork at office$75.7k - $136.3k
...solve complex challenges? Do you have a passion for automation and building systems that scale? Join our highly skilled Site Reliability Engineering team! Our team designs, develops, and manages applications and infrastructure that support Akamai Cloud's products and...SuggestedWork experience placementWork at office- ...Sr. Site Reliability Engineer Our client, a top tier IT Consulting firm is looking for several qualified Site Reliability Engineers to join a Top-Tier Investment Bank. Essential Requirements and Responsibilities: Proficiency in designing, deploying, and maintaining...Suggested
$130k - $160k
...About the Role The Site Reliability Engineering team at iCapital is fundamental to ensuring our platform delivers consistent, reliable service to our client base. As a Site Reliability Engineer, you'll work at the intersection of software engineering and operations...Full timeWork at officeRemote work$81.1k - $187k
...Job Description We are looking for a Site Reliability Engineer 3 to support mission-critical cloud services and production operations. The role focuses on improving service reliability, reducing operational risk, automating repetitive tasks, and driving faster detection...Temporary workImmediate startFlexible hoursShift work$136.2k - $214.01k
...outcomes Visionary in future focused problem-solving Exceptional in execution and impact The Role As a Senior Site Reliability Engineer at Proofpoint you will develop a deep understanding of the various services and applications that come together to...Full timeFlexible hours- ...Traders, Tower Research, PDT Partners, SIG, and more. We're looking for a midlevel or senior IC to join our Backend Engineering team as a Site Reliability Engineer. You'll own the uptime, performance, and observability of our platform, and help set the standard for how...Full timeLocal areaRemote work
$95k - $171k
.... Opportunities exist to focus on GPU infrastructure, Kubernetes, and ensuring reliability for AI workloads within Akamai's serverless inference platform. As an Site Reliability Engineer II, you will be responsible for: Building and maintaining dashboards, alerts...Permanent employmentWork experience placementWork at officeRemote workWork from homeWorldwideFlexible hours$55k - $151.47k
...ApplicableSpecialismIFS - Internal Firm Services - OtherManagement LevelSenior AssociateJob Description & SummaryThe OpportunityAs a Site Reliability Engineer - Senior Associate, you will play a pivotal role in enhancing the reliability, scalability, and performance of our...Full timeH1b- ...adeptlynavigate the intersection of technology, market needs, and business objectives. Collaborating closely with stakeholders and engineering teams, the Technical Product Manager defines and communicatesthe technical product vision, orchestrates the product roadmap, and...Full timeWork at officeRemote work
$169.3k - $304.7k
...in building and maintaining fast, efficient, scalable, and reliable routing software and infrastructure that is responsible... ...growth and stability of our global platform. As a Principal Site Reliability Engineer - Network, you will be responsible for: Architecting,...Work experience placementWork at office- ...office in SLC three days a week. We are not able to provide work sponsorship (i.e. OPT or H-1B visa) for this role *_ Site Reliability Engineers keep the platforms behind Extra Space Storage's applications available, fast, and secure. SRE owns the AWS environments those...Full timeInternshipH1bWork at office3 days per week
- ...Overview Join to apply for the Staff Site Reliability Engineer role at Addepar . Addepar is a global technology and data company that helps investment professionals provide the most informed, precise guidance for their clients. Hundreds of thousands of users...Full timeFlexible hours
- ...As the Manager of Site Reliability Engineering, you will lead the strategy, execution, and evolution of reliability for our world-class employee recognition platform. You will build, mentor, and empower a team of Site Reliability Engineers while partnering closely with...Shift work
$113k - $141.53k
...leader in global energy. Senior Solutions Engineer - Systems Integration serves as a... ...functionally in the field, ensuring safe, reliable, and performant operation across diverse... ...Willingness to travel to factories and project sites (25%).Preferred QualificationsMaster’s degree...Full timeFor contractorsLocal areaRemote workWorldwideFlexible hours- Reliability EngineerRio Tinto Kennecott | Salt Lake City, Utah | On-Site at the SmelterImprove asset reliability, maintenance strategies and equipment performance across... ...About the roleWe are looking for a Reliability Engineer to improve asset performance by identifying and...Full timeLocal areaWork from home
- ...contacts internal and external experts as required.• Utilizes reliability tools such as reliability analytics, failure evaluations, and... ....MINIMUM QUALIFICATIONS:• Bachelor’s Degree in Mechanical Engineering required.• Zero (0) years or more of experience required.As an...Full timeLocal area
$113k - $141.53k
...economic growth due to the availability of reliable, affordable electric power. The role... ...of an AES Clean Energy Senior Reliability Engineer will include, but is not limited to:Analyzing... ...operational sitesEvaluating operational sites for repowering and retrofit opportunities...Full timeContract workFor contractorsFor subcontractorWork at officeWorldwide$114.6k - $234.6k
...creates durable fixes and preventive controls. Designs performance, reliability, and fault-tolerance improvements for drivers, services, and... ...development lifecycle; provides guidance and coaching to engineers to drive improvements. Utilizes advanced knowledge to develop...Temporary workFlexible hoursShift work$25 per hour
...Distributed Systems Software Engineer, Python / Go Join to apply for the Distributed Systems Software Engineer, Python / Go role... ...automated testing approaches and infrastructure for validating reliability, performance, and resilience of cloud orchestration tools and...Full timeLocal areaRemote workWorldwide- ...in the interest of national security. Job Title: Lead Systems Engineer Job Code: 45046 Job Location: Salt Lake City, Utah Job... ...test, system integration, verification, qualification, or customer-site test events. Strong knowledge of requirements management,...For subcontractorLocal area
- ...the interest of national security. Job Title: Lead, Systems Engineer Job Code: 42191 Job Location: Salt Lake City- UT Job... ...to travel 20-40% of the time to customer and subcontractor sites in support of technical meetings and verification activities...For subcontractorLocal area
$80.52 - $85.52 per hour
...Pay Range: $80.52hr - $85.52hr Job Overview: The organization is seeking a Systems Engineer Lead to provide technical leadership across the full lifecycle of advanced data link and SATCOM systems, from concept development and capture through design, integration, verification...Temporary workLocal area$77k - $202k
...ApplicableSpecialismData, Analytics & AIManagement LevelSenior AssociateJob Description & SummaryThe OpportunityAs a GenAI Python Systems Engineer - Senior Associate, you will play a pivotal role in transforming raw data into actionable insights, enabling informed decision-...Full timeH1b$63k - $140k
...ApplicableSpecialismData, Analytics & AIManagement LevelAssociateJob Description & SummaryThe OpportunityAs a GenAI Python Systems Engineer - Experienced Associate, you will leverage advanced technologies and techniques to design and develop robust data solutions for...Full timeH1b- Salt Lake City, UtahSales - Sales Engineer /Full-time /On-siteFilevine is a Legal AI company delivering Legal Operating Intelligence for the future of legal work. Grounded in a singular system of truth, Filevine brings together data, documents, workflows, and teams into...Full timeContract workTemporary workWork at office2 days per week3 days per week
- ...discovering lasting solutions for unmet patient needs. Our Senior Quality Engineer, Software Validation position is a unique career opportunity... ...position will also represent the Draper facility on a cross-site council focused on global software-validation standards and...Full time
$152.5k - $205k
...stakeholder.What you’ll be responsible for:The Senior Software Engineer is responsible for extending Circle's in-house blockchain systems... ...and owning scalable microservices that are responsible for reliable and secure APIs that transfer value and assets across all blockchain...Permanent employmentRemote workFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
- site reliability engineer sre Salt Lake City, UT
- site reliability engineer Salt Lake City, UT
- official site Salt Lake City, UT
- site services specialist Salt Lake City, UT
- construction site safety Salt Lake City, UT
- IT site lead Salt Lake City, UT
- site leader Salt Lake City, UT
- site safety Salt Lake City, UT
- historic site Salt Lake City, UT
- junior website developer Salt Lake City, UT


