Site Reliability Engineer
GMI Cloud
About GMI
GMI Cloud is a fast-growing, AI-native infrastructure company delivering high-performance GPU compute, inference services, and infrastructure for AI agents.
Following 8x ARR growth, GMI Cloud continues to scale rapidly across the U.S. and APAC. As a Reference Platform NVIDIA Cloud Partner (NCP) and a validated leading NCP across both markets, we power production AI for leading AI-native companies including Fireworks AI, Cartesia, Reflection, and OpenRouter.
From large-scale compute to optimized inference and agentic workloads, GMI Cloud gives AI teams the infrastructure they need to build, deploy, and scale on one unified cloud.
One cloud for compute, inference, and agents.
Role Overview
We are seeking a skilled Site Reliability Engineer to join the GMI Global Infrastructure team. This role is hands-on and critical to ensuring the stability, efficiency, and reliability of the large-scale high performance AI/ML clusters in our data center. The ideal candidate will bring expertise in system-level troubleshooting, AI cluster maintenance, and operational excellence to ensure maximum performance for our infrastructure. Experience with large-scale infrastructure automation is considered a strong plus.
Responsibilities
- Design, implement and maintain scalable AI/ML infrastructure solutions.
- Proactively monitor GPU cluster health, performance and troubleshoot issues across compute, accelerator, and storage systems.
- Automate deployment, configuration and management of infrastructure resources.
- Manage GPU node lifecycle workflows, including provisioning, scaling, maintenance, decommissioning and upgrades of GPU nodes.
- Implement CI/CD pipelines for infrastructure deployment and orchestration.
- Ensure security, compliance and best practices across infrastructure.
- Manage incident response related to Infrastructure resources (GPU, CPU, Storage, Network).
- Handle customer provisioning requests for GPU resources, including onboarding, configuration and troubleshooting; resolve customer service requests related to GPU infrastructure, ensuring high customer satisfaction.
- Stay current with emerging GPU hardware and software technologies, integrating improvements as appropriate.
- Regional/international travel to GMI data center locations.
Qualifications
- Bachelor’s degree in Computer Science or related field.
- Over 3+ years of experience in data center operations, infrastructure, or systems engineering.
- Proven experience in site reliability engineering and infrastructure automation (e.g. Ansible, Terraform)
- Familiarity with containers orchestration platform (e.g. Kubernetes, Nvidia GPU operator, Nvidia Network operator, CNI, CSI) and job scheduling systems (e.g. Slurm).
- Familiarity with Linux system administration and scripting (Python, Bash).
- Familiarity with logging and monitoring tools such as Prometheus, Grafana, Loki.
- Good knowledge of GPU architecture, Nvidia CUDA, NCCL, or related AI/ML frameworks - added advantage.
- Strong troubleshooting skills and ability to analyze system logs and performance metrics.
- Excellent communication and teamwork abilities.
Meeting every qualification is not required—if you’re excited about this role, we’d love to hear from you. We believe diverse perspectives and experiences strengthen our team.
- Qualifications: 8+ years of Software Engineering experience, or equivalent demonstrated through... ...implement and maintain scalable and reliable infrastructure on Google Cloud Platform... ...vendor resources Willingness to work on-site at stated location in the job openingDepartment...SuggestedContract workFor contractorsWork experience placement
- ...their SAP and Oracle landscapes, whether running on-premises, in the cloud, or in a hybrid environment. We are looking for a Site Reliability Engineer II to join our global engineering team. In this role, you will apply software engineering principles to operations,...SuggestedWork at officeFlexible hoursShift work2 days per week
- ...ll be building the future of financial infrastructure. As part of our global expansion, we're looking for a hands-on Site Reliability Engineer (SRE) to design, scale, and safeguard the reliability of our next-generation financial platforms. This is a high-impact role...SuggestedRemote work
- Senior Talent Acquisition Specialist @ Centraprise Job Role: SRE Developer Job Type: Full time/ Permanent Location: Irving, TX Job Description: Must Have Technical/Functional Skills Proven experience in managing Windows and Linux servers. Proficiency...SuggestedPermanent employmentFull time
- ...on one unified cloud. One cloud for compute, inference, and agents. Role Overview We are seeking a skilled Site Reliability Engineer to join the GMI Global Infrastructure team. This role is hands-on and critical to ensuring the stability, efficiency, and...Suggested
- ...generative AI and cloud-native platforms to advanced release engineering practices, our teams are redefining how financial technology... ...championing best practices in coding, testing, and automation. Reliability Engineering: Establish service level objectives (SLOs) and...H1bWork at officeRemote workVisa sponsorshipFlexible hours2 days per week3 days per week
$140k - $150k
...your skills and experience — talk with your recruiter to learn more. Base pay range $140,000.00/yr - $150,000.00/yr Site Reliability Engineer II | 6-month Contract to Hire | Hybrid (Irving, TX) | 2x onsite per week Optomi, in partnership with a leading...Full timeContract work$72.1k - $158.62k
...person, one family and one community at a time. Position Summary We are seeking a highly skilled Software Development Engineer, Site Reliability Engineering (SRE), for Retail and Pharmacy platforms to drive reliability, scalability, and operational excellence. The...Hourly payFull timeTemporary workLocal area- ...Job Title: Site Reliability Engineer Location: Dallas TX (HYBRID) Duration :Full Time Job Description: Skill: Site Reliability Engineer • Ensures supported applications are functioning and available by minimizing downtime and maximizing performance...Full timeWork at office
- ...ensure applications are highly available, reliable, and performant at a global scale.... ...Bachelor of Computer Science or related Engineering field required. Master's Degree preferred... ...Minimum of 1 year of lead experience of site reliability engineering team required....Contract workWork at office
- ...and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems. As a Site Reliability Engineer III at JPMorgan Chase within the Chief Data & Analytics Office (CDAO) AI/ML & Data Platforms team,, you will solve complex...Work at office
- ...Ethos Group is seeking a talented and proactive Site Reliability Engineer (SRE) to join our growing technology team. This role is ideal for an engineer who is passionate about building highly available, scalable, and reliable systems while partnering closely with development...
- Mandatory Skills: AWS/Azure/GCP (GCP is not used very much ). Kubernetes /Helm,Docker,Gitlab,Grafana,Cyberark/Hashicorp Vault, Terraform etc. Experience utilizing Java, Perl, Python, Go and scripting experience in Shell and Perl to automate reports and monitor enterprise...
- ...healthcare fintech innovator, we’re transforming the patient journey and redefining what’s possible in dental care. This role: Site Reliability Engineer (SRE) with deep expertise in monitoring, debugging, and optimizing Azure App Services. This position is critical to...Full timeWork at office3 days per week
- ...improving platform infrastructure and applications with high reliability, resiliency, performance & quality, and faster time-to-market... ...documentation, including runbooks/playbooks; and, Using Chaos Engineering to test the robustness of the systems and applications....
- ...Senior Site Reliability Engineer Come join a growing bank at the heart of the innovation, technology, green tech and life sciences space. We continue to expand our global footprint and our banking technology is at the core of everything we do. As a Senior Site Reliability...
$136.2k - $214.01k
...outcomes Visionary in future focused problem-solving Exceptional in execution and impact The Role As a Senior Site Reliability Engineer at Proofpoint you will develop a deep understanding of the various services and applications that come together to...Full timeFlexible hours- ...Site Reliability Engineer We are looking for a Site Reliability Engineer for our client location in Dallas TX with the following skills: Java Spring Boot, Kubernetes, and eCommerce experience required. Key responsibilities include working with the applications, engineering...Work at office
- ...exclusive features. Our client is looking for a highly skilled Site Reliability Engineerto deploy, configure, and support our carrier-grade... ...customers. Work closely with customers, network engineers, and software developers to ensure seamless product integration...Full timeShift work
- ...automate them. # Experience in Implementing AI/ML-based monitoring and self-healing solutions. # Experience in Implementing Chaos Engineering/testing. Seniority level Mid-Senior level Employment type Full-time Job function Consulting, Analyst, and...Full time
- ...Role: Site Reliability Engineer 6+ months Contract role Remote About the Role We are looking for a dynamic and accomplished Site Reliability Engineer (SRE) who excels at solving complex reliability challenges and thrives in high-impact environments....Contract workRemote work
- ...Sr. Site Reliability Engineer Our client, a top tier IT Consulting firm is looking for several qualified Site Reliability Engineers to join a Top-Tier Investment Bank. Essential Requirements and Responsibilities: Proficiency in designing, deploying, and maintaining...
$174k - $252k
...systems by pushing for changes that improve reliability and velocity.Practice sustainable... ...:Bachelor’s degree in Computer Science, Engineering, a related field, or equivalent practical... ...degree in Computer Science or Engineering.Site Reliability Engineering (SRE) is what you...$125.7k - $203.1k
...collaborative team of systems and cloud engineers who thrive in a fast-paced environment built... ...maintain efficient platform uptime and reliability. • Lead end-to-end incident response and... ...steps.• Collaborate closely with Site Reliability Engineering (SRE) and Product...Permanent employmentFull timeTemporary workApprenticeshipWork experience placementLocal areaWorldwideFlexible hoursNight shift$147k - $210k
...product or system development code.Review code developed by other engineers and provide feedback to ensure best practices (e.g., style... ..., and troubleshooting large-scale distributed systems. Site Reliability Engineering (SRE) is what you get when you treat operations...$77.5k - $179k
Site Reliability Engineer I - Sales OperationsThis role has been designed as 'Hybrid' with a requirement that you will work on average 2 days per week from an HPE office.Who We Are:Hewlett Packard Enterprise is the global edge-to-cloud company advancing the way people live...Full timeWork experience placementInternshipWork at officeLocal areaImmediate start2 days per week$55k - $151.47k
...ApplicableSpecialismIFS - Internal Firm Services - OtherManagement LevelSenior AssociateJob Description & SummaryThe OpportunityAs a Site Reliability Engineer - Senior Associate, you will play a pivotal role in enhancing the reliability, scalability, and performance of our...Full timeH1b$172k - $300k
Job DescriptionGM Vehicle Autonomy is forming a centralized Site Reliability Engineering team to make reliability a measurable, engineered property of the systems used to build, validate, release, and operate autonomous-vehicle software.As one of our founding SREs, you...Full timeWork at officeLocal areaRemote workWork from homeRelocationRelocation packageFlexible hours$192.4k - $275.8k
...CloudOps— the team that keeps Splunk Cloud running for some of the world's most demanding enterprise customers, blending Site Reliability Engineering, Systems Engineering, and Service Engineering disciplines at a scale very few teams ever get to operate at. When the...Full timeTemporary workLocal areaFlexible hours- ...Position - Senior/Lead Site Reliability Engineer Observability Location - 100% Remote Experience - 8+ Years Type - Full Time Technology Stack - Splunk Enterprise, Splunk Cloud, Elasticsearch, ELK, Kibana, Prometheus, Grafana, Grafana Tempo, OpenTelemetry...Full timeRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!



