Site Reliability Engineer (SRE) - K8S/Slurm
GMI Cloud
About the Company
GMI Cloud is a fast-growing, AI-native infrastructure company delivering high-performance GPU compute, inference services, and infrastructure for AI agents.
Following 8x ARR growth, GMI Cloud continues to scale rapidly across the U.S. and APAC. As a Reference Platform NVIDIA Cloud Partner (NCP) and a validated leading NCP across both markets, we power production AI for leading AI-native companies including Fireworks AI, Cartesia, Reflection, and OpenRouter.
From large-scale compute to optimized inference and agentic workloads, GMI Cloud gives AI teams the infrastructure they need to build, deploy, and scale on one unified cloud.
One cloud for compute, inference, and agents.
Role Overview
We are seeking a skilled Site Reliability Engineer to join the GMI Global Infrastructure team. This role is hands-on and critical to ensuring the stability, efficiency, and reliability of the large-scale high performance AI/ML clusters in our data center. The ideal candidate will bring expertise in system-level troubleshooting, AI cluster maintenance, and operational excellence to ensure maximum performance for our infrastructure. Experience with large-scale infrastructure automation is considered a strong plus.
Responsibilities
- Design, implement and maintain scalable AI/ML infrastructure solutions.
- Proactively monitor GPU cluster health, performance and troubleshoot issues across compute, accelerator, and storage systems.
- Automate deployment, configuration and management of infrastructure resources.
- Manage GPU node lifecycle workflows, including provisioning, scaling, maintenance, decommissioning and upgrades of GPU nodes.
- Implement CI/CD pipelines for infrastructure deployment and orchestration.
- Ensure security, compliance and best practices across infrastructure.
- Manage incident response related to Infrastructure resources (GPU, CPU, Storage, Network).
- Handle customer provisioning requests for GPU resources, including onboarding, configuration and troubleshooting; resolve customer service requests related to GPU infrastructure, ensuring high customer satisfaction.
- Stay current with emerging GPU hardware and software technologies, integrating improvements as appropriate.
- Regional/international travel to GMI data center locations.
Qualifications
- Bachelor’s degree in Computer Science or related field.
- Over 3+ years of experience in data center operations, infrastructure, or systems engineering.
- Proven experience in site reliability engineering and infrastructure automation (e.g. Ansible, Terraform)
- Familiarity with containers orchestration platform (e.g. Kubernetes, Nvidia GPU operator, Nvidia Network operator, CNI, CSI) and job scheduling systems (e.g. Slurm).
- Familiarity with Linux system administration and scripting (Python, Bash).
- Familiarity with logging and monitoring tools such as Prometheus, Grafana, Loki.
- Good knowledge of GPU architecture, Nvidia CUDA, NCCL, or related AI/ML frameworks - added advantage.
- Strong troubleshooting skills and ability to analyze system logs and performance metrics.
- Excellent communication and teamwork abilities.
Meeting every qualification is not required—if you’re excited about this role, we’d love to hear from you. We believe diverse perspectives and experiences strengthen our team.
- ...our communities.This is a Software Engineering position at Director level, which is... ...role is for an experienced and driven Site Reliability Engineer (SRE) to join our AI Platform team to... ...distributed GPU cluster scheduling (e.g. Slurm, Kubernetes GPU scheduling)...Suggested
- ...Primary Location: Louisville, Kentucky V-Soft Consulting is currently hiring for a SRE-Site Reliability Engineer for our premier client in Louisville, Kentucky . Education And Experience » ~6+ years in IT infrastructure, with 3+ years focused on site reliability...SuggestedPermanent employmentFull timeWork experience placementCurrently hiringLocal area
- ...We are seeking an experienced Site Reliability Engineer (SRE) to support and maintain production systems hosted on AWS. The role focuses on production support, incident management, monitoring, observability, troubleshooting, and improving system reliability and availability...SuggestedContract work
- ...prioritize a diverse F5 community where each individual can thrive.Role SummaryWe are seeking a proactive and detail-oriented Site Reliability Engineer II (SRE II) to join our 24/7 Operations team in a hybrid capacity. In this role, you will provide round-the-clock, eyes-on-...SuggestedFull timeLocal areaImmediate startShift workNight shiftAfternoon shiftWeekday work
$70.8k - $131.4k
Job DescriptionThomson Reuters is strengthening its Site Reliability Engineering capability to help engineering and operations teams build, operate... ...review is needed.Key ResponsibilitiesSupport and maintain SRE operational tooling, including dashboards, alerts, runbooks...SuggestedFull timeWork at officeLocal areaFlexible hours- ...unwavering security to responsibly propel the global lottery industry ever forward.Position SummaryWe are looking for a skilled Site Reliability Engineer (SRE) to enhance the stability, performance, and reliability of our production systems. The SRE will work closely with...Permanent employmentFull timeWork experience placementLocal area
- ...the selected candidate for this role to work on site in the specified location(s).As a Senior Reliability Engineer, you will help shape the reliability, scalability... ...teams, you will apply Site Reliability Engineering (SRE) principles to improve system availability,...Full timeWork at office
$175k - $215k
...experiences — and we’re constantly looking for new ways to enhance these exciting experiences.Sr. Manager, Site Reliability Engineer provides strategic leadership across multiple SRE teams and their managers, ensuring alignment with organizational priorities and functional...$135k - $155k
...buyers at Fortune 1000 companies to tap into global manufacturing capacity.Xometry is seeking a Site Reliability Engineer II to join our Site Reliability Engineering (SRE) Organization. In this role as an individual contributor, you will guide the reliability and performance...Flexible hours- ...Period of performance: Up to 2 years in duration MUST HAVES: Minimum of 8 years of experience as a Site Reliability Engineer with a strong understanding of SRE principles for highly scalable and reliable systems Possess a bachelor's degree Experience working...Local areaRelocation package3 days per week
$100k - $200k
...OPPO US Research Center is seeking a skilled and proactive Site Reliability Engineer (SRE) to join our team. In this role, you will be responsible for ensuring the stability, scalability, and performance of our application systems. The ideal candidate is passionate about...Full time- ...We're seeking an SRE to ensure the reliability and performance of our clients' critical systems. You'll work on observability, incident response... ...management Nice to have Experience with chaos engineering Knowledge of distributed systems Background in high...Remote workFlexible hours
- ...We’re seeking a highly skilled Site Reliability Engineer (SRE) to join our engineering team and help ensure the reliability, scalability, and performance of our systems. As an SRE, you’ll blend software engineering with systems engineering to build and maintain resilient...Temporary workInterim roleRemote workFlexible hours
- ...,000 companies and 100,000 electronics engineers worldwide use Altium We are growing,... ...industry Role Overview: Senior Site Reliability Engineerensures the reliability, availability... ...shared responsibility model where the SRE team collaborates with and educates...Worldwide
$170k - $230k
...Site Reliability Engineer (SRE) Palo Alto / San Francisco Bay Area About Mithril Mithril is an AI infrastructure platform built to make GPU compute more accessible and affordable for the world's leading enterprises, AI startups, and the AI research community,...Work at officeLocal area1 day per week$165k - $225k
...Sr. Site Reliability Engineer (SRE) Chicago, IL or Remote Moonlite delivers high-performance AI infrastructure for organizations running intensive computational research, large-scale model training, and demanding data processing workloads. We provide infrastructure...Remote workFlexible hours- ...Site Reliability Engineer (SRE) Location: North Little Rock AR (onsite) Duration: Contract Required/Desired Skills: • Strong web development skills with a strong focus in C#/.NET • Someone who currently works in a hybrid skillset of BOTH.Net development AND...Contract work
$142.3k - $263.3k
...recently the new LIDAR iPad sensor. We are looking for the right Site Reliability Engineer to help us take our efforts to the next level. In this role,... ...Computer Vision Organization. As a main contributor to our SRE team you will develop and maintain infrastructure, tooling,...Work experience placementRelocation- ...globe. Join us on this journey to redefine resource management—and change lives along the way. The Role As a Site Reliability Engineer (SRE) at Air Apps, you will be responsible for ensuring the reliability, availability, and scalability of our systems. You...Temporary workWorldwide
- ...A Site Reliability Engineer (SRE) is responsible for ensuring the reliability, scalability, and performance of an organization's software systems and cloud infrastructure. The role combines software engineering with IT operations to automate processes, monitor system...
$100k - $180k
...Site Reliability Engineer (SRE) - Remote Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States. This is a fantastic opportunity to join an established...Full timeH1bLocal areaImmediate startRemote workVisa sponsorship$120k - $175k
...Senior Site Reliability Engineer (SRE) Atlanta, GA preferred, Remote At PrizePicks, we are the fastest-growing sports company in North America, as recognized by Inc. 5000. As the leading platform for Daily Fantasy Sports, we cover a diverse range of sports leagues...Full timeRemote workWork visaFlexible hours- ...Senior Site Reliability Engineer (SRE) Our client is a global technology consulting and digital solutions company that enables enterprises across industries to reimagine business models, accelerate innovation, and maximize growth by harnessing digital technologies....Local area
- ...Purple Drive Site Reliability Engineer (SRE) Contractual Atlanta, GA Key Highlights: Proven expertise in Google Cloud Platform (GCP) services, including BigQuery, Cloud Logging, IAM, and Service Accounts. Strong background in provisioning, monitoring, and...
$110k - $120k
...documentation that you are a U.S. citizen to qualify. Summary: The Site Reliability Engineer owns the day-to-day health, security, and usability of the organization's ICAM environment while applying SRE practices to keep delivery platforms reliable, observable, and...Contract workTemporary workFor contractorsWork at officeLocal areaFlexible hours- ...Role: Site Reliability Engineer (SRE) Location: Brentwood, TN (Onsite) Contract Experience: 6-8+ years Role Description: Combines software engineering and IT operations to ensure the reliability, scalability, and performance of systems, with...Contract work
- ...We are seeking a Site Reliability Engineer (SRE) to improve the reliability, scalability, and performance of cloud-based applications. The ideal candidate will automate infrastructure, monitor production systems, troubleshoot incidents, and collaborate with software engineering...
$40 - $80 per hour
...Job Title: Site Reliability Engineer (SRE) Duration (Contract): 12 Months Client Location: Southlake, TX Location Preference: Onsite Job Description: As a Site Reliability Engineer (SRE) , you will be responsible for improving the reliability, scalability...Hourly payContract work- ...Open role Site Reliability Engineer (SRE) San Francisco, CA (On-site) Responsibilities Develop and maintain advanced monitoring, alerting, and self-healing mechanisms that detect and address issues before they impact customers. Perform regular capacity...
- ...in Computer Science, Information Technology, Engineering, or equivalent field ~3-5 years of experience in Site Reliability Engineering, Production Support, Platform Engineering... ...application health ~ Understanding of SRE principles, including observability,...Remote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer (SRE) - K8S/Slurm. Be the first to apply!
- site reliability engineering manager United States
- site reliability engineer sre United States
- site reliability engineer United States
- site reliability engineer remote United States
- official site United States
- site merchandiser United States
- site services specialist United States
- construction site safety United States
- IT site lead United States
- site recruiter United States




