Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer (SRE) - K8S/Slurm

Full-time

GMI Cloud

About the Company

GMI Cloud is a fast-growing, AI-native infrastructure company delivering high-performance GPU compute, inference services, and infrastructure for AI agents.

Following 8x ARR growth, GMI Cloud continues to scale rapidly across the U.S. and APAC. As a Reference Platform NVIDIA Cloud Partner (NCP) and a validated leading NCP across both markets, we power production AI for leading AI-native companies including Fireworks AI, Cartesia, Reflection, and OpenRouter.

From large-scale compute to optimized inference and agentic workloads, GMI Cloud gives AI teams the infrastructure they need to build, deploy, and scale on one unified cloud.

One cloud for compute, inference, and agents.

Role Overview

We are seeking a skilled Site Reliability Engineer to join the GMI Global Infrastructure team. This role is hands-on and critical to ensuring the stability, efficiency, and reliability of the large-scale high performance AI/ML clusters in our data center. The ideal candidate will bring expertise in system-level troubleshooting, AI cluster maintenance, and operational excellence to ensure maximum performance for our infrastructure. Experience with large-scale infrastructure automation is considered a strong plus.

Responsibilities

  1. Design, implement and maintain scalable AI/ML infrastructure solutions.
  2. Proactively monitor GPU cluster health, performance and troubleshoot issues across compute, accelerator, and storage systems.
  3. Automate deployment, configuration and management of infrastructure resources.
  4. Manage GPU node lifecycle workflows, including provisioning, scaling, maintenance, decommissioning and upgrades of GPU nodes.
  5. Implement CI/CD pipelines for infrastructure deployment and orchestration.
  6. Ensure security, compliance and best practices across infrastructure.
  7. Manage incident response related to Infrastructure resources (GPU, CPU, Storage, Network).
  8. Handle customer provisioning requests for GPU resources, including onboarding, configuration and troubleshooting; resolve customer service requests related to GPU infrastructure, ensuring high customer satisfaction.
  9. Stay current with emerging GPU hardware and software technologies, integrating improvements as appropriate.
  10. Regional/international travel to GMI data center locations.

Qualifications

  1. Bachelor’s degree in Computer Science or related field.
  2. Over 3+ years of experience in data center operations, infrastructure, or systems engineering.
  3. Proven experience in site reliability engineering and infrastructure automation (e.g. Ansible, Terraform)
  4. Familiarity with containers orchestration platform (e.g. Kubernetes, Nvidia GPU operator, Nvidia Network operator, CNI, CSI) and job scheduling systems (e.g. Slurm).
  5. Familiarity with Linux system administration and scripting (Python, Bash).
  6. Familiarity with logging and monitoring tools such as Prometheus, Grafana, Loki.
  7. Good knowledge of GPU architecture, Nvidia CUDA, NCCL, or related AI/ML frameworks - added advantage.
  8. Strong troubleshooting skills and ability to analyze system logs and performance metrics.
  9. Excellent communication and teamwork abilities.

Meeting every qualification is not required—if you’re excited about this role, we’d love to hear from you. We believe diverse perspectives and experiences strengthen our team.

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer (SRE) - K8S/Slurm in United States vacancy
  •  ...our communities.This is a Software Engineering position at Director level, which is...  ...role is for an experienced and driven Site Reliability Engineer (SRE) to join our AI Platform team to...  ...distributed GPU cluster scheduling (e.g. Slurm, Kubernetes GPU scheduling)... 
    Suggested

    Morgan Stanley

    Alpharetta, GA
    3 days ago
  •  ...Primary Location: Louisville, Kentucky V-Soft Consulting is currently hiring for a SRE-Site Reliability Engineer for our premier client in Louisville, Kentucky . Education And Experience » ~6+ years in IT infrastructure, with 3+ years focused on site reliability... 
    Suggested
    Permanent employment
    Full time
    Work experience placement
    Currently hiring
    Local area

    V-Soft Consulting Group, Inc.

    Louisville, KY
    10 hours ago
  •  ...We are seeking an experienced Site Reliability Engineer (SRE) to support and maintain production systems hosted on AWS. The role focuses on production support, incident management, monitoring, observability, troubleshooting, and improving system reliability and availability... 
    Suggested
    Contract work

    2T Consulting

    Atlanta, GA
    a month ago
  •  ...prioritize a diverse F5 community where each individual can thrive.Role SummaryWe are seeking a proactive and detail-oriented Site Reliability Engineer II (SRE II) to join our 24/7 Operations team in a hybrid capacity. In this role, you will provide round-the-clock, eyes-on-... 
    Suggested
    Full time
    Local area
    Immediate start
    Shift work
    Night shift
    Afternoon shift
    Weekday work

    F5 Networks

    Reston, VA
    1 hour ago
  • $70.8k - $131.4k

    Job DescriptionThomson Reuters is strengthening its Site Reliability Engineering capability to help engineering and operations teams build, operate...  ...review is needed.Key ResponsibilitiesSupport and maintain SRE operational tooling, including dashboards, alerts, runbooks... 
    Suggested
    Full time
    Work at office
    Local area
    Flexible hours

    Thomson Reuters

    Saint Paul, MN
    2 days ago
  •  ...unwavering security to responsibly propel the global lottery industry ever forward.Position SummaryWe are looking for a skilled Site Reliability Engineer (SRE) to enhance the stability, performance, and reliability of our production systems. The SRE will work closely with... 
    Permanent employment
    Full time
    Work experience placement
    Local area

    Scientific Games Corporation

    Alpharetta, GA
    2 days ago
  •  ...the selected candidate for this role to work on site in the specified location(s).As a Senior Reliability Engineer, you will help shape the reliability, scalability...  ...teams, you will apply Site Reliability Engineering (SRE) principles to improve system availability,... 
    Full time
    Work at office

    The Charles Schwab Corporation

    Austin, TX
    3 days ago
  • $175k - $215k

     ...experiences — and we’re constantly looking for new ways to enhance these exciting experiences.Sr. Manager, Site Reliability Engineer provides strategic leadership across multiple SRE teams and their managers, ensuring alignment with organizational priorities and functional... 

    Disney Interactive

    Orlando, FL
    4 days ago
  • $135k - $155k

     ...buyers at Fortune 1000 companies to tap into global manufacturing capacity.Xometry is seeking a Site Reliability Engineer II to join our Site Reliability Engineering (SRE) Organization. In this role as an individual contributor, you will guide the reliability and performance... 
    Flexible hours

    Thomas

    Denver, CO
    1 day ago
  •  ...Period of performance: Up to 2 years in duration MUST HAVES: Minimum of 8 years of experience as a Site Reliability Engineer with a strong understanding of SRE principles for highly scalable and reliable systems Possess a bachelor's degree Experience working... 
    Local area
    Relocation package
    3 days per week

    Beyond SOF

    Vienna, VA
    4 days ago
  • $100k - $200k

     ...OPPO US Research Center is seeking a skilled and proactive Site Reliability Engineer (SRE) to join our team. In this role, you will be responsible for ensuring the stability, scalability, and performance of our application systems. The ideal candidate is passionate about... 
    Full time

    OPPO

    Palo Alto, CA
    2 days ago
  •  ...We're seeking an SRE to ensure the reliability and performance of our clients' critical systems. You'll work on observability, incident response...  ...management Nice to have Experience with chaos engineering Knowledge of distributed systems Background in high... 
    Remote work
    Flexible hours

    ACI Infotech

    Seattle, WA
    4 days ago
  •  ...We’re seeking a highly skilled Site Reliability Engineer (SRE) to join our engineering team and help ensure the reliability, scalability, and performance of our systems. As an SRE, you’ll blend software engineering with systems engineering to build and maintain resilient... 
    Temporary work
    Interim role
    Remote work
    Flexible hours

    OutSolve - Beyond Compliance

    Mission, KS
    1 day ago
  •  ...,000 companies and 100,000 electronics engineers worldwide use Altium We are growing,...  ...industry Role Overview: Senior Site Reliability Engineerensures the reliability, availability...  ...shared responsibility model where the SRE team collaborates with and educates... 
    Worldwide

    Renesas

    San Diego, CA
    1 day ago
  • $170k - $230k

     ...Site Reliability Engineer (SRE) Palo Alto / San Francisco Bay Area About Mithril Mithril is an AI infrastructure platform built to make GPU compute more accessible and affordable for the world's leading enterprises, AI startups, and the AI research community,... 
    Work at office
    Local area
    1 day per week

    Mithril

    Palo Alto, CA
    1 day ago
  • $165k - $225k

     ...Sr. Site Reliability Engineer (SRE) Chicago, IL or Remote Moonlite delivers high-performance AI infrastructure for organizations running intensive computational research, large-scale model training, and demanding data processing workloads. We provide infrastructure... 
    Remote work
    Flexible hours

    Moonlite AI

    Chicago, IL
    10 hours ago
  •  ...Site Reliability Engineer (SRE) Location: North Little Rock AR (onsite) Duration: Contract Required/Desired Skills: • Strong web development skills with a strong focus in C#/.NET • Someone who currently works in a hybrid skillset of BOTH.Net development AND... 
    Contract work

    Software Technology Inc

    North Little Rock, AR
    10 hours ago
  • $142.3k - $263.3k

     ...recently the new LIDAR iPad sensor. We are looking for the right Site Reliability Engineer to help us take our efforts to the next level. In this role,...  ...Computer Vision Organization. As a main contributor to our SRE team you will develop and maintain infrastructure, tooling,... 
    Work experience placement
    Relocation

    Apple

    San Diego, CA
    2 days ago
  •  ...globe. Join us on this journey to redefine resource management—and change lives along the way. The Role As a Site Reliability Engineer (SRE) at Air Apps, you will be responsible for ensuring the reliability, availability, and scalability of our systems. You... 
    Temporary work
    Worldwide

    airapps

    San Francisco, CA
    3 hours ago
  •  ...A Site Reliability Engineer (SRE) is responsible for ensuring the reliability, scalability, and performance of an organization's software systems and cloud infrastructure. The role combines software engineering with IT operations to automate processes, monitor system... 

    Mybridge

    Seattle, WA
    1 day ago
  • $100k - $180k

     ...Site Reliability Engineer (SRE) - Remote Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States. This is a fantastic opportunity to join an established... 
    Full time
    H1b
    Local area
    Immediate start
    Remote work
    Visa sponsorship

    Bright Vision Technologies

    United States
    1 day ago
  • $120k - $175k

     ...Senior Site Reliability Engineer (SRE) Atlanta, GA preferred, Remote At PrizePicks, we are the fastest-growing sports company in North America, as recognized by Inc. 5000. As the leading platform for Daily Fantasy Sports, we cover a diverse range of sports leagues... 
    Full time
    Remote work
    Work visa
    Flexible hours

    PrizePicks

    Atlanta, GA
    1 day ago
  •  ...Senior Site Reliability Engineer (SRE) Our client is a global technology consulting and digital solutions company that enables enterprises across industries to reimagine business models, accelerate innovation, and maximize growth by harnessing digital technologies.... 
    Local area

    E-Solutions

    New York, NY
    1 day ago
  •  ...Purple Drive Site Reliability Engineer (SRE) Contractual Atlanta, GA Key Highlights: Proven expertise in Google Cloud Platform (GCP) services, including BigQuery, Cloud Logging, IAM, and Service Accounts. Strong background in provisioning, monitoring, and... 

    Purple Drive

    Atlanta, GA
    1 day ago
  • $110k - $120k

     ...documentation that you are a U.S. citizen to qualify. Summary: The Site Reliability Engineer owns the day-to-day health, security, and usability of the organization's ICAM environment while applying SRE practices to keep delivery platforms reliable, observable, and... 
    Contract work
    Temporary work
    For contractors
    Work at office
    Local area
    Flexible hours

    Cherokee Federal

    Boulder, CO
    3 days ago
  •  ...Role: Site Reliability Engineer (SRE) Location: Brentwood, TN (Onsite) Contract Experience: 6-8+ years Role Description: Combines software engineering and IT operations to ensure the reliability, scalability, and performance of systems, with... 
    Contract work

    AceStack LLC

    Brentwood, TN
    10 hours ago
  •  ...We are seeking a Site Reliability Engineer (SRE) to improve the reliability, scalability, and performance of cloud-based applications. The ideal candidate will automate infrastructure, monitor production systems, troubleshoot incidents, and collaborate with software engineering... 

    Mybridge

    Austin, TX
    1 day ago
  • $40 - $80 per hour

     ...Job Title: Site Reliability Engineer (SRE) Duration (Contract): 12 Months Client Location: Southlake, TX Location Preference: Onsite Job Description: As a Site Reliability Engineer (SRE) , you will be responsible for improving the reliability, scalability... 
    Hourly pay
    Contract work

    Smart IMS

    Southlake, TX
    1 day ago
  •  ...Open role Site Reliability Engineer (SRE) San Francisco, CA (On-site) Responsibilities Develop and maintain advanced monitoring, alerting, and self-healing mechanisms that detect and address issues before they impact customers. Perform regular capacity... 

    Methodic

    San Francisco, CA
    1 day ago
  •  ...in Computer Science, Information Technology, Engineering, or equivalent field ~3-5 years of experience in Site Reliability Engineering, Production Support, Platform Engineering...  ...application health ~ Understanding of SRE principles, including observability,... 
    Remote work

    Anveta

    United States
    10 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer (SRE) - K8S/Slurm. Be the first to apply!