Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer

GMI Cloud

About GMI

GMI Cloud is a fast-growing, AI-native infrastructure company delivering high-performance GPU compute, inference services, and infrastructure for AI agents.

Following 8x ARR growth, GMI Cloud continues to scale rapidly across the U.S. and APAC. As a Reference Platform NVIDIA Cloud Partner (NCP) and a validated leading NCP across both markets, we power production AI for leading AI-native companies including Fireworks AI, Cartesia, Reflection, and OpenRouter.

From large-scale compute to optimized inference and agentic workloads, GMI Cloud gives AI teams the infrastructure they need to build, deploy, and scale on one unified cloud.

One cloud for compute, inference, and agents.

Role Overview

We are seeking a skilled Site Reliability Engineer to join the GMI Global Infrastructure team. This role is hands-on and critical to ensuring the stability, efficiency, and reliability of the large-scale high performance AI/ML clusters in our data center. The ideal candidate will bring expertise in system-level troubleshooting, AI cluster maintenance, and operational excellence to ensure maximum performance for our infrastructure. Experience with large-scale infrastructure automation is considered a strong plus.

Responsibilities

  • Design, implement and maintain scalable AI/ML infrastructure solutions.
  • Proactively monitor GPU cluster health, performance and troubleshoot issues across compute, accelerator, and storage systems.
  • Automate deployment, configuration and management of infrastructure resources.
  • Manage GPU node lifecycle workflows, including provisioning, scaling, maintenance, decommissioning and upgrades of GPU nodes.
  • Implement CI/CD pipelines for infrastructure deployment and orchestration.
  • Ensure security, compliance and best practices across infrastructure.
  • Manage incident response related to Infrastructure resources (GPU, CPU, Storage, Network).
  • Handle customer provisioning requests for GPU resources, including onboarding, configuration and troubleshooting; resolve customer service requests related to GPU infrastructure, ensuring high customer satisfaction.
  • Stay current with emerging GPU hardware and software technologies, integrating improvements as appropriate.
  • Regional/international travel to GMI data center locations.

Qualifications

  • Bachelor’s degree in Computer Science or related field.
  • Over 3+ years of experience in data center operations, infrastructure, or systems engineering.
  • Proven experience in site reliability engineering and infrastructure automation (e.g. Ansible, Terraform)
  • Familiarity with containers orchestration platform (e.g. Kubernetes, Nvidia GPU operator, Nvidia Network operator, CNI, CSI) and job scheduling systems (e.g. Slurm).
  • Familiarity with Linux system administration and scripting (Python, Bash).
  • Familiarity with logging and monitoring tools such as Prometheus, Grafana, Loki.
  • Good knowledge of GPU architecture, Nvidia CUDA, NCCL, or related AI/ML frameworks - added advantage.
  • Strong troubleshooting skills and ability to analyze system logs and performance metrics.
  • Excellent communication and teamwork abilities.

Meeting every qualification is not required—if you’re excited about this role, we’d love to hear from you. We believe diverse perspectives and experiences strengthen our team.

Vacancy posted 19 hours ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in Houston, TX vacancy
  • The Senior Site Reliability Engineer is responsible for improving the reliability, availability, scalability, and operational excellence of our critical infrastructure platforms and services. This role partners closely with Engineering, Security, and Infrastructure teams... 
    Suggested
    Full time
    Work at office
    Local area

    Castleton Commodities International

    Houston, TX
    3 days ago
  • As a Site Reliability Engineer, you will be responsible for: Operational Excellence & Incident Management- Maintain and monitor production systems for availability, latency, and performance.- Lead incident response efforts, including communication, resolution, and postmortem... 
    Suggested
    Permanent employment

    National Oilwell Varco

    Houston, TX
    4 days ago
  • Reliability Engineering Design, implement, and operate scalable, resilient, and highly available systems on Google Cloud Platform. Improve service...  ...Skills, and Abilities Three or more years of experience in Site Reliability Engineering, platform engineering, DevOps, cloud... 
    Suggested
    Remote work

    Patterson-UTI

    Houston, TX
    2 days ago
  •  ...and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Corporate Technology, Risk Technology team, you will solve complex and broad business problems... 
    Suggested

    JP Morgan Chase

    Houston, TX
    1 day ago
  •  ...Title:Site Reliability Engineer Location: Houston, TX 77002 (Hybrid: 3 days onsite / 2 days remote) Duration: Contract to Hire Work Requirements:U.S.Citizen, GC Holders,or Authorized to Work in the U.S. Job Description The Site Reliability Engineer is a founding... 
    Suggested
    Contract work
    Remote work
    Flexible hours

    INSPYR Solutions

    Houston, TX
    3 days ago
  • $136.2k - $214.01k

     ...outcomes Visionary in future focused problem-solving Exceptional in execution and impact The Role As a Senior Site Reliability Engineer at Proofpoint you will develop a deep understanding of the various services and applications that come together to... 
    Full time
    Flexible hours

    Proofpoint

    Houston, TX
    1 day ago
  •  ...Site Reliability Engineer The SRE manages and maintains infrastructure which supports cloud-based prototype applications as they transition into a production environment. The SRE assists in defining and measuring service level agreements and non-functional requirements... 

    ClifyX

    Houston, TX
    2 days ago
  •  ...applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems. As a Site Reliability Engineer III at JPMorgan Chase within the Corporate Technology, Corporate Know Your Customer (KYC) team, you will solve complex and... 

    JPMorgan Chase Bank, N.A.

    Houston, TX
    2 days ago
  •  ...As a Site Reliability Engineer, you will be responsible for: Operational Excellence & Incident Management Maintain and monitor production systems for availability, latency, and performance. Lead incident response efforts, including communication, resolution... 
    Permanent employment

    Talentify.io

    Houston, TX
    3 days ago
  • $55k - $151.47k

     ...ApplicableSpecialismIFS - Internal Firm Services - OtherManagement LevelSenior AssociateJob Description & SummaryThe OpportunityAs a Site Reliability Engineer - Senior Associate, you will play a pivotal role in enhancing the reliability, scalability, and performance of our... 
    Full time
    H1b

    PwC

    Houston, TX
    3 days ago
  • As an Entry-Level DevOps Site Reliability Engineer, you will join a team responsible for continuous improvement and support of customer facing products. Responsibilities will include collecting system requirements; improving existing tools and processes through scripting... 
    Work from home
    2 days per week

    Reynolds & Reynolds

    Houston, TX
    3 days ago
  •  ...ENGINEERLocation: HOUSTON, TXFLSA Class: EXEMPTResponsible to: Directo of Software EngineeringPosition Summary: DevOps / Site Reliability Engineer to implement and evolve the infrastructure, deployment pipelines, and reliability posture of our systems. You'll work closely... 
    Full time
    Local area

    VoltaGrid

    Houston, TX
    1 day ago
  •  ...globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability. As a Lead Site Reliability Engineer at JPMorgan Chase within the Corporate & Investment Bank (CIB) Management and Support Functions Digital & Platform... 

    JP Morgan Chase

    Houston, TX
    4 days ago
  • As a Lead Site Reliability Engineer at JPMorgan Chase within the Corporate Know Your Customer (KYC) Technology group, you hold a leadership role in your team, demonstrate strong knowledge across multiple technical domains, and advise others on the technical and business... 

    JP Morgan Chase

    Houston, TX
    13 hours ago
  •  ...globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability. As a Lead Site Reliability Engineer at JPMorgan Chase within the Corporate & Investment Bank (CIB) Management and Support Functions Digital &... 

    Talentify.io

    Houston, TX
    3 days ago
  •  ...contributing to revolutionary projects. You've discovered the perfect environment to have a major impact. As a Principal Site Reliability Engineer at JPMorgan Chase within the Corporate Technology Team, you draw upon your advanced knowledge to identify new... 

    Next Frontier Capital

    Houston, TX
    4 days ago
  • $213.1k - $300k

    Lead a team of engineers to maintain service uptime while managing global on-call rotations...  ...improve operational practices to drive reliability, maintainability, and stakeholder alignment...  ...or in a Manager, Software Engineer, Site Reliability Engineering-related occupation... 
    Full time
    Work at office

    Google

    Houston, TX
    13 hours ago
  •  ...JPMorgan Chase is seeking a Lead Site Reliability Engineer to shape the future of reliability for a globally recognized firm within the Corporate & Investment Bank domain. The role emphasizes leadership across multiple technical domains and mentoring peers. You will... 

    Jobleads-US

    Houston, TX
    3 days ago
  • $61k - $101k

     ...formal training or certification in software engineering concepts, along with 5+ years of applied...  .... We need deep expertise in reliability, scalability, performance, security, enterprise...  ...architecture, toil reduction, and other site reliability practices, with the ability... 
    Full time

    J.P. Morgan

    Houston, TX
    a month ago
  •  ...globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability.   As a Lead Site Reliability Engineer at JPMorgan Chase within the Corporate & Investment Bank (CIB) Management and Support Functions Digital &... 

    JPMorgan Chase & Co.

    Houston, TX
    3 days ago
  •  ...Talentify is seeking a Site Reliability Engineer in Houston to own reliability, observability, and performance across distributed systems. You will lead incident response, design health checks, and drive postmortems while collaborating with developers to evolve architecture... 

    Jobleads-US

    Houston, TX
    4 days ago
  •  ...Please extend your support for this role. Local candidate will get 1st preference. Job Title: SRE Engineer Location: Houston, TX and Jersey City, NJ - 3 Days Onsite Role FTE role with Mphasis Client: Mphasis H1B transfer will work... 
    Work experience placement
    H1b
    Local area

    Trinity Technology Solutions

    Houston, TX
    3 days ago
  • $113k - $141.53k

     ...leader in global energy. Senior Solutions Engineer - Systems Integration serves as a...  ...functionally in the field, ensuring safe, reliable, and performant operation across diverse...  ...Willingness to travel to factories and project sites (25%).Preferred QualificationsMaster’s degree... 
    Full time
    For contractors
    Local area
    Remote work
    Worldwide
    Flexible hours

    AES

    Houston, TX
    2 days ago
  • $94k - $112.63k

     ...global energy. AES is seeking a Solutions Engineer - Systems Integration to contribute to...  ...interface correctly and operate safely, reliably, and as intended across project environments...  ...to factories, laboratories, and project sites to support inspections, testing, commissioning... 
    Full time
    For contractors
    Local area
    Worldwide

    AES

    Houston, TX
    2 days ago
  • Position: Software Engineer- Flight & Ground Systems Location: Houston, TX Remote Status: On-Site Job Id: 866 # of Openings...  ...requirements and translate them into reliable software solutions.Self-motivated and... 
    Permanent employment
    Full time
    Remote work
    Relocation package

    Aegis Aerospace

    Houston, TX
    2 days ago
  • $76k - $155.7k

     ...Required: Up to 10%Type of Travel: Continental US* * *The Opportunity:CACI is looking for an experienced Flight Software Systems Engineer to support NASA’s Moon Base flight software development at the Johnson Space Center. The Moon Base is humanity’s first permanent lunar... 
    Permanent employment
    Contract work
    For contractors
    Work experience placement
    Flexible hours

    CACI International

    Houston, TX
    4 days ago
  • Lead ETRM Systems Developer:On behalf of our Energy client, Procom is searching for a Lead ETRM Systems Developer for a permanent role. This position is onsite at our client’s Houston, Texas office. Lead ETRM Systems Developer - Job Description:This role involves supporting...
    Permanent employment
    Work at office
    Immediate start

    ProCom

    Houston, TX
    2 days ago
  • $51 - $61 per hour

     ...onsite at the project, significantly reducing and/or eliminating the demands to travel. Key Responsibilities: As a Release Train Engineer, you will be responsible for facilitating Agile Release Train events and processes including communicating with stakeholders escalating... 
    Hourly pay
    Live in
    Work at office
    Local area
    Immediate start
    Flexible hours
    Shift work

    Accenture

    Houston, TX
    13 hours ago
  •  ...Release Train Engineer 4 Months- Contract To Hire Pay- $65-$70 W2 Onsite Houston, TX Job Description The Release Train...  ...sure all team activity, dashboards, and metrics are visible and reliable. Coach teams and Scrum Masters on agile and Scrum practices,... 
    Contract work
    Work at office

    Anveta

    Houston, TX
    13 hours ago
  •  ...data platforms that support high-volume, data-intensive workflows.The team works across backend engineering, infrastructure, and data systems, collaborating to deliver reliable, high-performance services in a modern cloud-native environment.Key Responsibilities-Backend... 
    Full time
    Flexible hours

    CGG

    Houston, TX
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!