Site Reliability Engineer
GMI Cloud
About GMI
GMI Cloud is a fast-growing, AI-native infrastructure company delivering high-performance GPU compute, inference services, and infrastructure for AI agents.
Following 8x ARR growth, GMI Cloud continues to scale rapidly across the U.S. and APAC. As a Reference Platform NVIDIA Cloud Partner (NCP) and a validated leading NCP across both markets, we power production AI for leading AI-native companies including Fireworks AI, Cartesia, Reflection, and OpenRouter.
From large-scale compute to optimized inference and agentic workloads, GMI Cloud gives AI teams the infrastructure they need to build, deploy, and scale on one unified cloud.
One cloud for compute, inference, and agents.
Role Overview
We are seeking a skilled Site Reliability Engineer to join the GMI Global Infrastructure team. This role is hands-on and critical to ensuring the stability, efficiency, and reliability of the large-scale high performance AI/ML clusters in our data center. The ideal candidate will bring expertise in system-level troubleshooting, AI cluster maintenance, and operational excellence to ensure maximum performance for our infrastructure. Experience with large-scale infrastructure automation is considered a strong plus.
Responsibilities
- Design, implement and maintain scalable AI/ML infrastructure solutions.
- Proactively monitor GPU cluster health, performance and troubleshoot issues across compute, accelerator, and storage systems.
- Automate deployment, configuration and management of infrastructure resources.
- Manage GPU node lifecycle workflows, including provisioning, scaling, maintenance, decommissioning and upgrades of GPU nodes.
- Implement CI/CD pipelines for infrastructure deployment and orchestration.
- Ensure security, compliance and best practices across infrastructure.
- Manage incident response related to Infrastructure resources (GPU, CPU, Storage, Network).
- Handle customer provisioning requests for GPU resources, including onboarding, configuration and troubleshooting; resolve customer service requests related to GPU infrastructure, ensuring high customer satisfaction.
- Stay current with emerging GPU hardware and software technologies, integrating improvements as appropriate.
- Regional/international travel to GMI data center locations.
Qualifications
- Bachelor’s degree in Computer Science or related field.
- Over 3+ years of experience in data center operations, infrastructure, or systems engineering.
- Proven experience in site reliability engineering and infrastructure automation (e.g. Ansible, Terraform)
- Familiarity with containers orchestration platform (e.g. Kubernetes, Nvidia GPU operator, Nvidia Network operator, CNI, CSI) and job scheduling systems (e.g. Slurm).
- Familiarity with Linux system administration and scripting (Python, Bash).
- Familiarity with logging and monitoring tools such as Prometheus, Grafana, Loki.
- Good knowledge of GPU architecture, Nvidia CUDA, NCCL, or related AI/ML frameworks - added advantage.
- Strong troubleshooting skills and ability to analyze system logs and performance metrics.
- Excellent communication and teamwork abilities.
Meeting every qualification is not required—if you’re excited about this role, we’d love to hear from you. We believe diverse perspectives and experiences strengthen our team.
- The Senior Site Reliability Engineer is responsible for improving the reliability, availability, scalability, and operational excellence of our critical infrastructure platforms and services. This role partners closely with Engineering, Security, and Infrastructure teams...SuggestedFull timeWork at officeLocal area
- As a Site Reliability Engineer, you will be responsible for: Operational Excellence & Incident Management- Maintain and monitor production systems for availability, latency, and performance.- Lead incident response efforts, including communication, resolution, and postmortem...SuggestedPermanent employment
- Reliability Engineering Design, implement, and operate scalable, resilient, and highly available systems on Google Cloud Platform. Improve service... ...Skills, and Abilities Three or more years of experience in Site Reliability Engineering, platform engineering, DevOps, cloud...SuggestedRemote work
- ...and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Corporate Technology, Risk Technology team, you will solve complex and broad business problems...Suggested
- ...Title:Site Reliability Engineer Location: Houston, TX 77002 (Hybrid: 3 days onsite / 2 days remote) Duration: Contract to Hire Work Requirements:U.S.Citizen, GC Holders,or Authorized to Work in the U.S. Job Description The Site Reliability Engineer is a founding...SuggestedContract workRemote workFlexible hours
$136.2k - $214.01k
...outcomes Visionary in future focused problem-solving Exceptional in execution and impact The Role As a Senior Site Reliability Engineer at Proofpoint you will develop a deep understanding of the various services and applications that come together to...Full timeFlexible hours- ...Site Reliability Engineer The SRE manages and maintains infrastructure which supports cloud-based prototype applications as they transition into a production environment. The SRE assists in defining and measuring service level agreements and non-functional requirements...
- ...applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems. As a Site Reliability Engineer III at JPMorgan Chase within the Corporate Technology, Corporate Know Your Customer (KYC) team, you will solve complex and...
- ...As a Site Reliability Engineer, you will be responsible for: Operational Excellence & Incident Management Maintain and monitor production systems for availability, latency, and performance. Lead incident response efforts, including communication, resolution...Permanent employment
$55k - $151.47k
...ApplicableSpecialismIFS - Internal Firm Services - OtherManagement LevelSenior AssociateJob Description & SummaryThe OpportunityAs a Site Reliability Engineer - Senior Associate, you will play a pivotal role in enhancing the reliability, scalability, and performance of our...Full timeH1b- As an Entry-Level DevOps Site Reliability Engineer, you will join a team responsible for continuous improvement and support of customer facing products. Responsibilities will include collecting system requirements; improving existing tools and processes through scripting...Work from home2 days per week
- ...ENGINEERLocation: HOUSTON, TXFLSA Class: EXEMPTResponsible to: Directo of Software EngineeringPosition Summary: DevOps / Site Reliability Engineer to implement and evolve the infrastructure, deployment pipelines, and reliability posture of our systems. You'll work closely...Full timeLocal area
- ...globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability. As a Lead Site Reliability Engineer at JPMorgan Chase within the Corporate & Investment Bank (CIB) Management and Support Functions Digital & Platform...
- As a Lead Site Reliability Engineer at JPMorgan Chase within the Corporate Know Your Customer (KYC) Technology group, you hold a leadership role in your team, demonstrate strong knowledge across multiple technical domains, and advise others on the technical and business...
- ...globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability. As a Lead Site Reliability Engineer at JPMorgan Chase within the Corporate & Investment Bank (CIB) Management and Support Functions Digital &...
- ...contributing to revolutionary projects. You've discovered the perfect environment to have a major impact. As a Principal Site Reliability Engineer at JPMorgan Chase within the Corporate Technology Team, you draw upon your advanced knowledge to identify new...
$213.1k - $300k
Lead a team of engineers to maintain service uptime while managing global on-call rotations... ...improve operational practices to drive reliability, maintainability, and stakeholder alignment... ...or in a Manager, Software Engineer, Site Reliability Engineering-related occupation...Full timeWork at office- ...JPMorgan Chase is seeking a Lead Site Reliability Engineer to shape the future of reliability for a globally recognized firm within the Corporate & Investment Bank domain. The role emphasizes leadership across multiple technical domains and mentoring peers. You will...
$61k - $101k
...formal training or certification in software engineering concepts, along with 5+ years of applied... .... We need deep expertise in reliability, scalability, performance, security, enterprise... ...architecture, toil reduction, and other site reliability practices, with the ability...Full time- ...globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability. As a Lead Site Reliability Engineer at JPMorgan Chase within the Corporate & Investment Bank (CIB) Management and Support Functions Digital &...
- ...Talentify is seeking a Site Reliability Engineer in Houston to own reliability, observability, and performance across distributed systems. You will lead incident response, design health checks, and drive postmortems while collaborating with developers to evolve architecture...
- ...Please extend your support for this role. Local candidate will get 1st preference. Job Title: SRE Engineer Location: Houston, TX and Jersey City, NJ - 3 Days Onsite Role FTE role with Mphasis Client: Mphasis H1B transfer will work...Work experience placementH1bLocal area
$113k - $141.53k
...leader in global energy. Senior Solutions Engineer - Systems Integration serves as a... ...functionally in the field, ensuring safe, reliable, and performant operation across diverse... ...Willingness to travel to factories and project sites (25%).Preferred QualificationsMaster’s degree...Full timeFor contractorsLocal areaRemote workWorldwideFlexible hours$94k - $112.63k
...global energy. AES is seeking a Solutions Engineer - Systems Integration to contribute to... ...interface correctly and operate safely, reliably, and as intended across project environments... ...to factories, laboratories, and project sites to support inspections, testing, commissioning...Full timeFor contractorsLocal areaWorldwide- Position: Software Engineer- Flight & Ground Systems Location: Houston, TX Remote Status: On-Site Job Id: 866 # of Openings... ...requirements and translate them into reliable software solutions.Self-motivated and...Permanent employmentFull timeRemote workRelocation package
$76k - $155.7k
...Required: Up to 10%Type of Travel: Continental US* * *The Opportunity:CACI is looking for an experienced Flight Software Systems Engineer to support NASA’s Moon Base flight software development at the Johnson Space Center. The Moon Base is humanity’s first permanent lunar...Permanent employmentContract workFor contractorsWork experience placementFlexible hours- Lead ETRM Systems Developer:On behalf of our Energy client, Procom is searching for a Lead ETRM Systems Developer for a permanent role. This position is onsite at our client’s Houston, Texas office. Lead ETRM Systems Developer - Job Description:This role involves supporting...Permanent employmentWork at officeImmediate start
$51 - $61 per hour
...onsite at the project, significantly reducing and/or eliminating the demands to travel. Key Responsibilities: As a Release Train Engineer, you will be responsible for facilitating Agile Release Train events and processes including communicating with stakeholders escalating...Hourly payLive inWork at officeLocal areaImmediate startFlexible hoursShift work- ...Release Train Engineer 4 Months- Contract To Hire Pay- $65-$70 W2 Onsite Houston, TX Job Description The Release Train... ...sure all team activity, dashboards, and metrics are visible and reliable. Coach teams and Scrum Masters on agile and Scrum practices,...Contract workWork at office
- ...data platforms that support high-volume, data-intensive workflows.The team works across backend engineering, infrastructure, and data systems, collaborating to deliver reliable, high-performance services in a modern cloud-native environment.Key Responsibilities-Backend...Full timeFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
- site reliability engineer sre Houston, TX
- site reliability engineer Houston, TX
- official site Houston, TX
- remote website tester Houston, TX
- site services specialist Houston, TX
- construction site safety Houston, TX
- IT site lead Houston, TX
- site recruiter Houston, TX
- site leader Houston, TX
- site safety Houston, TX



