Site Reliability Engineer
GMI Cloud
About GMI
GMI Cloud is a fast-growing, AI-native infrastructure company delivering high-performance GPU compute, inference services, and infrastructure for AI agents.
Following 8x ARR growth, GMI Cloud continues to scale rapidly across the U.S. and APAC. As a Reference Platform NVIDIA Cloud Partner (NCP) and a validated leading NCP across both markets, we power production AI for leading AI-native companies including Fireworks AI, Cartesia, Reflection, and OpenRouter.
From large-scale compute to optimized inference and agentic workloads, GMI Cloud gives AI teams the infrastructure they need to build, deploy, and scale on one unified cloud.
One cloud for compute, inference, and agents.
Role Overview
We are seeking a skilled Site Reliability Engineer to join the GMI Global Infrastructure team. This role is hands-on and critical to ensuring the stability, efficiency, and reliability of the large-scale high performance AI/ML clusters in our data center. The ideal candidate will bring expertise in system-level troubleshooting, AI cluster maintenance, and operational excellence to ensure maximum performance for our infrastructure. Experience with large-scale infrastructure automation is considered a strong plus.
Responsibilities
- Design, implement and maintain scalable AI/ML infrastructure solutions.
- Proactively monitor GPU cluster health, performance and troubleshoot issues across compute, accelerator, and storage systems.
- Automate deployment, configuration and management of infrastructure resources.
- Manage GPU node lifecycle workflows, including provisioning, scaling, maintenance, decommissioning and upgrades of GPU nodes.
- Implement CI/CD pipelines for infrastructure deployment and orchestration.
- Ensure security, compliance and best practices across infrastructure.
- Manage incident response related to Infrastructure resources (GPU, CPU, Storage, Network).
- Handle customer provisioning requests for GPU resources, including onboarding, configuration and troubleshooting; resolve customer service requests related to GPU infrastructure, ensuring high customer satisfaction.
- Stay current with emerging GPU hardware and software technologies, integrating improvements as appropriate.
- Regional/international travel to GMI data center locations.
Qualifications
- Bachelor’s degree in Computer Science or related field.
- Over 3+ years of experience in data center operations, infrastructure, or systems engineering.
- Proven experience in site reliability engineering and infrastructure automation (e.g. Ansible, Terraform)
- Familiarity with containers orchestration platform (e.g. Kubernetes, Nvidia GPU operator, Nvidia Network operator, CNI, CSI) and job scheduling systems (e.g. Slurm).
- Familiarity with Linux system administration and scripting (Python, Bash).
- Familiarity with logging and monitoring tools such as Prometheus, Grafana, Loki.
- Good knowledge of GPU architecture, Nvidia CUDA, NCCL, or related AI/ML frameworks - added advantage.
- Strong troubleshooting skills and ability to analyze system logs and performance metrics.
- Excellent communication and teamwork abilities.
Meeting every qualification is not required—if you’re excited about this role, we’d love to hear from you. We believe diverse perspectives and experiences strengthen our team.
- ...and best in class outcomesVisionary in future focused problem-solvingExceptional in execution and impactThe RoleAs a Senior Site Reliability Engineer at Proofpoint you will develop a deep understanding of the various services and applications that come together to deliver...SuggestedFull timeFlexible hours
- ...If you want to be part of an inclusive, adaptable, and forward-thinking organization, apply now.We are currently seeking a Site Reliability Engineer to join our team in Westlake, Texas (US-TX), United States (US).Position OverviewWe are seeking a highly skilled Site...SuggestedTemporary workWork at officeRemote workFlexible hours
- ...and foster a dynamic work environment where new ideas thrive. Are you ready to join our team and make an impact?As a Senior Site Reliability Engineer at TeamViewer, you’ll be a key player in ensuring the reliability, scalability, and performance of our Azure-based SaaS...SuggestedTemporary workCasual workWorldwide
$127k - $249k
The TeamPlatform Engineering sits within SRE and builds the core infrastructure powering MongoDB... ...plays a pivotal role in engineering the reliable, globally connected, multi-cloud network... ...are seeking a talented Senior Site Reliability Engineer (SRE) with a strong...SuggestedLocal areaRemote workWorldwideFlexible hours$98.58k - $138.02k
...Northern California / Silicon Valley Region / Denver, COProduct Engineering - DevOps /Full Time /HybridRestaurant365 is a SaaS company... ...office locations: Austin, TX; Irvine, CA; or Akron, OH. The Site Reliability Engineer II will be responsible for supporting, enhancing,...SuggestedFull timeWork at office$152k - $241.5k
...infrastructure platforms for automated host lifecycle management, fleet reliability/auto-healing, E2E observability or data-driven operations (... ...languages such as Python, Go, Perl, or Ruby.Mentored other engineers and influenced technical direction through design reviews,...Full time- ...Dimensional leverages the rapidly evolving state of the art to engineer scalable, innovative, and research driven solutions to improve... ...each of the developer tooling ecosystemsOwn the operational reliability of developer tooling ecosystems, including Python toolchains (...Full timeLocal area
$112.11k - $190.66k
...connected world more secure.This is for a hybrid role in Austin, TX.Position SummaryWe are looking for an experienced SRE - Site Reliability Engineer to work with our North American Team. Your responsibility will be to help design and build tools and infrastructure that...Full timeWork at officeLocal areaMonday to Friday- ...cutting-edge technology, including AI/ML, to enhance the reliability and performance of its critical applications. By joining this... ...power financial success. Unleash Your Expertise as a Site Reliability Engineer Are you a skilled engineer passionate about combining software...
- ...Operations and Maintenance Referrals increase your chances of interviewing at Infosys by 2x Get notified about new Site Reliability Engineer jobs in Austin, TX . Austin, TX $170,000.00-$190,000.00 1 day ago Site Reliability Engineer (SRE, Remote US)...Full timeRemote workRelocationAll shifts
- ...Job Title: Site Reliability Engineer Skill Level: Mid Level Employment Type: Full-Time Position Department: Engineering Reports To: Chief Technology Officer (CTO) Location: Austin, TX (HQ) (Optional, Based On Performance Remote & Hybrid) Overview...Full timeRemote work
$109.65k - $182.76k
...data to make the connected world more secure. Austin, TX - Hybrid (3 days a week) Position Summary We are seeking a Site Reliability Engineer to ensure the high level of service and operation excellence for the development of the innovative and ambitious...Full timeLocal area3 days per week- ...Role: Site Reliability Engineer Location: Southlake / Austin, TX - Onsite 4 days weekly Duration: 12 Months Job Summary We are seeking a motivated Site Reliability Engineer (Contractor) with 3 to 5 years of experience in automation, cloud infrastructure,...For contractors
$120k - $165k
...We provide tools, resources and support to enable users to reach their health goals. We are looking for a Software Engineer III - Site Reliability to join the MyFitnessPal PEAS team. As a member of the PEAS team, you'll have the opportunity to positively impact MyFitnessPal...Full timeContract workFor contractorsFor subcontractorWork at officeFlexible hours$121.4k - $218.6k
...complex challenges? Do you have a passion for automation and building systems that scale? Join our highly skilled Site Reliability Engineering team! Our team designs, develops, and manages applications and infrastructure that support Akamai Cloud's products...Work experience placementWork at office- ...configurationsKey Responsibilities:Build and operate scalable and reliable infrastructure.Collaborate with development teams to... ...insuranceVision insurance401(k)Get notified about new Site Reliability Engineer jobs in Austin, Texas Metropolitan AreaSite Reliability Engineer...Remote work
$136.2k - $214.01k
...outcomes Visionary in future focused problem-solving Exceptional in execution and impact The Role As a Senior Site Reliability Engineer at Proofpoint you will develop a deep understanding of the various services and applications that come together to...Full timeFlexible hours$130k - $153k
...our customers, and in our growing commitment to land stewardship and recreational access. WHAT YOU WILL DO onX is seeking a Site Reliability Engineer to build and maintain the infrastructure that enables our developers to ship reliably at scale. You\'ll manage onX\'s...Full timeLocal areaRemote workFlexible hours- ...Responsibilities Kforce has a client seeking a remote Site Reliability Engineer to join their team. We are seeking a highly motivated Site Reliability Engineer (SRE) to help build, scale, and maintain cloud infrastructure, CI/CD pipelines, and deployment automation...Hourly payContract workWork experience placementRemote work
$168k - $200k
...is passionate about creating transformative change in healthcare. What We're Looking For We're looking for a Senior Site Reliability Engineer to join our Data & ML Platform team. You'll be at the forefront of building and operating a resilient, observable, and...- ...Job Description Insight Global is seeking a Site Reliability Engineer to support one of our local government clients. This individual will serve as a subject matter expert for the Microsoft Power Platform, overseeing the governance, administration, security, automation...Local area
- ...safety, commercialization, and mass production to change the world for the better. JOB SUMMARY We are seeking an experienced Site Reliability Engineer to own and maintain the deployment of our cloud-based infrastructure to customer sites. In this role, you will work...Full timeLocal area
$75.7k - $136.3k
...solve complex challenges? Do you have a passion for automation and building systems that scale? Join our highly skilled Site Reliability Engineering team! Our team designs, develops, and manages applications and infrastructure that support Akamai Cloud's products and...Work experience placementWork at office- ...Job Summary We are seeking a Site Reliability Engineer to support and administer the agency's Microsoft Power Platform environment, with a primary focus on platform administration, reliability, governance, and optimization. This role is more heavily focused on hands...
$140k
...Title: Site Reliability Engineer SRE - ML platform Location: Austin, TX OR Sunnyvale, CA Type: FTE Salary/Rate : $140K Title: Site Reliability Engineer SRE - ML platform Responsibilities - Continuous Deployment using...- ...territory, and do work that genuinely matters, Future Secure AI is the place for you. About the Role We are looking for a Site Reliability Engineer to help design, build, and operate the platforms that power AI Co‑Workers. This is a hands‑on role for an engineer who...Flexible hours
- ...workplace embraces diversity and inclusion – it’s a place where you can grow, belong and thrive. Your day at NTT DATA The Site Reliability Engineer (SRE) is a seasoned subject matter expert, responsible for ensuring the reliability, availability, and performance of...Remote work
- ...Are you passionate about Site Reliability Engineering, automation, observability, and AI/ML-driven operations ? We’re looking for an experienced engineer who can help transform how mission-critical enterprise applications are deployed, monitored, and supported....
- ...Title: Site Reliability Engineer (SRE) Location: Austin, TX Description: We're searching for a driven Site Reliability Engineer (SRE) to join our innovative team. As an SRE, you'll be a cornerstone of our production software, ensuring our systems...Work experience placement
- ...Role: Site Reliability Engineer Rate: Location: Austin, TX (Hybrid 2 days onsite in a week, locals to TX) Visa: USC/GC/EAD/OPT Duration: 12+ months Client: Must have public sector (state client) at least on 1 project & focus on industry exp...Local area2 days per week
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
- site reliability engineer sre Austin, TX
- site reliability engineer Austin, TX
- site reliability engineer remote Austin, TX
- official site Austin, TX
- site merchandiser Austin, TX
- site services specialist Austin, TX
- construction site safety Austin, TX
- IT site lead Austin, TX
- site recruiter Austin, TX
- site leader Austin, TX


