Site Reliability Engineer
GMI Cloud
About GMI
GMI Cloud is a fast-growing, AI-native infrastructure company delivering high-performance GPU compute, inference services, and infrastructure for AI agents.
Following 8x ARR growth, GMI Cloud continues to scale rapidly across the U.S. and APAC. As a Reference Platform NVIDIA Cloud Partner (NCP) and a validated leading NCP across both markets, we power production AI for leading AI-native companies including Fireworks AI, Cartesia, Reflection, and OpenRouter.
From large-scale compute to optimized inference and agentic workloads, GMI Cloud gives AI teams the infrastructure they need to build, deploy, and scale on one unified cloud.
One cloud for compute, inference, and agents.
Role Overview
We are seeking a skilled Site Reliability Engineer to join the GMI Global Infrastructure team. This role is hands-on and critical to ensuring the stability, efficiency, and reliability of the large-scale high performance AI/ML clusters in our data center. The ideal candidate will bring expertise in system-level troubleshooting, AI cluster maintenance, and operational excellence to ensure maximum performance for our infrastructure. Experience with large-scale infrastructure automation is considered a strong plus.
Responsibilities
- Design, implement and maintain scalable AI/ML infrastructure solutions.
- Proactively monitor GPU cluster health, performance and troubleshoot issues across compute, accelerator, and storage systems.
- Automate deployment, configuration and management of infrastructure resources.
- Manage GPU node lifecycle workflows, including provisioning, scaling, maintenance, decommissioning and upgrades of GPU nodes.
- Implement CI/CD pipelines for infrastructure deployment and orchestration.
- Ensure security, compliance and best practices across infrastructure.
- Manage incident response related to Infrastructure resources (GPU, CPU, Storage, Network).
- Handle customer provisioning requests for GPU resources, including onboarding, configuration and troubleshooting; resolve customer service requests related to GPU infrastructure, ensuring high customer satisfaction.
- Stay current with emerging GPU hardware and software technologies, integrating improvements as appropriate.
- Regional/international travel to GMI data center locations.
Qualifications
- Bachelor’s degree in Computer Science or related field.
- Over 3+ years of experience in data center operations, infrastructure, or systems engineering.
- Proven experience in site reliability engineering and infrastructure automation (e.g. Ansible, Terraform)
- Familiarity with containers orchestration platform (e.g. Kubernetes, Nvidia GPU operator, Nvidia Network operator, CNI, CSI) and job scheduling systems (e.g. Slurm).
- Familiarity with Linux system administration and scripting (Python, Bash).
- Familiarity with logging and monitoring tools such as Prometheus, Grafana, Loki.
- Good knowledge of GPU architecture, Nvidia CUDA, NCCL, or related AI/ML frameworks - added advantage.
- Strong troubleshooting skills and ability to analyze system logs and performance metrics.
- Excellent communication and teamwork abilities.
Meeting every qualification is not required—if you’re excited about this role, we’d love to hear from you. We believe diverse perspectives and experiences strengthen our team.
$80k - $140k
Job DescriptionRBC Wealth Management Technology is seeking a Senior Site Reliability Engineer to join its Wealth Management SRE Team. This team is responsible for ensuring the performance, availability, resilience, and operational excellence of critical applications and...SuggestedFull timeFlexible hoursShift work$55 - $60 per hour
...experiences. The candidate will be responsible for ensuring the reliability, availability, performance, and scalability of cloud-based... ...Canada. We deliver agile, scalable talent solutions across IT, engineering, life sciences, clinical, and professional staffing, powered...SuggestedTemporary workLocal area- ...Site Reliability Engineer We are seeking a highly skilled Site Reliability Engineer (SRE) to support and enhance the reliability, scalability, and performance of enterprise applications and infrastructure. The ideal candidate will have strong experience in cloud environments...Suggested
$112.3k - $160.6k
...and organization with whom we partner and serve. Does this opportunity interest you? Western National is seeking a Site Reliability Engineer III to join our team! The individual in this role will have the opportunity to design, implement, manage, and monitor on...SuggestedFull timeWork at officeLocal areaRemote workWork visaFlexible hours$114.4k - $124.8k
...workforce solutions to Fortune 500 companies, government agencies, and enterprises across the United States and Canada. We serve IT, engineering, life sciences, clinical, and professional staffing needs across North America and Asia. As a certified Minority Business...SuggestedHourly payFull time- ...delivering speed, resilience, and choice to meet evolving marketplace needs. We are seeking an experienced Lead Site Reliability Engineer to join our engineering team and drive the reliability, scalability, and performance of our critical systems. This role combines...Full timePart time
$70.8k - $131.4k
Job DescriptionThomson Reuters is strengthening its Site Reliability Engineering capability to help engineering and operations teams build, operate, and improve reliable production services.The Site Reliability Engineer will support the tools, processes, and operational...Full timeWork at officeLocal areaFlexible hours$141.8k - $195k
...their best work, grow fast, and bring their full selves to the herd.Why You'll Love This RoleCribl Inc is seeking a Senior Site Reliability Engineer to join our mission where you will unlock the value of all observability data, as we expand our team in the U.S. Cribl...Remote work$95k - $171k
.... Opportunities exist to focus on GPU infrastructure, Kubernetes, and ensuring reliability for AI workloads within Akamai's serverless inference platform. As an Site Reliability Engineer II, you will be responsible for: Building and maintaining dashboards, alerts...Permanent employmentWork experience placementWork at officeRemote workWork from homeWorldwideFlexible hours$168k - $200k
...is passionate about creating transformative change in healthcare. What We're Looking For We're looking for a Senior Site Reliability Engineer to join our Data & ML Platform team. You'll be at the forefront of building and operating a resilient, observable, and...$81.1k - $187k
...Job Description We are looking for a Site Reliability Engineer 3 to support mission-critical cloud services and production operations. The role focuses on improving service reliability, reducing operational risk, automating repetitive tasks, and driving faster detection...Temporary workImmediate startFlexible hoursShift work$75.7k - $136.3k
...solve complex challenges? Do you have a passion for automation and building systems that scale? Join our highly skilled Site Reliability Engineering team! Our team designs, develops, and manages applications and infrastructure that support Akamai Cloud's products and...Work experience placementWork at office$121.4k - $218.6k
...solve complex challenges? Do you have a passion for automation and building systems that scale? Join our highly skilled Site Reliability Engineering team! Our team designs, develops, and manages applications and infrastructure that support Akamai Cloud's products and...Work experience placementWork at office$158.9k - $295.1k
Are you passionate about leading global engineering teams that keep mission-critical enterprise products reliable, scalable, and resilient? Join Thomson Reuters as Director, Site Reliability Engineering (SRE) within Service Management in our CIO organization, supporting...Full timeWork at officeLocal areaFlexible hours$169.3k - $304.7k
...in building and maintaining fast, efficient, scalable, and reliable routing software and infrastructure that is responsible... ...growth and stability of our global platform. As a Principal Site Reliability Engineer - Network, you will be responsible for: Architecting...Work experience placementWork at office$140k - $210k
...Our Mission As the world’s number 1 job site*, our mission is to help people get jobs. We strive to cultivate an... ...Comscore, Total Visits, March 2026) Day to Day As an Engineering Manager in Site Reliability Engineering at Indeed, you will manage and grow a team that...Work experience placementLocal area- ...Release Train Engineer (RTE) At Xebia we are always looking for talented people to deliver value to our clients and their customers. We help our customers to become data-driven by building their self-service data platform in the cloud and therefore deliver direct value...
- We are seeking experienced Reliability Engineers to support manufacturing and quality initiatives for CRM products, including pacemakers and leads... ...role, you will provide technical expertise across multiple sites, focusing on manufacturing support, project and program...
$114.7k - $195k
...applications, data, cloud environments, and AI agents. For engineers joining Veza today, this means the scale and resources of an... ...industry.Job DescriptionWe are seeking an exceptional Staff Site Reliability Engineer to lead critical infrastructure initiatives and drive...Work at officeRemote workFlexible hours$102.4k - $179k
...the hands-on design, development, and execution of AI Quality Engineering initiatives supporting Vizient's enterprise AI transformation... ...validation, AI-assisted testing, runtime observability, reliability engineering, and modern AI Quality Engineering practices.ResponsibilitiesDesign...Full time$88.9k - $155.5k
...will lead the day-to-day execution and delivery of AI Quality Engineering initiatives supporting Vizient's enterprise AI transformation... ..., and cross-functional delivery coordination to ensure the reliability, performance, and governance of AI-enabled systems.ResponsibilitiesLead...Full timeFor contractors$84.8k - $127.2k
...in purpose, growth, and impact. A Day in the Life The Reliability Engineer II will support key technical activities within the Electrophysiology... ...hybrid work model requiring a minimum of 4 days per week on-site for in-person collaboration and project activities. In...Full timeWork experience placementH1bWork at officeLocal areaImmediate startFlexible hours- Mission of the RoleThis position may be based in either our Minneapolis, MN or Canton, MA office.As a Systems Application Engineer II, you will play a pivotal role in shaping intelligent automation solutions across North America. You’ll partner with sales, industry teams...Full timeWork at officeMonday to FridayFlexible hoursWeekend work
$135.2k - $236.6k
...and in the future.Summary:In this role, you will lead Quality Engineering and product validation for an AI-enabled healthcare... ...will serve as a hands-on technical leader, ensuring accuracy, reliability, traceability, and usability across the product lifecycle, from...Full timeShift work$118.65k - $160k
...about the work and each other.AT A GLANCEThe Senior Software Engineer is a crucial role within our organization, requiring work in various... ...defects and performance bottlenecks to ensure optimal system reliability and user experience. Conduct comprehensive testing and...Full timeRemote workMonday to Friday- ...product development. We are seeking a motivated, skilled Software Engineer for our High-Performance Networking Platform team. The ideal... ...engineers and internal and outsourced development partners to develop reliable, cost effective and high quality solutions for assigned systems...Full timeWork experience placementWork at officeLocal areaImmediate start2 days per week
$102.4k - $179k
...organization. You will collaborate with Enterprise Architecture, Data Engineering, Product Management, Cloud Engineering, and Data Science... ...delivery.Optimize platform scalability, performance, reliability, and cost efficiency across cloud environments.Collaborate with...Full time- ...and collaborative environment, where your ideas and contributions will be valued and respected.Job SummaryAs a Principal Systems Engineer on the WaferSense team, you will provide technical leadership for the development and lifecycle support of a product line used by...Hourly pay16 hoursFull timeLocal areaWorldwide
$98.59k - $173.58k
OverviewGIS Solution Engineers on our commercial team are highly technical, trusted advisors to customers across many different commercial markets working with some of the largest and most complex companies around the world. They inspire customers supporting mission-critical...$79.25k - $130.73k
OverviewAs a GIS subject matter expert, you’re a natural at identifying the right analysis tools for the problem at hand. Not only do you create innovative solutions, you talk about solutions in ways that get others excited about the power of GIS technology. Join an account...Local area
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
- official site Minneapolis, MN
- remote website tester Minneapolis, MN
- site services specialist Minneapolis, MN
- construction site safety Minneapolis, MN
- IT site lead Minneapolis, MN
- site leader Minneapolis, MN
- site safety Minneapolis, MN
- historic site Minneapolis, MN
- junior website developer Minneapolis, MN
- website coordinator Minneapolis, MN



