Site Reliability Engineer
GMI Cloud
About GMI
GMI Cloud is a fast-growing, AI-native infrastructure company delivering high-performance GPU compute, inference services, and infrastructure for AI agents.
Following 8x ARR growth, GMI Cloud continues to scale rapidly across the U.S. and APAC. As a Reference Platform NVIDIA Cloud Partner (NCP) and a validated leading NCP across both markets, we power production AI for leading AI-native companies including Fireworks AI, Cartesia, Reflection, and OpenRouter.
From large-scale compute to optimized inference and agentic workloads, GMI Cloud gives AI teams the infrastructure they need to build, deploy, and scale on one unified cloud.
One cloud for compute, inference, and agents.
Role Overview
We are seeking a skilled Site Reliability Engineer to join the GMI Global Infrastructure team. This role is hands-on and critical to ensuring the stability, efficiency, and reliability of the large-scale high performance AI/ML clusters in our data center. The ideal candidate will bring expertise in system-level troubleshooting, AI cluster maintenance, and operational excellence to ensure maximum performance for our infrastructure. Experience with large-scale infrastructure automation is considered a strong plus.
Responsibilities
- Design, implement and maintain scalable AI/ML infrastructure solutions.
- Proactively monitor GPU cluster health, performance and troubleshoot issues across compute, accelerator, and storage systems.
- Automate deployment, configuration and management of infrastructure resources.
- Manage GPU node lifecycle workflows, including provisioning, scaling, maintenance, decommissioning and upgrades of GPU nodes.
- Implement CI/CD pipelines for infrastructure deployment and orchestration.
- Ensure security, compliance and best practices across infrastructure.
- Manage incident response related to Infrastructure resources (GPU, CPU, Storage, Network).
- Handle customer provisioning requests for GPU resources, including onboarding, configuration and troubleshooting; resolve customer service requests related to GPU infrastructure, ensuring high customer satisfaction.
- Stay current with emerging GPU hardware and software technologies, integrating improvements as appropriate.
- Regional/international travel to GMI data center locations.
Qualifications
- Bachelor’s degree in Computer Science or related field.
- Over 3+ years of experience in data center operations, infrastructure, or systems engineering.
- Proven experience in site reliability engineering and infrastructure automation (e.g. Ansible, Terraform)
- Familiarity with containers orchestration platform (e.g. Kubernetes, Nvidia GPU operator, Nvidia Network operator, CNI, CSI) and job scheduling systems (e.g. Slurm).
- Familiarity with Linux system administration and scripting (Python, Bash).
- Familiarity with logging and monitoring tools such as Prometheus, Grafana, Loki.
- Good knowledge of GPU architecture, Nvidia CUDA, NCCL, or related AI/ML frameworks - added advantage.
- Strong troubleshooting skills and ability to analyze system logs and performance metrics.
- Excellent communication and teamwork abilities.
Meeting every qualification is not required—if you’re excited about this role, we’d love to hear from you. We believe diverse perspectives and experiences strengthen our team.
$104.9k - $174.7k
...automation, AI, and platform transformation to improve efficiency, accuracy, and cost effectiveness.About the RoleSenior Site Reliability Engineer IIWe're looking for a Senior Site Reliability Engineer II to serve as a go-to technical resource for our Life Sciences SRE...SuggestedFull timeFor contractorsWork at officeLocal areaImmediate startFlexible hours- IXL Learning, developer of personalized learning products used by millions of people globally, is seeking a Senior Site Reliability Engineer to join our team, and help maintain the reliability and optimal performance of our products. We are seeking engineers with a passion...SuggestedWork at officeImmediate start
- ...Fluency: English (Required)Work Shift:1st shift (United States of America)Please review the following job description:The Site Reliability Engineer role focuses on enhancing the reliability and operational excellence of enterprise platforms across hybrid cloud and on-premises...SuggestedPermanent employmentFull timePart timeH1bWork at officeLocal areaImmediate startWork visaMonday to FridayShift workDay shift
$139.3k - $203.6k
...is a hybrid role***Meet the TeamThe Customer Experience (CX) engineering team is a group of extraordinary technical guides whose main... ...and employee satisfaction scores at Cisco.Your ImpactAs a Site Reliability Engineer (SRE) on this CX engineering team, you will play a...SuggestedFull timeTemporary workLocal areaFlexible hours$95k - $171k
.... Opportunities exist to focus on GPU infrastructure, Kubernetes, and ensuring reliability for AI workloads within Akamai's serverless inference platform. As an Site Reliability Engineer II, you will be responsible for: Building and maintaining dashboards, alerts...SuggestedPermanent employmentWork experience placementWork at officeRemote workWork from homeWorldwideFlexible hours- ...Site Reliability Engineer Job Overview: Site Reliability Engineer role focuses on reliability, scalability, and performance of enterprise platforms across cloud and on-prem environments. Position requires hands-on engineering with automation, observability, and cross...
- Role Profile:We are evolving our Site Reliability Engineering capabilities to strengthen reliability, observability, security, and operational excellence across our Markets and Risk Intelligence division.As a Senior SRE, you will be a senior hands‑on technical person help...Full timeShift work
$55k - $151.47k
...ApplicableSpecialismIFS - Internal Firm Services - OtherManagement LevelSenior AssociateJob Description & SummaryThe OpportunityAs a Site Reliability Engineer - Senior Associate, you will play a pivotal role in enhancing the reliability, scalability, and performance of our...Full timeH1b$84.9k - $209.5k
...Description Designs and architects infrastructure and service to ensure reliability and functionality. Forecasts demands and responds to capacity... ...new tools and develops and maintains advanced knowledge of site reliability trends. Responsibilities Key Responsibilities...Temporary workImmediate startFlexible hoursShift work$121.4k - $218.6k
...thrives in a dynamic environment? Join our highly skilled Site Reliability team! Our team designs, develops, and manages applications... ...and performance tuning. As a Senior Lead Site Reliability Engineer, you will be responsible for: Defining requirements as part...Work experience placementWork at office- ...Site Reliability Engineer (SRE) - Security Infrastructure Position Summary We are seeking an SRE to support reliability, scalability, and operational excellence for a large-scale network security transformation initiative. This role will focus on monitoring...
- ...Site Reliability Engineer Number of Position: 2 Only Fulltime I, Abhishek, would like to share a job opportunity as Site Reliability Engineer in Jacksonville, FL, Cary, NC or New York, NY (Onsite) location for a Fulltime position. *** In case, if you are not...Full timeWork visa
$125k - $185k
...Job Title: CaaS Private Site Reliability Engineer Corporate Title: Vice President Location: Cary, NC Who we are: In short – an essential part of Deutsche Bank’s technology solution, developing applications for key business areas. Our Technologists drive...Full timeWork at officeWork from homeShift work$65 - $70 per hour
...MatchPoint Solutions is a fast-growing global IT and Engineering services firm delivering innovative technology solutions to leading... ...solutions in a collaborative, high-growth environment. Site Reliability Engineer (SRE) – Security Infrastructure Location: Onsite...Local area3 days per week- ...technologies to enable scalable, secure, and reliable business operations. Applies strong... ...infrastructure.3. Manages infrastructure engineering projects and processes aligned with... ...benefit plans, please visit our Benefits site. Depending on the position and division,...Permanent employmentFull timePart timeWork experience placementH1bRemote workWork visaShift workWeekend workDay shift
$250k - $300k
We're looking for a hands-on engineering leader to build and own the Release Engineering, SecDevOps, and Site Reliability Engineering (SRE) functions for Infinia. This is a foundational role: you'll define how our software is built, secured, released, and kept running...Remote work- ...Hogan Software Engineer | 100% Remote (EST Hours) | W2 Contract (12 Months) About the Opportunity Optomi, in partnership with a... ...troubleshooting production issues, and ensuring application performance and reliability. The ideal candidate will have a strong background in...Contract workRemote work
$144k - $198k
...BaxterJoin our dynamic team as a Senior Principal Software Systems Engineer in the R&D/Software organization, where you will play a pivotal... ..., please speak with your recruiter or visit our Benefits site: Benefits | BaxterEqual Employment OpportunityBaxter is an equal...Full timeTemporary workLocal areaWork visaRelocation packageFlexible hours$130k - $180k
Position OverviewPower your future with Qualus as a Lead Relay Settings Engineer. In this role you will perform Protective Relay Design & Coordination: Design, specify, calculate settings, and coordinate protective relays and relay control schemes. Do you have 7+ years...Temporary workFlexible hours$136.09k - $168.11k
...looking to move fast and make a significant impact in an exciting space, you're in the right place!We are seeking a Lead Solution Engineer to join our North America GTM team, specifically supporting our Industry Verticals organization across SLED (State, Local & Education...Local areaFlexible hours$184k - $287.5k
NVIDIA is growing a senior engineering team focused on making our compute software stack first-class on NVIDIA CPU platforms. The team turns modern toolchains, build and code-health practices, performance-analysis workflows, and optimization techniques into repeatable improvements...Full timeRemote work$184k - $287.5k
...inspired to do their best work. Come join the team and see how you can make a lasting impact on the world.We are looking for a dedicated engineer for the Senior Systems Software Engineer role, focusing on GPU Performance at Scale. At NVIDIA, this role is uniquely positioned...Full timeRemote work- ...Reliability Engineer - Hydraulics & Pneumatics The Enviva team is driven by our shared vision for a renewable energy future. We are a fast... ...capital projects to include large greenfield and brownfield sites and small capital projects. Safeguard all work, including...For contractorsLocal area
- ...mobile, RF, networking, industrial, business equipment, and automotive. ~ Signal degradation over life Position Reliability Engineer Location Raleigh, NC Responsibilities Test, Validation & Qualification Develop and execute reliability and...Flexible hours
$100k - $153k
...Position Overview J ob Title CaaS Private Site Reliability Engineer Corporate Title Assistant Vice President Location Cary, NC Who we are: In short – an essential part of Deutsche Bank’s technology solution, developing applications for key business...Full timeWork at officeWork from homeShift work- ...AI infrastructure, working with server, cloud, and platform engineering teams.Operationalize machine learning workflows and support AI... ...implement system enhancements to improve performance, scalability, reliability, and cost efficiency.Collaborate across divisions to support...Full timeWork at officeRemote work
$175k - $200k
...desire to positively impact the environment and lives of others in a refreshing, vibrant, and inclusive culture.As a Senior AI Systems Engineer on the Advanced Systems team, you will architect and guide AI-enabled engineering systems that accelerate advanced aircraft...Full timeTemporary workWork at officeLocal areaRemote work$272k - $431.25k
...world.At NVIDIA, as a Principal Rack Scale Systems Infrastructure Engineer, you will build and guide the development of software systems.... ...with real-world deployment and integration needs. Establish reliability, security, validation, and left-shift strategies that reduce...Full timeRemote workShift work$120k - $160k
...PLATFORM ENGINEER LOCATION: Raleigh, NC or Beaverton, OR SALARY: $120K- $160K Full - Time Employee Hands-on role focuses on designing, building, and implementing scalable, secure infrastructure across commercial and CMMC Level 2 environments. The ideal...Full time$130k - $170k
About Us:BW Design Group is a fully integrated architecture, engineering, construction, system integration, and consulting firm committed to helping our clients realize their most critical goals from Strategy to Commercialization. As the only firm born from a manufacturing...Full timeFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!




