Site Reliability Engineer
GMI Cloud
About GMI
GMI Cloud is a fast-growing, AI-native infrastructure company delivering high-performance GPU compute, inference services, and infrastructure for AI agents.
Following 8x ARR growth, GMI Cloud continues to scale rapidly across the U.S. and APAC. As a Reference Platform NVIDIA Cloud Partner (NCP) and a validated leading NCP across both markets, we power production AI for leading AI-native companies including Fireworks AI, Cartesia, Reflection, and OpenRouter.
From large-scale compute to optimized inference and agentic workloads, GMI Cloud gives AI teams the infrastructure they need to build, deploy, and scale on one unified cloud.
One cloud for compute, inference, and agents.
Role Overview
We are seeking a skilled Site Reliability Engineer to join the GMI Global Infrastructure team. This role is hands-on and critical to ensuring the stability, efficiency, and reliability of the large-scale high performance AI/ML clusters in our data center. The ideal candidate will bring expertise in system-level troubleshooting, AI cluster maintenance, and operational excellence to ensure maximum performance for our infrastructure. Experience with large-scale infrastructure automation is considered a strong plus.
Responsibilities
- Design, implement and maintain scalable AI/ML infrastructure solutions.
- Proactively monitor GPU cluster health, performance and troubleshoot issues across compute, accelerator, and storage systems.
- Automate deployment, configuration and management of infrastructure resources.
- Manage GPU node lifecycle workflows, including provisioning, scaling, maintenance, decommissioning and upgrades of GPU nodes.
- Implement CI/CD pipelines for infrastructure deployment and orchestration.
- Ensure security, compliance and best practices across infrastructure.
- Manage incident response related to Infrastructure resources (GPU, CPU, Storage, Network).
- Handle customer provisioning requests for GPU resources, including onboarding, configuration and troubleshooting; resolve customer service requests related to GPU infrastructure, ensuring high customer satisfaction.
- Stay current with emerging GPU hardware and software technologies, integrating improvements as appropriate.
- Regional/international travel to GMI data center locations.
Qualifications
- Bachelor’s degree in Computer Science or related field.
- Over 3+ years of experience in data center operations, infrastructure, or systems engineering.
- Proven experience in site reliability engineering and infrastructure automation (e.g. Ansible, Terraform)
- Familiarity with containers orchestration platform (e.g. Kubernetes, Nvidia GPU operator, Nvidia Network operator, CNI, CSI) and job scheduling systems (e.g. Slurm).
- Familiarity with Linux system administration and scripting (Python, Bash).
- Familiarity with logging and monitoring tools such as Prometheus, Grafana, Loki.
- Good knowledge of GPU architecture, Nvidia CUDA, NCCL, or related AI/ML frameworks - added advantage.
- Strong troubleshooting skills and ability to analyze system logs and performance metrics.
- Excellent communication and teamwork abilities.
Meeting every qualification is not required—if you’re excited about this role, we’d love to hear from you. We believe diverse perspectives and experiences strengthen our team.
- ...candidate will think beyond support operations and approach reliability as an engineering discipline, using automation, resilience, observability,... ...Required Qualifications ~4-8 years of experience in Site Reliability Engineering, Production Engineering, Platform...SuggestedFull time
- ...This Site Reliability Engineer job is a 9+ Months Contract with a client located in Charlotte, NC (Hybrid 3 Days onsite every week). Pay Range: $65/hr - $68/hr on W2. The rate may be negotiable based on experience, education, geographic location, and other factors...SuggestedContract workTemporary workLocal area3 days per week
$80k - $90k
...Role: Site Reliability Engineer Location: Charlotte, NC We are At Synechron, we believe in the power of digital to transform businesses for the better. Our global consulting firm combines creativity and innovative technology to deliver industry-leading digital...SuggestedFull timeTemporary workFlexible hours- ...Position Summary We are seeking a Lead Site Reliability & Environment Monitoring Engineer to establish and evolve our enterprise observability and monitoring strategy across cloud and application platforms. This is a full-time leadership role responsible for owning...SuggestedFull timeShift work
- ...home!Where you’ll be:This position will be based at our Corporate Headquarters located in Charlotte, NC.About the Role:The Site Reliability Engineer plays a critical role in designing, building, and maintaining scalable, secure, and highly available cloud infrastructure...SuggestedFull timeFlexible hours
$91.2k - $136.8k
Reliability Engineer - IE08GEWe’re determined to make a difference and are proud to be an insurance company that goes well beyond coverages and... ...field.3+ years of experience in Infrastructure Engineering, Site Reliability Engineering (SRE), or DevOps.Hands-on experience...Full timeTemporary workWork at office3 days per week- ...Fluency: English (Required)Work Shift:1st shift (United States of America)Please review the following job description:The Site Reliability Engineer role focuses on enhancing the reliability and operational excellence of enterprise platforms across hybrid cloud and on-premises...Permanent employmentFull timePart timeH1bWork at officeLocal areaImmediate startWork visaMonday to FridayShift workDay shift
- ...Evaluate applications, platforms, and vendors to assess resiliency, reliability, and operational risk. Design and implement processes that... ...reliability tooling. Actively participate in reliability engineering and resilience communities of practice, contributing to...
- ...Site Reliability Engineer (SRE, Terraform, AWS, Dynatrace) Optomi, in partnership with a Fortune 500 digital platform leader, is seeking a Site Reliability Engineer to join their Digital SRE team! In this role, the Site Reliability Engineer will focus on improving the...
- ...Position: Observability Engineer Duration: 6 + Moths Location: Charlotte NC– Hybrid – Candidate must be local Pay Rate: 55/hr W2 Key Requirements Must Have AppDynamics Application Performance Monitoring Splunk Plus Automation PowerBI and...Work at officeLocal area
$57.51 per hour
...Job Title: Site Reliability Engineer II Location: Charlotte, NC Duration: Contract - 12 months Pay Range: $57.51/hr (W2) Job ID: 410454 About BCforward BCforward is a leading global IT consulting and workforce solutions firm providing services...Contract work- ...based on experience Introduction We are seeking a highly skilled and experienced professional to join our team as a Site Reliability Engineer III. This role involves collaborating with cross-functional teams to ensure the reliability and performance of critical...Work experience placementImmediate startRemote workFlexible hours
- ...SR for Technical Consultant (SRE, MS Dynamics) Job Summary The Senior Support Lead in Site Reliability engineering (SRE) will be responsible for overseeing the support and reliability operations within the organization. This role will focus on ensuring the stability...
- ...Job Title: Senior Site Reliability Engineer Duration: 18 months (possibility to extend or convert to FTE) Location: Charlotte, NC - Hybrid Role (3 days onsite in a week) Interview process: 2 rounds #1-hour virtual panel #1 hour on site technical panel...Shift work3 days per week
- ...track their careers in technology or operations within prestigious global organizations. Responsibilities: Platform & Reliability Engineering Embed SRE and production engineering principles into Payments Modernization from design through early life support...Full timeWorldwideVisa sponsorshipWork visa
- ...Site Reliability Engineer Job Location: San Francisco, CA or Charlotte, NC. Job Type: Contract Work with local API development squads, platform teams, product owners, scrum masters, and architects. The SRE ensures that both our internally critical and our externally...Contract workLocal area
$53 - $57 per hour
...A client of Innova Solutions is immediately hiring for the Site Reliability Engineer II Position Type: Full time, Hybrid Onsite Contract --- (***Hybrid Position - 3 days onsite and 2 days remote work in a week!!!) Duration: 12-18 Months contract Location:...Hourly payFull timeContract workTemporary workWork experience placementLocal areaImmediate startRemote workWorldwideFlexible hours- ...Senior Site Reliability Engineer Jersey City, New Jersey;Charlotte, North Carolina; Plano, Texas To proceed with your application, you must be at least 18 years of age. Acknowledge ( Bank of America employees are required to meet all posting eligibility requirements...Work at officeShift workDay shift
$60 - $65 per hour
...retail industries. Rate Range: $60-$65/Hr Job Description: The Client Document Generation team is seeking a Senior Software Engineer ( IT Onshore Band 4) to participate in the full system development lifecycle (SDLC) of enterprise applications that support high-...Immediate start- ...Senior Site Reliability Engineer (SRE) – PagerDuty / Moogsoft Migration Position Overview Lead a critical observability and incident-management transformation by migrating from the legacy Moogsoft AIOps platform to PagerDuty. This role focuses on redesigning incident management...
$57 - $62 per hour
...Site Reliability Engineer - Contract - $57-62 per HR This role involves leading reliability engineering efforts in a hybrid setting, focusing on automation and resilience to enhance system reliability. The ideal candidate will drive an engineering-led approach to...Contract work- ...PFB the JD must have skills Hands on SRE Engineer with good analytical skills Good Exposure to both incident and Problem Management Must have worked in AWS Must have good knowledge on Middleware components/services (Servers,Load Balancer/Trace Logs) and have...Permanent employmentFull time
$119.62k
...~ Proven leadership in SRE strategy, reliability by design, and observability. ~ Demonstrated... ...response, and capacity and reliability engineering. ~ Expertise in resilient engineering,... ..., and integrity. We are seeking a Site Reliability Engineer II to join our team...Hourly payFull timeContract work- ...delivering speed, resilience, and choice to meet evolving marketplace needs. We are seeking an experienced Lead Site Reliability Engineer to join our engineering team and drive the reliability, scalability, and performance of our critical systems. This role combines...Full timePart time
$127.6k - $191.4k
Staff Reliability Engineer - IE07KEWe’re determined to make a difference and are proud to be an insurance company that goes well beyond coverages... ...the future. The Hartford is seeking a dedicated Staff Data Site Reliability Engineer (DRE) to focus specifically on the integrity...Full timeTemporary workWork at office3 days per week$73.67 per hour
...Job Description Job Description Job Title: Site Reliability Engineer with Python Location: Pennington, NJ Duration: Contract - 9 months Pay Range: $73.67/hr (W2) Job ID: 410592 About BCforward BCforward is a leading global IT consulting and workforce...Contract work$153.23k
...administration and troubleshooting Infrastructure engineering background with proven automation... ...strongly preferred Knowledge of SRE, reliability engineering, and DevOps practices... ...collaboration, and integrity. We are seeking a Site Reliability Engineer with Python for our...Hourly payFull timeContract work- ...build a successful career with opportunities to learn, grow, and make an impact. Join us! Position Summary: The IKCP Site Reliability Engineer Lead is responsible for ensuring the reliability, scalability, performance, security, and operational excellence of the enterprise...Work at officeFlexible hoursShift workDay shift
- ...BC forward is seeking a highly motivated SRE Transformation Engineer for an opportunity in Charlotte, NC Hybrid/ Onsite! Position... ...experience in leading SRE strategy, automation, observability, and reliability by design across banking and payments and a proven ability to...Daily paidContract workTemporary workImmediate start
- ...Job Description Mainframe Systems Programmer / Infrastructure Engineer - 2 open - Contract \n \n Looking for someone that has held the title of Mainframe Systems Programmer for 5-10 years \n \n 6 month contract - with likely extension \n \n Qualifications...Contract work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
- site reliability engineer sre Charlotte, NC
- site reliability engineer Charlotte, NC
- official site Charlotte, NC
- remote website tester Charlotte, NC
- site services specialist Charlotte, NC
- construction site safety Charlotte, NC
- IT site lead Charlotte, NC
- site recruiter Charlotte, NC
- site leader Charlotte, NC
- site safety Charlotte, NC





