Site Reliability Engineer
GMI Cloud
About GMI
GMI Cloud is a fast-growing, AI-native infrastructure company delivering high-performance GPU compute, inference services, and infrastructure for AI agents.
Following 8x ARR growth, GMI Cloud continues to scale rapidly across the U.S. and APAC. As a Reference Platform NVIDIA Cloud Partner (NCP) and a validated leading NCP across both markets, we power production AI for leading AI-native companies including Fireworks AI, Cartesia, Reflection, and OpenRouter.
From large-scale compute to optimized inference and agentic workloads, GMI Cloud gives AI teams the infrastructure they need to build, deploy, and scale on one unified cloud.
One cloud for compute, inference, and agents.
Role Overview
We are seeking a skilled Site Reliability Engineer to join the GMI Global Infrastructure team. This role is hands-on and critical to ensuring the stability, efficiency, and reliability of the large-scale high performance AI/ML clusters in our data center. The ideal candidate will bring expertise in system-level troubleshooting, AI cluster maintenance, and operational excellence to ensure maximum performance for our infrastructure. Experience with large-scale infrastructure automation is considered a strong plus.
Responsibilities
- Design, implement and maintain scalable AI/ML infrastructure solutions.
- Proactively monitor GPU cluster health, performance and troubleshoot issues across compute, accelerator, and storage systems.
- Automate deployment, configuration and management of infrastructure resources.
- Manage GPU node lifecycle workflows, including provisioning, scaling, maintenance, decommissioning and upgrades of GPU nodes.
- Implement CI/CD pipelines for infrastructure deployment and orchestration.
- Ensure security, compliance and best practices across infrastructure.
- Manage incident response related to Infrastructure resources (GPU, CPU, Storage, Network).
- Handle customer provisioning requests for GPU resources, including onboarding, configuration and troubleshooting; resolve customer service requests related to GPU infrastructure, ensuring high customer satisfaction.
- Stay current with emerging GPU hardware and software technologies, integrating improvements as appropriate.
- Regional/international travel to GMI data center locations.
Qualifications
- Bachelor’s degree in Computer Science or related field.
- Over 3+ years of experience in data center operations, infrastructure, or systems engineering.
- Proven experience in site reliability engineering and infrastructure automation (e.g. Ansible, Terraform)
- Familiarity with containers orchestration platform (e.g. Kubernetes, Nvidia GPU operator, Nvidia Network operator, CNI, CSI) and job scheduling systems (e.g. Slurm).
- Familiarity with Linux system administration and scripting (Python, Bash).
- Familiarity with logging and monitoring tools such as Prometheus, Grafana, Loki.
- Good knowledge of GPU architecture, Nvidia CUDA, NCCL, or related AI/ML frameworks - added advantage.
- Strong troubleshooting skills and ability to analyze system logs and performance metrics.
- Excellent communication and teamwork abilities.
Meeting every qualification is not required—if you’re excited about this role, we’d love to hear from you. We believe diverse perspectives and experiences strengthen our team.
- Edmond, OKYouVersion - YouVersion Engineering /Full-Time/ Salary /On-siteThe YouVersion Senior Site Reliability Engineer is responsible for ensuring the integrity, performance, reliability, and cost-effectiveness of the cloud-based infrastructure and related systems supporting...SuggestedFull timeContract workTemporary workWork experience placementCasual workInternshipLocal areaWorldwide
$141.8k - $195k
...their best work, grow fast, and bring their full selves to the herd.Why You'll Love This RoleCribl Inc is seeking a Senior Site Reliability Engineer to join our mission where you will unlock the value of all observability data, as we expand our team in the U.S. Cribl...SuggestedRemote work$121.4k - $218.6k
...solve complex challenges? Do you have a passion for automation and building systems that scale? Join our highly skilled Site Reliability Engineering team! Our team designs, develops, and manages applications and infrastructure that support Akamai Cloud's products and...SuggestedWork experience placementWork at office$168k - $200k
...is passionate about creating transformative change in healthcare. What We're Looking For We're looking for a Senior Site Reliability Engineer to join our Data & ML Platform team. You'll be at the forefront of building and operating a resilient, observable, and...Suggested$95k - $171k
.... Opportunities exist to focus on GPU infrastructure, Kubernetes, and ensuring reliability for AI workloads within Akamai's serverless inference platform. As an Site Reliability Engineer II, you will be responsible for: Building and maintaining dashboards, alerts...SuggestedPermanent employmentWork experience placementWork at officeRemote workWork from homeWorldwideFlexible hours$140k - $210k
...Our Mission As the world’s number 1 job site*, our mission is to help people get jobs. We strive to cultivate an... ...Comscore, Total Visits, March 2026) Day to Day As an Engineering Manager in Site Reliability Engineering at Indeed, you will manage and grow a team that...Work experience placementLocal area$169.3k - $304.7k
...in building and maintaining fast, efficient, scalable, and reliable routing software and infrastructure that is responsible... ...growth and stability of our global platform. As a Principal Site Reliability Engineer - Network, you will be responsible for: Architecting,...Work experience placementWork at office- ...Job Title Senior Design Release Engineer About Canoo Canoo’s mission is to bring EVs to Everyone and build a world-class team to deploy this sustainable mobility revolution. We have developed breakthrough electric vehicles that are reinventing...Casual workLocal areaFlexible hours
$109.65k - $148.35k
Software Agile Release Train Engineer (RTE) (Lead or Experienced)Company:The Boeing CompanyBoeing Defense, Space & Security (BDS) has an exciting opportunity for a Software Agile Release Train Engineer (RTE) to support the B-1B Software team in Tinker AFB, OK. Our teams...Permanent employmentFull timeWork experience placementInterim roleCurrently hiringImmediate startRelocationVisa sponsorshipWork visaFlexible hoursShift work$111.35k - $150.65k
...Join to apply for the Lead Systems Engineer role at Boeing 3 days ago Be among the first 25 applicants Join to apply for the... ...architecture Perform analyses for affordability, safety, reliability, maintainability, testability, human systems integration, susceptibility...Permanent employmentFull timeInterim roleRelocationVisa sponsorshipWork visaFlexible hoursShift workDay shift$63k - $140k
...ApplicableSpecialismData, Analytics & AIManagement LevelAssociateJob Description & SummaryThe OpportunityAs a GenAI Python Systems Engineer - Experienced Associate, you will leverage advanced technologies and techniques to design and develop robust data solutions for...Full timeH1b- ...Experience: \tEducation/experience typically acquired through advanced technical education from an accredited course of study in engineering, computer science, mathematics, physics or chemistry (e.g. Bachelor) and typically 9 or more years' related work experience or an...Work experience placementWork at officeImmediate startRemote work
$206.4k - $283.8k
...teammates to share knowledge, improve demos, and strengthen the broader SE organizationCuriosity and Hands-On MindsetWe are looking for engineers who enjoy learning by building, experimenting, and digging into systems.You enjoy exploring how systems work by testing...Full timeRemote workVisa sponsorshipWork visa- DescriptionSAIC is currently seeking a Software Safety Engineer in Huntsville, Alabama to support to U.S. Army Integrated Fires Mission Command (IFMC) System of Systems (SoS) program within the Program Executive Office (PEO) for Missiles and Space, including technical...Work at officeRemote workRelocation2 days per week1 day per week
$95.3k - $158.8k
...users and stakeholders to translate ambiguous, unstructured problems into measurable AI solutions. Requirements:5+ years of software engineering experience, including 2+ years building and shipping LLM-based or ML-based systems to production BS in Engineering/Computer...Full timeWork experience placementLocal area- A well-respected Oklahoma City organization seeking a Senior Salesforce Developer to lead the design, development, implementation, and ongoing support of Salesforce solutions that support critical business operations.This position offers the opportunity to work across a...Full time
- Titan Professional Resources is looking for a skilled Sr. Software Engineer to join an exciting company here in the OKC area! This is a fully remote position offering weekly pay and great benefits! If this is something that interests you, apply today!Responsibilities:Ongoing...Weekly payRemote work
- ...continuous evolving technical needs of business Responsible for the in-house design, development, implementation, and support of engineering applications used for competitive advantageOther tasks as assignedSkills and Experience:Must have at least 7 years’ experience working...Full time
- Edmond, OKYouVersion - YouVersion Engineering /Full-Time/ Salary /On-siteAt YouVersion, we build technology that helps people around the... ...ThriveStrong software engineering experience with the ability to deliver reliable, scalable solutions.Comfort owning technical decisions and...Full timeTemporary workWork experience placementCasual workInternshipLocal area
$124k - $280k
...ApplicableSpecialismData, Analytics & AIManagement LevelSenior ManagerJob Description & SummaryAt PwC, our people in data and analytics engineering focus on leveraging advanced technologies and techniques to design and develop robust data solutions for clients. They play a...Full timeH1b- ...Platform Engineer Canoo's mission is to bring EVs to Everyone and build a world-class team to deploy this sustainable mobility revolution. We have developed breakthrough electric vehicles that are reinventing the automotive landscape with pioneering technologies, award...Casual workLocal areaFlexible hours
$83k - $166.1k
...maintaining backend systems, and ensuring the long-term scalability, reliability, and sustainability of the platform. The ideal candidate... ...Bachelor's degree in Computer Science, Information Systems, Engineering, Healthcare Informatics, or a related field. ~ Master's...Temporary workWork experience placementImmediate startFlexible hours- ...Preferred*** ***US Citizenship Required*** ***NO VISA SUPPORT or C2C*** ***On-site in OKC Required*** Primary Function: Long Wave Inc. is seeking a Reliability and Maintainability (R&M) Engineer to support Product Support for Department of Defense (DoD) programs. This...Full timeWork at officeVisa sponsorshipMonday to Thursday
$155k - $247k
...Job Title Director, Released Product Engineering Job Description The Director of Released Product Engineering is responsible for managing... .... Lean Manufacturing/Six Sigma Certification or Design for Reliability and Manufacturability Certification also preferred. You must...Full timeWork at officeWork visaRelocation package3 days per week- ...are, join our team.KPMG is currently seeking a Manager, Network Engineer - F5 Technical SME to join our Digital Nexus organization.Responsibilities... ...benefits can be found towards the bottom of our KPMG US Careers site at Benefits & How We Work.Follow this link to obtain salary...H1bLocal area
- ...the defense industry today!ABOUT THE JOBAs a Principal Software Engineer V (UAV Embedded), you will architect and implement mission-... ...applications that ensure deterministic control and absolute system reliability. Working at the intersection of hardware and software, you...Temporary workFor contractorsWork at officeLocal area
$120k - $140k
...Forward Deployment Engineer Fractal Analytics is a strategic AI partner to Fortune 500 companies with a vision to power every human... ...Ensure solutions meet enterprise requirements for security, reliability, maintainability, quality, and production readiness. Troubleshoot...Hourly payFull timeLocal areaRemote work$118.15k - $159.85k
Lead Software Engineer (Test & Verification)Company:The Boeing CompanyBoeing Defense, Space & Security (BDS) has an exciting opportunity... ...test and verification of software products to ensure quality, reliability, and functionalityPartners with stakeholders and leads the...Permanent employmentFull timeWork experience placementInterim roleRelocationVisa sponsorshipWork visaFlexible hoursShift work- Edmond, OKDigital Product - Digital Product Engineering LC /Full-Time/ Salary /On-siteAt Life.Church, our Digital Product team builds technology... ...used across our campuses and ministries, helping deliver reliable, scalable tools that serve people every week.The Digital...Full timeTemporary workWork experience placementCasual workInternship
- Company DescriptionWe are Olsson. We engineer and design solutions that improve the world around us. As a company, we promise to always be responsive, transparent, and focused on results - for our people, our clients, and our company.We’re a people-centric firm, so it’...Full timeWork at officeFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
- official site Oklahoma City, OK
- site services specialist Oklahoma City, OK
- construction site safety Oklahoma City, OK
- IT site lead Oklahoma City, OK
- site leader Oklahoma City, OK
- site safety Oklahoma City, OK
- historic site Oklahoma City, OK
- junior website developer Oklahoma City, OK
- website coordinator Oklahoma City, OK
- on-site clinical research associate (traveling/remote) Oklahoma City, OK




