Site Reliability Engineer
GMI Cloud
About GMI
GMI Cloud is a fast-growing, AI-native infrastructure company delivering high-performance GPU compute, inference services, and infrastructure for AI agents.
Following 8x ARR growth, GMI Cloud continues to scale rapidly across the U.S. and APAC. As a Reference Platform NVIDIA Cloud Partner (NCP) and a validated leading NCP across both markets, we power production AI for leading AI-native companies including Fireworks AI, Cartesia, Reflection, and OpenRouter.
From large-scale compute to optimized inference and agentic workloads, GMI Cloud gives AI teams the infrastructure they need to build, deploy, and scale on one unified cloud.
One cloud for compute, inference, and agents.
Role Overview
We are seeking a skilled Site Reliability Engineer to join the GMI Global Infrastructure team. This role is hands-on and critical to ensuring the stability, efficiency, and reliability of the large-scale high performance AI/ML clusters in our data center. The ideal candidate will bring expertise in system-level troubleshooting, AI cluster maintenance, and operational excellence to ensure maximum performance for our infrastructure. Experience with large-scale infrastructure automation is considered a strong plus.
Responsibilities
- Design, implement and maintain scalable AI/ML infrastructure solutions.
- Proactively monitor GPU cluster health, performance and troubleshoot issues across compute, accelerator, and storage systems.
- Automate deployment, configuration and management of infrastructure resources.
- Manage GPU node lifecycle workflows, including provisioning, scaling, maintenance, decommissioning and upgrades of GPU nodes.
- Implement CI/CD pipelines for infrastructure deployment and orchestration.
- Ensure security, compliance and best practices across infrastructure.
- Manage incident response related to Infrastructure resources (GPU, CPU, Storage, Network).
- Handle customer provisioning requests for GPU resources, including onboarding, configuration and troubleshooting; resolve customer service requests related to GPU infrastructure, ensuring high customer satisfaction.
- Stay current with emerging GPU hardware and software technologies, integrating improvements as appropriate.
- Regional/international travel to GMI data center locations.
Qualifications
- Bachelor’s degree in Computer Science or related field.
- Over 3+ years of experience in data center operations, infrastructure, or systems engineering.
- Proven experience in site reliability engineering and infrastructure automation (e.g. Ansible, Terraform)
- Familiarity with containers orchestration platform (e.g. Kubernetes, Nvidia GPU operator, Nvidia Network operator, CNI, CSI) and job scheduling systems (e.g. Slurm).
- Familiarity with Linux system administration and scripting (Python, Bash).
- Familiarity with logging and monitoring tools such as Prometheus, Grafana, Loki.
- Good knowledge of GPU architecture, Nvidia CUDA, NCCL, or related AI/ML frameworks - added advantage.
- Strong troubleshooting skills and ability to analyze system logs and performance metrics.
- Excellent communication and teamwork abilities.
Meeting every qualification is not required—if you’re excited about this role, we’d love to hear from you. We believe diverse perspectives and experiences strengthen our team.
$95k - $171k
.... Opportunities exist to focus on GPU infrastructure, Kubernetes, and ensuring reliability for AI workloads within Akamai's serverless inference platform. As an Site Reliability Engineer II, you will be responsible for: Building and maintaining dashboards, alerts...SuggestedPermanent employmentWork experience placementWork at officeRemote workWork from homeWorldwideFlexible hours$71.6k - $119.4k
...support application teams. Our services provide applications with reliability, security, and better customer experiences. About the Job:... ...automation, troubleshoot issues, and work closely with senior engineers to learn and apply best practices. You’ll gain exposure to a...SuggestedFull timeTemporary workInternshipLocal areaWork from home- ...Senior Site Reliability Engineer Location: West Lake, CA or Carrolton, TX (ONSITE) FTE ONLY Must Have Technical/Functional Skills ~5-7 years of professional experience in a Site Reliability, DevOps, or Systems Engineering role. ~3-5 years of hands-on experience...SuggestedPermanent employment
$75.7k - $136.3k
...solve complex challenges? Do you have a passion for automation and building systems that scale? Join our highly skilled Site Reliability Engineering team! Our team designs, develops, and manages applications and infrastructure that support Akamai Cloud's products and...SuggestedWork experience placementWork at office$121.4k - $218.6k
...solve complex challenges? Do you have a passion for automation and building systems that scale? Join our highly skilled Site Reliability Engineering team! Our team designs, develops, and manages applications and infrastructure that support Akamai Cloud's products and...SuggestedWork experience placementWork at office$55k - $151.47k
...ApplicableSpecialismIFS - Internal Firm Services - OtherManagement LevelSenior AssociateJob Description & SummaryThe OpportunityAs a Site Reliability Engineer - Senior Associate, you will play a pivotal role in enhancing the reliability, scalability, and performance of our...Full timeH1b- ...footprint spans across three continents with deployments in 13 different countries. We are looking for a Manager of Site Reliability Engineering to join our team. The role requires a highly engaged leader who is obsessed with software engineering, reliability, monitoring...Casual workFlexible hours
$169.3k - $304.7k
...in building and maintaining fast, efficient, scalable, and reliable routing software and infrastructure that is responsible... ...growth and stability of our global platform. As a Principal Site Reliability Engineer - Network, you will be responsible for: Architecting,...Work experience placementWork at office- ...Release Train Engineer Lead Agile Delivery Across a Complex Government Transformation Are you an experienced Release Train Engineer (... ...coordination, execution, and release outcomes. Location: Client site Sacramento, CA Work Model: Onsite at the client site four days per...Monday to Thursday
- ...Distributed Systems Software Engineer, Python / Go Join to apply for the Distributed Systems Software Engineer, Python / Go role... ...automated testing approaches and infrastructure for validating reliability, performance, and resilience of cloud orchestration tools and...Full timeLocal areaRemote workWorldwide
$130k - $170k
About Us:BW Design Group is a fully integrated architecture, engineering, construction, system integration, and consulting firm committed to helping our clients realize their most critical goals from Strategy to Commercialization. As the only firm born from a manufacturing...Full timeFlexible hours$98.59k - $173.58k
...interpersonal, and listening skillsAbility to travel domestically or internationally 25-50%Bachelor’s degree in geography, computer science, engineering, or a related fieldVisa sponsorship is not available for this posting. Applicants must be authorized to work for any employer in...Local area$77k - $202k
...ApplicableSpecialismData, Analytics & AIManagement LevelSenior AssociateJob Description & SummaryThe OpportunityAs a GenAI Python Systems Engineer - Senior Associate, you will play a pivotal role in transforming raw data into actionable insights, enabling informed decision-...Full timeH1b$63k - $140k
...ApplicableSpecialismData, Analytics & AIManagement LevelAssociateJob Description & SummaryThe OpportunityAs a GenAI Python Systems Engineer - Experienced Associate, you will leverage advanced technologies and techniques to design and develop robust data solutions for...Full timeH1b$255k - $300k
...Francisco, California / Sacramento, California / San Ramon, California / San Jose, California( Client Services / Solutions ) - Sales Engineering /Full Time /HybridAHEAD builds platforms for digital business. By weaving together advances in cloud infrastructure, automation...Full timeWork at office- Senior Software Engineer (C#/.NET/React)Company OverviewWe are a healthcare technology company located in the Davis, CA area. We have proudly... ...healthcare systems and standards (HL7, FHIR, EPIC) to enable reliable data exchange and workflows.Author unit, integration, and end-...Flexible hours
$144.43k - $183.12k
...Applications /Exempt /HybridBerkshire Hathaway Homestate Companies, Workers Compensation Division, has an opening for a Software Engineer 3. The Software Engineer 3 designs, develops, and maintains software solutions in a hybrid development environment, encompassing on...Work visa2 days per week$107.6k - $198.4k
...strategists, data scientists, operators, creatives, designers, engineers, and architects. Our team balances business strategy,... ...and delivery tasksLead other engineers in the use of secure, reliable, and scalable AI-assisted development practices, including agent...Local area- ...across large-scale enterprise environments. As a member of our AI engineering team, you’ll play a critical role in designing and deploying... ...pipelines to observability and evaluation layers that ensure reliability, accountability, and performance. You’ll be responsible not...Flexible hours
$150k - $236k
...are at the heart of how we innovate and grow. The Solutions Engineering Team The Solutions Engineering (SE) team at Logitech for... ...capabilities to customer outcomes and business value. Professional reliability, responsiveness, judgment, and follow-through in a customer-...Work at officeImmediate startRemote workWork from homeHome officeFlexible hoursNight shift$86.4k
...contractual/access requirements) The Platform Engineer designs, implements, administers, and... ...to deliver secure, scalable, and reliable solutions that support business objectives... ...regularly from the office to various work sites or from site-to-site Occasionally Works...For contractorsWork at officeLocal area$120k - $140k
...Forward Deployment Engineer Fractal Analytics is a strategic AI partner to Fortune 500 companies with a vision to power every human... ...Ensure solutions meet enterprise requirements for security, reliability, maintainability, quality, and production readiness. Troubleshoot...Hourly payFull timeLocal areaRemote work$65 - $90 per hour
...healthcare system to hire for The Developer & Employee Experience Team, enabling the client's journey to the cloud. The Platform Engineer, Consultant, will report to the Sr. Manager of Technical Engineering. In this role you will design, build, and maintain the Internal...Contract workRemote work$95k - $155k
OverviewCarollo Engineers is a leading engineering firm dedicated exclusively to water. For over 90 years, we've specialized in the planning... ...entry and analysis, participate in field activities such as site investigations, and pilot testingQualificationsBachelor’s...Full timeFlexible hours$131.2k - $218.6k
Position Summary Lead Integration Engineer II supports the design, build, and delivery of integration solutions that connect platforms... ...experience, and the ability to translate business needs into reliable technical solutions. Work you'll do As a Lead Integration...Local areaVisa sponsorship$146k
...customer-first mindset, and a genuine drive to deliver happiness through transformative AI solutions. About the Team Our Solutions Engineering team acts as the strategic and technical backbone for complex enterprise deal cycles. We partner closely with Account Executives...Work at officeRemote work- ...L3Harris Technologies, Inc. seeks a Lead, Mechanical Engineering in Sacramento, CA to lead design efforts for solid propulsion systems and components. The role emphasizes independent work, objective-driven design, and leadership within a product team. Responsibilities...
- ...Senior Customer Solutions Engineer Join Intel's Client Computing and Physical AI Group and help enable the next generation of mobile... ...driving customer success and influencing products from development through launch. This role will require an on-site presence....
$100k - $150k
About Us:BW Design Group is a fully integrated architecture, engineering, construction, system integration, and consulting firm committed to helping our clients realize their most critical goals from Strategy to Commercialization. As the only firm born from a manufacturing...Full timeContract workFlexible hours$173.62k
...Software Engineer Participate in full software development lifecycle (SDLC): Engage in the complete software development lifecycle,... ...industry methodologies (e.g., Agile, Scrum) to develop scalable and reliable applications. Design, develop, and deploy software solutions:...Relocation
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
- official site Sacramento, CA
- site services specialist Sacramento, CA
- construction site safety Sacramento, CA
- IT site lead Sacramento, CA
- site leader Sacramento, CA
- site safety Sacramento, CA
- historic site Sacramento, CA
- junior website developer Sacramento, CA
- website coordinator Sacramento, CA
- on-site clinical research associate (traveling/remote) Sacramento, CA


