Site Reliability Engineer
GMI Cloud
About GMI
GMI Cloud is a fast-growing, AI-native infrastructure company delivering high-performance GPU compute, inference services, and infrastructure for AI agents.
Following 8x ARR growth, GMI Cloud continues to scale rapidly across the U.S. and APAC. As a Reference Platform NVIDIA Cloud Partner (NCP) and a validated leading NCP across both markets, we power production AI for leading AI-native companies including Fireworks AI, Cartesia, Reflection, and OpenRouter.
From large-scale compute to optimized inference and agentic workloads, GMI Cloud gives AI teams the infrastructure they need to build, deploy, and scale on one unified cloud.
One cloud for compute, inference, and agents.
Role Overview
We are seeking a skilled Site Reliability Engineer to join the GMI Global Infrastructure team. This role is hands-on and critical to ensuring the stability, efficiency, and reliability of the large-scale high performance AI/ML clusters in our data center. The ideal candidate will bring expertise in system-level troubleshooting, AI cluster maintenance, and operational excellence to ensure maximum performance for our infrastructure. Experience with large-scale infrastructure automation is considered a strong plus.
Responsibilities
- Design, implement and maintain scalable AI/ML infrastructure solutions.
- Proactively monitor GPU cluster health, performance and troubleshoot issues across compute, accelerator, and storage systems.
- Automate deployment, configuration and management of infrastructure resources.
- Manage GPU node lifecycle workflows, including provisioning, scaling, maintenance, decommissioning and upgrades of GPU nodes.
- Implement CI/CD pipelines for infrastructure deployment and orchestration.
- Ensure security, compliance and best practices across infrastructure.
- Manage incident response related to Infrastructure resources (GPU, CPU, Storage, Network).
- Handle customer provisioning requests for GPU resources, including onboarding, configuration and troubleshooting; resolve customer service requests related to GPU infrastructure, ensuring high customer satisfaction.
- Stay current with emerging GPU hardware and software technologies, integrating improvements as appropriate.
- Regional/international travel to GMI data center locations.
Qualifications
- Bachelor’s degree in Computer Science or related field.
- Over 3+ years of experience in data center operations, infrastructure, or systems engineering.
- Proven experience in site reliability engineering and infrastructure automation (e.g. Ansible, Terraform)
- Familiarity with containers orchestration platform (e.g. Kubernetes, Nvidia GPU operator, Nvidia Network operator, CNI, CSI) and job scheduling systems (e.g. Slurm).
- Familiarity with Linux system administration and scripting (Python, Bash).
- Familiarity with logging and monitoring tools such as Prometheus, Grafana, Loki.
- Good knowledge of GPU architecture, Nvidia CUDA, NCCL, or related AI/ML frameworks - added advantage.
- Strong troubleshooting skills and ability to analyze system logs and performance metrics.
- Excellent communication and teamwork abilities.
Meeting every qualification is not required—if you’re excited about this role, we’d love to hear from you. We believe diverse perspectives and experiences strengthen our team.
$150k - $200k
...nationwide healthcare organization, creating unique engineering challenges around scale, reliability, security, real-time communication, and healthcare infrastructure... ...platform. NOCD is looking for a Senior Site Reliability Engineer (SRE) to help shape the...SuggestedFull timeWork at office- ...We are seeking a Staff Site Reliability Engineer to serve as the foundational Technical Lead for our Platform Engineering SRE organization. In this role, you will be the primary architect and visionary for the core technology foundations. As the technical lead for all...SuggestedFull time
$92.52k - $138.79k
...insights you need to drive results. FreeWheel’s platform makes TV and video advertising work.Job DescriptionWe're looking for a Site Reliability Engineer to own cloud infrastructure, system reliability, and observability for the Freewheel BuyerCloud and Revenue Science teams....SuggestedFull timeWorldwide$158.5k - $172k
...exceptional value they deserve.About The OpportunityAs a Senior Engineer on the Runtime Automation team, you will design, automate, and... .... This is a high-impact position driving continuous reliability, deep system optimization, and automation across our entire technology...SuggestedFull timeTemporary workWork at officeFlexible hours3 days per week- ...StatesIndustry: Trading FirmPosted: 2026-08-17Contact: Ethan HudsonEmail: ****@*****.***: (***) ***-****Job Title: Site Reliability Engineer (Infrastructure & Systems)Location: Chicago, IL (Greater Metro Area)About the OpportunityJoin a premier financial...SuggestedLocal area
- Qualifications: 8+ years of Software Engineering experience, or equivalent demonstrated through... ...implement and maintain scalable and reliable infrastructure on Google Cloud Platform... ...vendor resources Willingness to work on-site at stated location in the job openingDepartment...Contract workFor contractorsWork experience placement
- Play a key role in ensuring system reliability at one of the world’s most iconic and largest financial institutions.As a Site Reliability Engineer II at JPMorgan Chase within the Commercial and Investment Banking and Payment Technology Team, you will use technology to solve...
$130k - $180k
...of both work styles in a workplace that is intentional about belonging, collaboration, and accomplishment.Being a Senior Site Reliability Engineer at iManage Means… You are an engineer, a builder, and a systems thinker. You’ll create middleware and platform guardrails...Work at officeLocal areaRemote workWorldwideMonday to FridayFlexible hours$106k - $130k
...for any employer, at the date of hire. This position is ineligible for employment Visa sponsorship.Role Summary The Senior Site Reliability Engineer applies software engineering and systems engineering practices to improve the reliability, resilience, scalability, and...Hourly payFull timeImmediate startVisa sponsorshipWork visaFlexible hours- ...and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Commercial and Investment Banking and Payment Technology Team, you will solve complex and...
$91.2k - $136.8k
Reliability Engineer - IE08GEWe’re determined to make a difference and are proud to be an insurance company that goes well beyond coverages and... ...field.3+ years of experience in Infrastructure Engineering, Site Reliability Engineering (SRE), or DevOps.Hands-on experience...Full timeTemporary workWork at office3 days per week$108.08k - $172.5k
Work with development and platform engineering teams to migrate and maintain applications in Google Cloud. Apply Observability concepts and applications to maintain services. Monitor metrics, system health and analyze reports. Provide on-call rotation support for production...Full timeRemote workWorldwide$100k - $120k
OverviewThe Site Reliability Engineer is a key force behind improving Origami’s time to resolution and advancing overall site reliability and scalability. This person participates in efforts to identify root causes during post-incident investigations, while also identifying...Full timeTemporary workWork experience placementFlexible hours$130k - $170k
...Senior Site Reliability Engineer About Us Founded in 2014, we offer the industry’s first and only cloud‑based, fully‑customisable, end‑to‑end software solution to automate securities‑based lending from origination through the life of the loan. By combining thought...Full timeFlexible hoursShift work- ...A senior Site Reliability Engineer will join an established infrastructure function responsible for highly available, security-conscious cloud systems supporting complex business-critical workloads. You’ll take significant ownership of reliability, scalability, and...Full timeRemote work
- ...Get AI-powered advice on this job and more exclusive features. Direct message the job poster from Algo Capital Group Senior Site Reliability Engineer - Observability and Automation A leading high-frequency trading firm is seeking a mid to senior-level Site Reliability...Full timeWork at officeFlexible hours
$114k - $155k
...applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems. As a Site Reliability Engineer III at JPMorgan Chase within the Commercial and Investment Banking and Payment Technology Team, you will solve complex and...Local area$180k - $200k
...Come join tastytrade, part of IG Group, as we build the reliability practice behind the brokerage platform that active options... ...equities traders rely on every market day. As our first Senior Site Reliability Engineer, you'll define what reliability means at tastytrade, from...Work at office3 days per week- ...electronic production — all running on AWS with zero downtime tolerance for firms in active litigation. We're looking for a Site Reliability Engineer to help maintain the reliability, scalability, and security posture of that platform as we expand our AI capabilities (...
- ...Senior Site Reliability Engineer About The Position We are looking for a Senior Reliability Engineer to join our Platform team. In this position, you will be responsible for maintaining, designing, implementing and upgrading our cloud infrastructure to support our...Temporary workFlexible hours
$250k - $350k
...and India, where quantitative researchers, engineers, traders, and operational teams work together... ...that boost stability, throughput, and reliability Qualifications Minimum of 3 years’ experience in production support, site reliability, or infrastructure operations in...Full time$150k - $200k
...the job poster from Selby Jennings Recruitment Consultant @ Selby Jennings | Financial Technology We are seeking a Site Reliability Engineer to join our team and assist with the design, development, and administration of our trading and research systems. This role...Full timeWork at office$99.75k - $125k
...Play a key role in ensuring system reliability at one of the world\'s most iconic and largest financial institutions. As a Site Reliability Engineer II at JPMorgan Chase within the Commercial and Investment Banking and Payment Technology Team, you will use technology...Local area- ...professionals for this role. JOB DESCRIPTION Play a key role in ensuring system reliability at one of the world's most iconic and largest financial institutions. As a Site Reliability Engineer II at JPMorgan Chase within the Commercial and Investment Banking and Payment...Local area
- ...As a Site Reliability Engineer, you will build and secure infrastructure supporting our AI platform with special attention to safeguarding US customer data and supporting the Aerospace and Defense Industrial Base. You'll have strong ownership of US operations while collaborating...Immediate start
$152k - $205k
....About the roleAre you a systems-minded engineer who is happiest when production tells you... ...wasn't designed for? Do you want to own reliability for a platform that answers millions of... ...opportunity for you.We're looking for a Senior Site Reliability Engineer to join our...Local areaRemote workWork from homeVisa sponsorship- ...Job title : OTM - Senior Site Reliability Engineer Location: Chicago, IL Experience: 10 Years Role: Senior Site Reliability Engineer Oracle Transportation Management (OTM) Design, implement, and maintain highly available, scalable, and reliable Oracle Transportation...
$95.1k - $122.55k
About Us At Clearwater, we are dedicated to provide world-class enterprise applications, ensuring their performance and availability to support our clients in the ever-evolving fintech landscape. Our Enterprise Application Support team plays a vital role in maintaining...Shift work$160k - $200k
...world and need other like-minded individuals to accelerate and expand our efforts. Chicago, IL (Hybrid 3X a week) Senior Site Reliability Engineer (SRE) Chicago, IL (Hybrid) Opportunity Overview NOCD is looking for a Senior Site Reliability Engineer (SRE) to...Full timeWork at office- ...building and running systems that must perform reliably under real-time market conditions. The culture is highly collaborative, engineering-driven, and focused on continuous... ...related field ~3+ years of experience in site reliability, systems engineering, or technical...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
- site reliability engineer sre Chicago, IL
- site reliability engineer Chicago, IL
- site reliability engineer remote Chicago, IL
- official site Chicago, IL
- site services specialist Chicago, IL
- construction site safety Chicago, IL
- IT site lead Chicago, IL
- site recruiter Chicago, IL
- site leader Chicago, IL
- site safety Chicago, IL



