Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer

GMI Cloud

About GMI

GMI Cloud is a fast-growing, AI-native infrastructure company delivering high-performance GPU compute, inference services, and infrastructure for AI agents.

Following 8x ARR growth, GMI Cloud continues to scale rapidly across the U.S. and APAC. As a Reference Platform NVIDIA Cloud Partner (NCP) and a validated leading NCP across both markets, we power production AI for leading AI-native companies including Fireworks AI, Cartesia, Reflection, and OpenRouter.

From large-scale compute to optimized inference and agentic workloads, GMI Cloud gives AI teams the infrastructure they need to build, deploy, and scale on one unified cloud.

One cloud for compute, inference, and agents.

Role Overview

We are seeking a skilled Site Reliability Engineer to join the GMI Global Infrastructure team. This role is hands-on and critical to ensuring the stability, efficiency, and reliability of the large-scale high performance AI/ML clusters in our data center. The ideal candidate will bring expertise in system-level troubleshooting, AI cluster maintenance, and operational excellence to ensure maximum performance for our infrastructure. Experience with large-scale infrastructure automation is considered a strong plus.

Responsibilities

  • Design, implement and maintain scalable AI/ML infrastructure solutions.
  • Proactively monitor GPU cluster health, performance and troubleshoot issues across compute, accelerator, and storage systems.
  • Automate deployment, configuration and management of infrastructure resources.
  • Manage GPU node lifecycle workflows, including provisioning, scaling, maintenance, decommissioning and upgrades of GPU nodes.
  • Implement CI/CD pipelines for infrastructure deployment and orchestration.
  • Ensure security, compliance and best practices across infrastructure.
  • Manage incident response related to Infrastructure resources (GPU, CPU, Storage, Network).
  • Handle customer provisioning requests for GPU resources, including onboarding, configuration and troubleshooting; resolve customer service requests related to GPU infrastructure, ensuring high customer satisfaction.
  • Stay current with emerging GPU hardware and software technologies, integrating improvements as appropriate.
  • Regional/international travel to GMI data center locations.

Qualifications

  • Bachelor’s degree in Computer Science or related field.
  • Over 3+ years of experience in data center operations, infrastructure, or systems engineering.
  • Proven experience in site reliability engineering and infrastructure automation (e.g. Ansible, Terraform)
  • Familiarity with containers orchestration platform (e.g. Kubernetes, Nvidia GPU operator, Nvidia Network operator, CNI, CSI) and job scheduling systems (e.g. Slurm).
  • Familiarity with Linux system administration and scripting (Python, Bash).
  • Familiarity with logging and monitoring tools such as Prometheus, Grafana, Loki.
  • Good knowledge of GPU architecture, Nvidia CUDA, NCCL, or related AI/ML frameworks - added advantage.
  • Strong troubleshooting skills and ability to analyze system logs and performance metrics.
  • Excellent communication and teamwork abilities.

Meeting every qualification is not required—if you’re excited about this role, we’d love to hear from you. We believe diverse perspectives and experiences strengthen our team.

Vacancy posted 17 hours ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in Dallas, TX vacancy
  • Qualifications: 8+ years of Software Engineering experience, or equivalent demonstrated through...  ...implement and maintain scalable and reliable infrastructure on Google Cloud Platform...  ...vendor resources Willingness to work on-site at stated location in the job openingDepartment... 
    Suggested
    Contract work
    For contractors
    Work experience placement

    Cedent Consulting

    Dallas, TX
    3 days ago
  •  ...their SAP and Oracle landscapes, whether running on-premises, in the cloud, or in a hybrid environment. We are looking for a Site Reliability Engineer II to join our global engineering team. In this role, you will apply software engineering principles to operations,... 
    Suggested
    Work at office
    Flexible hours
    Shift work
    2 days per week

    Onapsis

    Dallas, TX
    3 hours ago
  •  ...ll be building the future of financial infrastructure. As part of our global expansion, we're looking for a hands-on Site Reliability Engineer (SRE) to design, scale, and safeguard the reliability of our next-generation financial platforms. This is a high-impact role... 
    Suggested
    Remote work

    Longbridge Singapore

    Dallas, TX
    3 days ago
  • Senior Talent Acquisition Specialist @ Centraprise Job Role: SRE Developer Job Type: Full time/ Permanent Location: Irving, TX Job Description: Must Have Technical/Functional Skills Proven experience in managing Windows and Linux servers. Proficiency...
    Suggested
    Permanent employment
    Full time

    Centraprise

    Irving, TX
    3 days ago
  •  ...on one unified cloud. One cloud for compute, inference, and agents. Role Overview We are seeking a skilled Site Reliability Engineer to join the GMI Global Infrastructure team. This role is hands-on and critical to ensuring the stability, efficiency, and... 
    Suggested

    GMI Cloud

    Syracuse, NY
    17 hours ago
  •  ...generative AI and cloud-native platforms to advanced release engineering practices, our teams are redefining how financial technology...  ...championing best practices in coding, testing, and automation. Reliability Engineering: Establish service level objectives (SLOs) and... 
    H1b
    Work at office
    Remote work
    Visa sponsorship
    Flexible hours
    2 days per week
    3 days per week

    GM Financial

    Irving, TX
    3 days ago
  • $140k - $150k

     ...your skills and experience — talk with your recruiter to learn more. Base pay range $140,000.00/yr - $150,000.00/yr Site Reliability Engineer II | 6-month Contract to Hire | Hybrid (Irving, TX) | 2x onsite per week Optomi, in partnership with a leading... 
    Full time
    Contract work

    Optomi

    Irving, TX
    3 days ago
  • $72.1k - $158.62k

     ...person, one family and one community at a time. Position Summary We are seeking a highly skilled Software Development Engineer, Site Reliability Engineering (SRE), for Retail and Pharmacy platforms to drive reliability, scalability, and operational excellence. The... 
    Hourly pay
    Full time
    Temporary work
    Local area

    CVS Health

    Richardson, TX
    1 day ago
  •  ...Job Title: Site Reliability Engineer Location: Dallas TX (HYBRID) Duration :Full Time Job Description: Skill: Site Reliability Engineer • Ensures supported applications are functioning and available by minimizing downtime and maximizing performance... 
    Full time
    Work at office

    Syntricate Technologies

    Dallas, TX
    16 hours ago
  •  ...ensure applications are highly available, reliable, and performant at a global scale....  ...Bachelor of Computer Science or related Engineering field required. Master's Degree preferred...  ...Minimum of 1 year of lead experience of site reliability engineering team required.... 
    Contract work
    Work at office

    3B Staffing LLC

    Irving, TX
    1 day ago
  •  ...and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems. As a Site Reliability Engineer III at JPMorgan Chase within the Chief Data & Analytics Office (CDAO) AI/ML & Data Platforms team,, you will solve complex... 
    Work at office

    J.P. Morgan

    Dallas, TX
    5 days ago
  •  ...Ethos Group is seeking a talented and proactive Site Reliability Engineer (SRE) to join our growing technology team. This role is ideal for an engineer who is passionate about building highly available, scalable, and reliable systems while partnering closely with development... 

    Ethos Group

    Irving, TX
    3 days ago
  • Mandatory Skills: AWS/Azure/GCP (GCP is not used very much ). Kubernetes /Helm,Docker,Gitlab,Grafana,Cyberark/Hashicorp Vault, Terraform etc. Experience utilizing Java, Perl, Python, Go and scripting experience in Shell and Perl to automate reports and monitor enterprise...

    Omni Inclusive

    Dallas, TX
    5 days ago
  •  ...healthcare fintech innovator, we’re transforming the patient journey and redefining what’s possible in dental care. This role: Site Reliability Engineer (SRE) with deep expertise in monitoring, debugging, and optimizing Azure App Services. This position is critical to... 
    Full time
    Work at office
    3 days per week

    Wellfit Technologies

    Irving, TX
    3 days ago
  •  ...improving platform infrastructure and applications with high reliability, resiliency, performance & quality, and faster time-to-market...  ...documentation, including runbooks/playbooks; and, Using Chaos Engineering to test the robustness of the systems and applications.... 

    Software Technology Inc

    Dallas, TX
    2 days ago
  •  ...Senior Site Reliability Engineer Come join a growing bank at the heart of the innovation, technology, green tech and life sciences space. We continue to expand our global footprint and our banking technology is at the core of everything we do. As a Senior Site Reliability... 

    Professional Recruiters

    Dallas, TX
    2 days ago
  • $136.2k - $214.01k

     ...outcomes Visionary in future focused problem-solving Exceptional in execution and impact The Role As a Senior Site Reliability Engineer at Proofpoint you will develop a deep understanding of the various services and applications that come together to... 
    Full time
    Flexible hours

    Proofpoint

    Dallas, TX
    4 days ago
  •  ...Site Reliability Engineer We are looking for a Site Reliability Engineer for our client location in Dallas TX with the following skills: Java Spring Boot, Kubernetes, and eCommerce experience required. Key responsibilities include working with the applications, engineering... 
    Work at office

    STIAOS Technologies

    Dallas, TX
    3 days ago
  •  ...exclusive features. Our client is looking for a highly skilled Site Reliability Engineerto deploy, configure, and support our carrier-grade...  ...customers. Work closely with customers, network engineers, and software developers to ensure seamless product integration... 
    Full time
    Shift work

    Felix Recruitment

    Dallas, TX
    3 days ago
  •  ...automate them. # Experience in Implementing AI/ML-based monitoring and self-healing solutions. # Experience in Implementing Chaos Engineering/testing. Seniority level Mid-Senior level Employment type Full-time Job function Consulting, Analyst, and... 
    Full time

    Infosys

    Richardson, TX
    3 days ago
  •  ...Role: Site Reliability Engineer 6+ months Contract role Remote About the Role We are looking for a dynamic and accomplished Site Reliability Engineer (SRE) who excels at solving complex reliability challenges and thrives in high-impact environments.... 
    Contract work
    Remote work

    ECHO IT SOLUTIONS INC .

    Farmers Branch, TX
    2 days ago
  •  ...Sr. Site Reliability Engineer Our client, a top tier IT Consulting firm is looking for several qualified Site Reliability Engineers to join a Top-Tier Investment Bank. Essential Requirements and Responsibilities: Proficiency in designing, deploying, and maintaining... 

    RIT Solutions

    Dallas, TX
    3 days ago
  • $174k - $252k

     ...systems by pushing for changes that improve reliability and velocity.Practice sustainable...  ...:Bachelor’s degree in Computer Science, Engineering, a related field, or equivalent practical...  ...degree in Computer Science or Engineering.Site Reliability Engineering (SRE) is what you... 

    Google

    Sunnyvale, TX
    4 days ago
  • $125.7k - $203.1k

     ...collaborative team of systems and cloud engineers who thrive in a fast-paced environment built...  ...maintain efficient platform uptime and reliability. • Lead end-to-end incident response and...  ...steps.• Collaborate closely with Site Reliability Engineering (SRE) and Product... 
    Permanent employment
    Full time
    Temporary work
    Apprenticeship
    Work experience placement
    Local area
    Worldwide
    Flexible hours
    Night shift

    CISCO Systems

    Richardson, TX
    1 day ago
  • $147k - $210k

     ...product or system development code.Review code developed by other engineers and provide feedback to ensure best practices (e.g., style...  ..., and troubleshooting large-scale distributed systems. Site Reliability Engineering (SRE) is what you get when you treat operations... 

    Google

    Sunnyvale, TX
    3 hours ago
  • $77.5k - $179k

    Site Reliability Engineer I - Sales OperationsThis role has been designed as 'Hybrid' with a requirement that you will work on average 2 days per week from an HPE office.Who We Are:Hewlett Packard Enterprise is the global edge-to-cloud company advancing the way people live... 
    Full time
    Work experience placement
    Internship
    Work at office
    Local area
    Immediate start
    2 days per week

    Hewlett Packard Enterprise

    Dallas, TX
    1 day ago
  • $55k - $151.47k

     ...ApplicableSpecialismIFS - Internal Firm Services - OtherManagement LevelSenior AssociateJob Description & SummaryThe OpportunityAs a Site Reliability Engineer - Senior Associate, you will play a pivotal role in enhancing the reliability, scalability, and performance of our... 
    Full time
    H1b

    PwC

    Dallas, TX
    3 days ago
  • $172k - $300k

    Job DescriptionGM Vehicle Autonomy is forming a centralized Site Reliability Engineering team to make reliability a measurable, engineered property of the systems used to build, validate, release, and operate autonomous-vehicle software.As one of our founding SREs, you... 
    Full time
    Work at office
    Local area
    Remote work
    Work from home
    Relocation
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, TX
    1 day ago
  • $192.4k - $275.8k

     ...CloudOps— the team that keeps Splunk Cloud running for some of the world's most demanding enterprise customers, blending Site Reliability Engineering, Systems Engineering, and Service Engineering disciplines at a scale very few teams ever get to operate at. When the... 
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    Richardson, TX
    4 days ago
  •  ...Position - Senior/Lead Site Reliability Engineer Observability Location - 100% Remote Experience - 8+ Years Type - Full Time Technology Stack - Splunk Enterprise, Splunk Cloud, Elasticsearch, ELK, Kibana, Prometheus, Grafana, Grafana Tempo, OpenTelemetry... 
    Full time
    Remote work

    vaaridatech

    Dallas, TX
    a month ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!