Site Reliability Engineer
$130k - $200kNscale
Site Reliability Engineer
Nscale is the GPU cloud built for AI. We run high-performance, cost-efficient infrastructure for AI-native startups and global enterprises, from bare metal up through the platform services teams actually build on. Our culture runs on ownership, accountability, and speed. We move with urgency, we tell each other the truth, and everyone here stays close to the infrastructure that makes AI work.
The Role
This is a career-level SRE role for someone who wants to own systems, not just watch them. You'll take real surface area: the automation and tooling other engineers depend on, and the reliability of production services running AI and GPU workloads at scale. You'll sit in the incident rotation, and you'll be expected to make the systems you touch quieter over time.
What You'll Do
• Build and own the automation and tooling that keeps the platform running; treat operational toil as a bug to be fixed, not a fact of life. • Define and maintain SLOs, SLIs, and the dashboards that make service health obvious at a glance. • Take point during incidents; troubleshoot under pressure, drive root cause analysis, and run post-incident reviews that actually change the system. • Investigate performance and reliability problems across Linux, networking, and distributed services, then fix them at the source. • Partner with Engineering, Networking, and Infrastructure teams to raise the reliability bar across the stack. • Improve availability, scalability, and efficiency through code, not manual effort.
What You'll Bring • 3-6 years in SRE, systems engineering, or software engineering, including time running production in a data center or cloud environment. • Strong programming skills (Python, Go, or similar) and a genuine bias toward automating the work away. • Solid command of Linux, networking fundamentals, and distributed systems. • A track record of troubleshooting live production issues and owning the fix through to the retro. • Fluency with monitoring and observability; metrics, logs, dashboards, and alerting. • Comfort in a fast-moving environment where priorities shift and you fill gaps without waiting to be asked.
Nice to Have
• Experience with AI or GPU workloads, or high-performance computing (HPC). • Familiarity with high-performance networking (InfiniBand, RDMA). • Kubernetes, plus virtualized or bare-metal environments.
On-Call and Pace A quick note on the shape of the job. This role sits close to production, so there is an on-call rotation, and some weeks are busier than others. We share it fairly, and we treat every page as a signal worth acting on rather than just an interruption. The goal is to make the systems quieter over time, so each rotation asks less of the person carrying it. If you take ownership of what you run and like leaving it in better shape than you found it, you'll do well here.
What We Offer • Competitive base plus equity, reviewed every 12 months. • Real scope early, and a progression plan built around the skills you want to sharpen. • Flexibility that treats you as an adult; we care that the work gets done, and we trust you to shape your day.
Salary Range $130,000 - $200,000 USD. Actual compensation varies with skill set, experience, and location, and the role may be eligible for bonus and equity.
Equal Opportunities Statement
At Nscale, we are committed to fostering an inclusive, diverse, and equitable workplace. We believe that a variety of perspectives enrich our work environment, and we encourage applications from candidates of all backgrounds, experiences, and abilities. We strongly encourage applications from people of colour, the LGBTQ+ community, people with disabilities, neurodivergent people, parents, carers, and people from lower socio-economic backgrounds.
If there's anything we can do to accommodate your specific situation, please let us know.
For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice.
- As a Site Reliability Engineer, you will be responsible for: Operational Excellence & Incident Management- Maintain and monitor production systems for availability, latency, and performance.- Lead incident response efforts, including communication, resolution, and postmortem...SuggestedPermanent employment
- The Senior Site Reliability Engineer is responsible for improving the reliability, availability, scalability, and operational excellence of our critical infrastructure platforms and services. This role partners closely with Engineering, Security, and Infrastructure teams...SuggestedFull timeWork at officeLocal area
- ...and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Corporate Technology, Risk Technology team, you will solve complex and broad business problems...Suggested
- Reliability Engineering Design, implement, and operate scalable, resilient, and highly available systems on Google Cloud Platform. Improve service... ...Skills, and Abilities Three or more years of experience in Site Reliability Engineering, platform engineering, DevOps, cloud...SuggestedRemote work
- ...Site Reliability Engineer The SRE manages and maintains infrastructure which supports cloud-based prototype applications as they transition into a production environment. The SRE assists in defining and measuring service level agreements and non-functional requirements...Suggested
- ...Site Reliability Engineer II PROS, Inc. is the leading offer management provider to the airline industry, helping airlines deliver seamless retail experiences designed to maximize revenue and margin growth. Powered by AI, the PROS Platform enables commercial teams to...Flexible hours
- ...applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems. As a Site Reliability Engineer III at JPMorgan Chase within the Corporate Technology, Corporate Know Your Customer (KYC) team, you will solve complex and...
- ...ENGINEERLocation: HOUSTON, TXFLSA Class: EXEMPTResponsible to: Directo of Software EngineeringPosition Summary: DevOps / Site Reliability Engineer to implement and evolve the infrastructure, deployment pipelines, and reliability posture of our systems. You'll work closely...Full timeLocal area
- As an Entry-Level DevOps Site Reliability Engineer, you will join a team responsible for continuous improvement and support of customer facing products. Responsibilities will include collecting system requirements; improving existing tools and processes through scripting...Work from home2 days per week
- As a Lead Site Reliability Engineer at JPMorgan Chase within the Corporate Know Your Customer (KYC) Technology group, you hold a leadership role in your team, demonstrate strong knowledge across multiple technical domains, and advise others on the technical and business...
- ...make an impact, and work with people who care, we'd love to meet you! ABOUT THE ROLE We are looking for a DevOps / Site Reliability Engineer to maintain operational resilience across Azure, AWS, and GCP in a 24x7 environment. This role blends platform...Local areaRemote workVisa sponsorshipWork visaFlexible hours
$213.1k - $300k
Lead a team of engineers to maintain service uptime while managing global on-call rotations... ...improve operational practices to drive reliability, maintainability, and stakeholder alignment... ...or in a Manager, Software Engineer, Site Reliability Engineering-related occupation...Full timeWork at office$61k - $101k
...formal training or certification in software engineering concepts, along with 5+ years of applied... .... We need deep expertise in reliability, scalability, performance, security, enterprise... ...architecture, toil reduction, and other site reliability practices, with the ability...- ...a company that values diversity, integrity, and growth. Role Overview PDI Technologies is looking for a Manager, Site Reliability Engineering to lead the SRE organization supporting Paylo, PDI’s payments, loyalty, and fuel-pricing product suite. This role owns the...
$132.23k - $176.31k
...future of AI‑ready connectivity, join us today. The Role We are seeking a highly skilled and proactive Senior Lead Site Reliability Engineer (SRE) to join our team, focusing on production support and performance optimization across our portal ecosystem. This role...Full timeTemporary workRemote work- ...JOB DESCRIPTION As a Site Reliability Engineer, you will be responsible for: Operational Excellence & Incident Management - Maintain and monitor production systems for availability, latency, and performance. - Lead incident response efforts, including communication...Permanent employmentFull time
- ...To Supervisor Analytics Cloud Services.Role OverviewThe Release Engineer is responsible for the deployment release and maintenance of... ...systems architects infrastructure and security teams to deliver reliable and scalable cloud solutions.Key ResponsibilitiesCI/CD Pipeline...
- ...accelerate autonomy development. We are seeking a software engineer with strong C++ expertise and a passion for building scalable simulation... ...in architecture and technical design discussions Build reliable, maintainable, and well-tested systems Contribute to code...Full time
- ...infrastructure, DevOps, SRE, and platform engineering. You will test AI-generated commands,... ...and deployment workflows for accuracy and reliability. Work with AWS, Azure, GCP,... ...Azure DevOps Cloud Infrastructure Site Reliability Engineering (SRE) Platform...Remote jobFor contractors
$113k - $141.53k
...leader in global energy. Senior Solutions Engineer - Systems Integration serves as a... ...functionally in the field, ensuring safe, reliable, and performant operation across diverse... ...Willingness to travel to factories and project sites (25%).Preferred QualificationsMaster’s degree...Full timeFor contractorsLocal areaWorldwideFlexible hours- Position: Software Engineer- Flight & Ground Systems Location: Houston, TX Remote Status: On-Site Job Id: 866 # of Openings... ...requirements and translate them into reliable software solutions.Self-motivated and...Permanent employmentFull timeRemote workRelocation package
- Position: Junior Software Engineer (Flight & Ground Systems) Location: Houston, TX Remote Status: On-Site Job Id: 867 # of Openings: 1 Junior Software Engineer (Flight & Ground Systems)Location...Permanent employmentFull timeInternshipRemote workRelocation package
- Reliability EngineerHouston, TXThe actual location of this job is in Houston, TX, US. Relocation... ...families, if needed.This is a fully site‑based role. Working together in person supports... ...environmentOpportunities to grow your engineering career in a global...Full timeRelocation package
$100k - $120k
...industry, join our team as we help shape a brighter way forward. JLL - Reliability Engineer (P3)Location: Spring, TX 77389Work Schedule: Hybrid (3 days onsite), M-F 8-4Travel requirements: 10% (other client sites in OR & WA)Reports to: Sr. Reliability EngineerEstimated base...Full timeWork at officeFlexible hours- ...together to shape what’s next. Whether you're engineering advanced materials, transforming... ...imagine everything.Role Impact Remote / Multi-Site - North America (Travel Required) Alternate Titles: Network Reliability Engineer | Plant Reliability Engineer | Maintenance...Contract workLocal areaRemote workWorldwideShift work
- ...infrastructure challenges. Job DetailsViridien is seeking a Platform Engineer - Infrastructure & Cloud Systems to design, build, and improve... ...observability tooling. This role focuses on building scalable, reliable systems and ensuring strong integration between infrastructure...Full timeRelocationFlexible hours
- ...Release Engineer - Azure Cloud Deployment Location: Houston, TX (Onsite) Job Type: Contract About the Role Join our dynamic... ...release engineering activities for cloud services to ensure secure, reliable deployments • Oversee application and infrastructure...Contract work
- ...data platforms that support high-volume, data-intensive workflows.The team works across backend engineering, infrastructure, and data systems, collaborating to deliver reliable, high-performance services in a modern cloud-native environment.Key Responsibilities-Backend...Full timeFlexible hours
- ...Release Engineer Visa status: U.S. Citizens and those authorized to work in the U.S. are encouraged to apply. Tax Terms: W2, 1099 Corp-Corp or 3rd Parties: Yes Technical skills and knowledge: Must be proficient with Source Control systems, like Git, to create...
$101.6k - $152.4k
We are seeking a talented Senior Engineer I, Digital Solutions to join our team and take charge of designing, developing, and deploying... ...alarms, and reports.Travel: Willingness to travel to customer sites as required. Travel is roughly expected to be around 25% but is...Full timeTemporary workImmediate startRemote workWork from homeFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
- site reliability engineer Houston, TX
- site reliability engineer sre Houston, TX
- site recruiter Houston, TX
- junior website developer Houston, TX
- on site coordinator Houston, TX
- construction site safety Houston, TX
- site services specialist Houston, TX
- website content developer Houston, TX
- website coordinator Houston, TX
- on-site clinical research associate (traveling/remote) Houston, TX




