Site Reliability Engineer
$100k - $170kNscale
Site Reliability Engineer
Houston; San Francisco; Seattle
About Nscale
Nscale is the GPU cloud built for AI. We run high-performance, cost-efficient infrastructure for AI-native startups and global enterprises, from bare metal up through the platform services teams actually build on. Our culture runs on ownership, accountability, and speed. We move with urgency, we tell each other the truth, and everyone here stays close to the infrastructure that makes AI work.
The Role
This is a career-level SRE role for someone who wants to own systems, not just watch them. You'll take real surface area: the automation and tooling other engineers depend on, and the reliability of production services running AI and GPU workloads at scale. You'll sit in the incident rotation, and you'll be expected to make the systems you touch quieter over time.
What You'll Do
- Build and own the automation and tooling that keeps the platform running; treat operational toil as a bug to be fixed, not a fact of life.
- Define and maintain SLOs, SLIs, and the dashboards that make service health obvious at a glance.
- Take point during incidents; troubleshoot under pressure, drive root cause analysis, and run post-incident reviews that actually change the system.
- Investigate performance and reliability problems across Linux, networking, and distributed services, then fix them at the source.
- Partner with Engineering, Networking, and Infrastructure teams to raise the reliability bar across the stack.
- Improve availability, scalability, and efficiency through code, not manual effort.
What You'll Bring
- 3-6 years in SRE, systems engineering, or software engineering, including time running production in a data center or cloud environment.
- Strong programming skills (Python, Go, or similar) and a genuine bias toward automating the work away.
- Solid command of Linux, networking fundamentals, and distributed systems.
- A track record of troubleshooting live production issues and owning the fix through to the retro.
- Fluency with monitoring and observability; metrics, logs, dashboards, and alerting.
- Comfort in a fast-moving environment where priorities shift and you fill gaps without waiting to be asked.
Nice to Have
- Experience with AI or GPU workloads, or high-performance computing (HPC).
- Familiarity with high-performance networking (InfiniBand, RDMA).
- Kubernetes, plus virtualized or bare-metal environments.
On-Call and Pace
A quick note on the shape of the job. This role sits close to production, so there is an on-call rotation, and some weeks are busier than others. We share it fairly, and we treat every page as a signal worth acting on rather than just an interruption. The goal is to make the systems quieter over time, so each rotation asks less of the person carrying it. If you take ownership of what you run and like leaving it in better shape than you found it, you'll do well here.
What We Offer
- Competitive base plus equity, reviewed every 12 months.
- Real scope early, and a progression plan built around the skills you want to sharpen.
- Flexibility that treats you as an adult; we care that the work gets done, and we trust you to shape your day.
Salary Range
$100,000 - $170,000 USD. Actual compensation varies with skill set, experience, and location, and the role may be eligible for bonus and equity.
Equal Opportunities Statement
At Nscale, we are committed to fostering an inclusive, diverse, and equitable workplace. We believe that a variety of perspectives enrich our work environment, and we encourage applications from candidates of all backgrounds, experiences, and abilities. We strongly encourage applications from people of color, the LGBTQ+ community, people with disabilities, neurodivergent people, parents, carers, and people from lower socio-economic backgrounds.
If there's anything we can do to accommodate your specific situation, please let us know.
The range below reflects the base salary for the position. Actual compensation may vary based on job-related factors such as skill set, experience, education, and location. In addition to base salary, this role may be eligible for bonus, equity, and/or commission programs. Nscale may offer a competitive benefits package including medical, dental, vision, flexible paid time off, parental leave, and retirement plan participation.
- As a Site Reliability Engineer, you will be responsible for: Operational Excellence & Incident Management- Maintain and monitor production systems for availability, latency, and performance.- Lead incident response efforts, including communication, resolution, and postmortem...SuggestedPermanent employment
- ...and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Corporate Technology, Risk Technology team, you will solve complex and broad business problems...Suggested
- The NexTier Technology team is looking for a Site Reliability Engineer (SRE) to help build, scale, and maintain highly reliable systems on Google Cloud Platform (GCP). This role blends software engineering with infrastructure expertise to ensure our services are performant...Suggested
- ...and AI agent a cryptographically secured identity, improving engineering velocity while maintaining security. We make trusted computing... ...problems that allow our customers to trust us for secure and reliable access to their infrastructure. Excellent security is table stakes...SuggestedWork at officeLocal areaRemote workSleeping nights
- ...Hynes & Khater is seeking a Cloud Support Engineer Lead to own the reliability, observability, and operational health of Azure-based applications. This is a hands-on leadership role; you will diagnose hard problems and define processes that keep systems running reliably...Suggested
- ...Site Reliability Engineer III There's nothing more exciting than being at the center of a rapidly growing field in technology and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems. As a Site Reliability...
$120k - $175k
...of sports fandom. Ready to reimagine the DFS industry together? We are seeking a highly skilled and experienced Senior Site Reliability Engineer to join our team. We are passionate about delivering cutting-edge solutions and pushing the boundaries of what's possible....Full timeRemote workWork visaFlexible hours$1,500 per month
...game worlds they inhabit. Our approach is centered around World Engine, our state-of-the-art onchain game server framework. World... ...architecture to keep our platform secure. Own delivery, scalability, and reliability of our backend infrastructure. Advise and collaborate with the...Full timeFlexible hours- As an Entry-Level DevOps Site Reliability Engineer, you will join a team responsible for continuous improvement and support of customer facing products. Responsibilities will include collecting system requirements; improving existing tools and processes through scripting...Work from home2 days per week
- ...ENGINEERLocation: HOUSTON, TXFLSA Class: EXEMPTResponsible to: Directo of Software EngineeringPosition Summary: DevOps / Site Reliability Engineer to implement and evolve the infrastructure, deployment pipelines, and reliability posture of our systems. You'll work closely...Full timeLocal area
- As a Lead Site Reliability Engineer at JPMorgan Chase within Market Risk Technology, you hold a leadership role in your team, demonstrate strong knowledge across multiple technical domains, and advise others on the technical and business issues facing them. Take lead and...
- ...Principal Site Reliability Engineer Join a globally recognized financial organization and advance your profession to new heights by contributing to revolutionary projects. You've discovered the perfect environment to have a major impact. As a Principal Site Reliability...
- ...JOB DESCRIPTION As a Site Reliability Engineer, you will be responsible for: Operational Excellence & Incident Management - Maintain and monitor production systems for availability, latency, and performance. - Lead incident response efforts, including communication...Permanent employmentFull time
- ...Please extend your support for this role. Local candidate will get 1st preference. Job Title: SRE Engineer Location: Houston, TX and Jersey City, NJ - 3 Days Onsite Role FTE role with Mphasis Client: Mphasis H1B transfer will work...Work experience placementH1bLocal area
$113k - $141.53k
...leader in global energy. Senior Solutions Engineer - Systems Integration serves as a... ...functionally in the field, ensuring safe, reliable, and performant operation across diverse... ...Willingness to travel to factories and project sites (25%).Preferred QualificationsMaster’s degree...Full timeFor contractorsLocal areaWorldwideFlexible hours- Reliability EngineerHouston, TXThe actual location of this job is in Houston, TX, US. Relocation... ...families, if needed.This is a fully site‑based role. Working together in person supports... ...environmentOpportunities to grow your engineering career in a global...Full timeRelocation package
$76k - $155.7k
...RegularPercentage of Travel Required: Up to 10%Type of Travel: Continental US* * *The Opportunity:CACI is seeking Software Systems Engineers to support the Artemis Next Generation Space Suit program at NASA Johnson Space Center. This position contributes to systems...Permanent employmentContract workFor contractorsWork experience placementImmediate startFlexible hours- ...infrastructure challenges. Job DetailsViridien is seeking a Platform Engineer - Infrastructure & Cloud Systems to design, build, and improve... ...observability tooling. This role focuses on building scalable, reliable systems and ensuring strong integration between infrastructure...Full timeRelocationFlexible hours
- ...data platforms that support high-volume, data-intensive workflows.The team works across backend engineering, infrastructure, and data systems, collaborating to deliver reliable, high-performance services in a modern cloud-native environment.Key Responsibilities-Backend...Full timeFlexible hours
$102.97k - $131.69k
...compliance with regulatory requirements and ATS policies and procedures. Partners with internal/external customer for engineered solutions to improve reliability and throughput. Identifies opportunities for Capital Expenditures for equipment replacement with supervision (...Work at office- ...accelerate autonomy development. We are seeking a software engineer with strong C++ expertise and a passion for building scalable simulation... ...in architecture and technical design discussions Build reliable, maintainable, and well-tested systems Contribute to code...Full time
- Position Title: Senior Software Engineer - Platform Location: Houston, TX onsiteFLSA Class: ExemptReports To: Manager of Software EngineeringPosition Summary:VoltaGrid is seeking a Senior Technical Solutions Engineer to join our Platform Team, responsible for designing...Full timeLocal area
- ...Role: Release Engineer Type: Contract Location: Houston, TX(5 days onsite) Release Engineer with CI/CD pipelines... ...systems architects infrastructure and security teams to deliver reliable and scalable cloud solutions. Key Responsibilities...Contract workShift work
- ...experience. Build software that matters alongside experienced engineers across software, firmware, and hardware domains. Grow into broader... ...scenarios to uncover edge cases, performance limits, and reliability opportunitiesDebug and resolve challenging issues that span multiple...Full timeWork experience placementWork at officeLocal areaImmediate start2 days per week
- ...Reliability Engineer Location: Houston, TX, US, 77049 LyondellBasell is a leader in the global chemical industry creating solutions for... ...supports capital projects with RAM analysis, and partners with site teams to develop and execute reliability and maintenance strategies...Local area
- ...Job Title The Reliability Engineer is primarily responsible for plant reliability measurement and improvement and the mechanical integrity program. Additional responsibilities include assisting the maintenance manager with maintenance systems and turn-around planning...Temporary workShift work
- ...Reliability Engineer Dynamis ElectroQuip Manufacturing - Houston, TX 77049 Overview Level: Experienced Position Type: Full Time Job... ...attendance; daily overtime required when on assignment at pad sites. EDUCATION/EXPERIENCE LEVEL Bachelor's degree in...Full timeWork at officeMonday to FridayShift workWeekend workAfternoon shiftEarly shift
- ...the transformation of low-Earth orbit into a global space marketplace. Our mission-driven team is seeking a bold and dynamic Reliability Engineer who is fueled by high ownership, execution horsepower, growth mindset, and driven to understand our world, science/...Permanent employmentWork at officeWeekend workAfternoon shift
- ...A leading engineering firm is seeking an Electrical Reliability Manager to oversee the electrical reliability program in their North American operations. The... ...driving improvements and optimizing the reliability of electrical systems across multiple sites. #J-18808-Ljbffr...
- ...innovative ways to help people? Do you like having the autonomy to build new solutions from the ground up? If so, being a Software Engineer III at Frost could be the job for you.At Frost, it’s about more than a job. It’s about having a flourishing career where you can...Full time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
- site reliability engineer Houston, TX
- site reliability engineer sre Houston, TX
- site services specialist Houston, TX
- construction site safety Houston, TX
- site leader Houston, TX
- official site Houston, TX
- website content developer Houston, TX
- remote website tester Houston, TX
- on site coordinator Houston, TX
- IT site lead Houston, TX



