Site Reliability Engineer
National Oilwell Varco
As a Site Reliability Engineer, you will be responsible for: Operational Excellence & Incident Management- Maintain and monitor production systems for availability, latency, and performance.- Lead incident response efforts, including communication, resolution, and postmortem documentation.- Design and implement health checks, alerting systems, and automated remediation workflows.- Drive root cause analysis and implement permanent resolutions for recurring issues.Observability & Insights- Set up and maintain full observability stacks (logging, metrics, tracing) using tools like Prometheus, Grafana, Datadog, OpenTelemetry, or ELK.- Analyze telemetry and logs to identify trends, anomalies, and opportunities for improvement.- Conduct post-incident reviews and use insights to inform future engineering investments.Performance & Systems Optimization- Tune and optimize distributed systems, including AKKA.NET actors, for performance and resource efficiency.- Work with developers to evolve architecture and improve system throughput, latency, and stability.- Optimize PostgreSQL performance, queries, and maintenance strategies.CI/CD & Automation- Design and maintain modern CI/CD pipelines using GitHub Actions, Azure Pipelines, or GitLab CI.- Automate deployment, testing, and rollback processes to reduce friction and increase deployment frequency.- Standardize infrastructure as code practices across environments.We’d love to talk to you if you have:- 5+ years of experience in SRE, DevOps, or Infrastructure Engineering roles.- Expertise in Kubernetes and container orchestration at scale.- Strong experience with AKKA.NET or similar actor-based frameworks.- Proficiency with scripting and automation (Bash, PowerShell, Python).- Experience with observability tools (Phobos,Datadog, Prometheus, Grafana, OpenTelemetry, ELK).- Hands-on experience with cloud platforms (AWS, Azure, or GCP).- Strong PostgreSQL knowledge—performance tuning, query optimization, maintenance.- Proven ability to lead incident management and drive postmortem processes.- A builder’s mindset with high standards for operational excellence and technical ownership.Preferred Tools & Ecosystem Experience- CI/CD: GitHub Actions, Azure Pipelines, GitLab CI- Infrastructure: Kubernetes, Docker, Terraform- Monitoring: Phobos (AKKA.NET), Datadog, Prometheus- Source Control: GitHub, GitLab, Azure DevOps- Programming: C#, Python, Bash, PowerShellEvery day, the oil and gas industry’s best minds put more than 150 years of experience to work to help our customers achieve lasting success.We Power the Industry that Powers the WorldThroughout every region in the world and across every area of drilling and production, our family of companies has provided the technical expertise, advanced equipment, and operational support necessary for success—now and in the future.Global FamilyWe are a global family of thousands of individuals, working as one team to create a lasting impact for ourselves, our customers, and the communities where we live and work. Purposeful InnovationThrough purposeful business innovation, product creation, and service delivery, we are driven to power the industry that powers the world better.Service Above AllThis drives us to anticipate our customers’ needs and work with them to deliver the finest products and services on time and on budget.CorporateOur family of companies is supported by our global Corporate teams, providing expert knowledge from functions including Human Resources, Information Technology, Compliance, Finance, QHSE, Marketing and Legal centers of expertise. We are structured to provide guidance and service above all to all our business operations.Full timePosting Date: 2026-06-23
- ...The Senior Site Reliability Engineer is responsible for improving the reliability, availability, scalability, and operational excellence of our critical infrastructure platforms and services. This role partners closely with Engineering, Security, and Infrastructure teams...SuggestedFull timeWork at officeLocal area
- ...Nscale, a GPU cloud for AI, seeks a senior SRE to raise the reliability bar across the platform. You will own the hardest problems, influence architectural decisions, and mentor others while maintaining an on-call rotation that becomes lighter over time. You’ll drive...Suggested
- As a Site Reliability Engineer, you will be responsible for: Operational Excellence & Incident Management- Maintain and monitor production systems for availability, latency, and performance.- Lead incident response efforts, including communication, resolution, and postmortem...SuggestedPermanent employment
$61k - $101k
...formal training or certification in software engineering concepts, along with 5+ years of applied... .... We need deep expertise in reliability, scalability, performance, security, enterprise... ...architecture, toil reduction, and other site reliability practices, with the ability...SuggestedFull time- ...globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability. As a Lead Site Reliability Engineer at JPMorgan Chase within the Corporate & Investment Bank (CIB) Management and Support Functions Digital &...Suggested
$213.1k - $300k
...Manager, Software Engineer, Site Reliability Engineering Share Manager, Software Engineer, Site Reliability Engineering Google Houston, TX, USA Advanced Experience owning outcomes and decision making, solving ambiguous problems and influencing stakeholders...Full timeWork at office- ...applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems. As a Site Reliability Engineer III at JPMorgan Chase within the Corporate Technology, Corporate Know Your Customer (KYC) team, you will solve complex and...
- ...and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Corporate Technology, Risk Technology team, you will solve complex and broad business problems...
- ...JOB DESCRIPTION As a Site Reliability Engineer, you will be responsible for: Operational Excellence & Incident Management - Maintain and monitor production systems for availability, latency, and performance. - Lead incident response efforts, including communication...Permanent employmentFull time
- Reliability Engineering Design, implement, and operate scalable, resilient, and highly available systems on Google Cloud Platform. Improve service... ...Skills, and Abilities Three or more years of experience in Site Reliability Engineering, platform engineering, DevOps, cloud...Remote work
- ...infrastructure, DevOps, SRE, and platform engineering. You will test AI-generated commands,... ...and deployment workflows for accuracy and reliability. Work with AWS, Azure, GCP,... ...Azure DevOps Cloud Infrastructure Site Reliability Engineering (SRE) Platform...Remote jobFor contractors
- ...accelerate autonomy development. We are seeking a software engineer with strong C++ expertise and a passion for building scalable simulation... ...in architecture and technical design discussions Build reliable, maintainable, and well-tested systems Contribute to code...Full time
- As an Entry-Level DevOps Site Reliability Engineer, you will join a team responsible for continuous improvement and support of customer facing products. Responsibilities will include collecting system requirements; improving existing tools and processes through scripting...Work from home2 days per week
- ...ENGINEERLocation: HOUSTON, TXFLSA Class: EXEMPTResponsible to: Directo of Software EngineeringPosition Summary: DevOps / Site Reliability Engineer to implement and evolve the infrastructure, deployment pipelines, and reliability posture of our systems. You'll work closely...Full timeLocal area
- As a Lead Site Reliability Engineer at JPMorgan Chase within the Corporate Know Your Customer (KYC) Technology group, you hold a leadership role in your team, demonstrate strong knowledge across multiple technical domains, and advise others on the technical and business...
- ...Release Train Engineer 4 Months- Contract To Hire Pay- $65-$70 W2 Onsite Houston, TX Job Description The Release Train... ...sure all team activity, dashboards, and metrics are visible and reliable. Coach teams and Scrum Masters on agile and Scrum practices,...Contract workWork at office
- ...Role: Release Engineer Type: Contract Location: Houston, TX(5 days onsite) Release Engineer with CI/CD pipelines... ...systems architects infrastructure and security teams to deliver reliable and scalable cloud solutions. Key Responsibilities...Contract workShift work
- ...Release Engineer Visa status: U.S. Citizens and those authorized to work in the U.S. are encouraged to apply. Tax Terms: W2, 1099 Corp-Corp or 3rd Parties: Yes Technical skills and knowledge: Must be proficient with Source Control systems, like Git, to create...
- ...Release Train Engineer The Release Train Engineer's primary purpose is to lead Agile Release Trains (ART) to success consistently in large environments. The RTE resolves and escalates impediments, manages risk, assures value delivery, and drives program level continuous...
$213.1k - $300k
...Google Houston, TX, USA is seeking a Manager, Software Engineer in Site Reliability Engineering to lead a team of engineers focused on uptime, availability, and scalable infrastructure across global services. The role emphasizes ownership and decision making, with...- ...a company that values diversity, integrity, and growth. Role Overview PDI Technologies is looking for a Manager, Site Reliability Engineering to lead the SRE organization supporting Paylo, PDI’s payments, loyalty, and fuel-pricing product suite. This role owns the...
- ...they're harder. Knows when to cut corners and when to build for the long haul. Bridges the gap between business goals and engineering reality without losing sight of either. Guides and grows other developers—teaching them how to think, not just how to code....
- ...transformation of low-Earth orbit into a global space marketplace. Our mission-driven team is seeking a bold and dynamic Software Systems Engineer who is fueled by high accountability, execution horsepower, and driven to understand our world, science/technology, and life...Permanent employmentFull timeWork at officeWeekend workAfternoon shift
$200k
...Senior Mechanical Reliability EngineerOur client is seeking a full-time Senior Mechanical Reliability Engineer. Hybrid. This individual may reside in Midland/Odessa, Houston, or... ...for operational improvement throughout the site and organization.Provide technical support...Full time- ...We are seeking a Lead Project Engineer (Power Systems) to serve as a key technical expert and contributor within our Engineering and Consulting team. In this role, you will execute complex power system studies, substation design and EPC bids, enhance technical capabilities...Full timeContract workFor subcontractor
- ...We are seeking a highly skilled and motivated Senior Software Engineer to architect, build, and operate the workflow orchestration platforms... ...training to data pipelines and CI/CD, our teams depend on reliable, scalable workflow systems to move fast. In this role, you will...Full timeTemporary work
$100k - $120k
...industry, join our team as we help shape a brighter way forward. JLL - Reliability Engineer (P3)Location: Spring, TX 77389Work Schedule: Hybrid (3 days onsite), M-F 8-4Travel requirements: 10%(other client sites in OR & WA)Reports to: Sr. Reliability EngineerEstimated base...Full timeWork at officeFlexible hours- ...Senior Reliability EngineerIndependently lead end-to-end Reliability Centered Maintenance (RCM) studies on rotating, static, electrical... ...client reliability/maintenance managers, operations, and engineering disciplines to validate assumptions, findings and recommendations...
- ...Senior Reliability EngineerAs a Senior Reliability Engineer, you are at the vanguard of the energy transition. You will architect the future of Power, Wind, and Electrification by driving cutting-edge systems integration, modeling, and simulation. From shaping data center...
- ...Venus Aerospace is revolutionizing rocket engine propulsion. With the first generational... ...building high‑quality software and ensuring its reliability through robust automated testing. We are... ...) validation systems. Location: On-site in Houston, TX Benefits: Venus...Permanent employmentFull timeFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
- site reliability engineer sre Houston, TX
- site reliability engineer Houston, TX
- site recruiter Houston, TX
- site services specialist Houston, TX
- junior website developer Houston, TX
- remote website tester Houston, TX
- official site Houston, TX
- on site coordinator Houston, TX
- site leader Houston, TX
- historic site Houston, TX




