Site Reliability Engineer
National Oilwell Varco
As a Site Reliability Engineer, you will be responsible for: Operational Excellence & Incident Management- Maintain and monitor production systems for availability, latency, and performance.- Lead incident response efforts, including communication, resolution, and postmortem documentation.- Design and implement health checks, alerting systems, and automated remediation workflows.- Drive root cause analysis and implement permanent resolutions for recurring issues.Observability & Insights- Set up and maintain full observability stacks (logging, metrics, tracing) using tools like Prometheus, Grafana, Datadog, OpenTelemetry, or ELK.- Analyze telemetry and logs to identify trends, anomalies, and opportunities for improvement.- Conduct post-incident reviews and use insights to inform future engineering investments.Performance & Systems Optimization- Tune and optimize distributed systems, including AKKA.NET actors, for performance and resource efficiency.- Work with developers to evolve architecture and improve system throughput, latency, and stability.- Optimize PostgreSQL performance, queries, and maintenance strategies.CI/CD & Automation- Design and maintain modern CI/CD pipelines using GitHub Actions, Azure Pipelines, or GitLab CI.- Automate deployment, testing, and rollback processes to reduce friction and increase deployment frequency.- Standardize infrastructure as code practices across environments.We’d love to talk to you if you have:- 5+ years of experience in SRE, DevOps, or Infrastructure Engineering roles.- Expertise in Kubernetes and container orchestration at scale.- Strong experience with AKKA.NET or similar actor-based frameworks.- Proficiency with scripting and automation (Bash, PowerShell, Python).- Experience with observability tools (Phobos,Datadog, Prometheus, Grafana, OpenTelemetry, ELK).- Hands-on experience with cloud platforms (AWS, Azure, or GCP).- Strong PostgreSQL knowledge—performance tuning, query optimization, maintenance.- Proven ability to lead incident management and drive postmortem processes.- A builder’s mindset with high standards for operational excellence and technical ownership.Preferred Tools & Ecosystem Experience- CI/CD: GitHub Actions, Azure Pipelines, GitLab CI- Infrastructure: Kubernetes, Docker, Terraform- Monitoring: Phobos (AKKA.NET), Datadog, Prometheus- Source Control: GitHub, GitLab, Azure DevOps- Programming: C#, Python, Bash, PowerShellEvery day, the oil and gas industry’s best minds put more than 150 years of experience to work to help our customers achieve lasting success.We Power the Industry that Powers the WorldThroughout every region in the world and across every area of drilling and production, our family of companies has provided the technical expertise, advanced equipment, and operational support necessary for success—now and in the future.Global FamilyWe are a global family of thousands of individuals, working as one team to create a lasting impact for ourselves, our customers, and the communities where we live and work. Purposeful InnovationThrough purposeful business innovation, product creation, and service delivery, we are driven to power the industry that powers the world better.Service Above AllThis drives us to anticipate our customers’ needs and work with them to deliver the finest products and services on time and on budget.CorporateOur family of companies is supported by our global Corporate teams, providing expert knowledge from functions including Human Resources, Information Technology, Compliance, Finance, QHSE, Marketing and Legal centers of expertise. We are structured to provide guidance and service above all to all our business operations.Full timePosting Date: 2026-06-23
- ...and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Corporate Technology, Risk Technology team, you will solve complex and broad business problems...Suggested
- The NexTier Technology team is looking for a Site Reliability Engineer (SRE) to help build, scale, and maintain highly reliable systems on Google Cloud Platform (GCP). This role blends software engineering with infrastructure expertise to ensure our services are performant...Suggested
- ...Hynes & Khater is seeking a Cloud Support Engineer Lead to own the reliability, observability, and operational health of Azure-based applications. This is a hands-on leadership role; you will diagnose hard problems and define processes that keep systems running reliably...Suggested
- ...and AI agent a cryptographically secured identity, improving engineering velocity while maintaining security. We make trusted computing... ...problems that allow our customers to trust us for secure and reliable access to their infrastructure. Excellent security is table stakes...SuggestedWork at officeLocal areaRemote workSleeping nights
$155k - $222.6k
...global cloud platform. As a team of six engineers distributed across the US, Canada, and the... ...with a strong focus on automation, reliability, and operational excellence. We are one... ...Qualifications ~2+ years of experience in Site Reliability Engineering, DevOps, Infrastructure...SuggestedPermanent employmentFull timeTemporary workLocal areaWorldwideFlexible hours- ...Site Reliability Engineer III There's nothing more exciting than being at the center of a rapidly growing field in technology and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems. As a Site Reliability...
- As a Lead Site Reliability Engineer at JPMorgan Chase within Market Risk Technology, you hold a leadership role in your team, demonstrate strong knowledge across multiple technical domains, and advise others on the technical and business issues facing them. Take lead and...
- ...ENGINEERLocation: HOUSTON, TXFLSA Class: EXEMPTResponsible to: Directo of Software EngineeringPosition Summary: DevOps / Site Reliability Engineer to implement and evolve the infrastructure, deployment pipelines, and reliability posture of our systems. You'll work closely...Full timeLocal area
- As an Entry-Level DevOps Site Reliability Engineer, you will join a team responsible for continuous improvement and support of customer facing products. Responsibilities will include collecting system requirements; improving existing tools and processes through scripting...Work from home2 days per week
$113.3k
...important to maintain our strong culture, achieve our goals, and thrive as #OneJamf. What you'll do at Jamf: As a Senior Site Reliability Engineer, you'll help us balance development velocity with the reliability our customers depend on. You'll partner with engineering...Work at officeRemote workWorldwideFlexible hours- ...JOB DESCRIPTION As a Site Reliability Engineer, you will be responsible for: Operational Excellence & Incident Management - Maintain and monitor production systems for availability, latency, and performance. - Lead incident response efforts, including communication...Permanent employmentFull time
- ...Please extend your support for this role. Local candidate will get 1st preference. Job Title: SRE Engineer Location: Houston, TX and Jersey City, NJ - 3 Days Onsite Role FTE role with Mphasis Client: Mphasis H1B transfer will work...Work experience placementH1bLocal area
- ...To Supervisor Analytics Cloud Services.Role OverviewThe Release Engineer is responsible for the deployment release and maintenance of... ...systems architects infrastructure and security teams to deliver reliable and scalable cloud solutions.Key ResponsibilitiesCI/CD Pipeline...
$113k - $141.53k
...leader in global energy. Senior Solutions Engineer - Systems Integration serves as a... ...functionally in the field, ensuring safe, reliable, and performant operation across diverse... ...Willingness to travel to factories and project sites (25%).Preferred QualificationsMaster’s degree...Full timeFor contractorsLocal areaWorldwideFlexible hours$76k - $155.7k
...RegularPercentage of Travel Required: Up to 10%Type of Travel: Continental US* * *The Opportunity:CACI is seeking Software Systems Engineers to support the Artemis Next Generation Space Suit program at NASA Johnson Space Center. This position contributes to systems...Permanent employmentContract workFor contractorsWork experience placementImmediate startFlexible hours- Reliability EngineerHouston, TXThe actual location of this job is in Houston, TX, US. Relocation... ...families, if needed.This is a fully site‑based role. Working together in person supports... ...environmentOpportunities to grow your engineering career in a global...Full timeRelocation package
- ...data platforms that support high-volume, data-intensive workflows.The team works across backend engineering, infrastructure, and data systems, collaborating to deliver reliable, high-performance services in a modern cloud-native environment.Key Responsibilities-Backend...Full timeFlexible hours
- ...infrastructure challenges. Job DetailsViridien is seeking a Platform Engineer - Infrastructure & Cloud Systems to design, build, and improve... ...observability tooling. This role focuses on building scalable, reliable systems and ensuring strong integration between infrastructure...Full timeRelocationFlexible hours
- ...accelerate autonomy development. We are seeking a software engineer with strong C++ expertise and a passion for building scalable simulation... ...in architecture and technical design discussions Build reliable, maintainable, and well-tested systems Contribute to code...Full time
- Position Title: Senior Software Engineer - Platform Location: Houston, TX onsiteFLSA Class: ExemptReports To: Manager of Software EngineeringPosition Summary:VoltaGrid is seeking a Senior Technical Solutions Engineer to join our Platform Team, responsible for designing...Full timeLocal area
- ...experience. Build software that matters alongside experienced engineers across software, firmware, and hardware domains. Grow into broader... ...scenarios to uncover edge cases, performance limits, and reliability opportunitiesDebug and resolve challenging issues that span multiple...Full timeWork experience placementWork at officeLocal areaImmediate start2 days per week
- ...A leading engineering firm is seeking an Electrical Reliability Manager to oversee the electrical reliability program in their North American operations. The... ...driving improvements and optimizing the reliability of electrical systems across multiple sites. #J-18808-Ljbffr...
$107.48k - $143.31k
...employees. What You'll Be Doing Lead and apply regional reliability engineering strategies to improve equipment performance, uptime, and... ...management skills with the ability to support multiple sites remotely. Willingness to travel to the plant and corporate...Temporary workRemote workFlexible hours- ...diversity is prized. Come find out how our people power modern life. About this Role We are seeking a Senior ITSM/ITAM Platform Engineer to lead the initial development and the ongoing administration of our enterprise IT Service Management and IT Asset Management...Full timeWork at officeLocal areaWorldwide
- Hybrid Remote • Houston, TXDescription The Platform Engineer is responsible for building, operating, and improving the shared infrastructure... ...teams use to deliver software. This role focuses on delivering reliable, secure, and well-documented platform services. This individual...Remote work
- ...innovative ways to help people? Do you like having the autonomy to build new solutions from the ground up? If so, being a Software Engineer III at Frost could be the job for you.At Frost, it’s about more than a job. It’s about having a flourishing career where you can...Full time
- ...cloud-native services, and infrastructure automation to improve reliability and speed to delivery. Our global team operates in a fast‑... ...operational excellence.About the Role:As a Senior Network Platform Engineer, you will be a key technical contributor responsible for...Full timeWork at officeFlexible hours
- ...we aim to ensure safety, security, and environmental stability for all.Job Summary:We are seeking a highly skilled Senior Systems Engineer to design, implement, and support enterprise infrastructure across on-premises and cloud environments.This role is hands-on and execution...Full time
- ...future. As a Senior Pre-Sales AI Solutions Engineer for Oil & Gas, you'll be the technical... ...engineering leaders, plant and reliability managers, HSE leadership, and OT/process... ...energy-sector events), and field/asset sites, which may include travel to operational...Full timeLocal areaRemote workWorldwideFlexible hoursShift work
- ...enterprise AD forest/domain and domain controllers: Group Policy, replication, sites and subnets, and account lifecycle in accordance with security, privacy, and regulatory best practices.Provide engineering support for Azure Virtual Desktop identity integration, including app...Full timeWork at officeLocal areaWorldwide
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
- site reliability engineer Houston, TX
- site services specialist Houston, TX
- construction site safety Houston, TX
- site leader Houston, TX
- official site Houston, TX
- website content developer Houston, TX
- remote website tester Houston, TX
- on site coordinator Houston, TX
- IT site lead Houston, TX
- site safety Houston, TX


