Site Reliability Engineer
NOV
Site Reliability Engineer
As a Site Reliability Engineer, you will be responsible for:
- Operational Excellence & Incident Management
- Maintain and monitor production systems for availability, latency, and performance.
- Lead incident response efforts, including communication, resolution, and postmortem documentation.
- Design and implement health checks, alerting systems, and automated remediation workflows.
- Drive root cause analysis and implement permanent resolutions for recurring issues.
- Observability & Insights
- Set up and maintain full observability stacks (logging, metrics, tracing) using tools like Prometheus, Grafana, Datadog, OpenTelemetry, or ELK.
- Analyze telemetry and logs to identify trends, anomalies, and opportunities for improvement.
- Conduct post-incident reviews and use insights to inform future engineering investments.
- Performance & Systems Optimization
- Tune and optimize distributed systems, including AKKA.NET actors, for performance and resource efficiency.
- Work with developers to evolve architecture and improve system throughput, latency, and stability.
- Optimize PostgreSQL performance, queries, and maintenance strategies.
- CI/CD & Automation
- Design and maintain modern CI/CD pipelines using GitHub Actions, Azure Pipelines, or GitLab CI.
- Automate deployment, testing, and rollback processes to reduce friction and increase deployment frequency.
- Standardize infrastructure as code practices across environments.
We'd love to talk to you if you have:
- 5+ years of experience in SRE, DevOps, or Infrastructure Engineering roles.
- Expertise in Kubernetes and container orchestration at scale.
- Strong experience with AKKA.NET or similar actor-based frameworks.
- Proficiency with scripting and automation (Bash, PowerShell, Python).
- Experience with observability tools (Phobos, Datadog, Prometheus, Grafana, OpenTelemetry, ELK).
- Hands-on experience with cloud platforms (AWS, Azure, or GCP).
- Strong PostgreSQL knowledge—performance tuning, query optimization, maintenance.
- Proven ability to lead incident management and drive postmortem processes.
- A builder's mindset with high standards for operational excellence and technical ownership.
Preferred Tools & Ecosystem Experience:
- CI/CD: GitHub Actions, Azure Pipelines, GitLab CI
- Infrastructure: Kubernetes, Docker, Terraform
- Monitoring: Phobos (AKKA.NET), Datadog, Prometheus
- Source Control: GitHub, GitLab, Azure DevOps
- Programming: C#, Python, Bash, PowerShell
About Us
Every day, the oil and gas industry's best minds put more than 150 years of experience to work to help our customers achieve lasting success. We Power the Industry that Powers the World Throughout every region in the world and across every area of drilling and production, our family of companies has provided the technical expertise, advanced equipment, and operational support necessary for success—now and in the future. Global Family We are a global family of thousands of individuals, working as one team to create a lasting impact for ourselves, our customers, and the communities where we live and work. Purposeful Innovation Through purposeful business innovation, product creation, and service delivery, we are driven to power the industry that powers the world better. Service Above All This drives us to anticipate our customers' needs and work with them to deliver the finest products and services on time and on budget.
About the Team
Corporate Our family of companies is supported by our global Corporate teams, providing expert knowledge from functions including Human Resources, Information Technology, Compliance, Finance, QHSE, Marketing and Legal centers of expertise. We are structured to provide guidance and service above all to all our business operations.
Job Info
- Job Identification 41952
- Job Category Technical
- Job Schedule Full time
- Job Shift Day
- Locations 10353 Richmond Avenue, Houston, TX, 77042, US (Hybrid)
- ...and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Corporate Technology, Risk Technology team, you will solve complex and broad business problems...Suggested
- Reliability Engineering Design, implement, and operate scalable, resilient, and highly available systems on Google Cloud Platform. Improve service... ...Skills, and Abilities Three or more years of experience in Site Reliability Engineering, platform engineering, DevOps, cloud...SuggestedRemote work
$74.1k - $148.3k
...Site Reliability Engineer Solve complex problems related to infrastructure cloud services and build automation to prevent problem recurrence. Design, write, and deploy software to improve the availability, scalability, and efficiency of Oracle products and services....SuggestedTemporary workImmediate startFlexible hours- ...Site Reliability Engineer II About PROS: PROS, Inc. is the leading offer management provider to the airline industry, helping airlines deliver seamless retail experiences designed to maximize revenue and margin growth. Powered by AI, the PROS Platform enables...SuggestedFlexible hours
- ...Senior Site Reliability Engineer The Senior Site Reliability Engineer is responsible for improving the reliability, availability, scalability, and operational excellence of our critical infrastructure platforms and services. This role partners closely with Engineering...SuggestedWork at officeLocal area
- ...Site Reliability Engineer III There's nothing more exciting than being at the center of a rapidly growing field in technology and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems. As a Site Reliability...
$104.9k - $174.7k
...58 Are you passionate about improving reliability, scalability, and resilience in complex... ...practices. Own prioritization of reliability engineering tasks within team backlogs. Lead... ...(IaaS). Background in DevOps, site reliability engineering practices, or related...Full timeLocal area- As a Lead Site Reliability Engineer at JPMorgan Chase within the Corporate Know Your Customer (KYC) Technology group, you hold a leadership role in your team, demonstrate strong knowledge across multiple technical domains, and advise others on the technical and business...
- As an Entry-Level DevOps Site Reliability Engineer, you will join a team responsible for continuous improvement and support of customer facing products. Responsibilities will include collecting system requirements; improving existing tools and processes through scripting...Work from home2 days per week
- ...ENGINEERLocation: HOUSTON, TXFLSA Class: EXEMPTResponsible to: Directo of Software EngineeringPosition Summary: DevOps / Site Reliability Engineer to implement and evolve the infrastructure, deployment pipelines, and reliability posture of our systems. You'll work closely...Full timeLocal area
- ...Please extend your support for this role. Local candidate will get 1st preference. Job Title: SRE Engineer Location: Houston, TX and Jersey City, NJ - 3 Days Onsite Role FTE role with Mphasis Client: Mphasis H1B transfer will work...Work experience placementH1bLocal area
- ...JOB DESCRIPTION As a Site Reliability Engineer, you will be responsible for: Operational Excellence & Incident Management - Maintain and monitor production systems for availability, latency, and performance. - Lead incident response efforts, including communication...Permanent employmentFull time
- ...To Supervisor Analytics Cloud Services.Role OverviewThe Release Engineer is responsible for the deployment release and maintenance of... ...systems architects infrastructure and security teams to deliver reliable and scalable cloud solutions.Key ResponsibilitiesCI/CD Pipeline...
- ...accelerate autonomy development. We are seeking a software engineer with strong C++ expertise and a passion for building scalable simulation... ...in architecture and technical design discussions Build reliable, maintainable, and well-tested systems Contribute to code...Full time
$113k - $141.53k
...leader in global energy. Senior Solutions Engineer - Systems Integration serves as a... ...functionally in the field, ensuring safe, reliable, and performant operation across diverse... ...Willingness to travel to factories and project sites (25%).Preferred QualificationsMaster’s degree...Full timeFor contractorsLocal areaWorldwideFlexible hours- ...together to shape what’s next. Whether you're engineering advanced materials, transforming... ...imagine everything.Role Impact Remote / Multi-Site - North America (Travel Required) Alternate Titles: Network Reliability Engineer | Plant Reliability Engineer | Maintenance...Contract workLocal areaRemote workWorldwideShift work
- Reliability EngineerHouston, TXThe actual location of this job is in Houston, TX, US. Relocation... ...families, if needed.This is a fully site‑based role. Working together in person supports... ...environmentOpportunities to grow your engineering career in a global...Full timeRelocation package
$76k - $155.7k
...RegularPercentage of Travel Required: Up to 10%Type of Travel: Continental US* * *The Opportunity:CACI is seeking Software Systems Engineers to support the Artemis Next Generation Space Suit program at NASA Johnson Space Center. This position contributes to systems...Permanent employmentContract workFor contractorsWork experience placementImmediate startFlexible hours- ...Release Engineer Visa status: U.S. Citizens and those authorized to work in the U.S. are encouraged to apply. Tax Terms: W2, 1099 Corp-Corp or 3rd Parties: Yes Technical skills and knowledge: Must be proficient with Source Control systems, like Git, to create...
- ...Release Train Engineer 4 Months- Contract To Hire Pay- $65-$70 W2 Onsite Houston, TX Job Description The Release Train... ...sure all team activity, dashboards, and metrics are visible and reliable. Coach teams and Scrum Masters on agile and Scrum practices,...Contract workWork at office
- ...infrastructure challenges. Job DetailsViridien is seeking a Platform Engineer - Infrastructure & Cloud Systems to design, build, and improve... ...observability tooling. This role focuses on building scalable, reliable systems and ensuring strong integration between infrastructure...Full timeRelocationFlexible hours
- ...data platforms that support high-volume, data-intensive workflows.The team works across backend engineering, infrastructure, and data systems, collaborating to deliver reliable, high-performance services in a modern cloud-native environment.Key Responsibilities-Backend...Full timeFlexible hours
- ...a company that values diversity, integrity, and growth. Role Overview PDI Technologies is looking for a Manager, Site Reliability Engineering to lead the SRE organization supporting Paylo, PDI’s payments, loyalty, and fuel-pricing product suite. This role owns the...
- Position Title: Senior Software Engineer - Platform Location: Houston, TX onsiteFLSA Class: ExemptReports To: Manager of Software EngineeringPosition Summary:VoltaGrid is seeking a Senior Technical Solutions Engineer to join our Platform Team, responsible for designing...Full timeLocal area
- ...experience. Build software that matters alongside experienced engineers across software, firmware, and hardware domains. Grow into broader... ...scenarios to uncover edge cases, performance limits, and reliability opportunitiesDebug and resolve challenging issues that span multiple...Full timeWork experience placementWork at officeLocal areaImmediate start2 days per week
- ...We are seeking a highly skilled and motivated Senior Software Engineer, Workflow Platforms to architect, build, and operate the workflow... ...training to data pipelines and CI/CD, our teams depend on reliable, scalable workflow systems to move fast. In this role, you will...Full timeTemporary work
$2,400 per month
...Description We’re seeking an experienced Software Engineer to maintain, enhance, and modernize a suite of .NET-based applications while developing new cross-platform, mobile, and distributed systems. This role bridges legacy modernization with next-generation engineering...Full time- ...Software System Engineer Here at The Exploration Company, we are developing, producing, and operating Nyx, a modular and reusable space orbital vehicle that can eventually be refuelled in orbit and that can carry cargo - and potentially humans in the longer run....Immediate startRelocationVisa sponsorshipRelocation package
$200k
...Senior Mechanical Reliability Engineer Our client is seeking a full-time Senior Mechanical Reliability Engineer. Hybrid. This individual may... ...leveraging opportunities for operational improvement throughout the site and organization. Provide technical support and expertise...Full time$107.48k - $143.31k
...employees. What You'll Be Doing Lead and apply regional reliability engineering strategies to improve equipment performance, uptime, and... ...management skills with the ability to support multiple sites remotely. Willingness to travel to the plant and corporate...Temporary workRemote workFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
- site reliability engineer Houston, TX
- site reliability engineer sre Houston, TX
- junior website developer Houston, TX
- website content developer Houston, TX
- on site coordinator Houston, TX
- after school site coordinator Houston, TX
- website coordinator Houston, TX
- site leader Houston, TX
- site recruiter Houston, TX
- historic site Houston, TX


