Observability Engineer / Site Reliability Engineer
Ontrac Solutions
About Ontrac Solutions
Ontrac Solutions is a leading technology consulting firm, specializing in cutting-edge solutions that drive business transformation. We partner with organizations to modernize their infrastructure, streamline processes, and deliver tangible results. By creating value beyond the hype, we help businesses modernize technology and build new strategies that fuel growth. Our team is committed to innovation, collaboration, and excellence, empowering our clients to succeed in an evolving digital landscape.
Role Overview
We are seeking an experienced Observability / Site Reliability Engineer (SRE) to design, scale, and maintain our enterprise monitoring and alerting ecosystems. In this role, you will bridge the gap between development and operations by ensuring high availability, performance tuning, and deep visibility across distributed multi-cloud and native systems. You will play a critical role in automating infrastructure and building robust observability pipelines using industry-leading cloud-native tools.
Key Responsibilities
- GCP & Cloud Management: Architect, optimize, and maintain observability frameworks across cloud environments, with a specific focus on implementing Google Cloud Platform (GCP) observability tools (Cloud Logging, Cloud Monitoring, Trace, and Profiler).
- Platform Management: Design, deploy, and maintain robust observability stacks across hybrid ecosystems, utilizing Prometheus, Grafana, and cloud-native integrations.
- Automation & IaC: Drive infrastructure-as-code (IaC) initiatives using Terraform and Ansible to ensure consistent, automated deployments of infrastructure and observability tooling.
- CI/CD Integration: Build, maintain, and optimize deployment workflows within Kubernetes and Google Kubernetes Engine (GKE) / OpenShift environments using GitHub, Harness, and other CI/CD pipelines.
- System Performance: Deeply analyze Linux/Unix system administration architectures, optimizing compute resource metrics and performance tuning across complex, distributed environments.
- SRE Evangelism: Implement SRE best practices, establishing meaningful Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets to ensure platform reliability.
Required Skills & Qualifications
- Cloud Infrastructure: Proven engineering experience within Google Cloud Platform (GCP) environments, particularly managing cloud-native monitoring and compute resources.
- Observability Tooling: Hands-on experience with Grafana, Prometheus, and Google Cloud Observability suites. Direct experience with GEM (Grafana Enterprise Metrics) is highly desirable.
- OS & Scripting: Expert-level knowledge of Linux/Unix operating systems paired with strong shell scripting skills for automation and systems management.
- Programming: Professional coding proficiency in at least one modern language (Python, Go, Java, Perl, or advanced Shell).
- Containers & Orchestration: Hands-on experience managing containerized applications on Kubernetes, GKE, and/or Red Hat OpenShift.
__________________________________
(Ontrac Solutions has partnered with PinpointVerify to help genuine applicants rise above the noise. Today, qualified candidates are too often overshadowed by fake and fraudulent applications. PinpointVerify gives our recruiters confidence that you are exactly who you say you are — and gives you a portable verification credential you can share with any employer.
Applicants who complete verification are prioritized over non-verified candidates with comparable experience. And if you're hired, Ontrac reimburses the full cost of your verification.
Get verified →
$159k - $272k
...opportunity to grow and make a difference in ways that matter to you. Role SummaryIn this role as Principal Site Reliability Engineer, Infrastructure Observability you will help formulate, develop, and implement a team of Site Reliability Engineers (SREs) focused on the...SuggestedFull timePrivate practiceLocal areaRemote workWork from home3 days per week- ...Information Technology group delivers secure, reliable technology solutions that enable... ...enterprise platforms.As a Principal Site Reliability Engineer (SRE), you will drive operational... ...reliability initiatives, champion observability and automation, drive major incident...SuggestedRemote workFlexible hours
- ...Job Description Job Description SRE Support Engineer - Observability While this position is not currently open, we are interviewing strong... ...support across Slack and tickets, improving monitoring reliability, and reducing incident impact through better triage, troubleshooting...SuggestedRemote work
$96k - $163k
...realize their greatest potential. Title and Summary Senior Site Reliability Engineer Who is Mastercard? At Mastercard technology, we... ...consistency and reliability in applying the skills. • Observability - Ability to use scripting and tooling to implement observability...SuggestedFull timePart timeWorldwideFlexible hours$76k - $127k
...governments realize their greatest potential. Title and Summary Site Reliability Engineer II Site Reliability Engineer II Who is Mastercard?... ...consistency and reliability in applying the skills. Observability - Ability to use scripting and tooling to implement...SuggestedFull timePart timeWorldwideFlexible hours$96k - $163k
...their greatest potential. Title and Summary Senior Site Reliability Engineer Overview The BizOps team is looking for a Senior... ...days, pro-rated based on date of hire; 10 annual paid U.S. observed holidays; 401k with a best-in-class company match; deferred...Full timePart timeWorldwideFlexible hoursShift work$76k - $127k
...realize their greatest potential. Title and Summary Site Reliability Engineer II The BizOps team at Mastercard is looking for a Site... ..., pro-rated based on date of hire; 10 annual paid U.S. observed holidays; 401k with a best-in-class company match; deferred...Full timePart timeWorldwideFlexible hoursEarly shift$158.5k - $172k
...deserve.About The OpportunityAs a Senior Engineer on the Runtime Automation team, you... ...system engineering, while ensuring top-tier observability and strict security across our mission-... ...-impact position driving continuous reliability, deep system optimization, and automation...Full timeWork at office3 days per week$104.9k - $174.7k
...link below, About the Role:We are hiring a hands-on Senior Site Reliability Engineer (SRE) to actively build, operate, and improve the... ...in designing infrastructure, writing Terraform, improving observability, and responding to real production incidents.If you live...Full timeWork at officeLocal areaRemote workWork from home- ...platforms to advanced release engineering practices, our teams are... ...behavior preferred Exposure to reliability engineering concepts such as... ...1#GMFjobsAbout The Role:The Site Reliability Engineer under... ...of SRE concepts, including observability, monitoring, incident response...Work experience placementH1bWork at officeRemote workVisa sponsorshipFlexible hoursShift work2 days per week
$130k - $180k
...belonging, collaboration, and accomplishment.Being a Senior Site Reliability Engineer at iManage Means… You are an engineer, a builder, and a... ...in on-call rotations. You’ll be a key voice in observability, change management, and service scalability, providing guidance...Work at officeLocal areaRemote workWorldwideMonday to FridayFlexible hours- ...home day is currently Tuesday.Engineering at Lambda is responsible for... ...teams to improve service reliability and deployment workflowsDeploy... ...network monitoring, observability, and management toolsImprove... ...rotationYouHave 5+ years of experience in Site Reliability Engineering,...Work at officeLocal areaWork from homeFlexible hours
$15k
...office, daily catered lunches, and more.As a Senior Cluster Site Reliability Engineer (SRE), you will help scale our research compute cluster... ...with engineering teamsDevelop robust metrics and observability for cluster health and use those metrics to inform your work...Work at officeLocal areaRemote work- Site Reliability Engineers are responsible for ensuring the availability, reliability, scalability, and performance of the firm’s most critical... ...operations, with a strong emphasis on building systems that are observable, resilient, and operable by default.This is an on-site...Local areaRemote workFlexible hoursShift work
- ...thinking organization, apply now.We are currently seeking a Site Reliability Engineer to join our team in Westlake, Texas (US-TX), United... ...Implement Site Reliability Engineering best practices including observability, incident management, capacity planning, and resiliency...Full timeTemporary workWork at officeRemote workFlexible hours
$134.25k - $214.8k
...that holds up in court. That's us.Axon's Platform team is the engine behind what hundreds of thousands of officers rely on every day... ...across Azure, AWS, or GCPHands on experience with observability platforms such as Grafana, Datadog, or New RelicFamiliarity with...Work experience placementWork at officeRemote work$140k - $150k
...OPTION: Remote_________________The NBA is hiring a Senior Site Reliability Engineer (SRE) - Messaging & Collaboration to ensure the availability... ...reason or sincerely held religious belief, practice, or observance.About the NBAThe National Basketball Association (NBA) is...Full timeTemporary workLocal areaRemote workWeekend work$67.2k - $100.8k
...functional partners (IT, Security, DevOps, Engineering) to improve operational health and apply SRE best practicesSupport the reliability, availability, scalability, and... ...configuration, patching, and releasesContribute to observability and monitoring (Dynatrace, Prometheus),...Temporary workH1bWork at officeRemote work$90k - $180k
...160 countries.JOB DESCRIPTION:About the RoleThis Senior Site Reliability Engineer position works on-site out of our Sylmar, CA or Sunnyvale... ...strategies, resource management, and cluster operations.Observability platform experience with tools such as Prometheus, Grafana...Remote workShift work$165k - $190k
...DevOps / SRE TeamThe DevOps/SRE team at Obsidian ensures that engineering excellence translates into stable, scalable, and high-... ...platformAddress complex challenges around scalability, reliability, observability, and cost efficiencyCollaborate with Engineering teams to...Work from home- Recognized as the No. 1 site trusted by real estate professionals, Realtor.com has... ...guidance.We are seeking a Senior Site Reliability Engineer to join our newly formed Operations... ...role will contribute to the reliability, observability, and operational excellence of our...Work at officeLocal area
$230k - $250k
GovCIO is hiring a Site Reliability Engineer with an active Secret clearance to ensure reliability, scalability, performance, and availability... ...manual processes.Develop monitoring, alerting, and observability solutions.Improve system performance, capacity, and resilience...Remote work$110k - $145k
...across the U.S., Canada, and India. We are seeking a Senior Site Reliability Engineer to own the reliability, scalability, performance, and... ...of service and system design to improve fault tolerance, observability and operational sustainability. Debug complex production...Contract workWork at officeWork from homeFlexible hours$210k - $230k
GovCIO is currently hiring for a Senior Site Reliability Engineer (SRE) to design, implement, and maintain highly available, scalable, and resilient... ...service mesh technologies (Istio, Linkerd)• Knowledge of observability platforms (Datadog, New Relic, Dynatrace)• Experience...Currently hiringRemote work$134.25k - $214.8k
...matters at a company where you matter.Your ImpactAs a Senior Site Reliability Engineer within the APX SRE organization, you’ll focus on... ...infrastructure, software builds, tests, and releases.Experience using observability tools such as APM, logging, and metrics to assist with...Work experience placementWork at officeRemote workFlexible hours$117k - $209.33k
...Position OverviewWant to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable,... ...as SLOs/SLIs, production readiness, incident management, observability, resilience testing, and toil reduction. Success in this...Full timeFor contractorsRemote work$152.13k - $162.13k
...about what’s next. Join us.General Summary:Unum Group seeks Site Reliability Engineers in Atlanta, GA.Applicants who are interested in this... ...Ref #66753) for consideration.Design, build, and maintain observability, monitoring, and alerting capabilities across consumer and...Full timeTemporary workWork at officeRemote work$104.43k - $156.65k
...Comcast prefers to have employees on-site collaborating unless the team has been... ...Paramount+, and many others.Our Site Reliability Engineering (SRE) team is at the heart of our... ...innovation with a focus on improving observability and reducing toil.Job Description*This...Permanent employmentFull timeWork at officeRemote workWorldwideFlexible hours$105.6k - $145.2k
Architect the Future as our Site Reliability Engineer!Are you ready to take your skills to the next level as a self-motivated and enthusiastic... ..., scalable cloud environments.Implement and enhance observability solutions using tools like New Relic, DataDog, Sumologic...Ongoing contractFull timeWork at officeLocal areaWorldwide$175k - $250k
...developed by our expert team of lawyers, engineers and research scientists. We’ve found... ...As a Software Engineer on the Site Reliability team at Harvey, you will ensure the reliability... ..., etc.). ~ Deep familiarity with observability tools (Datadog, Sentry, etc.) and incident...Full timeRelocation package
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Observability Engineer / Site Reliability Engineer. Be the first to apply!
- site reliability engineer remote Remote
- site reliability engineer sre Remote
- site reliability engineer Remote
- on-site clinical research associate (traveling/remote) Remote
- website coordinator Remote
- junior website developer Remote
- site leader Remote
- historic site Remote
- website content developer Remote
- construction site safety Remote



