Observability Engineer / Site Reliability Engineer
Ontrac Solutions
About Ontrac Solutions
Ontrac Solutions is a leading technology consulting firm, specializing in cutting-edge solutions that drive business transformation. We partner with organizations to modernize their infrastructure, streamline processes, and deliver tangible results. By creating value beyond the hype, we help businesses modernize technology and build new strategies that fuel growth. Our team is committed to innovation, collaboration, and excellence, empowering our clients to succeed in an evolving digital landscape.
Role Overview
We are seeking an experienced Observability / Site Reliability Engineer (SRE) to design, scale, and maintain our enterprise monitoring and alerting ecosystems. In this role, you will bridge the gap between development and operations by ensuring high availability, performance tuning, and deep visibility across distributed multi-cloud and native systems. You will play a critical role in automating infrastructure and building robust observability pipelines using industry-leading cloud-native tools.
Key Responsibilities
- GCP & Cloud Management: Architect, optimize, and maintain observability frameworks across cloud environments, with a specific focus on implementing Google Cloud Platform (GCP) observability tools (Cloud Logging, Cloud Monitoring, Trace, and Profiler).
- Platform Management: Design, deploy, and maintain robust observability stacks across hybrid ecosystems, utilizing Prometheus, Grafana, and cloud-native integrations.
- Automation & IaC: Drive infrastructure-as-code (IaC) initiatives using Terraform and Ansible to ensure consistent, automated deployments of infrastructure and observability tooling.
- CI/CD Integration: Build, maintain, and optimize deployment workflows within Kubernetes and Google Kubernetes Engine (GKE) / OpenShift environments using GitHub, Harness, and other CI/CD pipelines.
- System Performance: Deeply analyze Linux/Unix system administration architectures, optimizing compute resource metrics and performance tuning across complex, distributed environments.
- SRE Evangelism: Implement SRE best practices, establishing meaningful Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets to ensure platform reliability.
Required Skills & Qualifications
- Cloud Infrastructure: Proven engineering experience within Google Cloud Platform (GCP) environments, particularly managing cloud-native monitoring and compute resources.
- Observability Tooling: Hands-on experience with Grafana, Prometheus, and Google Cloud Observability suites. Direct experience with GEM (Grafana Enterprise Metrics) is highly desirable.
- OS & Scripting: Expert-level knowledge of Linux/Unix operating systems paired with strong shell scripting skills for automation and systems management.
- Programming: Professional coding proficiency in at least one modern language (Python, Go, Java, Perl, or advanced Shell).
- Containers & Orchestration: Hands-on experience managing containerized applications on Kubernetes, GKE, and/or Red Hat OpenShift.
__________________________________
(Ontrac Solutions has partnered with PinpointVerify to help genuine applicants rise above the noise. Today, qualified candidates are too often overshadowed by fake and fraudulent applications. PinpointVerify gives our recruiters confidence that you are exactly who you say you are — and gives you a portable verification credential you can share with any employer.
Applicants who complete verification are prioritized over non-verified candidates with comparable experience. And if you're hired, Ontrac reimburses the full cost of your verification.
Get verified →
$158.5k - $172k
...deserve.About The OpportunityAs a Senior Engineer on the Runtime Automation team, you... ...system engineering, while ensuring top-tier observability and strict security across our mission-... ...-impact position driving continuous reliability, deep system optimization, and automation...SuggestedFull timeTemporary workWork at officeFlexible hours3 days per week$145k - $175k
...designed to help you gain your full potential. Job OverviewThe Site Reliability Engineer supports deployments, cloud infrastructure, and... ...operations across our Kubernetes clusters, AWS environments, and observability platforms.We are hiring an experienced Site Reliability...SuggestedFull timeWork at officeLocal areaFlexible hours3 days per week$100.7k - $167.8k
Job SummaryThe Site Reliability Engineer III is a pivotal architect of stability for CME Clearing & Risk. You will engineer secure, scalable... ...defending SLIs, SLOs, and SLAs to maintain system health.Observability Mindset: Experience navigating the telemetry landscape using...SuggestedFull timeWorldwide- ...the world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Commercial and... ...discipline (e.g., Cloud, AI, Android, etc.)Experience in observability such as white and black box monitoring, service level objective...Suggested
- Play a key role in ensuring system reliability at one of the world’s most iconic and... ...largest financial institutions.As a Site Reliability Engineer II at JPMorgan Chase within the Commercial... ...application codeUnderstands observability patterns and strives to implement and...Suggested
$130k - $180k
...belonging, collaboration, and accomplishment.Being a Senior Site Reliability Engineer at iManage Means… You are an engineer, a builder, and a... ...in on-call rotations. You’ll be a key voice in observability, change management, and service scalability, providing guidance...Work at officeLocal areaRemote workWorldwideMonday to FridayFlexible hours$100k - $120k
OverviewThe Site Reliability Engineer is a key force behind improving Origami’s time to resolution and advancing overall site reliability and... ...delivery.Cross trains colleagues on how to best leverage observability tools during incident and performance investigations....Full timeTemporary workWork experience placementFlexible hours$130k - $225k
...consensus.The Algorithmic Trading Team is looking for a Site Reliability Engineer for our Chicago office. The SRE team is critical to the success... ..., trading and infrastructure - someone who creates observability that surfaces issues before they ever cause a problem, owns...Temporary workWork at officeFlexible hours$130k - $150k
...technologies is essential for this role. Position OverviewThe Site Reliability Engineer (SRE) helps ensure CRA’s critical business services are... ...reduce manual toil through automation, improve service observability, and strengthen incident response. The SRE partners...Work at officeWork from home3 days per week$127k - $249k
The TeamPlatform Engineering is the department within SRE that is responsible for a range... ...edge and internal service mesh), and observability and alerting systems.The Fleet Management... ...components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager...Work at officeLocal areaRemote workWorldwideFlexible hours$125.04k - $187.56k
...Digital and E-commerce, Technology and more. Overview The Site Reliability Engineer (SRE) III is responsible for ensuring the scalability,... ...and performance of production systems through automation, observability, incident response, and infrastructure engineering. This...Full timeWork at officeRemote workFlexible hours$190.8k - $267.1k
...influential and trafficked corners of the internet. As a Senior Site Reliability Engineer on Reddit’s Infrastructure SRE team, you’ll use your... ...will work very closely with the Compute, Traffic, and Observability infrastructure teams. They will own a suite of tools for...Work experience placementHome officeFlexible hours- ...Lenovo in the United States is seeking a Senior Site Reliability Engineer to strengthen reliability, observability, and operations for the Qira platform. You will work across device, edge, and cloud, collaborating with AI/ML, firmware, and platform teams to design scalable...Remote work
- ...leading quantitative trading firm is seeking a Head of Site Reliability Engineering to help scale one of its most critical infrastructure organisations... ...production infrastructure. Lead initiatives spanning observability, automation, incident management, platform engineering,...Immediate start
$165k - $225k
...demanding AI workloads with enterprise-grade reliability and compliance. Your Role: You will... ...core. Working closely with our systems engineers, network engineers, and platform... ...reliability while establishing the automation, observability, and operational practices. Job...Remote workFlexible hours$150k - $155k
...Site Reliability Engineer Hybrid (3 days onsite, 2 days remote) full‑time. No visa sponsorship. Base pay: $150,000 – $155,000 per year, subject... ...company seeks a Site Reliability Engineer focused on observation, logging, and capacity planning. The role requires experience...Full timeWork experience placementRemote workVisa sponsorship- ...Site Reliability Engineer II Play a key role in ensuring system reliability at one of the world's most iconic and largest financial institutions... ...engineering or updating application code Understands observability patterns and strives to implement and improve service...Local area
$180k - $200k
...Company Name: tastytrade Role: Senior Site Reliability Engineer Location: Chicago, IL (Hybrid, 3 days/week in office) Role Summary... ...and market data delivery to the error-budget policy and observability standards that guide every engineering team that follows....Work at office3 days per week- ...Edward Jones Site Reliability Engineer 100% remote Initial contract is 6 months, but will be a multi year engagement. Position Overview... ...with development teams to integrate reliability and observability best practices into the software development lifecycle....Contract workRemote work
- ...Job Description As a Senior DevOps / SRE Engineer on contract, you will be embedded with the Central Technology AI enablement team... ...or LangSmith Fleet, agent registries). · Experience with observability platforms (i.e. Langsmith) and cost-attribution patterns for...Contract workImmediate start
- ...powered advice on this job and more exclusive features. Direct message the job poster from Algo Capital Group Senior Site Reliability Engineer - Observability and Automation A leading high-frequency trading firm is seeking a mid to senior-level Site Reliability Engineer...Full timeWork at officeFlexible hours
$112.5k - $187.5k
...TransUnion, this role will report to a DevOps Director. The Site Reliability Engineering team drives reliability strategy, elevates engineering... ...and built for scale.Expert-level command of monitoring, observability, and alerting platforms (e.g., Datadog, Prometheus,...Full timeTemporary workWork experience placementWork at officeFlexible hours2 days per week$194k - $267k
...-educate on new concepts and tools.Position Overview:The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes... ...provide service-to-service communication, security, and observability within the Kubernetes clusters. Enable fine-grained...Permanent employmentWork at officeLocal areaWorldwideFlexible hours$132.1k - $220.1k
We're looking for a Staff Site Reliability Engineer to join our team, focusing on the core systems that power global financial markets. This... ...automate solutions at a global scale.Spearheadthe adoption of observability and performance testing, guiding teams to a "build with...Full timeWork at officeWorldwide2 days per week- ...and companies, alikeKlover’s engineering team powers one of the... ...grade systems that prioritize reliability, security, and performance,... ...candidateAbout the RoleAs a Senior/Staff Site Reliability Engineer, you... ..., a strong dedication to observability, and a laser-focus on...Work at officeImmediate startRemote work
$194k - $267k
...:We are seeking a highly technical StaffObservabilitySite Reliability Engineer with a specialty in Splunk to own and evolve our Splunk ecosystem... ...to delivering a world class, comprehensive, scalable Observability Platform that enables our SRE teams and business partners....Permanent employmentWork at officeLocal areaWorldwideFlexible hours- ...world's most complex and mission-critical systems. As a Site Reliability Engineer III at JPMorgan Chase within the Corporate Technology team... ...tools such as Python, Ansible and Terraform.Experience in observability including white and black box monitoring, service level...
$147k - $202k
...are too, let's talk.The Auth0 Platform Observability team owns the observability tooling... ...and we are looking for an Observability Engineer to help ensure that our Product and Platform... .... If you have experience within the Site Reliability Engineering (SRE) field or working as...Local areaWorldwideFlexible hours$130k - $165k
...Job Title: Senior Software Engineer Company: Snapsheet Job Location: USA, Remote... ...Department: Technology Team : Site Reliability Engineering About Snapsheet:... ...Build and operate our core internal observability platform Monitor our systems for capacity...Full timeTemporary workLocal areaRemote workVisa sponsorshipWork visaFlexible hours$250k - $350k
...where quantitative researchers, engineers, traders, and operational... ...stability, throughput, and reliability Qualifications Minimum of 3... ...experience in production support, site reliability, or... ...production setting Familiarity with observability tools (e.g., Prometheus), and...Full time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Observability Engineer / Site Reliability Engineer. Be the first to apply!
- site reliability engineer remote Chicago, IL
- site reliability engineer sre Chicago, IL
- site reliability engineer Chicago, IL
- on-site clinical research associate (traveling/remote) Chicago, IL
- website coordinator Chicago, IL
- junior website developer Chicago, IL
- site leader Chicago, IL
- historic site Chicago, IL
- website content developer Chicago, IL
- construction site safety Chicago, IL


