Observability Engineer / Site Reliability Engineer
Ontrac Solutions Inc
Observability / Site Reliability Engineer (SRE)Ontrac Solutions is a leading technology consulting firm, specializing in cutting-edge solutions that drive business transformation. We partner with organizations to modernize their infrastructure, streamline processes, and deliver tangible results. By creating value beyond the hype, we help businesses modernize technology and build new strategies that fuel growth. Our team is committed to innovation, collaboration, and excellence, empowering our clients to succeed in an evolving digital landscape.Role OverviewWe are seeking an experienced Observability / Site Reliability Engineer (SRE) to design, scale, and maintain our enterprise monitoring and alerting ecosystems. In this role, you will bridge the gap between development and operations by ensuring high availability, performance tuning, and deep visibility across distributed multi-cloud and native systems. You will play a critical role in automating infrastructure and building robust observability pipelines using industry-leading cloud-native tools.Key ResponsibilitiesGCP & Cloud Management: Architect, optimize, and maintain observability frameworks across cloud environments, with a specific focus on implementing Google Cloud Platform (GCP) observability tools (Cloud Logging, Cloud Monitoring, Trace, and Profiler).Platform Management: Design, deploy, and maintain robust observability stacks across hybrid ecosystems, utilizing Prometheus, Grafana, and cloud-native integrations.Automation & IaC: Drive infrastructure-as-code (IaC) initiatives using Terraform and Ansible to ensure consistent, automated deployments of infrastructure and observability tooling.CI/CD Integration: Build, maintain, and optimize deployment workflows within Kubernetes and Google Kubernetes Engine (GKE) / OpenShift environments using GitHub, Harness, and other CI/CD pipelines.System Performance: Deeply analyze Linux/Unix system administration architectures, optimizing compute resource metrics and performance tuning across complex, distributed environments.SRE Evangelism: Implement SRE best practices, establishing meaningful Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets to ensure platform reliability.Required Skills & QualificationsCloud Infrastructure: Proven engineering experience within Google Cloud Platform (GCP) environments, particularly managing cloud-native monitoring and compute resources.Observability Tooling: Hands-on experience with Grafana, Prometheus, and Google Cloud Observability suites. Direct experience with GEM (Grafana Enterprise Metrics) is highly desirable.OS & Scripting: Expert-level knowledge of Linux/Unix operating systems paired with strong shell scripting skills for automation and systems management.Programming: Professional coding proficiency in at least one modern language (Python, Go, Java, Perl, or advanced Shell).Containers & Orchestration: Hands-on experience managing containerized applications on Kubernetes, GKE, and/or Red Hat OpenShift.Ontrac Solutions has partnered with PinpointVerify to help genuine applicants rise above the noise. Today, qualified candidates are too often overshadowed by fake and fraudulent applications. PinpointVerify gives our recruiters confidence that you are exactly who you say you are — and gives you a portable verification credential you can share with any employer.Applicants who complete verification are prioritized over non-verified candidates with comparable experience. And if you're hired, Ontrac reimburses the full cost of your verification.
$158.5k - $172k
...deserve.About The OpportunityAs a Senior Engineer on the Runtime Automation team, you... ...system engineering, while ensuring top-tier observability and strict security across our mission-... ...-impact position driving continuous reliability, deep system optimization, and automation...SuggestedFull timeTemporary workWork at officeFlexible hours3 days per week$108.08k - $172.5k
Work with development and platform engineering teams to migrate and maintain applications in Google Cloud. Apply Observability concepts and applications to maintain services. Monitor metrics, system health and analyze reports. Provide on-call rotation support for production...SuggestedFull timeRemote workWorldwide- ...the world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Commercial and... ...discipline (e.g., Cloud, AI, Android, etc.)Experience in observability such as white and black box monitoring, service level objective...Suggested
$180k - $200k
...Company Name: tastytrade Role: Senior Site Reliability Engineer Location: Chicago, IL (Hybrid, 3 days/week in office) Role Summary... ...and market data delivery to the error-budget policy and observability standards that guide every engineering team that follows....SuggestedFull timeWork at office3 days per week$100k - $120k
OverviewThe Site Reliability Engineer is a key force behind improving Origami’s time to resolution and advancing overall site reliability and... ...delivery.Cross trains colleagues on how to best leverage observability tools during incident and performance investigations....SuggestedFull timeTemporary workWork experience placementFlexible hours- Play a key role in ensuring system reliability at one of the world’s most iconic and... ...largest financial institutions.As a Site Reliability Engineer II at JPMorgan Chase within the Commercial... ...application codeUnderstands observability patterns and strives to implement and...
$130k - $180k
...belonging, collaboration, and accomplishment.Being a Senior Site Reliability Engineer at iManage Means… You are an engineer, a builder, and a... ...in on-call rotations. You’ll be a key voice in observability, change management, and service scalability, providing guidance...Work at officeLocal areaRemote workWorldwideMonday to FridayFlexible hours$100.7k - $167.8k
Job SummaryThe Site Reliability Engineer III is a pivotal architect of stability for CME Clearing & Risk. You will engineer secure, scalable... ...defending SLIs, SLOs, and SLAs to maintain system health.Observability Mindset: Experience navigating the telemetry landscape using...Full timeWorldwide- ...Site Reliability Engineer (SRE) Immediate need for a talented Site Reliability Engineer (SRE). This is a 12+ months contract opportunity... ...Function Apps, Logic Apps - Must ~8+ years Monitoring & Observability tools of which 3+ years with Grafana, Prometheus, KQL, AppInsights...Contract workLocal areaImmediate start
$86k - $105k
...infrastructure and to be responsible for reliability, automation and scalability using and... .... Implement and evangelize Observability and monitoring systems to proactively... ...Minimum of 2 years prior DevOps, software engineering or related experience. Must be able...Hourly payWork at officeImmediate startVisa sponsorshipWork visaFlexible hours- ...powered advice on this job and more exclusive features. Direct message the job poster from Algo Capital Group Senior Site Reliability Engineer - Observability and Automation A leading high-frequency trading firm is seeking a mid to senior-level Site Reliability Engineer...Full timeWork at officeFlexible hours
- ...SRE Engineer We are seeking a highly capable engineer to join our dynamic SRE team.... ...file transfer teams to ensure secure and reliable operations. .NET Logging &... ...duties. ~ Demonstrated expertise in observability, monitoring, alerting, and troubleshooting...
$250k - $350k
...where quantitative researchers, engineers, traders, and operational... ...stability, throughput, and reliability Qualifications Minimum of 3... ...experience in production support, site reliability, or... ...production setting Familiarity with observability tools (e.g., Prometheus), and...Full time- ...Site Reliability Engineer As a Site Reliability Engineer, you will build and secure infrastructure supporting our AI platform with special... ...Monitoring & Incident Response: Implement comprehensive observability strategies using Prometheus, Grafana, and ELK. Define SLOs...
- ...Job Title: Site Reliability Engineer Location: Chicago, IL FTE Only Job Description Must Have Technical/Functional Skills ~ We are looking... ...with deep experience in AWS infrastructure, automation, observability, and production support. As an SRE, you will ensure our cloud...
$130k - $170k
...Senior Site Reliability Engineer About Us Founded in 2014, we offer the industry’s first and only cloud‑based, fully‑customisable, end‑to‑end... ...This position will lead the design and implementation of observability tools, incident response processes, and resilience...Full timeFlexible hoursShift work- ...Site Reliability Engineer II Play a key role in ensuring system reliability at one of the world's most iconic and largest financial institutions... ...engineering or updating application code Understands observability patterns and strives to implement and improve service...Local area
$127k - $249k
The TeamPlatform Engineering is the department within SRE that is responsible for a range... ...edge and internal service mesh), and observability and alerting systems.The Fleet Management... ...components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager...Work at officeLocal areaRemote workWorldwideFlexible hours$130k - $150k
...technologies is essential for this role. Position OverviewThe Site Reliability Engineer (SRE) helps ensure CRA’s critical business services are... ...reduce manual toil through automation, improve service observability, and strengthen incident response. The SRE partners...Work at officeWork from home3 days per week$106.9k - $200.6k
...opportunity We are seeking an AI Systems Engineer to own the delivery, model-serving, routing, and observability layer of EY’s AI-native platform. These are the... ...filtering, so every signal is captured and routed reliably. Automate GitOps-based delivery and...Full timeSummer holidayFlexible hours$130k - $225k
...consensus.The Algorithmic Trading Team is looking for a Site Reliability Engineer for our Chicago office. The SRE team is critical to the success... ..., trading and infrastructure - someone who creates observability that surfaces issues before they ever cause a problem, owns...Temporary workWork at officeFlexible hours$98.18k - $115.5k
...each other. Job Description Responsibilities The Reliability Observability Engineer 3 is responsible for enabling reliable, measurable, and... ...Partner with Product Owners , Application Engineering , Site Reliability Engineering (SRE) , and Operations Teams to...Full timeTemporary workWork experience placementLocal area3 days per week$174k - $267k
...new concepts and tools. Position Overview: The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes... ...provide service-to-service communication, security, and observability within the Kubernetes clusters. Enable fine-grained...Permanent employmentFull timeWork at officeLocal areaWorldwideFlexible hours$194k - $267k
...:We are seeking a highly technical StaffObservabilitySite Reliability Engineer with a specialty in Splunk to own and evolve our Splunk ecosystem... ...to delivering a world class, comprehensive, scalable Observability Platform that enables our SRE teams and business partners....Permanent employmentWork at officeLocal areaWorldwideFlexible hours- ...and companies, alikeKlover’s engineering team powers one of the... ...grade systems that prioritize reliability, security, and performance,... ...candidateAbout the RoleAs a Senior/Staff Site Reliability Engineer, you... ..., a strong dedication to observability, and a laser-focus on...Work at officeImmediate startRemote work
$120.6k - $150.9k
...are looking for a highly motivated, high-potential Staff Site Reliability Engineer (SRE) to join our team as a technical leader and drive transformative... ..., and platforms. Ensuring these systems are scalable, observable, and resilient is critical to unlocking business value and...Full timeFlexible hours$112.5k - $187.5k
...TransUnion, this role will report to a DevOps Director. The Site Reliability Engineering team drives reliability strategy, elevates engineering... ...and built for scale.Expert-level command of monitoring, observability, and alerting platforms (e.g., Datadog, Prometheus,...Full timeTemporary workWork experience placementWork at officeFlexible hours2 days per week$160k - $210k
...you'll do:Join our Platform Engineering team, where you'll ensure the... ...mentoring engineers across reliability initiativesAnalyze, troubleshoot... ...of experience in DevOps, Site Reliability Engineering, or... ...environmentsKnowledge of monitoring and observability tools such as Prometheus,...Work at officeWorldwideMonday to FridayFlexible hours$132.1k - $220.1k
...We're looking for a Staff Site Reliability Engineer to join our team, focusing on the core systems that power global financial markets. This... ...solutions at a global scale. Spearhead the adoption of observability and performance testing, guiding teams to a "build with...Full timeWork at officeWorldwide2 days per week- ...world's most complex and mission-critical systems. As a Site Reliability Engineer III at JPMorgan Chase within the Corporate Technology team... ...tools such as Python, Ansible and Terraform.Experience in observability including white and black box monitoring, service level...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Observability Engineer / Site Reliability Engineer. Be the first to apply!
- site reliability engineer remote Chicago, IL
- site reliability engineer Chicago, IL
- site reliability engineer sre Chicago, IL
- junior website developer Chicago, IL
- website content developer Chicago, IL
- on site coordinator Chicago, IL
- website coordinator Chicago, IL
- site leader Chicago, IL
- site recruiter Chicago, IL
- historic site Chicago, IL

