Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Observability Engineer / Site Reliability Engineer

Ontrac Solutions Inc

Observability / Site Reliability Engineer (SRE)

Ontrac Solutions is a leading technology consulting firm, specializing in cutting-edge solutions that drive business transformation. We partner with organizations to modernize their infrastructure, streamline processes, and deliver tangible results. By creating value beyond the hype, we help businesses modernize technology and build new strategies that fuel growth. Our team is committed to innovation, collaboration, and excellence, empowering our clients to succeed in an evolving digital landscape.

Role Overview

We are seeking an experienced Observability / Site Reliability Engineer (SRE) to design, scale, and maintain our enterprise monitoring and alerting ecosystems. In this role, you will bridge the gap between development and operations by ensuring high availability, performance tuning, and deep visibility across distributed multi-cloud and native systems. You will play a critical role in automating infrastructure and building robust observability pipelines using industry-leading cloud-native tools.

Key Responsibilities
  • GCP & Cloud Management: Architect, optimize, and maintain observability frameworks across cloud environments, with a specific focus on implementing Google Cloud Platform (GCP) observability tools (Cloud Logging, Cloud Monitoring, Trace, and Profiler).
  • Platform Management: Design, deploy, and maintain robust observability stacks across hybrid ecosystems, utilizing Prometheus, Grafana, and cloud-native integrations.
  • Automation & IaC: Drive infrastructure-as-code (IaC) initiatives using Terraform and Ansible to ensure consistent, automated deployments of infrastructure and observability tooling.
  • CI/CD Integration: Build, maintain, and optimize deployment workflows within Kubernetes and Google Kubernetes Engine (GKE) / OpenShift environments using GitHub, Harness, and other CI/CD pipelines.
  • System Performance: Deeply analyze Linux/Unix system administration architectures, optimizing compute resource metrics and performance tuning across complex, distributed environments.
  • SRE Evangelism: Implement SRE best practices, establishing meaningful Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets to ensure platform reliability.
Required Skills & Qualifications
  • Cloud Infrastructure: Proven engineering experience within Google Cloud Platform (GCP) environments, particularly managing cloud-native monitoring and compute resources.
  • Observability Tooling: Hands-on experience with Grafana, Prometheus, and Google Cloud Observability suites. Direct experience with GEM (Grafana Enterprise Metrics) is highly desirable.
  • OS & Scripting: Expert-level knowledge of Linux/Unix operating systems paired with strong shell scripting skills for automation and systems management.
  • Programming: Professional coding proficiency in at least one modern language (Python, Go, Java, Perl, or advanced Shell).
  • Containers & Orchestration: Hands-on experience managing containerized applications on Kubernetes, GKE, and/or Red Hat OpenShift.

Ontrac Solutions has partnered with PinpointVerify to help genuine applicants rise above the noise. Today, qualified candidates are too often overshadowed by fake and fraudulent applications. PinpointVerify gives our recruiters confidence that you are exactly who you say you are — and gives you a portable verification credential you can share with any employer.

Applicants who complete verification are prioritized over non-verified candidates with comparable experience. And if you're hired, Ontrac reimburses the full cost of your verification.

Vacancy posted 20 hours ago
Similar jobs that could be interesting for youBased on the Observability Engineer / Site Reliability Engineer in United States vacancy
  • $138.4k - $173k

     ...infrastructure as well as help improve the reliability, quality of services and overall observability patterns. Along with your team,...  ...’ll collaborate or embed with engineering teams, helping them to improve...  ...our locations by visiting our site.Compensation & BenefitsThe base... 
    Suggested
    Full time
    Flexible hours

    AppFolio

    Dallas, TX
    3 days ago
  •  ...serves customers in more than 35 countries worldwide.Position OverviewWe are seeking a highly experienced Senior Site Reliability Engineer (Unified Observability) to lead the design, implementation, and operational maturity of the F1 Next Generation Customer Unified... 
    Suggested
    Full time
    Worldwide
    Flexible hours

    NCR

    Atlanta, GA
    2 days ago
  • $159k - $272k

     ...opportunity to grow and make a difference in ways that matter to you. Role SummaryIn this role as Principal Site Reliability Engineer, Infrastructure Observability you will help formulate, develop, and implement a team of Site Reliability Engineers (SREs) focused on the... 
    Suggested
    Full time
    Private practice
    Local area
    Remote work
    Work from home
    3 days per week

    T. Rowe Price

    Owings Mills, MD
    5 days ago
  •  ...Role: Site Reliability Engineer Observability (AWS) Location : Remote Duration : 12+ months (Alternates depending on seniority/framing: Cloud Observability Engineer, Sr. SRE Monitoring & Observability, or DevOps Engineer Observability Platform... 
    Suggested
    Remote work

    Conch Technologies Inc

    New York, NY
    3 days ago
  •  ...Information Technology group delivers secure, reliable technology solutions that enable...  ...enterprise platforms.As a Principal Site Reliability Engineer (SRE), you will drive operational...  ...reliability initiatives, champion observability and automation, drive major incident... 
    Suggested
    Remote work
    Flexible hours

    DTCC- The Depository Trust & Clearing Corporation

    Jersey City, NJ
    4 days ago
  • $49.96 - $55 per hour

    job summary: Join an enterprise Service Health Platforms team as a Site Reliability Engineer managing large-scale network observability systems hosted in Azure. In this role, you will ensure high availability, security compliance, and optimal performance for internal... 
    Hourly pay
    Contract work
    Temporary work
    Work experience placement
    Remote work
    Washington DC
    9 days ago
  • OneTrust is seeking a Senior Software Engineer in Atlanta, Georgia. The role involves designing and maintaining a reliable application platform, collaborating with engineering...  ...enhancing customer experiences through observability tools. The ideal candidate will have a... 

    OneTrust

    Atlanta, GA
    3 days ago
  •  ...Job Description Job Description SRE Support Engineer - Observability While this position is not currently open, we are interviewing strong...  ...support across Slack and tickets, improving monitoring reliability, and reducing incident impact through better triage, troubleshooting... 
    Remote work

    Virtasant

    Austin, TX
    1 day ago
  • $143k - $191k

     ...capabilities to our customers. System Deployment Engineers work in complex environments with...  ...demonstrations and exercisesWork with site reliability engineers to provide and refine...  ...the security of our candidates. We've observed a rise in sophisticated phishing and fraudulent... 
    Full time
    Temporary work
    Work experience placement
    Immediate start

    Anduril Industries

    Costa Mesa, CA
    3 days ago
  •  ...world thrive at work.Location: Salt Lake City, UTAs a Senior Site Reliability Engineer, you will help define the future of reliability for our...  ..., and engineering best practices.Build and evolve observability platforms using OpenTelemetry, Datadog, Coralogix, or similar... 
    Full time
    Shift work

    O.C. Tanner

    Salt Lake City, UT
    3 days ago
  • $166k - $220k

     ...mission partners and operators to deploy reliable and robust capabilities on operationally...  ...must be reached. As a Senior Software Engineer, you will re-imagine the infrastructure...  ...and the security of our candidates. We've observed a rise in sophisticated phishing and fraudulent... 
    Full time
    Work experience placement
    Immediate start

    Anduril Industries

    Costa Mesa, CA
    3 days ago
  • $170k - $220k

     ...We're Looking ForWe’re looking for a hands-on, high-agency Site Reliability Engineer to help shape and scale the reliability layer of our stack...  ...and hotfix coordination. Build safe, repeatable, and observable workflows.GitHub Operations: Manage GitHub branching strategies... 

    Supio

    Seattle, WA
    3 days ago
  • $152.5k - $205k

     ...everyone is a stakeholder.What you’ll be responsible forThe Site Reliability Engineer builds and maintains shared platform capabilities, common...  ...workflows, and ensure appropriate reliability, observability, access controls, auditability, and cost management. You... 
    Flexible hours

    Circle

    San Francisco, CA
    2 days ago
  • $104.9k - $174.7k

     ...link below, About the Role:We are hiring a hands-on Senior Site Reliability Engineer (SRE) to actively build, operate, and improve the...  ...in designing infrastructure, writing Terraform, improving observability, and responding to real production incidents.If you live... 
    Full time
    Work at office
    Local area
    Remote work
    Work from home

    LexisNexis Risk Solutions Group

    Alpharetta, GA
    2 days ago
  • $166k - $220k

     ...expectations. Our systems integration engineers internalize the nuances of each deployment...  ...ship.ABOUT THE JOBWe are looking for a Site Reliability Engineer (SRE) to join AGD, our rapidly...  ...the security of our candidates. We've observed a rise in sophisticated phishing and... 
    Full time
    Work experience placement
    Immediate start

    Anduril Industries

    Washington DC
    2 days ago
  • $86.6k - $144.4k

     ...automation to solve complex security and reliability challenges?Do you enjoy shaping the...  ...of security operations through observability, compliance-as-code, and AI-driven...  ...LexisNexis Risk at our TeamOur Site Reliability Engineering (SRE) team plays a critical role in... 
    Full time
    Local area

    LexisNexis Risk Solutions Group

    Alpharetta, GA
    1 day ago
  • $81.1k - $187k

     ...infrastructure and service to ensure reliability and functionality. Forecasts...  ...and develops knowledge of site reliability trends.Only...  ...and mentorship to junior engineers. Communicate status, risks,...  ...alerting, centralized logging, observability, and reliability practices.... 
    Temporary work
    Flexible hours

    Oracle Corporation

    Nashville, TN
    20 hours ago
  • Recognized as the No. 1 site trusted by real estate professionals, Realtor.com has...  ...guidance.We are seeking a Senior Site Reliability Engineer to join our newly formed Operations...  ...role will contribute to the reliability, observability, and operational excellence of our... 
    Work at office
    Local area

    Realtor.com

    Austin, TX
    4 days ago
  • Reliability Engineering Design, implement, and operate scalable, resilient, and highly available systems...  ..., and reliable software delivery.Observability and Incident Management Develop...  ...Three or more years of experience in Site Reliability Engineering, platform engineering... 
    Remote work

    Patterson-UTI

    Houston, TX
    1 day ago
  • $165k - $225.6k

     ...we partner across functions to drive scale, reliability, and innovation through technology.The Senior Site Reliability Engineer OpportunityReporting to the Manager, Site...  ...GitLab, GitHub Actions).Containerization & Observability: Strong hands-on experience managing container... 
    Permanent employment
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    3 days ago
  • $158.5k - $172k

     ...deserve.About The OpportunityAs a Senior Engineer on the Runtime Automation team, you...  ...system engineering, while ensuring top-tier observability and strict security across our mission-...  ...-impact position driving continuous reliability, deep system optimization, and automation... 
    Full time
    Temporary work
    Work at office
    Flexible hours
    3 days per week

    GrubHub

    Chicago, IL
    3 days ago
  • $230k - $250k

    GovCIO is hiring a Site Reliability Engineer with an active Secret clearance to ensure reliability, scalability, performance, and availability...  ...manual processes.Develop monitoring, alerting, and observability solutions.Improve system performance, capacity, and resilience... 
    Remote work

    Govcio

    Arlington, VA
    3 days ago
  • $128.6k - $184.9k

     ...platform. As a team of six engineers distributed across the US, Canada...  ...strong focus on automation, reliability, and operational excellence....  ...7+ years of experience in Site Reliability Engineering,...  ...environments.Knowledge of monitoring, observability, and reliability engineering... 
    Permanent employment
    Full time
    Temporary work
    Local area
    Worldwide
    Flexible hours

    CISCO Systems

    Richardson, TX
    2 days ago
  •  ...the world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Corporate Technology,...  ...discipline (e.g., Cloud, AI, Android, etc.)Experience in observability such as white and black box monitoring, service level objective... 

    JP Morgan Chase

    Houston, TX
    1 day ago
  • $130k - $200k

     ...used by millions of people globally, is seeking a Senior Site Reliability Engineer to join our team, and help maintain the reliability and optimal...  ...message queues, load balancers, and databasesFluency in observability tools and methodologiesFlexibility to work with a variety... 
    Full time
    Work at office
    Immediate start

    IXL Learning

    San Mateo, CA
    1 day ago
  • $165k - $190k

     ...DevOps / SRE TeamThe DevOps/SRE team at Obsidian ensures that engineering excellence translates into stable, scalable, and high-...  ...platformAddress complex challenges around scalability, reliability, observability, and cost efficiencyCollaborate with Engineering teams to... 
    Work from home

    Obsidian Security

    Palo Alto, CA
    5 days ago
  • $45 - $85 per hour

    DescriptionThe Resy Site Reliability Engineering groups goal is to ensure Resy Customers can always use the service reliably. We're looking...  ...incident response and post-incident reviews - Applying observability engineering to our applications to ensure we can proactively... 
    Contract work
    Temporary work

    TEKsystems

    Phoenix, AZ
    2 days ago
  • Inspire Brands is hiring two Senior Site Reliability Engineers to help build and scale reliable, resilient, and observable systems supporting high-traffic, customer-facing digital platforms. These role blends software engineering, systems thinking, and operational excellence... 
    Worldwide

    Inspire Brands

    Atlanta, GA
    2 days ago
  • $104.9k - $174.7k

    About the role:A FinOps Site Reliability Engineer (SRE) bridges the gap between engineering, operations, and financial governance by embedding...  ...strong experience in Azure, AWS, Kubernetes, Terraform, observability, automation, AI platforms, and cloud financial management... 
    Full time
    Local area

    LexisNexis Risk Solutions Group

    Boca Raton, FL
    2 days ago
  • $130k - $153k

     ...and recreational access.WHAT YOU WILL DOonX is seeking a Site Reliability Engineer to build and maintain the infrastructure that enables our...  ...onX's infrastructure platform, deployment automation, and observability through infrastructure-as-code—keeping systems reliable... 
    Full time
    Part time
    Work at office

    onXmaps

    Bozeman, MT
    20 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Observability Engineer / Site Reliability Engineer. Be the first to apply!