Observability Engineer / Site Reliability Engineer [Remote]
Ontrac Solutions
- Remote job
About Ontrac Solutions
Ontrac Solutions is a leading technology consulting firm, specializing in cutting-edge solutions that drive business transformation. We partner with organizations to modernize their infrastructure, streamline processes, and deliver tangible results. By creating value beyond the hype, we help businesses modernize technology and build new strategies that fuel growth. Our team is committed to innovation, collaboration, and excellence, empowering our clients to succeed in an evolving digital landscape.
Role Overview
We are seeking an experienced Observability / Site Reliability Engineer (SRE) to design, scale, and maintain our enterprise monitoring and alerting ecosystems. In this role, you will bridge the gap between development and operations by ensuring high availability, performance tuning, and deep visibility across distributed multi-cloud and native systems. You will play a critical role in automating infrastructure and building robust observability pipelines using industry-leading cloud-native tools.
Key Responsibilities
- GCP & Cloud Management: Architect, optimize, and maintain observability frameworks across cloud environments, with a specific focus on implementing Google Cloud Platform (GCP) observability tools (Cloud Logging, Cloud Monitoring, Trace, and Profiler).
- Platform Management: Design, deploy, and maintain robust observability stacks across hybrid ecosystems, utilizing Prometheus, Grafana, and cloud-native integrations.
- Automation & IaC: Drive infrastructure-as-code (IaC) initiatives using Terraform and Ansible to ensure consistent, automated deployments of infrastructure and observability tooling.
- CI/CD Integration: Build, maintain, and optimize deployment workflows within Kubernetes and Google Kubernetes Engine (GKE) / OpenShift environments using GitHub, Harness, and other CI/CD pipelines.
- System Performance: Deeply analyze Linux/Unix system administration architectures, optimizing compute resource metrics and performance tuning across complex, distributed environments.
- SRE Evangelism: Implement SRE best practices, establishing meaningful Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets to ensure platform reliability.
Required Skills & Qualifications
- Cloud Infrastructure: Proven engineering experience within Google Cloud Platform (GCP) environments, particularly managing cloud-native monitoring and compute resources.
- Observability Tooling: Hands-on experience with Grafana, Prometheus, and Google Cloud Observability suites. Direct experience with GEM (Grafana Enterprise Metrics) is highly desirable.
- OS & Scripting: Expert-level knowledge of Linux/Unix operating systems paired with strong shell scripting skills for automation and systems management.
- Programming: Professional coding proficiency in at least one modern language (Python, Go, Java, Perl, or advanced Shell).
- Containers & Orchestration: Hands-on experience managing containerized applications on Kubernetes, GKE, and/or Red Hat OpenShift.
__________________________________
(Ontrac Solutions has partnered with PinpointVerify to help genuine applicants rise above the noise. Today, qualified candidates are too often overshadowed by fake and fraudulent applications. PinpointVerify gives our recruiters confidence that you are exactly who you say you are — and gives you a portable verification credential you can share with any employer.
Applicants who complete verification are prioritized over non-verified candidates with comparable experience. And if you're hired, Ontrac reimburses the full cost of your verification.
Get verified →
$194k - $267k
...career-defining work. We're all in on this mission. If you are too, let's talk. We are seeking a highly technical Observability Site Reliability Engineer with a specialty in Google Cloud, to own and expand our Observability ecosystem into GCP. In this role, you will...SuggestedPermanent employmentLocal areaWorldwideFlexible hours$92.4k - $148.8k
...Senior Site Reliability Engineer San Diego About SHEIN SHEIN is a global online fashion and lifestyle retailer, offering SHEIN branded... ..., cross-functional teams to design, build, and evolve observability and operational tooling—including metrics, logs, traces,...SuggestedTemporary workWork at officeWorldwideFlexible hours$64 - $68 per hour
...Akkodis is seeking a Site Reliability Engineer for a Contract with a client in Sunnyvale, CA/Austin, TX (Hybrid). The ideal candidate... ...Implement and support monitoring, logging, alerting, and observability platforms to ensure system reliability, performance, and...SuggestedHourly payContract workTemporary workLocal area$170k - $250k
...Site Reliability Engineer (SRE) Location: San Francisco, CA / Palo Alto, CA Company Stage of Funding: Growth-Stage AI Infrastructure Company... ...in reliability engineering to build the automation, observability, and platform infrastructure that powers their multi-cloud...SuggestedWork at officeVisa sponsorshipFlexible hours$100k - $170k
...Site Reliability Engineer Houston; San Francisco; Seattle About Nscale Nscale is the GPU cloud built for AI. We run high-performance... ...through to the retro. ~ Fluency with monitoring and observability; metrics, logs, dashboards, and alerting. ~ Comfort in...SuggestedFlexible hoursShift work$150k
...Site Reliability Engineer San Francisco, CA About The Role We are seeking an experienced Site Reliability Engineer (SRE) with a strong... ...external APIs; implement alerting and dashboards using observability tooling (e.g., CloudWatch, Datadog, Grafana). Lead...$98.58k - $138.02k
...Site Reliability Engineer II Restaurant365 is a SaaS company disrupting the restaurant industry! Our cloud-based platform provides a unique... ...and evolve monitoring tools and platforms to improve observability. Promote and apply best practices for reliability, scalability...Work at office$140k - $180k
...services. Our team of passionate engineers are constantly innovating,... ...simpler, smarter, and more reliable connectivity. We're looking... ...passionate and experienced Senior Site Reliability Engineer to join... ...of microservices. Build Observability for Microservices and cloud...Local areaWorldwideWeekend work$95k - $171k
...infrastructure, Kubernetes, and ensuring reliability for AI workloads within Akamai's... ...inference platform. As an Site Reliability Engineer II, you will be responsible for:... ...inference workloads using Akamai's existing observability platform Writing automation and tooling...Permanent employmentWork experience placementWork at officeRemote workWork from homeWorldwideFlexible hours- ...Senior Site Reliability Engineer Latitude AI is building the future of Ford's autonomy roadmap to make travel safer, less stressful, and... ...and release/deployment teams to provide uniform service observability and incident response. What you'll do: Build monitoring...Work at officeImmediate start
- ...redefine computing. About the Role We're seeking a Site Reliability Engineer to ensure Hyperbolic's GPU marketplace and AI... ...flags, and automated rollback mechanisms Proficient in observability tools and practices including metrics, logging, tracing,...
$86k - $105k
...infrastructure and to be responsible for reliability, automation and scalability using and... .... Implement and evangelize Observability and monitoring systems to proactively... ...Minimum of 2 years prior DevOps, software engineering or related experience. Must be able...Hourly payWork at officeImmediate startVisa sponsorshipWork visaFlexible hours$152.5k - $219.2k
...solutions that improve the reliability, scalability, and operational... ...environments. Partner with other engineering teams, product management,... ...~2+ years of experience in Site Reliability Engineering,... ...Knowledge of monitoring, observability, and reliability engineering...Full timeTemporary workLocal areaWorldwideFlexible hours$170k - $230k
...Site Reliability Engineer (SRE) Palo Alto / San Francisco Bay Area About Mithril Mithril is an AI infrastructure platform built to... ...keep the lights on' role — you will build the automation, observability, and tooling that allows Mithril to coordinate advanced compute...Work at officeLocal area1 day per week$146.4k - $263.6k
...Do you want to shape reliability practices for a new AI inference platform... ...decisions with product engineering teams, and shape SRE... ...at scale. As a Senior II Site Reliability Engineer, you will... ...for: Taking ownership of observability strategy for the serverless...Work experience placementWork at office$170k - $200k
...Site Reliability Engineer We are seeking a talented and motivated Site Reliability Engineer to join our engineering team. You will be responsible... ...servers, network devices, databases etc,.) Improve observability with logging, monitoring, alerting, and tracing tools (e....Full timeWorldwide$165k - $190k
...Site Reliability Engineer Palo Alto, California, USA Obsidian Security is the leading SaaS security platform, trusted by global enterprises... ...complex challenges around scalability, reliability, observability, and cost efficiency Collaborate with Engineering teams...Work from homeFlexible hours- ...attract incredibly creative scientists and engineers from leading academic institutions and... ...Summary We are looking for a Site Reliability Engineer to own the digital... ...predictably. Operational Rigor: You value observability, reproducibility, and clear operational...Visa sponsorship
$126.5k - $182k
...with a global team of software engineers and SREs responsible for... ...operations partners to ensure reliability and performance. Webex is... ...Your impact As a Site Reliability Engineer (SRE) supporting... ...improvements. Use observability data to guide capacity planning...Permanent employmentFull timeTemporary workLocal areaWorldwideFlexible hoursShift work- ...small team of former Google and Stripe engineers, including the founding team of Google... ...looking for a skilled and passionate Site Reliability Engineer to join our team. As a SRE,... ...ll be responsible for the reliability, observability, performance, and security of our core...Remote work1 day per week
$166k - $220k
...Senior Site Reliability Engineer Costa Mesa, California, United States Anduril Industries is a defense technology company with a mission... ...acquisition process and the security of our candidates. We've observed a rise in sophisticated phishing and fraudulent schemes...Full timeWork experience placementImmediate startRemote work- ...Join the engineering teams that bring OpenAI’s ideas safely to the world!! The Applied Engineering... ...About the Role We’re building the observability product for OpenAI—from scalable... ...tools to make OpenAI's production systems reliable, performant, and observable. What...Full time
- ...Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE Baseten is... ...We are hiring a Lead Software Engineer to build a first-class observability and root-cause analysis system for GPU fabrics. This is a hard...Full timeFlexible hours
- ...Join the engineering teams that bring OpenAI’s ideas safely to the world!! The Applied Engineering... ...ensuring that they are performant and reliable. You will work in a deeply iterative... ...play a crucial role in ensuring the observability, reliability, scalability, and...Full timeWork experience placementRelocation package
- ...our new San Francisco headquarters. About the Role As a Site Reliability Engineer (SRE) at Mercor, you’ll own production reliability across... ...so they are stable, resource‑efficient, isolated, and well‑observed. Introduce and champion modern SRE practices (e.g., incident...
$210k - $240k
Join to apply for the Senior Site Reliability Engineer role at Alembic Technologies This range is provided by Alembic Technologies. Your actual... ...(SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You’ll partner...Full time$150k - $195k
...team is growing, and we are looking for engineers with passion for automation. You will... ...network/connectivity, workload management, observability, and storage services. We build tooling... ...teams to improve the scalability and reliability of internal processes. Participate in...Full timeWorldwide$175k - $250k
...Senior Cloud Infrastructure Engineer Location: San Francisco, CA.... ...Remote unavailable. Modality: On-Site only. Must live within... ...scalability, performance, and reliability across environments. What... ...systems for orchestration, observability, distributed storage, and networking...Full timeRemote workRelocationRelocation package$145k - $165k
...About the role Bolt Graphics is seeking a highly experienced Site Reliability Engineer (SRE) to design, build, and operate highly reliable... ...or Microsoft Azure environments. Experience implementing observability, monitoring, and alerting using tools such as Prometheus...Work at officeImmediate start$175k - $240k
...intelligent agents ubiquitous. We build the foundation for agent engineering in the real world, helping developers move from prototypes to... ...Fullstack Engineer for our commercial product LangSmith, an observability and evals platform. In this role, you'll have the opportunity...Full timeWork at officeFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Observability Engineer / Site Reliability Engineer [Remote]. Be the first to apply!


