Staff Site Reliability Engineer - Observability GCP
$194k - $267kOkta
Secure Every Identity, from AI to Human
Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk.We are seeking a highly technical Observability Site Reliability Engineer with a specialty in Google Cloud, to own and expand our Observability ecosystem into GCP. In this role, you will move beyond simple monitoring to delivering a world class, comprehensive, scalable Observability Platform that enables our SRE teams and business partners. You will treat infrastructure as code —utilizing Terraform and strong coding proficiency in Go, Python, or Ruby —to automate the deployment of agents and collectors across complex distributed systems.
Key Responsibilities
- Automated Infrastructure: Design, build, and maintain scalable observability infrastructure using tools like Terraform.
- GCP Observabilty Engineering: Optimize the collection, processing, and storage of Observabilty data to ensure high reliability and low latency of our Splunk and Grafana services
- Incident Response: Participate in on-call rotations and lead post-incident reviews to drive systemic improvements and "observability-driven development."
- Automation: Eliminate "toil" by automating the deployment and scaling of observability agents and collectors.
Required Skills & Experience (The Essentials)
GKE: Minimum 5+ Experience scaling and managing observability in a Google Cloud platform. Visualization: Expertise in creating intuitive, actionable Splunk or Grafana dashboards that correlate data across multiple sources. SRE Mindset: Minimum 3+ years of experience in an SRE, DevOps, or Systems Engineering role with a focus on high-availability systems.
- Programming Proficiency: Strong coding skills in Python , Go for building internal tools and automating workflows.
- Distributed Systems: Deep understanding of Linux internals, networking (TCP/IP, DNS, Load Balancing), and container orchestration (Kubernetes/GKE).
- Problem Solving: A data-driven approach to debugging complex, cross-service performance bottlenecks.
Bonus Skills (The "Nice-to-Haves")
- Telemetry Standards: Hands-on experience with OpenTelemetry (OTel), Vector, or similar frameworks for instrumenting applications.
- Grafana Loki: Experience in migrating Splunk to Grafana Loki
Other Cloud Platforms: Experience managing observability native tools within AWS.
Additional requirements:
- This position requires the ability to access federal environments and/or have access to protected federal data. As a condition of employment for this position, the successful candidate must be able to submit documentation establishing U.S. Person status (e.g. a U.S. Citizen, National, Lawful Permanent Resident, Refugee, or Asylee. 22 CFR 120.15) upon hire.
#LI-MM
#LI-Hybrid
P24517_3387022
The annual base salary range for this position for candidates located in the San Francisco Bay area is between: $194,000—$267,000 USD Below is the annual base salary range for candidates located in California (excluding San Francisco Bay Area), Colorado, Illinois, New York and Washington. Your actual base salary will depend on factors such as your skills, qualifications, experience, and work location. In addition, Okta offers equity (where applicable), bonus, and benefits, including health, dental and vision insurance, 401(k), flexible spending account, and paid leave (including PTO and parental leave) in accordance with our applicable plans and policies. To learn more about our Total Rewards program please visit: .
The Okta Experience
We are intentional about connection. Our global community, spanning over 20 offices worldwide, is united by a drive to innovate. Your journey begins with an immersive, in-person onboarding experience designed to accelerate your impact and connect you to our mission and team from day one.
Okta is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, ancestry, marital status, age, physical or mental disability, or status as a protected veteran. We also consider for employment qualified applicants with arrest and convictions records, consistent with applicable laws. If reasonable accommodation is needed to complete any part of the job application, interview process, or onboarding please use this Form to request an accommodation. Notice for New York City Applicants & Employees: Okta may use Automated Employment Decision Tools (AEDT), as defined by New York City Local Law 144, that use artificial intelligence, machine learning, or other automated processes to assist in our recruitment and hiring process. In accordance with NYC Local Law 144, if you are an applicant or employee residing in New York City, please click here to view our full NYC AEDT Notice.$185k - $210k
...innovation. About the role The Observability department plays a pivotal role in CoreWeave... ...We are seeking senior observability engineers with specializations in logging and... ...Improve the performance, security, reliability, and scalability of observability services...SuggestedFull timeTemporary workCasual workWork at officeRemote workFlexible hours$149.4k - $202k
...Senior Software Engineer- Site Reliability Engineering (SRE) DC, MD, VA, CA The Site Reliability Engineering discipline at Noctua Technology... ...as Code (IaC), monitor it through advanced observability stacks, and protect it by engineering for failure. We work...SuggestedRemote work- ...Signature IT World Inc Production support expertise with SRE Observability experience : Proactive issue identification using... ...new job is posted. Sign in to set job alerts for “Senior Site Reliability Engineer” roles. Bellevue, WA $204,000.00-$259,000.00 1 day ago Seattle...SuggestedContract workRemote work
$175k - $250k
...Senior Cloud Infrastructure Engineer Location: San Francisco, CA.... ...Remote unavailable. Modality: On‑Site only. Must live within... ...scalability, performance, and reliability across environments. What... ...systems for orchestration, observability, distributed storage, and networking...SuggestedFull timeRemote workRelocationRelocation package$103.5k - $150k
...self. The Role and Team The Site Reliability Engineering organization at Medallia brings together... ...availability, and performance using observability and alerting platforms.... ...infrastructure platforms such as AWS, OCI, or GCP. ~ Demonstrated experience with Linux...SuggestedTemporary workWork experience placementLocal area3 days per week$121.4k - $218.6k
...building systems that scale? Join our highly skilled Site Reliability Engineering team! Our team designs, develops, and manages applications... ...of our infrastructure. Designing and implementing observability solutions, including monitoring, logging, alerting, and...Work experience placementWork at office$153k - $185k
...Senior Site Reliability Engineer El Segundo, California, United States About Varda Low Earth... ...this goal, with leadership and staff comprised of veterans from SpaceX, Blue... ...as Terraform Implement and operate observability systems (metrics, logging, tracing) and...Permanent employmentFull timeImmediate startRelocation packageFlexible hoursWeekend work$135k - $150k
Senior Site Reliability Engineer Job number: 884 This is a remote position. Ad Hoc is a technology company that empowers organizations... ...the platform toward them Designing and operating observability across metrics, logging, tracing, and alerting Leading...Remote workFlexible hours$95k - $171k
...infrastructure, Kubernetes, and ensuring reliability for AI workloads within Akamai's... ...inference platform. As an Site Reliability Engineer II, you will be responsible for:... ...inference workloads using Akamai's existing observability platform Writing automation and tooling...Permanent employmentWork experience placementWork at officeRemote workWork from homeWorldwideFlexible hours$165k - $230k
...the ultimate goal of enabling human life on Mars. SR. SITE RELIABILITY ENGINEER (STARSHIELD) Starshield leverages SpaceX’s Starlink technology... ...for government use, with an initial focus on earth observation, communications, and hosted payloads. SpaceX’s satellite programs...Permanent employmentTemporary workImmediate startWeekend work$147k - $202k
...are too, let's talk. The Auth0 Platform Observability team owns the observability tooling... ...and we are looking for an Observability Engineer to help ensure that our Product and... ...stability. If you have experience within the Site Reliability Engineering (SRE) field or as a...Local areaFlexible hours- ...including, with hands-on Development and Systems engineering background ~3-5 years of experience in a Site Reliability Engineering role ~ Experience with... ...Twistlock, Sysdig, Aqua etc. Platform Monitoring, Observability, & Performance Tools: Nginx, New Relic, AppDynamics...Temporary workImmediate start
$106.3k - $221.1k
...missions and the government forward! Job Description The Site Reliability Engineer will ensure the reliability, performance, and scalability... ...and need an accommodation for a disability or religious observance during the interview process or for the job you are...Live inWork at officeLocal area$51.9 per hour
...is responsible for the reliability, availability, and... ...role blends software engineering, clinical engineering,... ...functionally with AHN site leaders and teams to navigate... ...management and staff productivity.Plan, organize... ...actions. Utilizes observability practices to gain deep...For contractorsLocal area$194k - $267k
...Position Overview: We are seeking a highly technical Staff Observability Site Reliability Engineer with a specialty in Splunk to own and evolve our... ...Experience managing observability native tools within AWS or GCP. Additional requirements: This position requires...Permanent employmentWork at officeLocal areaWorldwideFlexible hours$144.2k - $288.4k
...-focused Principal Software Development Engineer to join our high-energy, growing team. In... ..., disaster recovery, security, and observability, ensuring solutions meet enterprise standards... ...solutions on Google Cloud Platform (GCP). 5+ years of experience designing modern...Hourly payFull timeTemporary workLocal areaFlexible hours$105.79k - $141.05k
...Role We are seeking a highly skilled and proactive Lead Site Reliability Engineer (SRE) to join our team, focusing on production support... ...our systems, with a strong emphasis on AWS infrastructure, observability, automation, and AI-assisted engineering practices....Full timeTemporary workRemote work$132.23k - $176.31k
...We are seeking a highly skilled and proactive Senior Lead Site Reliability Engineer (SRE) to join our team, focusing on production support... ...our systems, with a strong emphasis on AWS infrastructure, observability, automation and AI-assisted engineering practices. The...Full timeTemporary workRemote work- ...Job Title: GCP Engineer Job ID: 2024-12689 Job Location: Mt Laurel, NJ or New York, NY or Toronto, ON or London, ON (2 days/week onsite) Job Travel Location(s): # Positions: 1 Employment Type: W2 Duration:Long term # of Layers:0 Work Eligibility:All Work Authorizations...Work at office2 days per week
- ...seeking a talented Backend Engineer to join our Streamlit... ...APIs for performance, reliability, and developer... ...Implement robust monitoring, observability, and error handling... ...platforms (AWS, Azure, GCP) and hands-on... ...the Snowflake Careers Site for salary and benefits...Full timeWorldwide
- ...Stack Cloud : Google Cloud Platform (GCP) Infrastructure as Code : Pulumi... ...another senior role. As our Senior Software Engineer, you will be working closely with the CTO... ...gathers dust). ~12 Paid Holidays: We observe all standard holidays to give the team time...Full timeWork at officeRemote workMonday to ThursdayFlexible hours
$196k - $294k
...~ We are seeking a Principal Software Engineer to lead architectural and strategic initiatives... ...platform API strategy. Own platform observability, product analytics frameworks, and... ...experience with cloud infrastructure (AWS, GCP, Azure). ~ Exceptional strategic...Full timeWork experience placementLocal areaRelocation package$166k - $220k
...Motion Planning, Hardware, Test Engineering, Space, Networking, and... ..., or Google Cloud Platform (GCP). Containerization and Orchestration... ...: Use Kubernetes to ensure reliable deployment and scaleability... ...Familiarity with observability concepts and tools. Knowledge...Full timeWork experience placementImmediate startRelocation package- ...distributed systems and database engine problems, care about... ...teams and external customers Observability, reliability, and operational tooling for... ...on AWS, Azure, or GCP Prior work on multi-tenant... ...posting on the Snowflake Careers Site for salary and benefits information...Full time
- ...and passionate Software Engineer for our Snowpark... ...Coding, Reviews, Testing, Observability, Tooling and On-Call support... .... ● Build highly reliable and fault-tolerant... ...● Ability to work on-site in our downtown Bellevue... ...support (AWS, Azure, GCP). ● Implementing multi...Full timeWork at office
$146k - $194k
...Anduril’s Cloud Infrastructure Engineering team is the foundation upon... ...cloud environments (AWS, Azure, GCP). Automate Everything:... ...critical data. Enable Scale and Reliability: Engineer systems and... ...security of our candidates. We've observed a rise in sophisticated phishing...Full timeWork experience placementImmediate start$112k - $176k
...interesting and challenging problems alongside a great team of engineers. * Developing new skills as you push your knowledge, and our technology... .... * Experience with cloud technologies, specifically AWS and GCP, including deploying, monitoring, and scaling applications. *...Full timeWorldwide- ...A leading security infrastructure firm in Washington, D.C. is seeking a hands-on Site Reliability Engineer (SRE) with expertise in Kubernetes and cloud infrastructure. The role emphasizes total ownership of security infrastructure while defending against advanced threats...
$131k - $227.13k
...Description: The 1LMX MES COE is seeking an engineer who will own infrastructure‑as‑code, cloud platform, and reliability for the Apriso environment on AWS. This role blends full‑stack development, DevOps, and Site Reliability Engineering (SRE) practices to deliver...Full timeTemporary workWork experience placementWork at officeRemote workRelocationFlexible hoursShift work3 days per week- ...SRE Engineer Location: Washington, DC (Onsite) Duration: 08-17-2026 - 07-30-2027 Key Responsibilities Observability & Monitoring: Standardize and automate Dynatrace installations... ...knowledge base articles. Reliability Engineering: Champion SRE metrics...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Staff Site Reliability Engineer - Observability GCP. Be the first to apply!
- senior staff engineer Washington DC
- engineering aide Washington DC
- software engineer staff Washington DC
- assistant engineer Washington DC
- technology administrator Washington DC
- senior staff systems engineer Washington DC
- staff engineer Washington DC
- site reliability engineer Washington DC
- site reliability engineer remote Washington DC
- site reliability engineer sre Washington DC



