Staff Site Reliability Engineer - Splunk
$194k - $267kOkta
Secure Every Identity, from AI to Human
Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk.Position Overview:
We are seeking a highly technical Staff Observability Site Reliability Engineer with a specialty in Splunk to own and evolve our Splunk ecosystem. In this role, you will move beyond simple monitoring to delivering a world class, comprehensive, scalable Observability Platform that enables our SRE teams and business partners. You will treat infrastructure as code —utilizing Terraform and strong coding proficiency in Go, Python, or Ruby —to automate the deployment of agents and collectors across complex distributed systems.
Key Responsibilities
- Automated Infrastructure: Design, build, and maintain scalable observability infrastructure using tools like Terraform.
- Splunk Engineering: Optimize the collection, processing, and storage of log data to ensure high reliability and low latency of our Splunk services
- Incident Response: Participate in on-call rotations and lead post-incident reviews to drive systemic improvements and "observability-driven development."
- Automation: Eliminate "toil" by automating the deployment and scaling of observability agents and collectors.
Required Skills & Experience (The Essentials)
Log Management: Minimum 5+ Experience scaling and managing Splunk Cloud at scale (1000+ SVCs), including Workload Management (WLM) and HEC optimization. Visualization: Expertise in creating intuitive, actionable Splunk dashboards that correlate data across multiple sources.
SRE Mindset: Minimum 5+ years of experience in an SRE, DevOps, or Systems Engineering role with a focus on high-availability systems.
- Programming Proficiency: Strong coding skills in SPL , Go for building internal tools and automating workflows.
- Distributed Systems: Deep understanding of Linux internals, networking (TCP/IP, DNS, Load Balancing), and container orchestration (Kubernetes/EKS).
- Problem Solving: A data-driven approach to debugging complex, cross-service performance bottlenecks.
Bonus Skills (The "Nice-to-Haves")
- Telemetry Standards: Hands-on experience with OpenTelemetry (OTel), Vector, or similar frameworks for instrumenting applications.
- Charge-back app: Experience in implementing Splunk charge-back app for usage reporting
Cloud Platforms: Experience managing observability native tools within AWS or GCP.
Additional requirements:
- This position requires the ability to access federal environments and/or have access to protected federal data. As a condition of employment for this position, the successful candidate must be able to submit documentation establishing U.S. Person status (e.g. a U.S. Citizen, National, Lawful Permanent Resident, Refugee, or Asylee. 22 CFR 120.15) upon hire.
- This person must attend in person onboarding in our San Francisco/Chicago office the first week of employment.
#LI-MM
#LI-Hybrid
P14596_3372199
Below is the annual base salary range for candidates located in California (excluding San Francisco Bay Area), Colorado, Illinois, New York and Washington. Your actual base salary will depend on factors such as your skills, qualifications, experience, and work location. In addition, Okta offers equity (where applicable), bonus, and benefits, including health, dental and vision insurance, 401(k), flexible spending account, and paid leave (including PTO and parental leave) in accordance with our applicable plans and policies. To learn more about our Total Rewards program please visit: .
The Okta Experience
We are intentional about connection. Our global community, spanning over 20 offices worldwide, is united by a drive to innovate. Your journey begins with an immersive, in-person onboarding experience designed to accelerate your impact and connect you to our mission and team from day one.
Okta is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, ancestry, marital status, age, physical or mental disability, or status as a protected veteran. We also consider for employment qualified applicants with arrest and convictions records, consistent with applicable laws. If reasonable accommodation is needed to complete any part of the job application, interview process, or onboarding please use this Form to request an accommodation. Notice for New York City Applicants & Employees: Okta may use Automated Employment Decision Tools (AEDT), as defined by New York City Local Law 144, that use artificial intelligence, machine learning, or other automated processes to assist in our recruitment and hiring process. In accordance with NYC Local Law 144, if you are an applicant or employee residing in New York City, please click here to view our full NYC AEDT Notice.- ...to SearchRemote in United States of America: New YorkSite Reliability Engineering& 4 othersLooking for something else?Find a vacancy that works... ...OP2 / Vector (Observability Pipelines Worker) alongside Splunk, GCP Pub/Sub sinks, and OpenTelemetry (OTel)AWS-Native Monitoring...Splunk
- ...Site Reliability Engineer Our Client, a multinational telecommunications technology company is seeking a Site Reliability Engineer (SRE I)... ...monitoring/observability tools (Grafana, Prometheus, ELK, Splunk, DataDog, or similar). •General knowledge of networking fundamentals...SplunkTemporary work
- ...Senior Site Reliability Engineer (SRE) Our client is seeking a Senior Site Reliability Engineer (SRE) with 10–15 years of experience to support... ...-on experience with observability and monitoring tools: Splunk, Grafana, Prometheus, Nagios, Datadog, OpenTelemetry, CloudWatch...Splunk
- ...Site Reliability Engineer (SRE) Job Title Site Reliability Engineer (SRE) Job Summary We are seeking a skilled Site... ...Prometheus, Grafana, ELK Stack, OpenTelemetry, Datadog, or Splunk. Familiarity with service mesh technologies such as Istio...SplunkFlexible hours
- ...DevOps Engineer Responsible for reliability and support of container platform on-prem and external clouds (Azure /AWS /Google) Monitor and troubleshoot... ..., Golang, and shell scripting Experience with Splunk/Dynatrace Experience with Java Full Stack...Splunk
$104k - $178k
...About the Role Role Overview You will join the Site Reliability Engineering (SRE) team within DoubleVerify's Technology organization.... ...monitoring and observability tools such as Prometheus, Grafana, Splunk, or Nagios. ~ Hands-on experience with Infrastructure-as...Splunk- ...significant effect in a realm tailored for top achievers in site reliability. As a Lead Site Reliability Engineer at JPMorgan Chase within the Commercial & Investment... ...and service management expertise (Dynatrace, Splunk, Geneos, Grafana; ITIL Incident/Problem/Change/...Splunk
- ...We are seeking a highly motivated Site Reliability Engineer (SRE) to join the Equity Trading Platform Engineering team, supporting critical trading... ...with monitoring and observability tools such as Splunk, Grafana, Prometheus, Nagios, Datadog, OpenTelemetry, CloudWatch...SplunkPermanent employmentWork at officeAfternoon shift
$139k - $257.55k
...Community CCM organization is seeking an outstanding Senior Site Reliability Engineer (SRE) to support innovation through machine learning,... ...bash, RubyCI/CD: Jenkins, Argo CDObservability: New Relic, Splunk, Grafana, PrometheusAbout AdobeAdobe empowers everyone to create...SplunkFull timeTemporary workLocal areaRemote workWorldwide$189.59k - $220k
Director, Site Reliability EngineeringNBCUniversal is one of the world's... ...NBCUniversal's Production Software Engineering team, responsible for... ...Engineers (Contractors and/or Staff), and third-party vendors.... ...like Grafana, Influx, Graylog/Splunk, Selenium, New Relic, and...SplunkFor contractorsRemote work$145k - $160k
...We are seeking a specialized Observability & Infrastructure Engineer to lead the monitoring, telemetry, and platform-as-code initiatives... ...using OP2 / Vector (Observability Pipelines Worker) alongside Splunk, GCP Pub/Sub sinks, and OpenTelemetry (OTel) AWS-Native...SplunkTemporary workH1bRemote workFlexible hours$86.13k - $127.19k
...sustainable, more inclusive world. Job Role: Job Role: Site Reliability Engineer Location: Boston, MA Duration: Fulltime Summary:... ...alerting distributed systems at scale using Datadog and Splunk. ~ Experience supporting DevOps practices for service delivery...SplunkFull timeLocal area- ...combines Strategy, Experience & Design, Engineering and Managed Services. We build digital... ...existing observability tools (e.g., Datadog, Splunk, Grafana).Tracing: Set up end-to-end... ...6+ years in SRE, platform reliability or observability engineering, with strong...Splunk
- ...The Role:GIPHY is seeking a highly experienced Site Reliability Engineer to join our SRE team. You will help design, build, operate, and evolve the infrastructure that powers GIPHY, including our cloud environment, Kubernetes clusters, and CI/CD platforms.You will also...Full timeWork experience placementRemote work
$207k - $300k
...systems by pushing for changes that improve reliability and velocity.Practice sustainable... ...:Master's degree in Computer Science or Engineering.Experience mentoring engineers and cultivating... ...consensus across cross-functional teams.Site Reliability Engineering (SRE) combines...$208.5k - $216.5k
...including its signal quality and its cost.Drive reliability improvements using SLOs and telemetry... ...To The TableExperience:Demonstrated Staff-level scope: has owned a production... ...years in SRE, DevOps or infrastructure engineering, though scope and impact weigh more than...Full timeTemporary workFlexible hours$235k - $250k
...along the way, come join the Broadridge team.Broadridge is Growing. We are seeking a strategic, hands-on Global Manager of Site Reliability Engineering to lead the reliability, release engineering, and operational evolution of its Integrated Platform. This event-driven...Permanent employmentFull timeLocal areaFlexible hours$120k - $150k
...allows each person to achieve personal success and add value to our teams and communities.We are currently looking for a Site Reliability Engineer to join our Platform Engineering team in New York, NY.About the RoleJoin our Platform Engineering team as aSite Reliability...Full time- ...our clients to succeed in an evolving digital landscape. Role Overview We are seeking an experienced Observability / Site Reliability Engineer (SRE) to design, scale, and maintain our enterprise monitoring and alerting ecosystems. In this role, you will bridge the...
$150k - $170k
...Senior Site Reliability Engineer – Zip CoJoin to apply for the Senior Site Reliability Engineer role at Zip CoAt Zip, we build cloud-native software applications that serve millions of customers and process billions of dollars in payments. We're looking for a seasoned...Casual workWork at officeRemote work$200k - $240k
...problems and help health systems deliver better care, we'd love to meet you! About the role We're looking for a Senior Site Reliability Engineer to join our Infrastructure Engineering team and get their hands directly into the systems that keep our healthcare...Work at office3 days per week- ...Plaid Inc is looking for a Staff Site Reliability Engineer to lead the reliability practices across product engineering. You will architect SLO and error-budget programs, ensuring new products are production-ready while promoting safety gates for high velocity. The...
$120k - $180k
...people, and works with high-profile manufacturers including leaders in space and defense. You will be the first dedicated Site Reliability Engineer and own critical infrastructure end to end. This is a greenfield opportunity to architect the path from AWS to on-premises...Permanent employmentFull timeRelocation package$80k - $95k
...s users, applications, and web-based product offerings. In this role, the Site Reliability Engineer (SRE) will play a key role in maintaining resources at peak efficiency to guarantee staff are able to perform their functions efficiently and to ensure a good customer...Remote workVisa sponsorshipWork visa- ...Job Description Job Description Location: New York, NY, USA Exp: 8-12 Years Client: Amex Job Description: SRE Engineer (This is not a Devops role, strictly need an SRE Engineer, who has great analytical skills and is a good incident manager as well)...
$191k - $226k
...— and using AI to scale that impact further and faster than anyone else can. About the role: We are seeking a Senior Site Reliability Engineer to own the reliability, performance, and resilience of the cloud infrastructure powering Garner's products and AI/ML workloads...Remote workWork visaFlexible hours- ...Job Title: Site Reliability Engineer (SRE) Job Location: New York, NY Job Type: Contract Job description : # Owning infrastructure end-to-end designing, building, and scaling cloud and on[1]prem systems as the sole SRE # Architecting the migration...Full timeContract work
$189k - $283.6k
...the SRE team, you will proactively and reactively improve the reliability of Block's platform and critical infrastructure. You are metrics... ...~ A strong desire to perform and grow as an engineer ~5+ years of software development experience Technologies...Full timeLocal areaRemote workRelocation packageFlexible hoursShift work- ...Job Title WHAT YOU'LL DO DAY-TO-DAY: The engineer will be responsible for developing and implementing new systems and services in the areas of infrastructure monitoring, configuration management, and automation. Additional responsibilities include upgrading and/or...
- ...Software Reliability Engineer Good software has to run where customers need it. For many of Retool's largest customers, that means running Retool in their own infrastructure, behind their own controls, with the reliability and operational clarity they would expect...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Staff Site Reliability Engineer - Splunk. Be the first to apply!
- software engineer staff New York, NY
- technology administrator New York, NY
- assistant engineer New York, NY
- staff data engineer New York, NY
- assistant chief engineer New York, NY
- staff engineer New York, NY
- staff security engineer New York, NY
- assistant engineering manager New York, NY
- senior staff systems engineer New York, NY
- staff design engineer New York, NY



