Senior Site Reliability Engineer - Observability
Dimensional Fund Advisors
Job Description:About the Role: We are looking for a Senior SRE to join our Platform Engineering team as the operations owner of our observability platforms. You’ll be responsible for the reliability, scalability, and continued evolution of the tools that give our engineering organization visibility into everything they build and run. The current observability platform is primarily comprised of on-premises ELK (Elasticsearch, Logstash, Kibana) Stack and Grafana, with some exposure to New Relic and SolarWinds. This is a hybrid role: roughly half your time will be spent on steady-state operations and platform support, and the other half on engineering projects that meaningfully advance the platforms you support. It’s a great fit for someone who is genuinely motivated by the pursuit of excellence – not just sustaining what works but relentlessly refining it. You take pride in the platforms you own, and that pride drives you to keep improving them, whether that means tightening an SLO, eliminating a source of toil, or building something that gives teams faster insight into their systems. What You’ll Work On: Operations & Reliability (~ 50%) Serve as a primary escalation point for production support involving the ELK Stack, Grafana, and New Relic Own platform health, capacity planning, and performance tuning for on-premises observability infrastructure – including Elasticsearch cluster management, index lifecycle policies, and retention strategies Monitor and maintain SLOs for the observability platforms, ensuring the tools engineers depend on are highly available and performant Support engineering teams in onboarding to observability platforms – helping teams instrument their applications, build dashboards, and define meaningful alerts Manage patching, upgrades, and configuration management across the observability stack Collaborate with security to harden platform configurations and manage software vulnerabilities Contribute to on-call rotations and maintain runbooks and escalation procedures Platform Engineering (~ 50%) Design and build tooling/automation to reduce toil and improve the experience for teams using observability platforms Lead or contribute to platform modernization initiatives – e.g., improving ingestion pipelines, scaling platform capacity, standardizing Grafana dashboard and alerting patterns, or evaluating new capabilities within the existing stack Develop and maintain infrastructure-as-code (Terraform, Helm, Ansible, etc.) for platform components Build and enforce standards around logging metrics and alerting that help engineering teams adopt observability best practices at scale Participate in design reviews and contribute to the overall platform roadmap What We’re Looking For: Bachelor’s degree in a technical field or equivalent practical experience 5+ years of experience in SRE, DevOps, or platform engineering roles Deep hands-on experience with the ELK Stack – Elasticsearch cluster operations, Logstash pipeline development, Kibana, and index lifecycle management Strong experience with Grafana, including data source integrations, dashboard design, and alerting Solid understanding of observability principles Experience operating on-premises infrastructure, including capacity planning, server management, and the operational tradeoffs with managed cloud services Proficiency in Python for automation and tooling; familiarity with shell scripting Strong Linux systems knowledge and comfort working with configuration management tools (e.g., Ansible, Chef, Puppet, etc.) Demonstrated ability to drive incidents to resolution and communicate clearly under pressure A bias toward automation and a low tolerance for repetitive manual work Nice to Have: Experience with Prometheus Experience with New Relic administration or APM instrumentation Familiarity with log shipping agents and pipeline tools such as Beats, Fluentd, or Fluent Bit Experience with distributed tracing tools like OpenTelemetry Exposure to cloud-based observability offerings and experience thinking through hybrid strategies Prior experience building or governing observability standards across a large engineering organization #LI-HybridDimensional offers a variety of programs to help take care of you, your family, and your career, including comprehensive benefits, educational initiatives, and special celebrations of our history, culture, and growth.It is the policy of the Company to provide equal opportunity for all employees and applicants. The Company recruits, hires, trains, promotes, compensates, and administers all personnel actions without regard to actual or perceived race, color, religion, religious practice, creed, sex, sex stereotyping, pregnancy (which includes pregnancy, childbirth, and medical conditions related to pregnancy, childbirth, or breastfeeding), caregiver status, gender, gender identity, gender expression, transgender identity, national origin, age, mental or physical disability, ancestry, medical condition, marital status, familial status, domestic partnership status, military or veteran status or service, unemployment status, citizenship status or alienage, sexual orientation, status as a victim of domestic violence, status as a victim of stalking, status as a victim of sex offenses, genetic information, political activities or recreational activities, arrest or conviction record, salary history, natural hairstyle or any other status protected by applicable law except as otherwise required or permitted by law or regulation applicable to the Company or its affiliates. SummaryLocation: Austin; CharlotteType: Full time
- ...that seriously!The RoleThe Senior SRE at 2K is a hands-on technical... ...partnering with network engineers, systems architects, and game... ...direction, influencing reliability from architecture review through... ...across game service deployments.Observability & ReliabilityBuild and run...Senior
$127k - $249k
The TeamPlatform Engineering is the department within SRE that is responsible for a range... ...edge and internal service mesh), and observability and alerting systems.The Fleet Management... ...components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager...SeniorWork at officeLocal areaRemote workWorldwideFlexible hours$152k - $241.5k
...intelligence.We’re looking for a Senior SRE to join our Compute Farm team... ...host lifecycle management, fleet reliability/auto-healing, E2E observability or data-driven operations (AIOps/ML... ...Go, Perl, or Ruby.Mentored other engineers and influenced technical direction...SeniorFull time$165k - $241.4k
...Security, Collaboration, and Observability portfolios Your ImpactThe... ....We’re looking for talented engineers with a software or operations... ...development teams to ensure the reliability, performance and security of... ...see the Cisco careers site to discover more benefits and...SeniorFull timeTemporary workWork at officeLocal areaFlexible hours1 day per week$184k - $287.5k
...Infrastructure organization is seeking a Senior System Software Engineer to lead the evolution of our next-generation Data & Observability Platform. We serve and collaborate... ...distributed pipelines, and ensure platform reliability.What you’ll be doing:Architect High-Performance...SeniorFull time- ...are an integrated product, engineering, strategy and risk team, all... ...we serve our clients. As a Senior Engineer on AI.x, you will... ...technology today.As a Senior AI Site Reliability Engineer you will support... ...implement comprehensive observability frameworks to minimize MTTD...SeniorFull time
$127k - $249k
We are looking for an experienced Senior or Staff Engineer for our SRE, InfraSec team, to guide the security of our cloud-based infrastructure... ...-focused areas, such as runtime scanning, security observability, CSPM, and moreCloud Expertise: Strong experience with at...SeniorLocal areaRemote workWorldwideFlexible hours$110.7k - $171.8k
...Participation in oncall rotation as a platform reliability escalation point Incident... ...requirements. Collaborate with engineering teams across the organization to influence... ...improvements. ~ Experience with observability concepts and practices (monitoring, logging...SeniorWork experience placementWork at officeLocal area$185k - $227k
...read on for more details. ROLE AND RESPONSIBILITIES: A Senior Site Reliability Engineer (SRE) is expected to own the operational stability and... ...workloads using EKS, GKE, ECS, Cloud Run with service mesh,observability, and security best practices Implement disaster...SeniorRemote work$178.42k - $230.5k
...maintaining the tools and services engineers here at GM use every day to... ...RoleWe are looking for a Senior Engineer with an extensive... ...delivering impact through observability frameworks and will evolve depending... ...and maintaining a robust, reliable, and cohesive inner loop...SeniorFull timeWork experience placementLocal areaWork from homeRelocation packageFlexible hours- ...selected candidate for this role to work on site in the specified location(s).We are seeking a Kafka Site Reliability Engineer to help build, operate, and continuously... ...operational excellence through automation, observability, Infrastructure as Code, and AIOps-driven capabilities...SeniorFull timeWork at office
- Recognized as the No. 1 site trusted by real estate professionals, Realtor.com... ...expert guidance.We are seeking a Senior Site Reliability Engineer to join our newly formed Operations... ...will contribute to the reliability, observability, and operational excellence of our platform...SeniorWork at officeLocal area
- ...the rapidly evolving state of the art to engineer scalable, innovative, and research... ...business. About the Role:We are looking for a Senior SRE to serve as the operations owner for... ...tooling ecosystemsOwn the operational reliability of developer tooling ecosystems, including...SeniorFull timeLocal area
$81.1k - $187k
.... You’ll partner with customer support, service owners, and engineering teams around the globe to ensure high-quality service for customers... ...posted.Career Level - IC3Escalation points for junior site reliability engineers during complex or high-impact incidents.Manage and...SeniorTemporary workMonday to FridayFlexible hoursShift workNight shift$127k - $249k
...Eastern or Central time zones. We are looking for an experienced Senior Engineer for our SRE, Atlas team to support, maintain and grow the... ...crucial workloads. Role OverviewWe are seeking a talented Site Reliability Engineer (SRE) with a strong infrastructure background....SeniorLocal areaRemote workWorldwideFlexible hours$98.58k - $138.02k
...Silicon Valley Region / Denver, COProduct Engineering - DevOps /Full Time /HybridRestaurant... ..., TX; Irvine, CA; or Akron, OH. The Site Reliability Engineer II will be responsible for... ...tools and platforms to improve observability. Promote and apply best practices for...Full timeWork at office$100k - $125k
...to join our team and make an impact?ResponsibilitiesAs a Site Reliability Engineer at TeamViewer, you’ll be a key player in ensuring the reliability... ...(e.g. MS SQL, Postgres).Experience with monitoring and observability tools (e.g. Datadog, Grafana, Prometheus).Understanding...Temporary workCasual workWorldwide$167.18k - $203.61k
...DescriptionCox Automotive Corporate Services, LLCLEAD SITE RELIABILITY ENGINEERJob Description: Lead Site Reliability Engineer positions offered by Cox Automotive Corporate... ...performance testing. Develop and maintain observability practices, including SLI, SLO, SLA, error...Full timeWork at officeRemote workFlexible hours$180k - $230k
...services and partner integrations that power reliable money movement across the region. This... ...Strong product mindset, connecting engineering decisions to customer outcomes, payout... ...software, predictable delivery, effective observability, and thoughtful operational ownership....SeniorFull time- ...METRIX IT SOLUTIONS INC is seeking a senior database engineer to design, deploy, and manage multi-region CockroachDB clusters in production... ...global deployments. You will monitor cluster health using observability tools, troubleshoot complex issues, and automate...
$198k - $247.5k
...transformation for the world’s largest enterprises.About Forward Deployed Engineering (FDE) at ClouderaCloudera’s Forward Deployed Engineering (FDE... ...of LLMOps/MLOps practices including evaluation, observability, serving, monitoring, and lifecycle managementExperience designing...SeniorFull timeWork from homeRelocation$131.6k - $210.3k
...a highly skilled and experiencedStaff Site Reliability Engineerto join our DevOps squad. This... ...technologies, as well as the ability to mentor engineers and contribute to the overall... ...Istio. * Proficiency in monitoring and observability tools like Grafana, Grafana Loki,...Work experience placementWork at officeLocal areaRemote work$184k - $287.5k
NVIDIA's Infrastructure Specialists team is hiring a Senior Solutions Architect - AI Factory Observability & Visualization! This remote role develops full-... ...equivalent experience in Computer Science, Mathematics, Engineering, Physics, or related field.6+ years of experience...SeniorFull timeRemote work$190k - $215k
...meaningful connections every day.As a Senior Software Engineer on our AI & Intelligent Systems team,... ...while ensuring they remain reliable, performant, and easy to evolve.This... ...cloud-native systems that are reliable, observable, secure, and resilient while balancing...SeniorPermanent employmentLive in$224k - $356.5k
...limits (memory bandwidth, FLOP/s, interconnect throughput) to observed inference throughput, latency, and utilizationOwn the model validation... ...we need to see:BS, MS, or PhD in Computer Science, Computer Engineering, Electrical Engineering, or equivalent experience.12+ years of...SeniorFull timeLocal area$119.2k - $175.45k
...business analytics platform engineering team is at the forefront of... ...everyone across the company.As a Senior Systems Engineer, you'll... ...develop scalable, secure, and reliable analytics platforms on Azure... ...mitigation.Establish robust observability solutions, implement...SeniorFull timeH1bLocal areaWork from homeRelocation packageFlexible hours- ...for this role to work on site in the specified location(... ...organization is seeking a Senior Kafka Platform Engineer with experience to lead the... ...enhance scalability, security, reliability, performance, stability,... ...for testing, reliability, observability, and operational readiness...SeniorFull timeWork at office
$92.5k - $209.5k
As a Senior Software Engineer on the Identity and Access Management (IAM) team, you’ll design, build, and maintain secure, reliable, and scalable cloud services that power our global authentication... ...Azure, or GCP).Experience with observability tools (metrics, logging, tracing...SeniorTemporary workFlexible hours$170k - $230k
Job DescriptionWe are hiring a Senior Platform Engineer to join the Autonomous Vehicle (AV) Cloud Engineering... ...and is motivated by building a reliable, easy-to-consume platform that helps... ..., strong service ownership, observability, and SLIs/SLOs. Take ownership of complex...SeniorFull timeWork experience placementWork at officeLocal areaWork from homeFlexible hours- ...DescriptionThe Role:General Motors is seeking a Senior Software Engineer to support, design, and improve... ...environments, strengthening system reliability, enabling scalable integrations, and... ...reliability, maintainability, observability, and supportabilityApply modern engineering...SeniorFull timeLocal areaWork from homeRelocation package
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Site Reliability Engineer - Observability. Be the first to apply!
- site reliability engineer Austin, TX
- site reliability engineer sre Austin, TX
- senior manufacturing manager Austin, TX
- senior business analyst Austin, TX
- senior cost estimator Austin, TX
- senior manager tax Austin, TX
- senior devops Austin, TX
- senior recruiter Austin, TX
- senior property manager Austin, TX
- senior paralegal Austin, TX

