Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer - Data Infrastructure

TikTok

The Data Infrastructure SRE team is responsible for the reliability, scalability, and efficiency of the core data services that power our products. We manage a massive, distributed environment built on technologies like Kubernetes, Redis, MySQL, and Message Queue. Our work is not about building features, but about engineering the resilience and performance of the underlying platform that all product teams depend on. We are the guardians of production, ensuring our data systems run smoothly, nonstop.This role includes participation in a rotational on-call schedule to ensure nonstop coverage for our critical data infrastructure. You will be expected to respond to, troubleshoot, and resolve production incidents. Our team collaborates across multiple time zones, and you will engage in rigorous change management and post-incident review processes to maintain system stability.Responsibilities:- Incident response and triage: Serve as a first responder for production alerts and incidents. Execute established runbooks to mitigate issues and escalate effectively when necessary.- Operational excellence and change management: Perform routine operational tasks, such as deployments, configuration changes, and system maintenance, following established change control processes to minimize production risk.- Iterative automation and AI augmentation: Identify and automate repetitive manual tasks using scripting (e.g., Python, Go, Bash) and AI Agents to reduce toil, improve operational consistency, and boost overall productivity.- Observability and monitoring: Improve our observability posture by refining monitoring dashboards, tuning alert thresholds, and ensuring that our systems are sufficiently instrumented to detect and diagnose problems.- Data Center and AI Infrastructure: Support the daily operations, construction, and maintenance of data center environments and AI infrastructure to meet the demands of large-scale data processing.Minimum Qualifications:- Bachelor’s degree in Computer Science, a related technical field, or equivalent practical experience.- 2+ years of experience in an SRE, DevOps, Systems Administration, or similar role.- Experience with at least one scripting language (e.g., Python, Bash, Go).- Solid understanding of Linux operating systems and networking concepts.Preferred Qualifications:- Familiarity with container technologies like Docker and Kubernetes.- Hands-on experience with at least one common data store (e.g., MySQL, Redis, PostgreSQL).- Experience with monitoring and observability tools (e.g., Prometheus, Grafana, ELK Stack).- A strong desire to learn, a proactive attitude toward problem-solving, and excellent communication skills.- Experience in the operation and construction of Data Centers is a big plus.Req ID: A32205

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer - Data Infrastructure in Seattle, WA vacancy
  • The Data Infrastructure SRE team is responsible for the reliability, scalability, and efficiency of the core data services that...  ...building features, but about engineering the resilience and performance...  ...maintain system stability.As a Site Reliability Engineer, you will... 
    Suggested

    TikTok

    Seattle, WA
    3 days ago
  • $127k - $249k

     ...looking for an experienced Senior or Staff Engineer for our SRE, InfraSec team, to guide the security of our cloud-based infrastructure. As a Staff SRE, you will be very hands-on...  ...speed of the market. We have redefined the data platform for the AI era, enabling builders... 
    Suggested
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    Seattle, WA
    2 days ago
  •  ...team supports all Big Data services and products...  ...are responsible for the reliability of all the company's...  ..., services, and query engines. We serve business needs...  ...-impacting events.- Infrastructure Automation: Automate infrastructure...  ...related to site reliability and... 
    Suggested

    TikTok

    Seattle, WA
    3 days ago
  • $165k - $225.6k

     ...building the trusted, neutral infrastructure that enables organizations...  ...s talk. THE TECHNOLOGY, DATA AND INTELLIGENCE TEAM...  ...functions to drive scale, reliability, and innovation through technology. THE SENIOR SITE RELIABILITY ENGINEER OPPORTUNITY Reporting to... 
    Suggested
    Permanent employment
    Full time
    Local area
    Worldwide
    Flexible hours

    Okta

    Bellevue, WA
    1 day ago
  • $159.2k - $301.6k

     ...on the cloud. In this reliability-focused role, you will...  ...building on top of AWS cloud infrastructure. You'll partner with the backend engineers building these APIs to...  ...reliable Own database data protection, backup, and...  ...years of experience in site reliability engineering... 
    Suggested
    Full time
    Temporary work
    Local area
    Worldwide

    Adobe Systems

    Seattle, WA
    3 days ago
  • $134.25k - $214.8k

     ...where you matter.Your ImpactAre you an engineer who gets excited about the challenge...  ...instrumenting them, but designing the infrastructure that makes traces, metrics, and logs useful...  ...the Observability team within Axon's Site Reliability organization — a focused team... 
    Work experience placement
    Work at office
    Remote work

    Axon

    Seattle, WA
    3 days ago
  • $166.9k - $203.9k

     ...be building the self-service infrastructure and guardrails that allow...  ...to own their own speed and reliability. You operate across the full...  ...— spanning service design, data access patterns, and delivery...  ...​5+ years in SRE, Platform Engineering, or Backend/Full-Stack engineering... 
    Minimum wage
    Full time
    Night shift
    Weekend work

    Rock central

    Seattle, WA
    2 days ago
  • $143k - $191k

     ...operating system that turns thousands of data streams into a realtime, 3D command and...  ...to our customers. System Deployment Engineers work in complex environments with shared...  ...demonstrations and exercisesWork with site reliability engineers to provide and refine requirements... 
    Full time
    Temporary work
    Work experience placement
    Immediate start

    Anduril Industries

    Seattle, WA
    1 day ago
  • $127k - $249k

    The TeamPlatform Engineering is the department within SRE that is responsible...  ...for a range of critical infrastructure and operational functions...  ...that ensure cluster reliability and security (e.g., CoreDNS,...  ...market. We have redefined the data platform for the AI era, enabling... 
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    Seattle, WA
    4 days ago
  • $170k - $220k

     ...'re Looking ForWe’re looking for a hands-on, high-agency Site Reliability Engineer to help shape and scale the reliability layer of our stack...  ..., weekly deploys, and hotfixes — while also automating infrastructure, monitoring systems, and GitHub workflows. This is a software... 

    Supio

    Seattle, WA
    2 days ago
  • $166k - $258k

     ...Seattle office a minimum of 4 days/week in order to be considered for this position.Nordstrom is looking for a Senior Engineer 2 to join our Site Reliability Engineering (SRE) team — and we think that person could be you.You'll help build the scalable, reliable, and... 
    Full time
    Work at office

    Nordstrom

    Seattle, WA
    9 hours ago
  • $140k - $205k

    Senior Technology Site Reliability EngineerCooley is seeking a Senior Site Reliability Engineer to join the Infrastructure & Development Operations team.Position summary: The Senior Technology...  ...working with advanced ETL data workflows including technologies such as... 
    Full time
    Temporary work
    Work at office
    Flexible hours
    Weekend work

    Cooley

    Seattle, WA
    2 days ago
  •  ...interactions. Working extensively with Windows including patching and certificate provisioning and renewals. Able to map out existing infrastructure flows and dependencies. Leading root cause analysis meetings. Responding to incidents. Understanding of logs and monitoring... 

    Comtech

    Seattle, WA
    4 days ago
  • $210.6k - $305.1k

     ...enterprise network telemetry data, ThousandEyes enables IT teams...  ...of FedRAMP-compliant infrastructure and systems, ensuring excellence...  ...led a distributed team of 5+ engineers, can demonstrate strong technical...  ...Please see the Cisco careers site to discover more benefits and... 
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    Seattle, WA
    4 days ago
  • $194k - $267k

     ...secures AI by building the trusted, neutral infrastructure that enables organizations to safely...  ...technical StaffObservabilitySite Reliability Engineer with a specialty in Splunk to own and...  ...collection, processing, and storage of log data to ensure high reliability and low... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta

    Bellevue, WA
    2 days ago
  • $163.62k - $212.71k

     ...advertising. We deal with BIG data, operating mainly in AWS...  ...processes that improve our engineering teams' productivity and streamline...  ...strategic Lead/Principal Site Reliability Engineer to drive the...  ...available, and reliable cloud infrastructure in AWS leveraging modern... 
    Full time
    Part time
    Work experience placement
    Work at office
    Local area
    Immediate start
    Remote work
    Work from home
    Flexible hours
    Shift work
    3 days per week
    1 day per week

    iSpot.tv

    Bellevue, WA
    1 day ago
  • $194k - $267k

     ...AI by building the trusted, neutral infrastructure that enables organizations to safely...  ...you are too, let's talk.The TeamThe Site Reliability team is dedicated to architecting and...  ...that maximize platform reliability and engineering velocity.The ideal candidate is someone... 
    Local area
    Worldwide
    Flexible hours

    Okta

    Bellevue, WA
    9 hours ago
  • $194k - $267k

     ...AI by building the trusted, neutral infrastructure that enables organizations to safely...  ...concepts and tools.Position Overview:The Site Reliability Engineer (SRE) will play a key role in...  .../or have access to protected federal data. As a condition of employment for this... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta

    Bellevue, WA
    4 days ago
  •  ...the world's most complex and mission-critical systems.As a Site Reliability Engineer III - DevOps Engineer at JPMorgan Chase within the...  ...simple and straightforward solutions. Through code and cloud infrastructure, you will configure, maintain, monitor, and optimize applications... 

    JP Morgan Chase

    Seattle, WA
    1 day ago
  •  ...Site Reliability Engineer Comtech is a woman-owned small business founded in 1998 and headquartered in Reston, VA....  ...project management, applications development, infrastructure, Cyber security, and enterprise content/data management services. We have developed our methodologies... 

    Comtech LLC

    Seattle, WA
    1 day ago
  •  ...Lululemon Site Reliability Engineering Engineer We are a yoga-inspired technical apparel company...  ...with internal investment work using data to inform your decisions Identify...  ...gaps in our observability tooling and infrastructure and recommend and implement appropriate... 

    Samprasoft

    Seattle, WA
    2 days ago
  • Technical/Functional Skills Windows Servers, Digital: Microsoft Azure Windows Powershell, Digital: DevOps Roles & Responsibilities Windows Server 2012 -2019 Administration Microsoft Azure Azure AAD DFSR, DHCP DNS, KMS, WSUS TCP/IP Hyper-V High Availability Clusters ...

    The Dignify Solutions, LLC

    Bellevue, WA
    13 hours ago
  •  ...Principal Site Reliability Engineer (IC4) As a Principal Site Reliability Engineer (IC4), you will...  ...combine software engineering with infrastructure expertise to improve service reliability...  ...optimize operational processes using data-driven insights. Incident... 

    Oracle

    Seattle, WA
    4 days ago
  • $140k - $200k

     ...office. These include frontend and backend engineers, AI research scientists, and others from...  ...Overview We're looking to hire for our Data side of our AI team at Speechify. This...  ...low cost through a tight integration of infrastructure, engineering, and research work. We are... 
    Remote job
    Full time
    Work at office
    Shift work

    Speechify

    Bellevue, WA
    13 hours ago
  • $109k - $160k

     ...CoreWeave combines superior infrastructure performance with deep technical...  .... About The Role: The Data Platforms Team serves as the...  ...We are seeking a senior engineer with specialization in database...  ...the performance, security, reliability, and scalability of our data... 
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Remote work
    Flexible hours

    Coreweave

    Bellevue, WA
    13 hours ago
  • $172.5k - $313.7k

     ...the TeamSlack is seeking experienced engineers to join its Core Infrastructure organization — the team responsible...  ...driving our systems toward greater reliability, performance, scalability, and...  ...pipelines that process and transform data for Slack's search infrastructure.Partner... 
    Full time

    Salesforce

    Seattle, WA
    2 days ago
  • $192k - $240k

     ...resources, and support you need to grow your career.Data at BrexOur Scientists and Engineers work together to make data — and insights derived from...  ...just crunching numbers. The Data team at Brex develops infrastructure, statistical models, and products using financial... 
    Work at office
    3 days per week

    Brex

    Seattle, WA
    3 days ago
  • About the TeamDoorDash is a data driven organization and relies on timely, accurate and reliable data to drive many business and...  ...organization owns all the infrastructure necessary to run an operationally...  ...simplify data workflows for engineers, analysts, and ML... 
    Hourly pay
    Work at office
    Local area
    Remote work
    Relocation
    Flexible hours

    Doordash

    Seattle, WA
    2 days ago
  • $197.3k - $313.7k

     ...future of Salesforce.Own the data and intelligence...  ...Services team owns the data infrastructure, telemetry pipelines, marketplace...  ...technical direction every data engineer and ML practitioner on the team...  ...engineering team, and pipeline reliability SLOs.Self-service analytics.... 
    Full time
    Contract work
    Live in
    Work at office

    Salesforce

    Seattle, WA
    1 day ago
  • $192k - $240k

     ...you need to grow your career.Engineering at BrexEngineering at Brex...  ...intention. Our teams span Software, Data, Security, and IT, and...  ...data, query, and search infrastructure that powers critical Brex product...  ...customizable, scalable, and reliable for finance teams, and that... 
    Work at office
    Remote work
    Work from home

    Brex

    Seattle, WA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer - Data Infrastructure. Be the first to apply!