Site Reliability Engineer - Data Infrastructure
TikTok
The Data Infrastructure SRE team is responsible for the reliability, scalability, and efficiency of the core data services that power our products. We manage a massive, distributed environment built on technologies like Kubernetes, Redis, MySQL, and Message Queue. Our work is not about building features, but about engineering the resilience and performance of the underlying platform that all product teams depend on. We are the guardians of production, ensuring our data systems run smoothly, nonstop.This role includes participation in a rotational on-call schedule to ensure nonstop coverage for our critical data infrastructure. You will be expected to respond to, troubleshoot, and resolve production incidents. Our team collaborates across multiple time zones, and you will engage in rigorous change management and post-incident review processes to maintain system stability.Responsibilities:- Incident response and triage: Serve as a first responder for production alerts and incidents. Execute established runbooks to mitigate issues and escalate effectively when necessary.- Operational excellence and change management: Perform routine operational tasks, such as deployments, configuration changes, and system maintenance, following established change control processes to minimize production risk.- Iterative automation and AI augmentation: Identify and automate repetitive manual tasks using scripting (e.g., Python, Go, Bash) and AI Agents to reduce toil, improve operational consistency, and boost overall productivity.- Observability and monitoring: Improve our observability posture by refining monitoring dashboards, tuning alert thresholds, and ensuring that our systems are sufficiently instrumented to detect and diagnose problems.- Data Center and AI Infrastructure: Support the daily operations, construction, and maintenance of data center environments and AI infrastructure to meet the demands of large-scale data processing.Minimum Qualifications:- Bachelor’s degree in Computer Science, a related technical field, or equivalent practical experience.- 2+ years of experience in an SRE, DevOps, Systems Administration, or similar role.- Experience with at least one scripting language (e.g., Python, Bash, Go).- Solid understanding of Linux operating systems and networking concepts.Preferred Qualifications:- Familiarity with container technologies like Docker and Kubernetes.- Hands-on experience with at least one common data store (e.g., MySQL, Redis, PostgreSQL).- Experience with monitoring and observability tools (e.g., Prometheus, Grafana, ELK Stack).- A strong desire to learn, a proactive attitude toward problem-solving, and excellent communication skills.- Experience in the operation and construction of Data Centers is a big plus.Req ID: A32205
- The Data Infrastructure SRE team is responsible for the reliability, scalability, and efficiency of the core data services that... ...building features, but about engineering the resilience and performance... ...maintain system stability.As a Site Reliability Engineer, you will...Suggested
$127k - $249k
...looking for an experienced Senior or Staff Engineer for our SRE, InfraSec team, to guide the security of our cloud-based infrastructure. As a Staff SRE, you will be very hands-on... ...speed of the market. We have redefined the data platform for the AI era, enabling builders...SuggestedLocal areaRemote workWorldwideFlexible hours- ...team supports all Big Data services and products... ...are responsible for the reliability of all the company's... ..., services, and query engines. We serve business needs... ...-impacting events.- Infrastructure Automation: Automate infrastructure... ...related to site reliability and...Suggested
$165k - $225.6k
...building the trusted, neutral infrastructure that enables organizations... ...s talk. THE TECHNOLOGY, DATA AND INTELLIGENCE TEAM... ...functions to drive scale, reliability, and innovation through technology. THE SENIOR SITE RELIABILITY ENGINEER OPPORTUNITY Reporting to...SuggestedPermanent employmentFull timeLocal areaWorldwideFlexible hours$159.2k - $301.6k
...on the cloud. In this reliability-focused role, you will... ...building on top of AWS cloud infrastructure. You'll partner with the backend engineers building these APIs to... ...reliable Own database data protection, backup, and... ...years of experience in site reliability engineering...SuggestedFull timeTemporary workLocal areaWorldwide$134.25k - $214.8k
...where you matter.Your ImpactAre you an engineer who gets excited about the challenge... ...instrumenting them, but designing the infrastructure that makes traces, metrics, and logs useful... ...the Observability team within Axon's Site Reliability organization — a focused team...Work experience placementWork at officeRemote work$166.9k - $203.9k
...be building the self-service infrastructure and guardrails that allow... ...to own their own speed and reliability. You operate across the full... ...— spanning service design, data access patterns, and delivery... ...5+ years in SRE, Platform Engineering, or Backend/Full-Stack engineering...Minimum wageFull timeNight shiftWeekend work$143k - $191k
...operating system that turns thousands of data streams into a realtime, 3D command and... ...to our customers. System Deployment Engineers work in complex environments with shared... ...demonstrations and exercisesWork with site reliability engineers to provide and refine requirements...Full timeTemporary workWork experience placementImmediate start$127k - $249k
The TeamPlatform Engineering is the department within SRE that is responsible... ...for a range of critical infrastructure and operational functions... ...that ensure cluster reliability and security (e.g., CoreDNS,... ...market. We have redefined the data platform for the AI era, enabling...Work at officeLocal areaRemote workWorldwideFlexible hours$170k - $220k
...'re Looking ForWe’re looking for a hands-on, high-agency Site Reliability Engineer to help shape and scale the reliability layer of our stack... ..., weekly deploys, and hotfixes — while also automating infrastructure, monitoring systems, and GitHub workflows. This is a software...$166k - $258k
...Seattle office a minimum of 4 days/week in order to be considered for this position.Nordstrom is looking for a Senior Engineer 2 to join our Site Reliability Engineering (SRE) team — and we think that person could be you.You'll help build the scalable, reliable, and...Full timeWork at office$140k - $205k
Senior Technology Site Reliability EngineerCooley is seeking a Senior Site Reliability Engineer to join the Infrastructure & Development Operations team.Position summary: The Senior Technology... ...working with advanced ETL data workflows including technologies such as...Full timeTemporary workWork at officeFlexible hoursWeekend work- ...interactions. Working extensively with Windows including patching and certificate provisioning and renewals. Able to map out existing infrastructure flows and dependencies. Leading root cause analysis meetings. Responding to incidents. Understanding of logs and monitoring...
$210.6k - $305.1k
...enterprise network telemetry data, ThousandEyes enables IT teams... ...of FedRAMP-compliant infrastructure and systems, ensuring excellence... ...led a distributed team of 5+ engineers, can demonstrate strong technical... ...Please see the Cisco careers site to discover more benefits and...Full timeTemporary workLocal areaFlexible hours$194k - $267k
...secures AI by building the trusted, neutral infrastructure that enables organizations to safely... ...technical StaffObservabilitySite Reliability Engineer with a specialty in Splunk to own and... ...collection, processing, and storage of log data to ensure high reliability and low...Permanent employmentWork at officeLocal areaWorldwideFlexible hours$163.62k - $212.71k
...advertising. We deal with BIG data, operating mainly in AWS... ...processes that improve our engineering teams' productivity and streamline... ...strategic Lead/Principal Site Reliability Engineer to drive the... ...available, and reliable cloud infrastructure in AWS leveraging modern...Full timePart timeWork experience placementWork at officeLocal areaImmediate startRemote workWork from homeFlexible hoursShift work3 days per week1 day per week$194k - $267k
...AI by building the trusted, neutral infrastructure that enables organizations to safely... ...you are too, let's talk.The TeamThe Site Reliability team is dedicated to architecting and... ...that maximize platform reliability and engineering velocity.The ideal candidate is someone...Local areaWorldwideFlexible hours$194k - $267k
...AI by building the trusted, neutral infrastructure that enables organizations to safely... ...concepts and tools.Position Overview:The Site Reliability Engineer (SRE) will play a key role in... .../or have access to protected federal data. As a condition of employment for this...Permanent employmentWork at officeLocal areaWorldwideFlexible hours- ...the world's most complex and mission-critical systems.As a Site Reliability Engineer III - DevOps Engineer at JPMorgan Chase within the... ...simple and straightforward solutions. Through code and cloud infrastructure, you will configure, maintain, monitor, and optimize applications...
- ...Site Reliability Engineer Comtech is a woman-owned small business founded in 1998 and headquartered in Reston, VA.... ...project management, applications development, infrastructure, Cyber security, and enterprise content/data management services. We have developed our methodologies...
- ...Lululemon Site Reliability Engineering Engineer We are a yoga-inspired technical apparel company... ...with internal investment work using data to inform your decisions Identify... ...gaps in our observability tooling and infrastructure and recommend and implement appropriate...
- Technical/Functional Skills Windows Servers, Digital: Microsoft Azure Windows Powershell, Digital: DevOps Roles & Responsibilities Windows Server 2012 -2019 Administration Microsoft Azure Azure AAD DFSR, DHCP DNS, KMS, WSUS TCP/IP Hyper-V High Availability Clusters ...
- ...Principal Site Reliability Engineer (IC4) As a Principal Site Reliability Engineer (IC4), you will... ...combine software engineering with infrastructure expertise to improve service reliability... ...optimize operational processes using data-driven insights. Incident...
$140k - $200k
...office. These include frontend and backend engineers, AI research scientists, and others from... ...Overview We're looking to hire for our Data side of our AI team at Speechify. This... ...low cost through a tight integration of infrastructure, engineering, and research work. We are...Remote jobFull timeWork at officeShift work$109k - $160k
...CoreWeave combines superior infrastructure performance with deep technical... .... About The Role: The Data Platforms Team serves as the... ...We are seeking a senior engineer with specialization in database... ...the performance, security, reliability, and scalability of our data...Permanent employmentFull timeTemporary workCasual workWork at officeRemote workFlexible hours$172.5k - $313.7k
...the TeamSlack is seeking experienced engineers to join its Core Infrastructure organization — the team responsible... ...driving our systems toward greater reliability, performance, scalability, and... ...pipelines that process and transform data for Slack's search infrastructure.Partner...Full time$192k - $240k
...resources, and support you need to grow your career.Data at BrexOur Scientists and Engineers work together to make data — and insights derived from... ...just crunching numbers. The Data team at Brex develops infrastructure, statistical models, and products using financial...Work at office3 days per week- About the TeamDoorDash is a data driven organization and relies on timely, accurate and reliable data to drive many business and... ...organization owns all the infrastructure necessary to run an operationally... ...simplify data workflows for engineers, analysts, and ML...Hourly payWork at officeLocal areaRemote workRelocationFlexible hours
$197.3k - $313.7k
...future of Salesforce.Own the data and intelligence... ...Services team owns the data infrastructure, telemetry pipelines, marketplace... ...technical direction every data engineer and ML practitioner on the team... ...engineering team, and pipeline reliability SLOs.Self-service analytics....Full timeContract workLive inWork at office$192k - $240k
...you need to grow your career.Engineering at BrexEngineering at Brex... ...intention. Our teams span Software, Data, Security, and IT, and... ...data, query, and search infrastructure that powers critical Brex product... ...customizable, scalable, and reliable for finance teams, and that...Work at officeRemote workWork from home
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer - Data Infrastructure. Be the first to apply!
- site reliability engineer Seattle, WA
- software data engineer Seattle, WA
- entry level big data engineer Seattle, WA
- big data developer Seattle, WA
- senior data quality engineer Seattle, WA
- sr data engineer Seattle, WA
- junior big data engineer Seattle, WA
- big data cloud engineer Seattle, WA
- data platform engineer Seattle, WA
- senior cloud data engineer Seattle, WA


