Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Lead Site Reliability Engineer

$198.36k - $416.1k

TikTok USDS Joint Venture

Responsibilities

Team Intro

TikTok video system is a world-leading video platform that provides multimedia storage, delivery, transcoding services. As part of the USDS, the Video Platform team is responsible for building the next generation video processing platform which provides excellent experiences for billions of users around the world.

Responsibilities
  • Provide technical leadership and mentorship to a team of Site Reliability Engineers focused on building observable, fault-tolerant video services with high availability
  • Drive architectural decisions for large-scale, globally distributed service mesh architectures
  • Establish and maintain production ownership models, incident response protocols, and service level agreements
  • Develop strategic roadmaps for observability and automation initiatives that enhance system reliability
  • Balance technical contributions with people management responsibilities, including career development, performance evaluations, and team growth
  • Foster a culture of trust, continuous learning and improvement, and knowledge sharing within your team and across the organization
  • Lead security initiatives to safeguard critical assets, partnering with security and compliance teams to implement robust protocols and enforcement that ensure data protection and regulatory compliance across all services
Qualifications

Minimum Qualifications:

  • 5+ years of experience and expertise in designing, analyzing, and troubleshooting large-scale distributed systems with various databases, caching solutions and other web service components
  • Experience in running high-availability web services at massive scale, with comprehensive knowledge of cloud-native architectures and advanced networking concepts
  • Experience in design/development or SRE experience in video streaming or related services
  • Previous experience leading a small to mid-size team while maintaining significant "hands-on" technical contributions
  • Strong understanding of Unix/Linux operating systems internals and networking fundamentals
  • Proficiency in writing production-grade code in Go, Python, Java or similar languages
  • Proven track record of establishing and implementing SRE best practices across engineering organizations
Preferred Qualifications
  • Previous experience in design/development or SRE in video processing applications
  • Experience in large scale Kubernetes systems
  • Deep expertise in algorithms, data structures, and systems design with proven ability to architect complex technical solutions
About USDS

TikTok USDS Joint Venture LLC is dedicated to the safety and security of millions of Americans who create, discover, and connect with what they love on the apps we operate. The Joint Venture has been established in compliance with the Executive Order signed by President Trump on September 25, 2025. Our foundation is a comprehensive data privacy and cybersecurity program we operate under defined safeguards to protect national security and secure U.S. user data, apps and the algorithm. We safeguard the U.S. content ecosystem, holding decision‑making authority for trust and safety policies and moderation. USDS Joint Venture helps ensure Americans can continue to express their creativity, discover new hobbies and interests, and build thriving communities and businesses on a global scale.

On‑site presence across teams allows the company to operate with greater speed, alignment, and agility — especially in areas like real‑time decision‑making, team development, and integrated execution. As such, the company is shifting from a hybrid work model to a fully in‑person schedule up to 5 days a week.

Why Join Us

Inspiring creativity is at the core of TikTok's mission. Our innovative product is built to help people authentically express themselves, discover and connect – and our global, diverse teams make that possible. Together, we create value for our communities, inspire creativity and bring joy - a mission we work towards every day.

We strive to do great things with great people. We lead with curiosity, humility, and a desire to make impact in a rapidly growing tech company. Every challenge is an opportunity to learn and innovate as one team. We're resilient and embrace challenges as they come. By constantly iterating and fostering an "Always Day 1" mindset, we achieve meaningful breakthroughs for ourselves, our company, and our users. When we create and grow together, the possibilities are limitless. Join us.

Diversity & Inclusion

TikTok is committed to creating an inclusive space where employees are valued for their skills, experiences, and unique perspectives. Our platform connects people from across the globe and so does our workplace. At TikTok, our mission is to inspire creativity and bring joy. To achieve that goal, we are committed to celebrating our diverse voices and to creating an environment that reflects the many communities we reach. We are passionate about this and hope you are too.

USDS Reasonable Accommodation

USDS is committed to providing reasonable accommodations in our recruitment processes for candidates with disabilities, pregnancy, sincerely held religious beliefs or other reasons protected by applicable laws. If you need assistance or a reasonable accommodation, please reach out to us at

Job Information

【For Pay Transparency】Compensation Description (Annually)

The base salary range for this position in the selected city is $198360 - $416100 annually.

Compensation may vary outside of this range depending on a number of factors, including a candidate’s qualifications, skills, competencies and experience, and location. Base pay is one part of the Total Package that is provided to compensate and recognize employees for their work, and this role may be eligible for additional discretionary bonuses/incentives, and restricted stock units.

Employees have day one access to medical, dental, and vision insurance, a 401(k) savings plan with company match, paid parental leave, short‑term and long‑term disability coverage, life insurance, wellbeing benefits, among others.

Employees also receive 10 paid holidays per year, 10 paid sick days per year and 17 days of Paid Personal Time (prorated upon hire with increasing accruals by tenure).

The Company reserves the right to modify or change these benefits programs at any time, with or without notice.

For Los Angeles County (unincorporated) Candidates:

  1. Interacting and occasionally having unsupervised contact with internal/external clients and/or colleagues;
  2. Appropriately handling and managing confidential information including proprietary and trade secret information and access to information technology systems;
  3. Exercising sound judgment.
#J-18808-Ljbffr
Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Lead Site Reliability Engineer in Seattle, WA vacancy
  • $167.7k - $245.2k

     ...also delivering AI-powered assurance insights within Cisco’s leading Networking, Security, Collaboration, and Observability...  ...maintaining our FedRAMP offering. Your ImpactAs a FedRAMP Site Reliability Engineer(SRE), you will lead the operations and architecture of our... 
    Suggested
    Full time
    Temporary work
    Work at office
    Local area
    Flexible hours
    2 days per week

    CISCO Systems

    Seattle, WA
    4 days ago
  • $134.25k - $214.8k

     ...matters at a company where you matter.Your ImpactAre you an engineer who gets excited about the challenge of making complex distributed...  ...it.You will be part of the Observability team within Axon's Site Reliability organization — a focused team responsible for Axon's metrics,... 
    Suggested
    Work experience placement
    Work at office
    Remote work

    Axon

    Seattle, WA
    4 days ago
  • $134.25k - $214.8k

     ...change. Constantly grow as you work hard for a mission that matters at a company where you matter.Your ImpactAs a Senior Site Reliability Engineer within the APX SRE organization, you’ll focus on delivering practical, scalable solutions to support the reliability and performance... 
    Suggested
    Work experience placement
    Work at office
    Remote work
    Flexible hours

    Axon

    Seattle, WA
    3 days ago
  • $170k - $220k

     ...We're Looking ForWe’re looking for a hands-on, high-agency Site Reliability Engineer to help shape and scale the reliability layer of our stack....  ...shipping process.You’ll work closely with engineers, product leads, and company leadership to ensure uptime, speed, and... 
    Suggested

    Supio

    Seattle, WA
    3 days ago
  •  ...Windows including patching and certificate provisioning and renewals. Able to map out existing infrastructure flows and dependencies. Leading root cause analysis meetings. Responding to incidents. Understanding of logs and monitoring tools (Splunk, Sumo Logic, New Relic,... 
    Suggested

    Comtech

    Seattle, WA
    12 hours ago
  • $147k - $202.4k

     ...let's talk.Our company is seeking a highly skilled Senior Site Reliability Engineer to join our team. We are a SaaS company specializing in securing...  ...and be a primary responder for critical incidents, leading root cause analysis and implementing preventative measures... 
    Work at office
    Local area
    Worldwide
    Flexible hours
    Shift work

    Okta

    Bellevue, WA
    4 days ago
  • $165k - $225.6k

     ...From core infrastructure to enterprise platforms, we partner across functions to drive scale, reliability, and innovation through technology.The Senior Site Reliability Engineer OpportunityReporting to the Manager, Site Reliability Engineering, this role will help build,... 
    Permanent employment
    Local area
    Worldwide
    Flexible hours

    Okta

    Bellevue, WA
    3 days ago
  • $127k - $249k

    The TeamPlatform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions...  ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper).... 
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    Seattle, WA
    12 hours ago
  • $143k - $191k

     ...mission critical capabilities to our customers. System Deployment Engineers work in complex environments with shared environmental...  ...customersParticipate in customer demonstrations and exercisesWork with site reliability engineers to provide and refine requirements for tooling and... 
    Full time
    Temporary work
    Work experience placement
    Immediate start

    Anduril Industries

    Seattle, WA
    3 days ago
  • $194k - $267k

     ...Overview:We are seeking a highly technical StaffObservabilitySite Reliability Engineer with a specialty in Splunk to own and evolve our Splunk...  ...servicesIncident Response: Participate in on-call rotations and lead post-incident reviews to drive systemic improvements and "... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta

    Bellevue, WA
    3 days ago
  • $194k - $267k

     ...do something more than once, automate it” and who can rapidly self-educate on new concepts and tools.Position Overview:The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta

    Bellevue, WA
    4 days ago
  •  ...adventure where you can push the limits of what's possible.As a Lead Software Engineer at JPMorganChase within the Enterprise Technology,...  ...These benefits include comprehensive health care coverage, on-site health and wellness centers, a retirement savings plan, backup... 

    JP Morgan Chase

    Seattle, WA
    12 hours ago
  • $204k - $306k

     ...mission. If you are too, let's talk.Manager, Site Reliability EngineeringSan Francisco,...  ...Francisco Office. The IDaaS Site Reliability Engineering GroupOkta authenticates, authorizes and...  ...capabilities, and robust self-healing patterns.Lead, mentor, and grow a high-performing... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours
    2 days per week

    Okta

    Bellevue, WA
    3 days ago
  • $194k - $267k

     ...let's talk.The TeamWe are looking for an experienced Staff Site Reliability Engineer to join Okta's Emerging Products Group (EPG). Our mission is...  ...technical challenges. You will serve as a key technical lead within the EPG SRE organization, partnering with software engineers... 
    Local area
    Worldwide
    Flexible hours

    Okta

    Bellevue, WA
    1 day ago
  •  ...together. We are responsible for the reliability of all the company's major...  ...products, services, and query engines. We serve business needs...  ...effectively.- Incident Management: Lead efforts to troubleshoot and...  ...emerging technologies related to site reliability and infrastructure... 

    TikTok

    Seattle, WA
    12 hours ago
  • $180.1k - $278.7k

     ...Staff Infrastructure Reliability EngineerThe Staff Infrastructure Reliability Engineer is responsible for the technical leadership of Redfin's production database...  ...reliability and maintainability. They will help lead the team's strategy as we expand the database and... 
    Minimum wage
    Immediate start

    Rocket Companies

    Seattle, WA
    4 days ago
  •  ...A leading data and AI infrastructure provider is seeking a Senior Staff Technical Program Manager for Reliability to enhance the reliability and performance of their multi-cloud infrastructure...  ...programs in partnership with senior engineering leaders, requiring over 10 years of... 

    Menlo Ventures

    Bellevue, WA
    3 days ago
  • $177.69k - $341.73k

     ...A global technology company based in Seattle is seeking engineers for their Site Reliability Engineering team. The role involves extensive work on data infrastructure and high-performance systems. Candidates should have a Bachelor's degree in Computer Science and at least... 

    ByteDance

    Seattle, WA
    3 days ago
  •  ...ByteDance’s Infrastructure Engineering team in Seattle designs, builds, and operates global infrastructure spanning public and private...  ...storage. Join a fast-paced, collaborative team focused on reliability, scalability, and continuous optimization, driving improvements... 

    ByteDance

    Seattle, WA
    3 days ago
  •  ...Amazon.com Services LLC is seeking an Industry Specialist within Surface Transportation to lead the Fleet Management program for Middle Mile assets, focusing on uptime, cost reduction, and a scalable Vendor Quality Audit Program. You will define vision and roadmaps, collaborate... 

    Amazon

    Bellevue, WA
    3 days ago
  • SRE / Sr. DevOps EngineerLocation – Seattle, WA (Hybrid - 03 days onsite) Duration – 12 months + Rate: DOE US Citizens and Green card holders are Preferred.Core Skills:Azure Cloud, AKS – Scalability, monitoring, deployment, check logs, ensure node and pod health.Databases...
    Work experience placement

    Georgia IT Inc

    Seattle, WA
    4 days ago
  • $127k - $249k

    We are looking for an experienced Senior or Staff Engineer for our SRE, InfraSec team, to guide the security of our cloud-based infrastructure...  ...Responsibilities:Cloud Security Design and Implementation: Help lead the design and deployment of security solutions for cloud... 
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    Seattle, WA
    3 days ago
  • $94k - $142.3k

     ...to level-up your career at the company leading workforce transformation in the...  ...and partner enablement, applications engineering, infrastructure, collaboration, enterprise...  ...globally, at scale, sustainably.As a Site Reliability Operations Engineer you'll be part of... 
    Full time
    Shift work

    Salesforce

    Seattle, WA
    3 days ago
  •  ...System Reliability Engineer (SRE)At T-Mobile, we invest in YOU! Our Total Rewards Package ensures that employees get the same big love we give...  ...Security and Access Management (ISAM) organization is leading a major transformation in logical access compliance. What was... 
    Full time
    Temporary work
    Part time
    Work experience placement
    Flexible hours

    T Mobile US

    Bellevue, WA
    4 days ago
  • SRE / DevOps EngineerSeattle based client. Seattle-WA (3 days onsite).U.S. Citizens and those authorized to work in the U.S. are encouraged to apply. We are unable to sponsor currently.Must Have SkillsF5 load balancerChef and Terraform - must have most critical skillVMware...

    Georgia IT Inc

    Seattle, WA
    4 days ago
  • Technical/Functional Skills Windows Servers, Digital: Microsoft Azure Windows Powershell, Digital: DevOps Roles & Responsibilities Windows Server 2012 -2019 Administration Microsoft Azure Azure AAD DFSR, DHCP DNS, KMS, WSUS TCP/IP Hyper-V High Availability Clusters ...

    The Dignify Solutions, LLC

    Bellevue, WA
    1 day ago
  • $198.36k - $416.1k

     ...Responsibilities Team Intro TikTok video system is a world-leading video platform that provides multimedia storage, delivery,...  ...Provide technical leadership and mentorship to a team of Site Reliability Engineers focused on building observable, fault-tolerant video... 
    Temporary work
    Shift work

    TikTok USDS Joint Venture

    Seattle, WA
    3 days ago
  •  ...Lululemon Site Reliability Engineering Engineer We are a yoga-inspired technical apparel company up to big things. The practice and philosophy...  ...kindness, and creates the space for others to do the same. Leads with courage, knowing the possibility of greatness is... 

    Samprasoft

    Seattle, WA
    3 days ago
  • Job Title Required Skills: CHEF experience - Must have most critical Azure Cloud – experience - Must have most critical AKS- Azure Kubernetes services - Must have most critical Kubernetes - Must have most critical NoSQL DB – Cassandra / Mongo DB ...

    Syntricate Technologies

    Seattle, WA
    3 days ago
  • $166k - $244k

    # Senior Software Engineer, Site Reliability EngineeringGoogle • onsite • 601 N 34th St, Seattle, WA 98103, USA • full\_timePay: USD 166000.00 - USD 244000.00 / unspecifiedBusinesses of all shapes and sizes rely on Google’s unparalleled advertising solutions to help them... 
    Temporary work

    Epic Games

    Seattle, WA
    5 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Lead Site Reliability Engineer. Be the first to apply!