Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Lead Site Reliability Engineer

$198.36k - $416.1k

TikTok USDS Joint Venture

Responsibilities Team Intro TikTok video system is a world-leading video platform that provides multimedia storage, delivery, transcoding services. As part of the USDS, the Video Platform team is responsible for building the next generation video processing platform which provides excellent experiences for billions of users around the world. Provide technical leadership and mentorship to a team of Site Reliability Engineers focused on building observable, fault-tolerant video services with high availability Drive architectural decisions for large-scale, globally distributed service mesh architectures Establish and maintain production ownership models, incident response protocols, and service level agreements Develop strategic roadmaps for observability and automation initiatives that enhance system reliability Balance technical contributions with people management responsibilities, including career development, performance evaluations, and team growth Foster a culture of trust, continuous learning and improvement, and knowledge sharing within your team and across the organization Lead security initiatives to safeguard critical assets, partnering with security and compliance teams to implement robust protocols and enforcement that ensure data protection and regulatory compliance across all services Qualifications Minimum Qualifications: 5+ years of experience and expertise in designing, analyzing, and troubleshooting large-scale distributed systems with various databases, caching solutions and other web service components Experience in running high-availability web services at massive scale, with comprehensive knowledge of cloud-native architectures and advanced networking concepts Experience in design/development or SRE experience in video streaming or related services Previous experience leading a small to mid-size team while maintaining significant "hands-on" technical contributions Strong understanding of Unix/Linux operating systems internals and networking fundamentals Proficiency in writing production-grade code in Go, Python, Java or similar languages Proven track record of establishing and implementing SRE best practices across engineering organizations Preferred Qualifications: Previous experience in design/development or SRE in video processing applications Experience in large scale Kubernetes systems Deep expertise in algorithms, data structures, and systems design with proven ability to architect complex technical solutions About USDS TikTok USDS Joint Venture LLC is dedicated to the safety and security of millions of Americans who create, discover, and connect with what they love on the apps we operate. The Joint Venture has been established in compliance with the Executive Order signed by President Trump on September 25, 2025. Our foundation is a comprehensive data privacy and cybersecurity program we operate under defined safeguards to protect national security and secure U.S. user data, apps and the algorithm. We safeguard the U.S. content ecosystem, holding decision-making authority for trust and safety policies and moderation. USDS Joint Venture helps ensure Americans can continue to express their creativity, discover new hobbies and interests, and build thriving communities and businesses on a global scale. On-site presence across teams allows the company to operate with greater speed, alignment, and agility — especially in areas like real-time decision-making, team development, and integrated execution. As such, the company is shifting from a hybrid work model to a fully in-person schedule up to 5 days a week. Why Join Us Inspiring creativity is at the core of TikTok's mission. Our innovative product is built to help people authentically express themselves, discover and connect and our global, diverse teams make that possible. Together, we create value for our communities, inspire creativity and bring joy - a mission we work towards every day. We strive to do great things with great people. We lead with curiosity, humility, and a desire to make impact in a rapidly growing tech company. Every challenge is an opportunity to learn and innovate as one team. We're resilient and embrace challenges as they come. By constantly iterating and fostering an "Always Day 1" mindset, we achieve meaningful breakthroughs for ourselves, our company, and our users. When we create and grow together, the possibilities are limitless. Join us. Diversity & Inclusion TikTok is committed to creating an inclusive space where employees are valued for their skills, experiences, and unique perspectives. Our platform connects people from across the globe and so does our workplace. At TikTok, our mission is to inspire creativity and bring joy. To achieve that goal, we are committed to celebrating our diverse voices and to creating an environment that reflects the many communities we reach. We are passionate about this and hope you are too. USDS Reasonable Accommodation USDS is committed to providing reasonable accommodations in our recruitment processes for candidates with disabilities, pregnancy, sincerely held religious beliefs or other reasons protected by applicable laws. If you need assistance or a reasonable accommodation, please reach out to us at Job Information 【For Pay Transparency】Compensation Description (Annually) The base salary range for this position in the selected city is $198360 - $416100 annually. Compensation may vary outside of this range depending on a number of factors, including a candidate’s qualifications, skills, competencies and experience, and location. Base pay is one part of the Total Package that is provided to compensate and recognize employees for their work, and this role may be eligible for additional discretionary bonuses/incentives, and restricted stock units. Benefits may vary depending on the nature of employment and the country work location. Employees have day one access to medical, dental, and vision insurance, a 401(k) savings plan with company match, paid parental leave, short-term and long-term disability coverage, life insurance, wellbeing benefits, among others. Employees also receive 10 paid holidays per year, 10 paid sick days per year and 17 days of Paid Personal Time (prorated upon hire with increasing accruals by tenure). The Company reserves the right to modify or change these benefits programs at any time, with or without notice. For Los Angeles County (unincorporated) Candidates: Interacting and occasionally having unsupervised contact with internal/external clients and/or colleagues; Appropriately handling and managing confidential information including proprietary and trade secret information and access to information technology systems; and Exercising sound judgment. #J-18808-Ljbffr

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Lead Site Reliability Engineer in Seattle, WA vacancy
  • $143k - $191k

     ...mission critical capabilities to our customers. System Deployment Engineers work in complex environments with shared environmental...  ...customersParticipate in customer demonstrations and exercisesWork with site reliability engineers to provide and refine requirements for tooling and... 
    Suggested
    Full time
    Temporary work
    Work experience placement
    Immediate start

    Anduril Industries

    Seattle, WA
    3 days ago
  • $166k - $258k

     ...considered for this position.Nordstrom is looking for a Senior Engineer 2 to join our Site Reliability Engineering (SRE) team — and we think that person could...  .... You'll set and maintain high operational standards, lead key projects from concept to execution, and bring fresh... 
    Suggested
    Full time
    Work at office

    Nordstrom

    Seattle, WA
    7 days ago
  • $134.25k - $214.8k

     ...matters at a company where you matter.Your ImpactAre you an engineer who gets excited about the challenge of making complex distributed...  ...it.You will be part of the Observability team within Axon's Site Reliability organization — a focused team responsible for Axon's metrics,... 
    Suggested
    Work experience placement
    Work at office
    Remote work

    Axon

    Seattle, WA
    4 days ago
  • $127k - $249k

    The TeamPlatform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions...  ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper).... 
    Suggested
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    Seattle, WA
    5 days ago
  • $165k - $225.6k

     ...From core infrastructure to enterprise platforms, we partner across functions to drive scale, reliability, and innovation through technology.The Senior Site Reliability Engineer OpportunityReporting to the Manager, Site Reliability Engineering, this role will help build,... 
    Suggested
    Permanent employment
    Local area
    Worldwide
    Flexible hours

    Okta

    Bellevue, WA
    3 days ago
  • $134.25k - $214.8k

     ...change. Constantly grow as you work hard for a mission that matters at a company where you matter.Your ImpactAs a Senior Site Reliability Engineer within the APX SRE organization, you’ll focus on delivering practical, scalable solutions to support the reliability and performance... 
    Work experience placement
    Work at office
    Remote work
    Flexible hours

    Axon

    Seattle, WA
    3 days ago
  • $170k - $220k

     ...We're Looking ForWe’re looking for a hands-on, high-agency Site Reliability Engineer to help shape and scale the reliability layer of our stack....  ...shipping process.You’ll work closely with engineers, product leads, and company leadership to ensure uptime, speed, and... 

    Supio

    Seattle, WA
    3 days ago
  •  ...Windows including patching and certificate provisioning and renewals. Able to map out existing infrastructure flows and dependencies. Leading root cause analysis meetings. Responding to incidents. Understanding of logs and monitoring tools (Splunk, Sumo Logic, New Relic,... 

    Comtech

    Seattle, WA
    5 days ago
  • $204k - $306k

     ...mission. If you are too, let's talk.Manager, Site Reliability EngineeringSan Francisco,...  ...Francisco Office. The IDaaS Site Reliability Engineering GroupOkta authenticates, authorizes and...  ...capabilities, and robust self-healing patterns.Lead, mentor, and grow a high-performing... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours
    2 days per week

    Okta

    Bellevue, WA
    7 days ago
  • $194k - $267k

     ...Overview:We are seeking a highly technical StaffObservabilitySite Reliability Engineer with a specialty in Splunk to own and evolve our Splunk...  ...servicesIncident Response: Participate in on-call rotations and lead post-incident reviews to drive systemic improvements and "... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta

    Bellevue, WA
    3 days ago
  • $194k - $267k

     ...let's talk.The TeamWe are looking for an experienced Staff Site Reliability Engineer to join Okta's Emerging Products Group (EPG). Our mission is...  ...technical challenges. You will serve as a key technical lead within the EPG SRE organization, partnering with software engineers... 
    Local area
    Worldwide
    Flexible hours

    Okta

    Bellevue, WA
    6 days ago
  • $194k - $267k

     ...do something more than once, automate it” and who can rapidly self-educate on new concepts and tools.Position Overview:The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta

    Bellevue, WA
    3 days ago
  •  ...together. We are responsible for the reliability of all the company's major...  ...products, services, and query engines. We serve business needs...  ...effectively.- Incident Management: Lead efforts to troubleshoot and...  ...emerging technologies related to site reliability and infrastructure... 

    TikTok

    Seattle, WA
    5 days ago
  • $127k - $249k

    We are looking for an experienced Senior or Staff Engineer for our SRE, InfraSec team, to guide the security of our cloud-based infrastructure...  ...Responsibilities:Cloud Security Design and Implementation: Help lead the design and deployment of security solutions for cloud... 
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    Seattle, WA
    3 days ago
  •  ...Oracle seeks a senior Oracle database engineer to own end-to-end management of the database stack, emphasizing security, resiliency, scale, and performance. You will collaborate with SRE and development teams to define and deliver premier capabilities across on-premises... 
    Shift work
    Weekend work

    Ll Oefentherapie

    Seattle, WA
    4 days ago
  •  ...A leading tech company is seeking a Site Reliability Engineer in Seattle to ensure seamless operation of physical infrastructure. Responsibilities include automation solutions, system monitoring, and collaboration with engineering teams. The ideal candidate has a degree... 

    Tik Tok

    Seattle, WA
    3 days ago
  •  ...Amazon.com Services LLC is seeking an Industry Specialist within Surface Transportation to lead the Fleet Management program for Middle Mile assets, focusing on uptime, cost reduction, and a scalable Vendor Quality Audit Program. You will define vision and roadmaps, collaborate... 

    Amazon

    Bellevue, WA
    3 days ago
  •  ...System Reliability Engineer (SRE)At T-Mobile, we invest in YOU! Our Total Rewards Package ensures that employees get the same big love we give...  ...Security and Access Management (ISAM) organization is leading a major transformation in logical access compliance. What was... 
    Full time
    Temporary work
    Part time
    Work experience placement
    Flexible hours

    T Mobile US

    Bellevue, WA
    4 days ago
  •  ...AWS/SRE leader to own and scale the cloud footprint. You will design, implement, and maintain AWS infrastructure, collaborate with engineering and product teams, and drive robust service management for deployments, monitoring, and disaster recovery. The role requires 10+... 

    Jobtailor

    Bellevue, WA
    9 hours ago
  • $180.1k - $278.7k

     ...Staff Infrastructure Reliability EngineerThe Staff Infrastructure Reliability Engineer is responsible for the technical leadership of Redfin's production database...  ...reliability and maintainability. They will help lead the team's strategy as we expand the database and... 
    Minimum wage
    Immediate start

    Rocket Companies

    Seattle, WA
    4 days ago
  •  ...A leading data and AI infrastructure provider is seeking a Senior Staff Technical Program Manager for Reliability to enhance the reliability and performance of their multi-cloud infrastructure...  ...programs in partnership with senior engineering leaders, requiring over 10 years of... 

    Menlo Ventures

    Bellevue, WA
    3 days ago
  • SRE / DevOps EngineerSeattle based client. Seattle-WA (3 days onsite).U.S. Citizens and those authorized to work in the U.S. are encouraged to apply. We are unable to sponsor currently.Must Have SkillsF5 load balancerChef and Terraform - must have most critical skillVMware...

    Georgia IT Inc

    Seattle, WA
    4 days ago
  • $177.69k - $341.73k

     ...A global technology company based in Seattle is seeking engineers for their Site Reliability Engineering team. The role involves extensive work on data infrastructure and high-performance systems. Candidates should have a Bachelor's degree in Computer Science and at least... 

    ByteDance

    Seattle, WA
    3 days ago
  • SRE / Sr. DevOps EngineerLocation – Seattle, WA (Hybrid - 03 days onsite) Duration – 12 months + Rate: DOE US Citizens and Green card holders are Preferred.Core Skills:Azure Cloud, AKS – Scalability, monitoring, deployment, check logs, ensure node and pod health.Databases...
    Work experience placement

    Georgia IT Inc

    Seattle, WA
    4 days ago
  • $198.36k - $416.1k

     ...Responsibilities Team Intro TikTok video system is a world-leading video platform that provides multimedia storage, delivery,...  ...Provide technical leadership and mentorship to a team of Site Reliability Engineers focused on building observable, fault-tolerant video services... 
    Temporary work
    Shift work

    TikTok USDS Joint Venture

    Seattle, WA
    3 days ago
  • $166k - $244k

    # Senior Software Engineer, Site Reliability EngineeringGoogle • onsite • 601 N 34th St, Seattle, WA 98103, USA • full\_timePay: USD 166000.00 - USD 244000.00 / unspecifiedBusinesses of all shapes and sizes rely on Google’s unparalleled advertising solutions to help them... 
    Temporary work

    Epic Games

    Seattle, WA
    16 hours ago
  • $232k - $319k

     ...scale the service with great people and reliable, cost-effective, and efficient infrastructure...  ...& tooling. What you’ll be doing Lead the Infra platform and shared services org...  ...serviceAccelerate the velocity of SRE and product engineering by developing robust platforms, powerful... 
    Permanent employment
    Local area
    Worldwide
    Flexible hours

    Okta

    Bellevue, WA
    6 days ago
  • The Data Infrastructure SRE team is responsible for the reliability, scalability, and efficiency of the core data services that power...  ...Message Queue. Our work is not about building features, but about engineering the resilience and performance of the underlying platform... 

    TikTok

    Seattle, WA
    5 days ago
  • $147k - $202.4k

     ...let's talk.Our company is seeking a highly skilled Senior Site Reliability Engineer to join our team. We are a SaaS company specializing in securing...  ...and be a primary responder for critical incidents, leading root cause analysis and implementing preventative measures... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours
    Shift work

    Okta

    Bellevue, WA
    3 days ago
  •  ...team is responsible for the reliability, scalability, and efficiency...  ...building features, but about engineering the resilience and performance...  ...maintain system stability.As a Site Reliability Engineer, you...  ...Capacity and cost optimization: Lead initiatives in capacity planning... 

    TikTok

    Seattle, WA
    5 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Lead Site Reliability Engineer. Be the first to apply!