Lead Site Reliability Engineer
$198.36k - $416.1kTikTok USDS Joint Venture
Responsibilities
Team Intro
TikTok video system is a world-leading video platform that provides multimedia storage, delivery, transcoding services. As part of the USDS, the Video Platform team is responsible for building the next generation video processing platform which provides excellent experiences for billions of users around the world.
Responsibilities
- Provide technical leadership and mentorship to a team of Site Reliability Engineers focused on building observable, fault-tolerant video services with high availability
- Drive architectural decisions for large-scale, globally distributed service mesh architectures
- Establish and maintain production ownership models, incident response protocols, and service level agreements
- Develop strategic roadmaps for observability and automation initiatives that enhance system reliability
- Balance technical contributions with people management responsibilities, including career development, performance evaluations, and team growth
- Foster a culture of trust, continuous learning and improvement, and knowledge sharing within your team and across the organization
- Lead security initiatives to safeguard critical assets, partnering with security and compliance teams to implement robust protocols and enforcement that ensure data protection and regulatory compliance across all services
Qualifications
Minimum Qualifications:
- 5+ years of experience and expertise in designing, analyzing, and troubleshooting large-scale distributed systems with various databases, caching solutions and other web service components
- Experience in running high-availability web services at massive scale, with comprehensive knowledge of cloud-native architectures and advanced networking concepts
- Experience in design/development or SRE experience in video streaming or related services
- Previous experience leading a small to mid-size team while maintaining significant "hands-on" technical contributions
- Strong understanding of Unix/Linux operating systems internals and networking fundamentals
- Proficiency in writing production-grade code in Go, Python, Java or similar languages
- Proven track record of establishing and implementing SRE best practices across engineering organizations
Preferred Qualifications
- Previous experience in design/development or SRE in video processing applications
- Experience in large scale Kubernetes systems
- Deep expertise in algorithms, data structures, and systems design with proven ability to architect complex technical solutions
About USDS
TikTok USDS Joint Venture LLC is dedicated to the safety and security of millions of Americans who create, discover, and connect with what they love on the apps we operate. The Joint Venture has been established in compliance with the Executive Order signed by President Trump on September 25, 2025. Our foundation is a comprehensive data privacy and cybersecurity program we operate under defined safeguards to protect national security and secure U.S. user data, apps and the algorithm. We safeguard the U.S. content ecosystem, holding decision‑making authority for trust and safety policies and moderation. USDS Joint Venture helps ensure Americans can continue to express their creativity, discover new hobbies and interests, and build thriving communities and businesses on a global scale.
On‑site presence across teams allows the company to operate with greater speed, alignment, and agility — especially in areas like real‑time decision‑making, team development, and integrated execution. As such, the company is shifting from a hybrid work model to a fully in‑person schedule up to 5 days a week.
Why Join Us
Inspiring creativity is at the core of TikTok's mission. Our innovative product is built to help people authentically express themselves, discover and connect – and our global, diverse teams make that possible. Together, we create value for our communities, inspire creativity and bring joy - a mission we work towards every day.
We strive to do great things with great people. We lead with curiosity, humility, and a desire to make impact in a rapidly growing tech company. Every challenge is an opportunity to learn and innovate as one team. We're resilient and embrace challenges as they come. By constantly iterating and fostering an "Always Day 1" mindset, we achieve meaningful breakthroughs for ourselves, our company, and our users. When we create and grow together, the possibilities are limitless. Join us.
Diversity & Inclusion
TikTok is committed to creating an inclusive space where employees are valued for their skills, experiences, and unique perspectives. Our platform connects people from across the globe and so does our workplace. At TikTok, our mission is to inspire creativity and bring joy. To achieve that goal, we are committed to celebrating our diverse voices and to creating an environment that reflects the many communities we reach. We are passionate about this and hope you are too.
USDS Reasonable Accommodation
USDS is committed to providing reasonable accommodations in our recruitment processes for candidates with disabilities, pregnancy, sincerely held religious beliefs or other reasons protected by applicable laws. If you need assistance or a reasonable accommodation, please reach out to us at
Job Information
【For Pay Transparency】Compensation Description (Annually)
The base salary range for this position in the selected city is $198360 - $416100 annually.
Compensation may vary outside of this range depending on a number of factors, including a candidate’s qualifications, skills, competencies and experience, and location. Base pay is one part of the Total Package that is provided to compensate and recognize employees for their work, and this role may be eligible for additional discretionary bonuses/incentives, and restricted stock units.
Employees have day one access to medical, dental, and vision insurance, a 401(k) savings plan with company match, paid parental leave, short‑term and long‑term disability coverage, life insurance, wellbeing benefits, among others.
Employees also receive 10 paid holidays per year, 10 paid sick days per year and 17 days of Paid Personal Time (prorated upon hire with increasing accruals by tenure).
The Company reserves the right to modify or change these benefits programs at any time, with or without notice.
For Los Angeles County (unincorporated) Candidates:
- Interacting and occasionally having unsupervised contact with internal/external clients and/or colleagues;
- Appropriately handling and managing confidential information including proprietary and trade secret information and access to information technology systems;
- Exercising sound judgment.
$167.7k - $245.2k
...also delivering AI-powered assurance insights within Cisco’s leading Networking, Security, Collaboration, and Observability... ...maintaining our FedRAMP offering. Your ImpactAs a FedRAMP Site Reliability Engineer(SRE), you will lead the operations and architecture of our...SuggestedFull timeTemporary workWork at officeLocal areaFlexible hours2 days per week$134.25k - $214.8k
...matters at a company where you matter.Your ImpactAre you an engineer who gets excited about the challenge of making complex distributed... ...it.You will be part of the Observability team within Axon's Site Reliability organization — a focused team responsible for Axon's metrics,...SuggestedWork experience placementWork at officeRemote work$134.25k - $214.8k
...change. Constantly grow as you work hard for a mission that matters at a company where you matter.Your ImpactAs a Senior Site Reliability Engineer within the APX SRE organization, you’ll focus on delivering practical, scalable solutions to support the reliability and performance...SuggestedWork experience placementWork at officeRemote workFlexible hours$170k - $220k
...We're Looking ForWe’re looking for a hands-on, high-agency Site Reliability Engineer to help shape and scale the reliability layer of our stack.... ...shipping process.You’ll work closely with engineers, product leads, and company leadership to ensure uptime, speed, and...Suggested- ...Windows including patching and certificate provisioning and renewals. Able to map out existing infrastructure flows and dependencies. Leading root cause analysis meetings. Responding to incidents. Understanding of logs and monitoring tools (Splunk, Sumo Logic, New Relic,...Suggested
$147k - $202.4k
...let's talk.Our company is seeking a highly skilled Senior Site Reliability Engineer to join our team. We are a SaaS company specializing in securing... ...and be a primary responder for critical incidents, leading root cause analysis and implementing preventative measures...Work at officeLocal areaWorldwideFlexible hoursShift work$165k - $225.6k
...From core infrastructure to enterprise platforms, we partner across functions to drive scale, reliability, and innovation through technology.The Senior Site Reliability Engineer OpportunityReporting to the Manager, Site Reliability Engineering, this role will help build,...Permanent employmentLocal areaWorldwideFlexible hours$127k - $249k
The TeamPlatform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions... ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper)....Work at officeLocal areaRemote workWorldwideFlexible hours$143k - $191k
...mission critical capabilities to our customers. System Deployment Engineers work in complex environments with shared environmental... ...customersParticipate in customer demonstrations and exercisesWork with site reliability engineers to provide and refine requirements for tooling and...Full timeTemporary workWork experience placementImmediate start$194k - $267k
...Overview:We are seeking a highly technical StaffObservabilitySite Reliability Engineer with a specialty in Splunk to own and evolve our Splunk... ...servicesIncident Response: Participate in on-call rotations and lead post-incident reviews to drive systemic improvements and "...Permanent employmentWork at officeLocal areaWorldwideFlexible hours$194k - $267k
...do something more than once, automate it” and who can rapidly self-educate on new concepts and tools.Position Overview:The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and...Permanent employmentWork at officeLocal areaWorldwideFlexible hours- ...adventure where you can push the limits of what's possible.As a Lead Software Engineer at JPMorganChase within the Enterprise Technology,... ...These benefits include comprehensive health care coverage, on-site health and wellness centers, a retirement savings plan, backup...
$204k - $306k
...mission. If you are too, let's talk.Manager, Site Reliability EngineeringSan Francisco,... ...Francisco Office. The IDaaS Site Reliability Engineering GroupOkta authenticates, authorizes and... ...capabilities, and robust self-healing patterns.Lead, mentor, and grow a high-performing...Permanent employmentWork at officeLocal areaWorldwideFlexible hours2 days per week$194k - $267k
...let's talk.The TeamWe are looking for an experienced Staff Site Reliability Engineer to join Okta's Emerging Products Group (EPG). Our mission is... ...technical challenges. You will serve as a key technical lead within the EPG SRE organization, partnering with software engineers...Local areaWorldwideFlexible hours- ...together. We are responsible for the reliability of all the company's major... ...products, services, and query engines. We serve business needs... ...effectively.- Incident Management: Lead efforts to troubleshoot and... ...emerging technologies related to site reliability and infrastructure...
$180.1k - $278.7k
...Staff Infrastructure Reliability EngineerThe Staff Infrastructure Reliability Engineer is responsible for the technical leadership of Redfin's production database... ...reliability and maintainability. They will help lead the team's strategy as we expand the database and...Minimum wageImmediate start- ...A leading data and AI infrastructure provider is seeking a Senior Staff Technical Program Manager for Reliability to enhance the reliability and performance of their multi-cloud infrastructure... ...programs in partnership with senior engineering leaders, requiring over 10 years of...
$177.69k - $341.73k
...A global technology company based in Seattle is seeking engineers for their Site Reliability Engineering team. The role involves extensive work on data infrastructure and high-performance systems. Candidates should have a Bachelor's degree in Computer Science and at least...- ...ByteDance’s Infrastructure Engineering team in Seattle designs, builds, and operates global infrastructure spanning public and private... ...storage. Join a fast-paced, collaborative team focused on reliability, scalability, and continuous optimization, driving improvements...
- ...Amazon.com Services LLC is seeking an Industry Specialist within Surface Transportation to lead the Fleet Management program for Middle Mile assets, focusing on uptime, cost reduction, and a scalable Vendor Quality Audit Program. You will define vision and roadmaps, collaborate...
- SRE / Sr. DevOps EngineerLocation – Seattle, WA (Hybrid - 03 days onsite) Duration – 12 months + Rate: DOE US Citizens and Green card holders are Preferred.Core Skills:Azure Cloud, AKS – Scalability, monitoring, deployment, check logs, ensure node and pod health.Databases...Work experience placement
$127k - $249k
We are looking for an experienced Senior or Staff Engineer for our SRE, InfraSec team, to guide the security of our cloud-based infrastructure... ...Responsibilities:Cloud Security Design and Implementation: Help lead the design and deployment of security solutions for cloud...Local areaRemote workWorldwideFlexible hours$94k - $142.3k
...to level-up your career at the company leading workforce transformation in the... ...and partner enablement, applications engineering, infrastructure, collaboration, enterprise... ...globally, at scale, sustainably.As a Site Reliability Operations Engineer you'll be part of...Full timeShift work- ...System Reliability Engineer (SRE)At T-Mobile, we invest in YOU! Our Total Rewards Package ensures that employees get the same big love we give... ...Security and Access Management (ISAM) organization is leading a major transformation in logical access compliance. What was...Full timeTemporary workPart timeWork experience placementFlexible hours
- SRE / DevOps EngineerSeattle based client. Seattle-WA (3 days onsite).U.S. Citizens and those authorized to work in the U.S. are encouraged to apply. We are unable to sponsor currently.Must Have SkillsF5 load balancerChef and Terraform - must have most critical skillVMware...
- Technical/Functional Skills Windows Servers, Digital: Microsoft Azure Windows Powershell, Digital: DevOps Roles & Responsibilities Windows Server 2012 -2019 Administration Microsoft Azure Azure AAD DFSR, DHCP DNS, KMS, WSUS TCP/IP Hyper-V High Availability Clusters ...
$198.36k - $416.1k
...Responsibilities Team Intro TikTok video system is a world-leading video platform that provides multimedia storage, delivery,... ...Provide technical leadership and mentorship to a team of Site Reliability Engineers focused on building observable, fault-tolerant video...Temporary workShift work- ...Lululemon Site Reliability Engineering Engineer We are a yoga-inspired technical apparel company up to big things. The practice and philosophy... ...kindness, and creates the space for others to do the same. Leads with courage, knowing the possibility of greatness is...
- Job Title Required Skills: CHEF experience - Must have most critical Azure Cloud – experience - Must have most critical AKS- Azure Kubernetes services - Must have most critical Kubernetes - Must have most critical NoSQL DB – Cassandra / Mongo DB ...
$166k - $244k
# Senior Software Engineer, Site Reliability EngineeringGoogle • onsite • 601 N 34th St, Seattle, WA 98103, USA • full\_timePay: USD 166000.00 - USD 244000.00 / unspecifiedBusinesses of all shapes and sizes rely on Google’s unparalleled advertising solutions to help them...Temporary work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Lead Site Reliability Engineer. Be the first to apply!
- lead network engineer Seattle, WA
- lead engineer Seattle, WA
- lead system engineer Seattle, WA
- lead operating engineer Seattle, WA
- lead infrastructure engineer Seattle, WA
- site reliability engineer sre Seattle, WA
- site reliability engineer Seattle, WA
- site recruiter Seattle, WA
- site services specialist Seattle, WA
- junior website developer Seattle, WA

