Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer, Cloud Infrastructure - USDS

$118.66k - $259.2k

TikTok

Site Reliability Engineer, Cloud Infrastructure - USDS

Get AI-powered advice on this job and more exclusive features.

Base Pay Range

$118,657.00/yr - $259,200.00/yr

Responsibilities

The Systems and Networking team is committed to ensuring the seamless operation of TikTok's US physical infrastructure. We handle the provisioning of physical servers and maintain the TikTok US physical network. Additionally, we engage in collaborative efforts with vendors such as OCI and Akamai to manage physical hardware, networks, and uphold assurance and compliance objectives. We also work closely with our colleagues around the world to build and support various platforms within our US region, including internal platforms that support our daily operations. Our primary goal is to ensure the uninterrupted functionality of TikTok's US Physical Infrastructure, thus facilitating other internal middleware teams to deliver essential intermediary services to internal business units such as Product, e-Commerce, Ads/Monetization, etc., all while strictly adhering to compliance standards.

In order to enhance collaboration and cross-functional partnerships, at this time, our organization follows a hybrid work schedule that requires employees to work in the office 3 days a week, or as directed by their manager/department. We regularly review our hybrid work model, and the specific requirements may change at any time.

Drive infrastructure automation and tooling: Design develop, and maintain solutions for efficient operation, optimization, and comprehensive monitoring of global infrastructure, minimizing manual intervention.

Collaborate on service lifecycle management: Partner with engineering teams to design, deploy, operate, and continuously improve robust and scalable systems and services, from inception to refinement.

Ensure service reliability and performance: Proactively monitor system health, conduct performance testing, and manage incidents to maximize uptime, availability, and adherence to defined SLAs/SLOs.

Execute core SRE practices: Perform on-call duties and production operations, including change management, capacity planning, and disaster recovery, while contributing to documentation and process improvements across teams.

Qualifications

Minimum Qualifications

  • Proficient in one or more programming languages (e.g., Python, Go, Java, C++).
  • Strong understanding of Linux operating systems and open-source technologies.
  • Experience in network architecture and troubleshooting, database modeling, cloud systems, and large-scale distributed systems.
  • Knowledge of monitoring tools and methodologies (such as Prometheus, Grafana), AIOPS, APM, Disaster Recovery.
  • Experience in designing, analyzing, and building automation and tools for large-scale systems.
  • Experience in building solutions with AWS, GCP, Azure, and other cloud services.

Preferred Qualifications

  • Expertise in any of these tech stacks: Kubernetes, ElasticSearch, ClickHouse, Message Queue, OpenTSDB, Service Mesh, MySQL, Redis, etc.
  • Master's degree in Computer Science, Engineering, or a related field.

As a condition of employment, all successful candidates must be able to establish authorization to work in the United States. For this position, the Company does not provide sponsorship for any immigration-related benefits.

About USDS

TikTok is the leading destination for short-form mobile video. Our mission is to inspire creativity and bring joy. U.S. Data Security ("USDS") is a subsidiary of TikTok in the U.S. This new, security-first division was created to bring heightened focus and governance to our data protection policies and content assurance protocols to keep U.S. users safe. Our focus is on providing oversight and protection of the TikTok platform and U.S. user data, so millions of Americans can continue turning to TikTok to learn something new, earn a living, express themselves creatively, or be entertained. The teams within USDS that deliver on this commitment daily span across Trust & Safety, Security & Privacy, Engineering, User & Product Ops, Corporate Functions and more.

Why Join Us

Inspiring creativity is at the core of TikTok's mission. Our innovative product is built to help people authentically express themselves, discover and connect – and our global, diverse teams make that possible. Together, we create value for our communities, inspire creativity and bring joy - a mission we work towards every day. We strive to do great things with great people. We lead with curiosity, humility, and a desire to make impact in a rapidly growing tech company. Every challenge is an opportunity to learn and innovate as one team. We're resilient and embrace challenges as they come. By constantly iterating and fostering an "Always Day 1" mindset, we achieve meaningful breakthroughs for ourselves, our company, and our users. When we create and grow together, the possibilities are limitless. Join us.

Diversity & Inclusion

TikTok is committed to creating an inclusive space where employees are valued for their skills, experiences, and unique perspectives. Our platform connects people from across the globe and so does our workplace. At TikTok, our mission is to inspire creativity and bring joy. To achieve that goal, we are committed to celebrating our diverse voices and to creating an environment that reflects the many communities we reach. We are passionate about this and hope you are too.

USDS Reasonable Accommodation

USDS is committed to providing reasonable accommodations in our recruitment processes for candidates with disabilities, pregnancy, sincerely held religious beliefs or other reasons protected by applicable laws. If you need assistance or a reasonable accommodation, please reach out to us at

Job Information

The base salary range for this position in the selected city is $118,657 - $259,200 annually. Compensation may vary outside of this range depending on a number of factors, including a candidate’s qualifications, skills, competencies and experience, and location. Base pay is one part of the Total Package that is provided to compensate and recognize employees for their work, and this role may be eligible for additional discretionary bonuses/incentives, and restricted stock units. Benefits may vary depending on the nature of employment and the country work location. Employees have day one access to medical, dental, and vision insurance, a 401(k) savings plan with company match, paid parental leave, short-term and long-term disability coverage, life insurance, wellbeing benefits, among others. Employees also receive 10 paid holidays per year, 10 paid sick days per year and 17 days of Paid Personal Time (prorated upon hire with increasing accruals by tenure). The Company reserves the right to modify or change these benefits programs at any time, with or without notice. For Los Angeles County (unincorporated) Candidates: Qualified applicants with arrest or conviction records will be considered for employment in accordance with all federal, state, and local laws including the Los Angeles County Fair Chance Ordinance for Employers and the California Fair Chance Act. Our company believes that criminal history may have a direct, adverse and negative relationship on the following job duties, potentially resulting in the withdrawal of the conditional offer of employment: 1. Interacting and occasionally having unsupervised contact with internal/external clients and/or colleagues; 2. Appropriately handling and managing confidential information including proprietary and trade secret information and access to information technology systems; and 3. Exercising sound judgment.

Seniority Level

Mid-Senior level

Employment Type

Full-time

Job Function

Engineering and Information Technology, Technology, Information and Internet

#J-18808-Ljbffr
Vacancy posted 7 hours ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer, Cloud Infrastructure - USDS in San Jose, CA vacancy
  • $118.66k - $259.2k

     ...Site Reliability Engineer, Cloud Infrastructure- USDS Site Reliability Engineer, Cloud Infrastructure- USDS Get AI-powered advice on this job and more exclusive features. Responsibilities The infrastructure team of US Tech Services Department at TikTok supports... 
    Suggested
    Full time
    Temporary work
    Internship
    Work at office
    Local area
    3 days per week

    TikTok

    San Jose, CA
    7 hours ago
  • $118.66k - $187.2k

     ...Site Reliability Engineer - Video Platform (Entry Level) - USDS Team Intro TikTok video system is a world‑leading video...  ...system stability and save infrastructure costs. Provide strong support...  ...AWS, Google, Azure and other cloud services is a plus. Passionate... 
    Suggested
    Temporary work
    Work experience placement

    TikTok

    San Jose, CA
    7 hours ago
  • $184k - $287.5k

    At NVIDIA, Site Reliability Engineering provides a rare chance to define, develop, and support large-scale production systems with high efficiency...  ...-tolerant, performant, and supportable.Background with infrastructure automation.Experience running critical services in... 
    Suggested
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • Site Reliability Engineer, Data Platform - USDS Responsibilities Engage in and improve the whole lifecycle of service...  ...-effective data, services and infrastructures. Maintain services once they are...  ...services at scale and understanding cloud native technologies and... 
    Suggested

    TikTok

    Mountain View, CA
    5 days ago
  • $118.66k - $259.2k

     ...Site Reliability Engineer - AML Global Recommendation - USDS About the Team: Site Reliability Engineering (SRE) of the AML (Applied Machine Learning) team combines system engineering and the art of machine learning to develop and run a massively distributed AI/ML recommendation... 
    Suggested
    Temporary work
    Work at office
    Shift work
    3 days per week

    TikTok

    San Jose, CA
    7 hours ago
  • $187.04k - $359.72k

     ...E-commerce AI Platform team (USDS). Our mission is to build the...  ...stakeholders to drive complex engineering efforts and align technical...  ...and novel ideas to keep our infrastructure at the cutting edge. Qualifications...  ...on a global scale. On-site presence across teams allows... 
    Temporary work
    Local area
    Flexible hours
    Shift work

    TikTok USDS Joint Venture

    San Jose, CA
    7 hours ago
  • $152k - $241.5k

     ...team is seeking a Senior System Software Engineer to help bring NVIDIA's autonomous...  ...expansion of hardware-in-the-loop (HIL) infrastructure to support the robust deployment and lifecycle...  ...(e.g., Jenkins, Docker, Kubernetes) in cloud-native or hybrid environments for mission... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

    We are seeking a Senior DevOps / Cloud Simulation Infrastructure Engineer to own the complete end-to-end cloud execution pipeline for SimReady assets...  ...observable, debuggable, and production-ready.Operational Reliability: Implement atomic update semantics and safe failure... 
    Full time
    Local area

    Nvidia

    Santa Clara, CA
    4 days ago
  • $168k - $264.5k

     ...NVIDIA's Digital Marketing Organization seeks a senior Site Reliability Engineer (SRE) to join our Santa Clara, CA team. As an SRE at NVIDIA...  ...across deployment pipelines, Akamai CDN, WAF, and cloud infrastructure.Author, test, and activate shared and non-shared Cloudlet... 
    Full time

    Nvidia

    Santa Clara, CA
    6 days ago
  • $280k - $380k

     ...an active-active, multi-cloud platform on AWS and GCP to...  ...internet scale. We focus on reliability and automation, engineering systems that perform...  ...can, and turning complex infrastructure into reliable, well‑documented...  ...experienced DevOps/SRE (Site Reliability Engineering)... 
    Work at office
    Local area
    Remote work
    Monday to Thursday
    Flexible hours

    Roku

    San Jose, CA
    2 days ago
  •  ...advance your career. THE ROLE:We are looking for a systems-minded engineer who lives at the intersection of large-scale model inference,...  .... This role focuses on post-training and inference infrastructure, with particular emphasis on P/D disaggregation, KV cache lifecycle... 

    AMD

    San Jose, CA
    2 days ago
  • $168k - $270.25k

     ...streamlines cluster provisioning, workload management, and infrastructure monitoring. It provides all the tools you need to deploy...  ...excellent, comprehensive support to our customers! ​Sr Site Reliability Engineer in this role will significantly impact and contribute to... 
    Full time
    Worldwide

    Nvidia

    Santa Clara, CA
    2 days ago
  • $124k - $271.2k

    What You Can ExpectAs a Lead Staff Site Reliability Engineer, you will be one of the technical leads for our DevOps Platforms organization...  ...group is responsible for DevOps Platforms including cloud infrastructure, physical data center orchestration, critical security... 
    Full time
    Work at office
    Remote work

    Zoom

    San Jose, CA
    18 hours ago
  • $176k - $276k

    Cloud Foundations Reliability (CFR) is part of NVIDIA’s Global Network Infrastructure (GNI) organization. We deploy, integrate, and operate the Kubernetes-based platform and...  ....We are looking for a hands-on senior engineer to own the lifecycle and automation of the... 
    Full time
    Remote work
    Weekend work

    Nvidia

    Santa Clara, CA
    5 days ago
  • Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI...  ...physical data centers. We are looking for a Senior Site Reliability Engineer to improve the reliability, scalability, and... 
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    2 days ago
  • $272k - $431.25k

    NVIDIA is looking for a Cloud Site Reliability Engineering Architect to work in IPP's (Infrastructure, Planning and Process) Cloud Infrastructure Team. IPP is a global organization within NVIDIA. This group works with various other groups within NVIDIA such as Graphics... 
    Full time
    Work experience placement
    Worldwide

    Nvidia

    Santa Clara, CA
    4 days ago
  • Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers...  ...home day is currently Tuesday.Engineering at Lambda is responsible for building...  ..., upgrades, and scaling.Own the reliability, performance, and security of... 
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    5 days ago
  • $178k - $321k

     ...markets. We are safe and reliable, backed by our Proof...  ...harness: a resilient cloud platform, the agentic...  ...governed data and AI infrastructure everything else depends...  ...This is a two-person engineering team: you deploy, debug...  ...internal or external careers site.Notice:All official... 

    OKX

    San Jose, CA
    4 days ago
  • $184k - $287.5k

     .... We are looking for a Senior Software Engineer to lead the bring-up, triage, benchmarking...  ...-art LLM workloads run efficiently and reliably at scale. You will lead deep...  ...debugging of large-scale AI clusters, infrastructure, and end-to-end workloads, setting the... 
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $100k

     ...all seniorities.Tenstorrent’s AI Software Infrastructure team builds the platforms that power...  ...You AreStrong backend or infrastructure engineer with experience building and operating platforms...  ...data center infrastructure differs from cloud-native environments at scale.How... 
    Permanent employment

    Tenstorrent

    Santa Clara, CA
    5 days ago
  • $116k - $189.75k

     ...large language model workloads. We are looking for a Software Engineer focused on bring-up, triage, benchmarking, analysis, and...  ...doing:Bring up, validate, and debug large-scale AI clusters, infrastructure, and end-to-end workloads.Bring up, tune, and benchmark AI pre... 
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    3 days ago
  • Technical Cloud Architect (Virtualization & Cloud Infrastructure)This role has been designed as 'Hybrid' with a requirement that you will work on average 2...  ...impact role that sits at the intersection of pre-sales engineering, solutions architecture, and product strategy. You... 
    Full time
    Work experience placement
    Work at office
    Local area
    Immediate start
    2 days per week

    Hewlett Packard Enterprise

    San Jose, CA
    2 days ago
  • Senior Site Reliability Engineer ILocationSan Jose, Costa Rica - RemoteSummary of roleOwn availability...  ..., increase efficiency in our use of cloud resources and our developer’s time,...  ...creation to production supportVersed in Infrastructure as Code practices using technologies... 
    Flexible hours

    Sumo Logic

    San Jose, CA
    5 days ago
  • $152k - $241.5k

     ...and network fabrics.Use IaC(Infrastructure‑as‑Code) and config management...  ...globally distributed, multi‑cloud hybrid environment - On‑prem...  ...lifecycle management, fleet reliability/auto-healing, E2E...  ...Perl, or Ruby.Mentored other engineers and influenced technical direction... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $128.6k - $184.9k

     ...scalability, and efficiency of the infrastructure that powers our global cloud platform. As a team of six engineers distributed across the US,...  ...strong focus on automation, reliability, and operational excellence....  ...7+ years of experience in Site Reliability Engineering,... 
    Permanent employment
    Full time
    Temporary work
    Local area
    Worldwide
    Flexible hours

    CISCO Systems

    Milpitas, CA
    1 day ago
  • $168k - $270.25k

    NVIDIA is looking for a Senior Site Reliability Engineer (SRE) to join its GeForce Now (GFN) team. SRE...  ...our internal and external-facing GPU cloud gaming services have reliability and...  ...to and mitigating high-severity infrastructure alerts and service degradations.Ways... 
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $267k - $356k

     ..., The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands...  ....Lambda's Storage Engineering team is the backbone behind...  ...the industry, which means reliability and performance aren't just...  ...across new and existing sites using tools such as... 
    Work experience placement
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    4 days ago
  • $230k - $250k

     ...autonomous networking, giving engineers and AI agents the ability...  ...across every major cloud and vendor environment.Global...  ...done.Forward is looking for a Site Reliability EngineerAbout the Role This...  ...closely with engineering, infrastructure, and product to ensure our... 
    Night shift

    Forward Networks

    Santa Clara, CA
    2 days ago
  • $168k - $270.25k

     ....Join our team at NVIDIA as a Senior Site reliability engineer focused on HPC storage and play a crucial...  ...while harnessing the power of cloud computing. You will be responsible for...  ...closely with engineering teams to align infrastructure with their evolving needs, document... 
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $101k - $161k

     ...in data-driven, client-to-cloud networking for large data center...  ...awards, such as Best Engineering Team, Best Company for Diversity...  ...Work WithWe’re looking for Site Reliability Engineers to join our...  ...microservices stack, monitoring infrastructure, and much more. What You'll... 

    Arista Networks

    Santa Clara, CA
    5 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer, Cloud Infrastructure - USDS. Be the first to apply!