Site Reliability Engineer, Cloud Infrastructure - USDS
$118.66k - $259.2kTikTok
Site Reliability Engineer, Cloud Infrastructure - USDS
Get AI-powered advice on this job and more exclusive features.
Base Pay Range
$118,657.00/yr - $259,200.00/yr
Responsibilities
The Systems and Networking team is committed to ensuring the seamless operation of TikTok's US physical infrastructure. We handle the provisioning of physical servers and maintain the TikTok US physical network. Additionally, we engage in collaborative efforts with vendors such as OCI and Akamai to manage physical hardware, networks, and uphold assurance and compliance objectives. We also work closely with our colleagues around the world to build and support various platforms within our US region, including internal platforms that support our daily operations. Our primary goal is to ensure the uninterrupted functionality of TikTok's US Physical Infrastructure, thus facilitating other internal middleware teams to deliver essential intermediary services to internal business units such as Product, e-Commerce, Ads/Monetization, etc., all while strictly adhering to compliance standards.
In order to enhance collaboration and cross-functional partnerships, at this time, our organization follows a hybrid work schedule that requires employees to work in the office 3 days a week, or as directed by their manager/department. We regularly review our hybrid work model, and the specific requirements may change at any time.
Drive infrastructure automation and tooling: Design develop, and maintain solutions for efficient operation, optimization, and comprehensive monitoring of global infrastructure, minimizing manual intervention.
Collaborate on service lifecycle management: Partner with engineering teams to design, deploy, operate, and continuously improve robust and scalable systems and services, from inception to refinement.
Ensure service reliability and performance: Proactively monitor system health, conduct performance testing, and manage incidents to maximize uptime, availability, and adherence to defined SLAs/SLOs.
Execute core SRE practices: Perform on-call duties and production operations, including change management, capacity planning, and disaster recovery, while contributing to documentation and process improvements across teams.
Qualifications
Minimum Qualifications
- Proficient in one or more programming languages (e.g., Python, Go, Java, C++).
- Strong understanding of Linux operating systems and open-source technologies.
- Experience in network architecture and troubleshooting, database modeling, cloud systems, and large-scale distributed systems.
- Knowledge of monitoring tools and methodologies (such as Prometheus, Grafana), AIOPS, APM, Disaster Recovery.
- Experience in designing, analyzing, and building automation and tools for large-scale systems.
- Experience in building solutions with AWS, GCP, Azure, and other cloud services.
Preferred Qualifications
- Expertise in any of these tech stacks: Kubernetes, ElasticSearch, ClickHouse, Message Queue, OpenTSDB, Service Mesh, MySQL, Redis, etc.
- Master's degree in Computer Science, Engineering, or a related field.
As a condition of employment, all successful candidates must be able to establish authorization to work in the United States. For this position, the Company does not provide sponsorship for any immigration-related benefits.
About USDS
TikTok is the leading destination for short-form mobile video. Our mission is to inspire creativity and bring joy. U.S. Data Security ("USDS") is a subsidiary of TikTok in the U.S. This new, security-first division was created to bring heightened focus and governance to our data protection policies and content assurance protocols to keep U.S. users safe. Our focus is on providing oversight and protection of the TikTok platform and U.S. user data, so millions of Americans can continue turning to TikTok to learn something new, earn a living, express themselves creatively, or be entertained. The teams within USDS that deliver on this commitment daily span across Trust & Safety, Security & Privacy, Engineering, User & Product Ops, Corporate Functions and more.
Why Join Us
Inspiring creativity is at the core of TikTok's mission. Our innovative product is built to help people authentically express themselves, discover and connect – and our global, diverse teams make that possible. Together, we create value for our communities, inspire creativity and bring joy - a mission we work towards every day. We strive to do great things with great people. We lead with curiosity, humility, and a desire to make impact in a rapidly growing tech company. Every challenge is an opportunity to learn and innovate as one team. We're resilient and embrace challenges as they come. By constantly iterating and fostering an "Always Day 1" mindset, we achieve meaningful breakthroughs for ourselves, our company, and our users. When we create and grow together, the possibilities are limitless. Join us.
Diversity & Inclusion
TikTok is committed to creating an inclusive space where employees are valued for their skills, experiences, and unique perspectives. Our platform connects people from across the globe and so does our workplace. At TikTok, our mission is to inspire creativity and bring joy. To achieve that goal, we are committed to celebrating our diverse voices and to creating an environment that reflects the many communities we reach. We are passionate about this and hope you are too.
USDS Reasonable Accommodation
USDS is committed to providing reasonable accommodations in our recruitment processes for candidates with disabilities, pregnancy, sincerely held religious beliefs or other reasons protected by applicable laws. If you need assistance or a reasonable accommodation, please reach out to us at
Job Information
The base salary range for this position in the selected city is $118,657 - $259,200 annually. Compensation may vary outside of this range depending on a number of factors, including a candidate’s qualifications, skills, competencies and experience, and location. Base pay is one part of the Total Package that is provided to compensate and recognize employees for their work, and this role may be eligible for additional discretionary bonuses/incentives, and restricted stock units. Benefits may vary depending on the nature of employment and the country work location. Employees have day one access to medical, dental, and vision insurance, a 401(k) savings plan with company match, paid parental leave, short-term and long-term disability coverage, life insurance, wellbeing benefits, among others. Employees also receive 10 paid holidays per year, 10 paid sick days per year and 17 days of Paid Personal Time (prorated upon hire with increasing accruals by tenure). The Company reserves the right to modify or change these benefits programs at any time, with or without notice. For Los Angeles County (unincorporated) Candidates: Qualified applicants with arrest or conviction records will be considered for employment in accordance with all federal, state, and local laws including the Los Angeles County Fair Chance Ordinance for Employers and the California Fair Chance Act. Our company believes that criminal history may have a direct, adverse and negative relationship on the following job duties, potentially resulting in the withdrawal of the conditional offer of employment: 1. Interacting and occasionally having unsupervised contact with internal/external clients and/or colleagues; 2. Appropriately handling and managing confidential information including proprietary and trade secret information and access to information technology systems; and 3. Exercising sound judgment.
Seniority Level
Mid-Senior level
Employment Type
Full-time
Job Function
Engineering and Information Technology, Technology, Information and Internet
#J-18808-Ljbffr$118.66k - $259.2k
...Site Reliability Engineer, Cloud Infrastructure- USDS Site Reliability Engineer, Cloud Infrastructure- USDS Get AI-powered advice on this job and more exclusive features. Responsibilities The infrastructure team of US Tech Services Department at TikTok supports...SuggestedFull timeTemporary workInternshipWork at officeLocal area3 days per week$118.66k - $187.2k
...Site Reliability Engineer - Video Platform (Entry Level) - USDS Team Intro TikTok video system is a world‑leading video... ...system stability and save infrastructure costs. Provide strong support... ...AWS, Google, Azure and other cloud services is a plus. Passionate...SuggestedTemporary workWork experience placement$184k - $287.5k
At NVIDIA, Site Reliability Engineering provides a rare chance to define, develop, and support large-scale production systems with high efficiency... ...-tolerant, performant, and supportable.Background with infrastructure automation.Experience running critical services in...SuggestedFull time- Site Reliability Engineer, Data Platform - USDS Responsibilities Engage in and improve the whole lifecycle of service... ...-effective data, services and infrastructures. Maintain services once they are... ...services at scale and understanding cloud native technologies and...Suggested
$118.66k - $259.2k
...Site Reliability Engineer - AML Global Recommendation - USDS About the Team: Site Reliability Engineering (SRE) of the AML (Applied Machine Learning) team combines system engineering and the art of machine learning to develop and run a massively distributed AI/ML recommendation...SuggestedTemporary workWork at officeShift work3 days per week$187.04k - $359.72k
...E-commerce AI Platform team (USDS). Our mission is to build the... ...stakeholders to drive complex engineering efforts and align technical... ...and novel ideas to keep our infrastructure at the cutting edge. Qualifications... ...on a global scale. On-site presence across teams allows...Temporary workLocal areaFlexible hoursShift work$152k - $241.5k
...team is seeking a Senior System Software Engineer to help bring NVIDIA's autonomous... ...expansion of hardware-in-the-loop (HIL) infrastructure to support the robust deployment and lifecycle... ...(e.g., Jenkins, Docker, Kubernetes) in cloud-native or hybrid environments for mission...Full time$184k - $287.5k
We are seeking a Senior DevOps / Cloud Simulation Infrastructure Engineer to own the complete end-to-end cloud execution pipeline for SimReady assets... ...observable, debuggable, and production-ready.Operational Reliability: Implement atomic update semantics and safe failure...Full timeLocal area$168k - $264.5k
...NVIDIA's Digital Marketing Organization seeks a senior Site Reliability Engineer (SRE) to join our Santa Clara, CA team. As an SRE at NVIDIA... ...across deployment pipelines, Akamai CDN, WAF, and cloud infrastructure.Author, test, and activate shared and non-shared Cloudlet...Full time$280k - $380k
...an active-active, multi-cloud platform on AWS and GCP to... ...internet scale. We focus on reliability and automation, engineering systems that perform... ...can, and turning complex infrastructure into reliable, well‑documented... ...experienced DevOps/SRE (Site Reliability Engineering)...Work at officeLocal areaRemote workMonday to ThursdayFlexible hours- ...advance your career. THE ROLE:We are looking for a systems-minded engineer who lives at the intersection of large-scale model inference,... .... This role focuses on post-training and inference infrastructure, with particular emphasis on P/D disaggregation, KV cache lifecycle...
$168k - $270.25k
...streamlines cluster provisioning, workload management, and infrastructure monitoring. It provides all the tools you need to deploy... ...excellent, comprehensive support to our customers! Sr Site Reliability Engineer in this role will significantly impact and contribute to...Full timeWorldwide$124k - $271.2k
What You Can ExpectAs a Lead Staff Site Reliability Engineer, you will be one of the technical leads for our DevOps Platforms organization... ...group is responsible for DevOps Platforms including cloud infrastructure, physical data center orchestration, critical security...Full timeWork at officeRemote work$176k - $276k
Cloud Foundations Reliability (CFR) is part of NVIDIA’s Global Network Infrastructure (GNI) organization. We deploy, integrate, and operate the Kubernetes-based platform and... ....We are looking for a hands-on senior engineer to own the lifecycle and automation of the...Full timeRemote workWeekend work- Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI... ...physical data centers. We are looking for a Senior Site Reliability Engineer to improve the reliability, scalability, and...Work at officeLocal areaWork from homeFlexible hours
$272k - $431.25k
NVIDIA is looking for a Cloud Site Reliability Engineering Architect to work in IPP's (Infrastructure, Planning and Process) Cloud Infrastructure Team. IPP is a global organization within NVIDIA. This group works with various other groups within NVIDIA such as Graphics...Full timeWork experience placementWorldwide- Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers... ...home day is currently Tuesday.Engineering at Lambda is responsible for building... ..., upgrades, and scaling.Own the reliability, performance, and security of...Work at officeLocal areaWork from homeFlexible hours
$178k - $321k
...markets. We are safe and reliable, backed by our Proof... ...harness: a resilient cloud platform, the agentic... ...governed data and AI infrastructure everything else depends... ...This is a two-person engineering team: you deploy, debug... ...internal or external careers site.Notice:All official...$184k - $287.5k
.... We are looking for a Senior Software Engineer to lead the bring-up, triage, benchmarking... ...-art LLM workloads run efficiently and reliably at scale. You will lead deep... ...debugging of large-scale AI clusters, infrastructure, and end-to-end workloads, setting the...Full timeRemote work$100k
...all seniorities.Tenstorrent’s AI Software Infrastructure team builds the platforms that power... ...You AreStrong backend or infrastructure engineer with experience building and operating platforms... ...data center infrastructure differs from cloud-native environments at scale.How...Permanent employment$116k - $189.75k
...large language model workloads. We are looking for a Software Engineer focused on bring-up, triage, benchmarking, analysis, and... ...doing:Bring up, validate, and debug large-scale AI clusters, infrastructure, and end-to-end workloads.Bring up, tune, and benchmark AI pre...Full timeRemote work- Technical Cloud Architect (Virtualization & Cloud Infrastructure)This role has been designed as 'Hybrid' with a requirement that you will work on average 2... ...impact role that sits at the intersection of pre-sales engineering, solutions architecture, and product strategy. You...Full timeWork experience placementWork at officeLocal areaImmediate start2 days per week
- Senior Site Reliability Engineer ILocationSan Jose, Costa Rica - RemoteSummary of roleOwn availability... ..., increase efficiency in our use of cloud resources and our developer’s time,... ...creation to production supportVersed in Infrastructure as Code practices using technologies...Flexible hours
$152k - $241.5k
...and network fabrics.Use IaC(Infrastructure‑as‑Code) and config management... ...globally distributed, multi‑cloud hybrid environment - On‑prem... ...lifecycle management, fleet reliability/auto-healing, E2E... ...Perl, or Ruby.Mentored other engineers and influenced technical direction...Full time$128.6k - $184.9k
...scalability, and efficiency of the infrastructure that powers our global cloud platform. As a team of six engineers distributed across the US,... ...strong focus on automation, reliability, and operational excellence.... ...7+ years of experience in Site Reliability Engineering,...Permanent employmentFull timeTemporary workLocal areaWorldwideFlexible hours$168k - $270.25k
NVIDIA is looking for a Senior Site Reliability Engineer (SRE) to join its GeForce Now (GFN) team. SRE... ...our internal and external-facing GPU cloud gaming services have reliability and... ...to and mitigating high-severity infrastructure alerts and service degradations.Ways...Full time$267k - $356k
..., The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands... ....Lambda's Storage Engineering team is the backbone behind... ...the industry, which means reliability and performance aren't just... ...across new and existing sites using tools such as...Work experience placementWork at officeLocal areaWork from homeFlexible hours$230k - $250k
...autonomous networking, giving engineers and AI agents the ability... ...across every major cloud and vendor environment.Global... ...done.Forward is looking for a Site Reliability EngineerAbout the Role This... ...closely with engineering, infrastructure, and product to ensure our...Night shift$168k - $270.25k
....Join our team at NVIDIA as a Senior Site reliability engineer focused on HPC storage and play a crucial... ...while harnessing the power of cloud computing. You will be responsible for... ...closely with engineering teams to align infrastructure with their evolving needs, document...Full time$101k - $161k
...in data-driven, client-to-cloud networking for large data center... ...awards, such as Best Engineering Team, Best Company for Diversity... ...Work WithWe’re looking for Site Reliability Engineers to join our... ...microservices stack, monitoring infrastructure, and much more. What You'll...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer, Cloud Infrastructure - USDS. Be the first to apply!
- site reliability engineer San Jose, CA
- site reliability engineer sre San Jose, CA
- senior cloud security engineer San Jose, CA
- senior cloud solutions architect San Jose, CA
- senior cloud data engineer San Jose, CA
- cloud engineering manager San Jose, CA
- informatica cloud developer San Jose, CA
- senior cloud engineer San Jose, CA
- cloud architect San Jose, CA
- aws cloud architect San Jose, CA

