Site Reliability Engineer - Compute Platform
TikTok
Team introductionOur Compute Platform SRE team supports all Big Data services and products across the company. We are a newly established team and waiting for talents like you to shape the team's future together. We are responsible for the reliability of all the company's major data warehouse products, services, and query engines. We serve business needs across domains within TikTok. We look forward to welcoming you to the team.Responsibilities:- Responsible for the reliability of all TikTok's major data warehouse products, services, and query engines, such as ClickHouse, Spark, Presto, Doris, etc.- Uphold Service Level Agreements (SLAs): Ensure that all service level objectives and agreements from ByteDance's Data Platform services are met. Respond promptly to any system outages or issues.- Continuous Performance Optimization: Analyze service performance and reliability patterns to identify potential performance bottlenecks. Implement proactive measures to prevent service disruptions. Work with development teams to optimize application performance, ensuring that services run efficiently and that resources are utilized effectively.- Incident Management: Lead efforts to troubleshoot and resolve service incidents and postmortems. Coordinate with cross-functional teams to manage and mitigate service-impacting events.- Infrastructure Automation: Automate infrastructure provisioning, scaling, and management processes to reduce manual interventions and improve service quality.- Collaboration: Engage with product and development teams to integrate reliability and performance considerations into the software lifecycle.- Capacity and Demand Planning: Assess and forecast infrastructure needs based on growth patterns and upcoming initiatives.- Stay Updated: Keep current with industry trends, best practices, and emerging technologies related to site reliability and infrastructure engineering.Minimal Qualifications:- Bachelor's Degree or above, in Computer Science, Engineering, or a related field. Passionate about computer science and Internet technology.- In-depth understanding of Linux, computer networking, and databases. Proficient in common SRE/DevOps open-source toolsets, system monitoring tools, and container orchestration platforms like Kubernetes.- Experience or familiarity with open-source or commercial technologies such as ClickHouse, Hadoop, Doris, Spark, Presto and Kubernetes.- Strong coding skills in at least one scripting or programming language, including but not limited to Python, Shell, Java, Go, etc.- Excellent problem-solving skills and the ability to think critically.Req ID: A237226
- A leading social media platform based in Seattle is seeking a Site Reliability Engineer for its U.S. Data Security division. The role involves developing automation... .... Candidates should hold a Bachelor's degree in Computer Science with 3+ years of experience and proficiency...Suggested
- ...Lululemon Site Reliability Engineering Engineer We are a yoga-inspired technical apparel company up... ...Qualifications ~ Bachelor’s degree in computer science/engineering or equivalent ~... ...solutions, log aggregation platforms, and distributed tracing frameworks...Suggested
$130k - $200k
...Site Reliability Engineer Nscale is the GPU cloud built for AI. We run high-performance, cost-efficient... ..., from bare metal up through the platform services teams actually build on. Our... ...GPU workloads, or high-performance computing (HPC). • Familiarity with high-performance...SuggestedShift work$160k - $250k
...DevOps And Systems Engineer Hive is the leading provider of cloud... ...high performance computing integrating GPUs. Even with... ...need to grow our DevOps and Site Reliability team to maintain the reliability... ...diverse array of technology platforms, following best practices and...Suggested$194k - $267k
...Staff Site Reliability Engineer - Kubernetes Important: if an employer asks you to log into their... ...in building and managing Kubernetes platforms that support cloud-native applications... ...principles. Bachelor’s degree in Computer Science, Engineering, or related field...SuggestedPermanent employmentWork at officeLocal areaWorldwideFlexible hours$160k - $210k
...buying with our Deep Learning Advertising Platform. Since 2015, we have harnessed the... ...! We are looking for a Senior Site Reliability Engineer to strengthen our AWS infrastructure... ...Evaluate our existing AWS architecture (compute, networking, security) and ensure we...Work at officeImmediate startRemote workWork from home$179k - $294k
Senior Software Engineer - High Performance Computing Design and implement improvements to Zoox's cutting‑edge... ..., ensuring high performance and reliability for our machine learning workloads... ...applications. Hands‑on experience with cloud platforms (AWS, GCP, Azure), using their...Temporary work$56 per hour
...Software Developer, Scientific Computer & DevOps Type: Long-term... ...approaches that support long-term reliability and sustainability. Design... ...steps. Collaborate with IT, engineering analysts, and software... ...Experience with relevant tools and platforms such as SLURM, AzHOP,...Long term contractLocal areaRemote work$204k - $306k
...If you are too, let's talk. Manager, Site Reliability Engineering San Francisco, California Secure... ...As the Manager of Infrastructure Platform and Shared Services, you will oversee... ...communication and interpersonal skills ~ Computer Science Degree or related degree or...Permanent employmentWork at officeLocal areaWorldwideFlexible hours2 days per week- C++ Engineer - High Performance Computing (HPC) About the Role What if your deep systems programming expertise could directly shape the infrastructure... ...annotation, validation, and quality control Improve reliability, performance, and safety across existing C++ codebases...Hourly payOngoing contractContract workFreelanceRemote workFlexible hours
- Overview As a Principal Software Engineer at JPMorganChase within the Core Foundational Platforms team, you provide expertise and... ...practices in high-performance computing (HPC) practices that intersect... ...comprehensive health care coverage, on-site health and wellness centers, a...
$405k
...Build core systems tracking datacenter sites, designs, bills of materials, equipment,... ...construction management, scheduling, and engineering systems. Design fine-grained permissions... ...users, and prioritize work that brings compute online sooner. Partner with research...Full timeWork at officeVisa sponsorshipFlexible hours- ...connect the orbital economy. As a Software Engineer, Reliability, you will ensure the software powering our avionics and high‑performance compute systems operates flawlessly in orbit.... ...frameworks and testing compute platforms under mission‑representative conditions...
- ...Arm is seeking an engineer on the AI Compute Infra team in Seattle to design, build, and operate large-scale infrastructure for AI training,... ...accelerator enablement, and high-performance networking to improve reliability and performance. Responsibilities include building and...
$262.7k - $355.4k
As a Principal Software Engineer on the AI Compute Platform team, you will design and build a secure, reliable, and easy-to-use platform for running AI workloads at scale. You will develop control-plane services, APIs, orchestration systems, and developer tools for distributed...Work at officeLocal areaVisa sponsorshipRelocation package$180.5k - $225.6k
...AI and data infrastructure platform so our customers can use deep... ...their business. Founded by engineers — and customer obsessed — we... ...getting started.The Serverless Compute Platform is the backbone of... ...with production-grade reliability spanning a range of use cases...Local areaRemote workWorldwide- SID Global Solutions in Bellevue, WA is seeking a Computer Vision Engineer to lead end-to-end cross-functional projects that integrate IIoT hardware... ...experience, and excellent program and stakeholder management across teams and sites. #J-18808-Ljbffr SID Global Solutions
$232k - $319k
...s talk. The Infrastructure Platform and Shared Services Team... ...service with great people and reliable, cost-effective, and efficient... ...velocity of SRE and product engineering by developing robust platforms... ...interpersonal skills ~ Computer Science Degree or related degree...Permanent employmentLocal areaWorldwideFlexible hours- ...provider is seeking a Senior Staff Technical Program Manager for Reliability to enhance the reliability and performance of their multi-... ...role involves leading programs in partnership with senior engineering leaders, requiring over 10 years of experience in cloud infrastructure...
$166k - $244k
# Senior Software Engineer, Site Reliability EngineeringGoogle • onsite • 601 N 34th St, Seattle, WA 98103, USA • full\_timePay: USD 166000.00 - USD 244000.00 / unspecifiedBusinesses of all shapes and sizes rely on Google’s unparalleled advertising solutions to help them...Temporary work$198.36k - $416.1k
...Team Intro TikTok video system is a world-leading video platform that provides multimedia storage, delivery, transcoding... ...Provide technical leadership and mentorship to a team of Site Reliability Engineers focused on building observable, fault-tolerant video services...Temporary workShift work- ...JPMorganChase in Seattle seeks a Lead Software Engineer to join the Enterprise Technology, Infrastructure Platforms team. You will act as a core technical contributor, delivering trusted, scalable technology across multiple domains, while guiding AI-assisted engineering...
- ...Nscale is seeking a Senior Site Reliability Engineer to own the reliability bar for AI infrastructure operations. You will tackle the hardest... ...challenges, mentor peers, and shape SRE practices across the platform. You will work closely with teams running AI workloads...
- ...Robinhood is seeking a Staff Software Engineer for its Storage Platform in Bellevue, WA. You will design and evolve core storage infrastructure—relational... ...databases, key‑value systems, and caching—while driving reliability, performance, and cost efficiency across multi‑region...
- ...ByteDance’s Infrastructure Engineering team in Seattle designs, builds, and operates global infrastructure spanning public and private... ...storage. Join a fast-paced, collaborative team focused on reliability, scalability, and continuous optimization, driving improvements...
- Job Title Required Skills: CHEF experience - Must have most critical Azure Cloud – experience - Must have most critical AKS- Azure Kubernetes services - Must have most critical Kubernetes - Must have most critical NoSQL DB – Cassandra / Mongo DB ...
- ...Sr. Site Reliability Engineer Comtech is a woman-owned small business founded in 1998 and headquartered in Reston, VA. We offer IT solutions... ...going into API and handling code /bugs /error etc. 2. Platform experience = Chef, Puppet, Azure, Ansible. Good hands on in...Local area
- ...Starbucks Senior Site Reliability Engineer (Cloud) This position contributes to Starbucks on their Data Platform Services team. This team maintains and improves the data platform that many Starbucks services are dependent on. When you order coffee with your rewards...Local areaWorldwide
$134.25k - $214.8k
...Sr. Site Reliability Engineer I Seattle, Washington, United States At Axon, we're on a mission to Protect Life. We're explorers, pursuing... ...of our mission-critical, cloud-native global Kubernetes platform and the services that run on it. You care deeply about system...Work experience placementWork at officeRemote workFlexible hours- Technical/Functional Skills Windows Servers, Digital: Microsoft Azure Windows Powershell, Digital: DevOps Roles & Responsibilities Windows Server 2012 -2019 Administration Microsoft Azure Azure AAD DFSR, DHCP DNS, KMS, WSUS TCP/IP Hyper-V High Availability Clusters ...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer - Compute Platform. Be the first to apply!
- site reliability engineer sre Seattle, WA
- site reliability engineer Seattle, WA
- data platform engineer Seattle, WA
- platform engineer Seattle, WA
- senior platform engineer Seattle, WA
- platform engineering manager Seattle, WA
- platform developer Seattle, WA
- client platform engineer Seattle, WA
- site recruiter Seattle, WA
- site services specialist Seattle, WA


