Site Reliability Engineer - Compute Platform
TikTok
Team introductionOur Compute Platform SRE team supports all Big Data services and products across the company. We are a newly established team and waiting for talents like you to shape the team's future together. We are responsible for the reliability of all the company's major data warehouse products, services, and query engines. We serve business needs across domains within TikTok. We look forward to welcoming you to the team.Responsibilities:- Responsible for the reliability of all TikTok's major data warehouse products, services, and query engines, such as ClickHouse, Spark, Presto, Doris, etc.- Uphold Service Level Agreements (SLAs): Ensure that all service level objectives and agreements from ByteDance's Data Platform services are met. Respond promptly to any system outages or issues.- Continuous Performance Optimization: Analyze service performance and reliability patterns to identify potential performance bottlenecks. Implement proactive measures to prevent service disruptions. Work with development teams to optimize application performance, ensuring that services run efficiently and that resources are utilized effectively.- Incident Management: Lead efforts to troubleshoot and resolve service incidents and postmortems. Coordinate with cross-functional teams to manage and mitigate service-impacting events.- Infrastructure Automation: Automate infrastructure provisioning, scaling, and management processes to reduce manual interventions and improve service quality.- Collaboration: Engage with product and development teams to integrate reliability and performance considerations into the software lifecycle.- Capacity and Demand Planning: Assess and forecast infrastructure needs based on growth patterns and upcoming initiatives.- Stay Updated: Keep current with industry trends, best practices, and emerging technologies related to site reliability and infrastructure engineering.Minimal Qualifications:- Bachelor's Degree or above, in Computer Science, Engineering, or a related field. Passionate about computer science and Internet technology.- In-depth understanding of Linux, computer networking, and databases. Proficient in common SRE/DevOps open-source toolsets, system monitoring tools, and container orchestration platforms like Kubernetes.- Experience or familiarity with open-source or commercial technologies such as ClickHouse, Hadoop, Doris, Spark, Presto and Kubernetes.- Strong coding skills in at least one scripting or programming language, including but not limited to Python, Shell, Java, Go, etc.- Excellent problem-solving skills and the ability to think critically.Req ID: A237226
- Overview Site Reliability Engineer, Compute - USDS TikTok is the leading destination for short-form mobile video. U.S. Data Security (USDS) is a subsidiary... ...Experience with containers and container orchestration platforms such as Docker, Kubernetes or equivalent. Candidates...SuggestedWork experience placement
- A leading social media platform based in Seattle is seeking a Site Reliability Engineer for its U.S. Data Security division. The role involves developing automation... .... Candidates should hold a Bachelor's degree in Computer Science with 3+ years of experience and proficiency...Suggested
$180.5k - $225.6k
...AI and data infrastructure platform so our customers can use deep... ...their business. Founded by engineers — and customer obsessed — we... ...getting started.The Serverless Compute Platform is the backbone of... ...with production-grade reliability spanning a range of use cases...SuggestedLocal areaRemote workWorldwide$253k - $416k
...DescriptionDistinguished Software Engineer, Systems Infrastructure - Compute, Deployment, Infrastructure as Code... ...of LinkedIn’s Core Infrastructure platform, which provides shared system... ...architecture, balancing scalability, reliability, and costHelp evolve LinkedIn’s private...SuggestedFor contractorsWork experience placementWork at officeFlexible hours$166k - $258k
...considered for this position.Nordstrom is looking for a Senior Engineer 2 to join our Site Reliability Engineering (SRE) team — and we think that person could... ....You own this if you have…A bachelor’s degree in computer science, Engineering, or equivalent experience10+ years...SuggestedFull timeWork at office$143k - $191k
...cutting-edge autonomy, AI, computer vision, sensor fusion, and networking... .... System Deployment Engineers work in complex environments... ...as a data and applications platform and be able to educate customersParticipate... ...and exercisesWork with site reliability engineers to provide and...Full timeTemporary workWork experience placementImmediate start$134.25k - $214.8k
...matter.Your ImpactAre you an engineer who gets excited about the... ...next-generation observability platform, enabling the entire... ...Observability team within Axon's Site Reliability organization — a focused team... ...QualificationsBachelor's Degree in Computer Science, Engineering, or an...Work experience placementWork at officeRemote work$184.9k - $250.2k
...evolving to support fault-tolerant quantum computing. We need a Senior SDE who can build... ...infrastructure for quantum workloads - the kind of engineer who thinks about latency budgets,... ...or architecture (design patterns, reliability and scaling) of new and existing...InternshipFlexible hours$143.7k - $194.4k
Serverless Compute ( is changing the way we think about computing in the cloud. Serverless... ...work with team to build the new generic platform by using latest AWS technologies. You... ...design or architecture (design patterns, reliability and scaling) of new and existing systems...InternshipFlexible hours$194k - $267k
...concepts and tools.Position Overview:The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications... ...containerization principles.Bachelor’s degree in Computer Science, Engineering, or related field (...Permanent employmentWork at officeLocal areaWorldwideFlexible hours- ...place.As a Principal Software Engineer at JPMorganChase within the Core Foundational Platforms team, you provide expertise and... ...practices in high-performance computing (HPC) practices that intersect... ...comprehensive health care coverage, on-site health and wellness centers, a...
$160k - $250k
...emphasis on distributed high performance computing integrating GPUs. Even with these data... ..., we also need to grow our DevOps and Site Reliability team to maintain the reliability of... ...Manage a diverse array of technology platforms, following best practices and procedures...$127k - $249k
...looking for an experienced Senior or Staff Engineer for our SRE, InfraSec team, to guide the... ...implement controls that reinforce the platform’s security posture.This is an SRE team,... ...AWS, Azure, GCP), including network and compute security, identity management, and cloud...Local areaRemote workWorldwideFlexible hours- ...are in. About this team Site Reliability Engineering We are looking for a motivated engineer... ...~ Bachelor's degree in computer science/engineering or equivalent... ...monitoring solutions, log aggregation platforms, and distributed tracing frameworks...
- Apple in Seattle seeks an Engineering Manager to lead the Compute Node Software team, shaping cloud compute platforms for containers, VMs, and related infrastructure. You will steer vision, architecture, and execution, balancing security, performance, and Apple’s privacy...
- ...Whoever deploys frontier compute infrastructure fastest... .... The Production Engineering Team Examples of... ...month build window: our platform covers burn-in,... ...production-ready" before a site goes live, not after.... ...on. Own end-to-end reliability, scalability, and operation...Local area
- ...is responsible for the reliability, scalability, and... ...building features, but about engineering the resilience and... ...performance of the underlying platform that all product teams... ...system stability.As a Site Reliability Engineer,... ...Bachelor’s degree in Computer Science, a related...
- ...team is responsible for the reliability, scalability, and efficiency... ...building features, but about engineering the resilience and performance of the underlying platform that all product teams depend... ...Qualifications:- Bachelor’s degree in Computer Science, a related technical...
$174k - $252k
Senior Software Engineer, Site Reliability Engineering corporate_fare Google place Seattle, WA, USA ;... ...Kirkland, WA, USA . Bachelor’s degree in Computer Science, Engineering, a related field... ...the next generation of Google platforms, we make Google's product portfolio possible...Temporary work$133.2k - $219.6k
Site Reliability Engineer of Container Service Direct message the job poster from Alibaba Cloud Global Talent Acquisition Talent Sourcer Job Description... ...of Linux systems, Alibaba Cloud services for computing, storage, and networking. Strong ownership and results-driven...Full time$129.96k - $246.24k
Site Reliability Engineer, Product - USDS 3 weeks ago Be among the first 25 applicants Responsibilities... ...Bachelor or above degree in Computer Science or a related technical discipline... ...oversight and protection of the TikTok platform and U.S. user data, so millions of...Full timeTemporary workWork at officeLocal area3 days per week$129.96k - $246.24k
Overview Site Reliability Engineer, Edge Services - USDS Base pay range: $129,960.00/yr - $246,240.00/yr Responsibilities Architect and implement... ...: Bachelor\'s degree with 2+ years of experience in Computer Engineering, Computer Science, or related fields, or equivalent...$56 per hour
...Software Developer, Scientific Computer & DevOps Type: Long-term... ...that support long-term reliability and sustainability. Design... ...steps. Collaborate with IT, engineering analysts, and software... ...Experience with relevant tools and platforms such as SLURM, AzHOP,...Long term contractTemporary workLocal areaRemote work- ...Lambda's mission is to make compute as ubiquitous as electricity... ...Tuesday. The Compute Software Engineering role plays a key part in the... ...contribute to and evolve the internal platform tooling, and workflows. As... ...software efficiently and reliably. The position works closely...Work at officeLocal areaWork from homeFlexible hours
$196k - $230k
..., and so are the rewards. The Software Platform team accelerates developer velocity and increases system reliability by building the foundational platforms and tools that power the company engineering. Within this group, Compute team focuses on building and operating a...Work at officeFlexible hoursShift work3 days per week$176k - $333.5k
We are looking for a Senior Software Engineer for our Autonomous Vehicle efforts within the... ...state of the art Machine Learning and Computer Vision techniques to generate a variety... ...background in software development. We build platforms, web services, and tools to ingest...- ...to add talented individuals to our team.Computer vision plays an indispensable role in modern... ..., amongst others. Computer vision engineers at Valve are working on all those areas to... ...planGenerous vacation and family leaveOn-site amenities in support of health and efficiencyFertility...Flexible hours
$179k - $294k
Senior Software Engineer - High Performance Computing Design and implement improvements to Zoox's cutting‑edge... ..., ensuring high performance and reliability for our machine learning workloads... ...applications. Hands‑on experience with cloud platforms (AWS, GCP, Azure), using their...Temporary work$124.9k - $228.9k
...Handling over 1 trillion queries per day, our platform operates at an unprecedented scale. We... .... About the Role The High-Performance Computing team powers core bidding platform by... ...as our platform evolves. We expect our engineers to be end-to-end owners. You will participate...Full timeTemporary workLocal area- Crusoe Cloud is revolutionizing HPC by offering sustainable, low-cost GPU compute power. As a Senior Cloud Support Engineer, you will be the primary technical support contact, helping customers leverage Crusoe Cloud to achieve their research goals and accelerate development...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer - Compute Platform. Be the first to apply!
- site reliability engineer remote Seattle, WA
- site reliability engineer Seattle, WA
- site reliability engineer sre Seattle, WA
- senior platform engineer Seattle, WA
- platform developer Seattle, WA
- platform engineer Seattle, WA
- client platform engineer Seattle, WA
- platform engineering manager Seattle, WA
- data platform engineer Seattle, WA
- junior website developer Seattle, WA


