Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer - Compute Platform

TikTok

Team introductionOur Compute Platform SRE team supports all Big Data services and products across the company. We are a newly established team and waiting for talents like you to shape the team's future together. We are responsible for the reliability of all the company's major data warehouse products, services, and query engines. We serve business needs across domains within TikTok. We look forward to welcoming you to the team.Responsibilities:- Responsible for the reliability of all TikTok's major data warehouse products, services, and query engines, such as ClickHouse, Spark, Presto, Doris, etc.- Uphold Service Level Agreements (SLAs): Ensure that all service level objectives and agreements from ByteDance's Data Platform services are met. Respond promptly to any system outages or issues.- Continuous Performance Optimization: Analyze service performance and reliability patterns to identify potential performance bottlenecks. Implement proactive measures to prevent service disruptions. Work with development teams to optimize application performance, ensuring that services run efficiently and that resources are utilized effectively.- Incident Management: Lead efforts to troubleshoot and resolve service incidents and postmortems. Coordinate with cross-functional teams to manage and mitigate service-impacting events.- Infrastructure Automation: Automate infrastructure provisioning, scaling, and management processes to reduce manual interventions and improve service quality.- Collaboration: Engage with product and development teams to integrate reliability and performance considerations into the software lifecycle.- Capacity and Demand Planning: Assess and forecast infrastructure needs based on growth patterns and upcoming initiatives.- Stay Updated: Keep current with industry trends, best practices, and emerging technologies related to site reliability and infrastructure engineering.Minimal Qualifications:- Bachelor's Degree or above, in Computer Science, Engineering, or a related field. Passionate about computer science and Internet technology.- In-depth understanding of Linux, computer networking, and databases. Proficient in common SRE/DevOps open-source toolsets, system monitoring tools, and container orchestration platforms like Kubernetes.- Experience or familiarity with open-source or commercial technologies such as ClickHouse, Hadoop, Doris, Spark, Presto and Kubernetes.- Strong coding skills in at least one scripting or programming language, including but not limited to Python, Shell, Java, Go, etc.- Excellent problem-solving skills and the ability to think critically.Req ID: A237226

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer - Compute Platform in Seattle, WA vacancy
  • Overview Site Reliability Engineer, Compute - USDS TikTok is the leading destination for short-form mobile video. U.S. Data Security (USDS) is a subsidiary...  ...Experience with containers and container orchestration platforms such as Docker, Kubernetes or equivalent. Candidates... 
    Suggested
    Work experience placement

    TikTok

    Seattle, WA
    4 days ago
  • A leading social media platform based in Seattle is seeking a Site Reliability Engineer for its U.S. Data Security division. The role involves developing automation...  .... Candidates should hold a Bachelor's degree in Computer Science with 3+ years of experience and proficiency... 
    Suggested

    TikTok

    Seattle, WA
    4 days ago
  • $180.5k - $225.6k

     ...AI and data infrastructure platform so our customers can use deep...  ...their business. Founded by engineers — and customer obsessed — we...  ...getting started.The Serverless Compute Platform is the backbone of...  ...with production-grade reliability spanning a range of use cases... 
    Suggested
    Local area
    Remote work
    Worldwide

    DataBricks

    Bellevue, WA
    1 day ago
  • $253k - $416k

     ...DescriptionDistinguished Software Engineer, Systems Infrastructure - Compute, Deployment, Infrastructure as Code...  ...of LinkedIn’s Core Infrastructure platform, which provides shared system...  ...architecture, balancing scalability, reliability, and costHelp evolve LinkedIn’s private... 
    Suggested
    For contractors
    Work experience placement
    Work at office
    Flexible hours

    Linkedin

    Bellevue, WA
    2 days ago
  • $166k - $258k

     ...considered for this position.Nordstrom is looking for a Senior Engineer 2 to join our Site Reliability Engineering (SRE) team — and we think that person could...  ....You own this if you have…A bachelor’s degree in computer science, Engineering, or equivalent experience10+ years... 
    Suggested
    Full time
    Work at office

    Nordstrom

    Seattle, WA
    3 days ago
  • $143k - $191k

     ...cutting-edge autonomy, AI, computer vision, sensor fusion, and networking...  .... System Deployment Engineers work in complex environments...  ...as a data and applications platform and be able to educate customersParticipate...  ...and exercisesWork with site reliability engineers to provide and... 
    Full time
    Temporary work
    Work experience placement
    Immediate start

    Anduril Industries

    Seattle, WA
    4 days ago
  • $134.25k - $214.8k

     ...matter.Your ImpactAre you an engineer who gets excited about the...  ...next-generation observability platform, enabling the entire...  ...Observability team within Axon's Site Reliability organization — a focused team...  ...QualificationsBachelor's Degree in Computer Science, Engineering, or an... 
    Work experience placement
    Work at office
    Remote work

    Axon

    Seattle, WA
    21 hours ago
  • $184.9k - $250.2k

     ...evolving to support fault-tolerant quantum computing. We need a Senior SDE who can build...  ...infrastructure for quantum workloads - the kind of engineer who thinks about latency budgets,...  ...or architecture (design patterns, reliability and scaling) of new and existing... 
    Internship
    Flexible hours

    Amazon

    Seattle, WA
    4 days ago
  • $143.7k - $194.4k

    Serverless Compute ( is changing the way we think about computing in the cloud. Serverless...  ...work with team to build the new generic platform by using latest AWS technologies. You...  ...design or architecture (design patterns, reliability and scaling) of new and existing systems... 
    Internship
    Flexible hours

    Amazon

    Seattle, WA
    2 days ago
  • $194k - $267k

     ...concepts and tools.Position Overview:The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications...  ...containerization principles.Bachelor’s degree in Computer Science, Engineering, or related field (... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta

    Bellevue, WA
    1 day ago
  •  ...place.As a Principal Software Engineer at JPMorganChase within the Core Foundational Platforms team, you provide expertise and...  ...practices in high-performance computing (HPC) practices that intersect...  ...comprehensive health care coverage, on-site health and wellness centers, a... 

    JP Morgan Chase

    Seattle, WA
    4 days ago
  • $160k - $250k

     ...emphasis on distributed high performance computing integrating GPUs. Even with these data...  ..., we also need to grow our DevOps and Site Reliability team to maintain the reliability of...  ...Manage a diverse array of technology platforms, following best practices and procedures... 

    Hive

    Seattle, WA
    3 days ago
  • $127k - $249k

     ...looking for an experienced Senior or Staff Engineer for our SRE, InfraSec team, to guide the...  ...implement controls that reinforce the platform’s security posture.This is an SRE team,...  ...AWS, Azure, GCP), including network and compute security, identity management, and cloud... 
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    Seattle, WA
    4 days ago
  •  ...are in. About this team Site Reliability Engineering We are looking for a motivated engineer...  ...~ Bachelor's degree in computer science/engineering or equivalent...  ...monitoring solutions, log aggregation platforms, and distributed tracing frameworks... 

    Kaav Inc.

    Seattle, WA
    2 days ago
  • Apple in Seattle seeks an Engineering Manager to lead the Compute Node Software team, shaping cloud compute platforms for containers, VMs, and related infrastructure. You will steer vision, architecture, and execution, balancing security, performance, and Apple’s privacy... 

    Apple Inc.

    Seattle, WA
    1 day ago
  •  ...Whoever deploys frontier compute infrastructure fastest...  .... The Production Engineering Team Examples of...  ...month build window: our platform covers burn-in,...  ...production-ready" before a site goes live, not after....  ...on. Own end-to-end reliability, scalability, and operation... 
    Local area

    Fluidstack

    Seattle, WA
    3 days ago
  •  ...is responsible for the reliability, scalability, and...  ...building features, but about engineering the resilience and...  ...performance of the underlying platform that all product teams...  ...system stability.As a Site Reliability Engineer,...  ...Bachelor’s degree in Computer Science, a related... 

    TikTok

    Seattle, WA
    1 day ago
  •  ...team is responsible for the reliability, scalability, and efficiency...  ...building features, but about engineering the resilience and performance of the underlying platform that all product teams depend...  ...Qualifications:- Bachelor’s degree in Computer Science, a related technical... 

    TikTok

    Seattle, WA
    1 day ago
  • $174k - $252k

    Senior Software Engineer, Site Reliability Engineering corporate_fare Google place Seattle, WA, USA ;...  ...Kirkland, WA, USA . Bachelor’s degree in Computer Science, Engineering, a related field...  ...the next generation of Google platforms, we make Google's product portfolio possible... 
    Temporary work

    Google Inc.

    Seattle, WA
    2 days ago
  • $133.2k - $219.6k

    Site Reliability Engineer of Container Service Direct message the job poster from Alibaba Cloud Global Talent Acquisition Talent Sourcer Job Description...  ...of Linux systems, Alibaba Cloud services for computing, storage, and networking. Strong ownership and results-driven... 
    Full time

    Alibaba Cloud

    Seattle, WA
    4 days ago
  • $129.96k - $246.24k

    Site Reliability Engineer, Product - USDS 3 weeks ago Be among the first 25 applicants Responsibilities...  ...Bachelor or above degree in Computer Science or a related technical discipline...  ...oversight and protection of the TikTok platform and U.S. user data, so millions of... 
    Full time
    Temporary work
    Work at office
    Local area
    3 days per week

    TikTok

    Seattle, WA
    4 days ago
  • $129.96k - $246.24k

    Overview Site Reliability Engineer, Edge Services - USDS Base pay range: $129,960.00/yr - $246,240.00/yr Responsibilities Architect and implement...  ...: Bachelor\'s degree with 2+ years of experience in Computer Engineering, Computer Science, or related fields, or equivalent... 

    TikTok

    Seattle, WA
    4 days ago
  • $56 per hour

     ...Software Developer, Scientific Computer & DevOps Type: Long-term...  ...that support long-term reliability and sustainability. Design...  ...steps. Collaborate with IT, engineering analysts, and software...  ...Experience with relevant tools and platforms such as SLURM, AzHOP,... 
    Long term contract
    Temporary work
    Local area
    Remote work

    System One

    Bellevue, WA
    17 days ago
  •  ...Lambda's mission is to make compute as ubiquitous as electricity...  ...Tuesday. The Compute Software Engineering role plays a key part in the...  ...contribute to and evolve the internal platform tooling, and workflows. As...  ...software efficiently and reliably. The position works closely... 
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Corporation

    Bellevue, WA
    3 days ago
  • $196k - $230k

     ..., and so are the rewards. The Software Platform team accelerates developer velocity and increases system reliability by building the foundational platforms and tools that power the company engineering. Within this group, Compute team focuses on building and operating a... 
    Work at office
    Flexible hours
    Shift work
    3 days per week

    United States Digital Space LLC

    Bellevue, WA
    21 hours ago
  • $176k - $333.5k

    We are looking for a Senior Software Engineer for our Autonomous Vehicle efforts within the...  ...state of the art Machine Learning and Computer Vision techniques to generate a variety...  ...background in software development. We build platforms, web services, and tools to ingest... 

    NVIDIA

    Seattle, WA
    4 days ago
  •  ...to add talented individuals to our team.Computer vision plays an indispensable role in modern...  ..., amongst others. Computer vision engineers at Valve are working on all those areas to...  ...planGenerous vacation and family leaveOn-site amenities in support of health and efficiencyFertility... 
    Flexible hours

    Valve Software

    Bellevue, WA
    1 day ago
  • $179k - $294k

    Senior Software Engineer - High Performance Computing Design and implement improvements to Zoox's cutting‑edge...  ..., ensuring high performance and reliability for our machine learning workloads...  ...applications. Hands‑on experience with cloud platforms (AWS, GCP, Azure), using their... 
    Temporary work

    jobs.frontdoordefense.com - Jobboard

    Seattle, WA
    3 days ago
  • $124.9k - $228.9k

     ...Handling over 1 trillion queries per day, our platform operates at an unprecedented scale. We...  .... About the Role The High-Performance Computing team powers core bidding platform by...  ...as our platform evolves. We expect our engineers to be end-to-end owners. You will participate... 
    Full time
    Temporary work
    Local area

    The Trade Desk, Inc.

    Bellevue, WA
    3 days ago
  • Crusoe Cloud is revolutionizing HPC by offering sustainable, low-cost GPU compute power. As a Senior Cloud Support Engineer, you will be the primary technical support contact, helping customers leverage Crusoe Cloud to achieve their research goals and accelerate development... 

    Crusoe

    Bellevue, WA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer - Compute Platform. Be the first to apply!