Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer - Compute Platform

TikTok

Team introductionOur Compute Platform SRE team supports all Big Data services and products across the company. We are a newly established team and waiting for talents like you to shape the team's future together. We are responsible for the reliability of all the company's major data warehouse products, services, and query engines. We serve business needs across domains within TikTok. We look forward to welcoming you to the team.Responsibilities:- Responsible for the reliability of all TikTok's major data warehouse products, services, and query engines, such as ClickHouse, Spark, Presto, Doris, etc.- Uphold Service Level Agreements (SLAs): Ensure that all service level objectives and agreements from ByteDance's Data Platform services are met. Respond promptly to any system outages or issues.- Continuous Performance Optimization: Analyze service performance and reliability patterns to identify potential performance bottlenecks. Implement proactive measures to prevent service disruptions. Work with development teams to optimize application performance, ensuring that services run efficiently and that resources are utilized effectively.- Incident Management: Lead efforts to troubleshoot and resolve service incidents and postmortems. Coordinate with cross-functional teams to manage and mitigate service-impacting events.- Infrastructure Automation: Automate infrastructure provisioning, scaling, and management processes to reduce manual interventions and improve service quality.- Collaboration: Engage with product and development teams to integrate reliability and performance considerations into the software lifecycle.- Capacity and Demand Planning: Assess and forecast infrastructure needs based on growth patterns and upcoming initiatives.- Stay Updated: Keep current with industry trends, best practices, and emerging technologies related to site reliability and infrastructure engineering.Minimal Qualifications:- Bachelor's Degree or above, in Computer Science, Engineering, or a related field. Passionate about computer science and Internet technology.- In-depth understanding of Linux, computer networking, and databases. Proficient in common SRE/DevOps open-source toolsets, system monitoring tools, and container orchestration platforms like Kubernetes.- Experience or familiarity with open-source or commercial technologies such as ClickHouse, Hadoop, Doris, Spark, Presto and Kubernetes.- Strong coding skills in at least one scripting or programming language, including but not limited to Python, Shell, Java, Go, etc.- Excellent problem-solving skills and the ability to think critically.Req ID: A237226

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer - Compute Platform in Seattle, WA vacancy
  • $209.1k - $282.9k

    As a software Engineer on the AI Compute Platform team, you will design and build a secure, reliable, and easy-to-use platform for running AI workloads at scale. You will develop control-plane services, APIs, orchestration systems, and developer tools for distributed training... 
    Suggested
    Work at office
    Local area
    Visa sponsorship
    Relocation package

    ARM

    Seattle, WA
    4 days ago
  • $180.5k - $225.6k

     ...AI and data infrastructure platform so our customers can use deep...  ...their business. Founded by engineers — and customer obsessed — we...  ...getting started.The Serverless Compute Platform is the backbone of...  ...with production-grade reliability spanning a range of use cases... 
    Suggested
    Local area
    Remote work
    Worldwide

    DataBricks

    Bellevue, WA
    3 days ago
  • $253k - $416k

     ...DescriptionDistinguished Software Engineer, Systems Infrastructure - Compute, Deployment, Infrastructure as Code...  ...of LinkedIn’s Core Infrastructure platform, which provides shared system...  ...architecture, balancing scalability, reliability, and costHelp evolve LinkedIn’s private... 
    Suggested
    For contractors
    Work experience placement
    Work at office
    Flexible hours

    Linkedin

    Bellevue, WA
    4 days ago
  •  ...MangoApps runs an enterprise SaaS platform that thousands of customers...  ...hiring a senior, hands-on engineer to own the reliability, availability, security,...  ...-on Cloud Operations and Site Reliability Engineering,...  ...speak in specifics about AWS compute, networking, IAM, EKS/Kubernetes... 
    Suggested
    Full time

    MangoApps

    Seattle, WA
    12 hours ago
  • $143k - $194k

     ...cutting-edge autonomy, AI, computer vision, sensor fusion, and networking...  .... System Deployment Engineers work in complex environments...  ...as a data and applications platform and be able to educate customersParticipate...  ...and exercisesWork with site reliability engineers to provide and... 
    Suggested
    Full time
    Temporary work
    Work experience placement
    Immediate start

    Anduril Industries

    Seattle, WA
    1 day ago
  • $55k - $151.47k

     ...Description & SummaryThe OpportunityAs a Site Reliability Engineer - Senior Associate, you will play a...  ...- Utilizing cloud infrastructure platforms such as AWS, Google Cloud Platform, and...  ...Business Administration/Management, Computer Science/Information Systems, Engineering... 
    Full time
    H1b

    PwC

    Seattle, WA
    1 day ago
  • $194k - $267k

     ...concepts and tools.Position Overview:The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications...  ...containerization principles.Bachelor’s degree in Computer Science, Engineering, or related field (... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta

    Bellevue, WA
    2 days ago
  • $232k - $319k

     ...s talk. The Infrastructure Platform and Shared Services Team...  ...service with great people and reliable, cost-effective, and efficient...  ...velocity of SRE and product engineering by developing robust platforms...  ...interpersonal skills ~ Computer Science Degree or related degree... 
    Permanent employment
    Local area
    Worldwide
    Flexible hours

    Okta

    Bellevue, WA
    9 days ago
  •  ...This is an engineering-first Senior SRE role. We’re looking for...  ...systems and/or distributed platforms Owned services end-to-end...  ...design → launch → on-call → reliability improvements) Led incident...  ...degree is mandatory: BS/MS in Computer Science, Computer Engineering... 

    Practice by Numbers

    Bellevue, WA
    1 day ago
  •  ...place.As a Principal Software Engineer at JPMorganChase within the Core Foundational Platforms team, you provide expertise and...  ...practices in high-performance computing (HPC) practices that intersect...  ...comprehensive health care coverage, on-site health and wellness centers, a... 

    JP Morgan Chase

    Seattle, WA
    1 day ago
  • $160k - $250k

     ...DevOps And Systems Engineer Hive is the leading provider of cloud...  ...high performance computing integrating GPUs. Even with...  ...need to grow our DevOps and Site Reliability team to maintain the reliability...  ...diverse array of technology platforms, following best practices and... 

    Hive

    Seattle, WA
    12 hours ago
  •  ...are in. About this team Site Reliability Engineering We are looking for a motivated engineer...  ...~ Bachelor's degree in computer science/engineering or equivalent...  ...monitoring solutions, log aggregation platforms, and distributed tracing frameworks... 

    Kaav Inc.

    Seattle, WA
    4 days ago
  •  ...A Site Reliability Engineer (SRE) is responsible for ensuring the reliability, scalability, and performance of an organization's software systems...  ...Problem-Solving Education Bachelor's degree in Computer Science, Information Technology, Software Engineering, or... 

    Mybridge

    Seattle, WA
    1 day ago
  •  ...Technical Support Engineer/Site Reliability Engineer At F5, we strive to bring a better digital...  ...and scale an AI Security Public SaaS platform, operating AI inference workloads at...  ...Qualifications ~ Bachelor's degree in Computer Science, Information Technology, or a... 

    F5

    Seattle, WA
    1 day ago
  • $209.1k - $282.9k

    As an engineer on the AI Compute Infra team, you will design, build, and operate large-scale infrastructure for AI training, fine-tuning, evaluation...  ...partnering with AI researchers and engineers to improve reliability, performance, scalability, and developer productivity.... 
    Work at office
    Local area
    Visa sponsorship
    Relocation package

    ARM

    Seattle, WA
    4 days ago
  •  ...Google LLC in Seattle, WA is seeking a Senior Software Engineer for the Infrastructure team within Google Cloud Compute. You will design, build, and scale distributed systems that power Google's cloud infrastructure and services. You have 5 years of experience with... 

    Jobleads-US

    Seattle, WA
    1 day ago
  • $174k - $252k

     ...Senior Software Engineer, Infrastructure, Google Cloud Compute Share Senior Software Engineer, Infrastructure...  ...building the next generation of Google platforms, we make Google's product...  ...unparalleled scale, efficiency, reliability and velocity. Our customers include... 
    Temporary work
    Worldwide

    Jobleads-US

    Seattle, WA
    1 day ago
  •  ...is responsible for the reliability, scalability, and...  ...building features, but about engineering the resilience and...  ...performance of the underlying platform that all product teams...  ...system stability.As a Site Reliability Engineer,...  ...Bachelor’s degree in Computer Science, a related... 

    TikTok

    Seattle, WA
    3 days ago
  •  ...team is responsible for the reliability, scalability, and efficiency...  ...building features, but about engineering the resilience and performance of the underlying platform that all product teams depend...  ...Qualifications:- Bachelor’s degree in Computer Science, a related technical... 

    TikTok

    Seattle, WA
    3 days ago
  •  ...to add talented individuals to our team.Computer vision plays an indispensable role in modern...  ..., amongst others. Computer vision engineers at Valve are working on all those areas to...  ...planGenerous vacation and family leaveOn-site amenities in support of health and efficiencyFertility... 
    Flexible hours

    Valve Software

    Bellevue, WA
    3 days ago
  •  ...Job Title : Site Reliability Engineer (Azure + Data) Work Mode : Bellevue, WA - Hybrid Job type...  ...reliability engineering cloud operations or platform engineering including strong hands-on...  ...Right size resources optimize compute storage networking SQL and Cosmos capacity... 
    Contract work

    VDart

    Bellevue, WA
    12 hours ago
  • $133.2k - $219.6k

     ...are the AI Inference Platform at Alibaba Group, committed...  ...innovation and engineering practices. Our team focuses...  ...and enterprise-level reliability. By doing so, we aim...  ...and technically skilled Site Reliability Engineer (...  .... Experience in cloud computing, AI infrastructure,... 
    Full time

    Alibaba Group

    Bellevue, WA
    3 hours ago
  • $175k - $308.5k

     ...found it.The Apple Service Engineering (ASE) team builds and provides...  ...Service Engineering (ASE)'s Compute team is seeking an...  ...components of Apple's Cloud Platform with an emphasis on VM lifecycle...  ...preferred. ~7+ years in a Site Reliability Engineering Infrastructure focused... 
    Relocation

    Jobleads-US

    Seattle, WA
    3 days ago
  •  ...Meta is seeking a software engineer to join our AI & Systems Co-Design team to drive the definition of our next-generation compute and storage architectures. As a key member of the...  ...work closely with internal software and platforms engineering teams to drive workload analysis... 

    Jobleads-US

    Bellevue, WA
    1 day ago
  •  ...production at scale. We own the platform end to end: backend systems,...  ...-on infrastructure role for engineers who want to work on deeply...  ...Improve performance, reliability, and operational excellence...  ...offer of employment: protect computer hardware entrusted to you from... 

    OpenAI

    Seattle, WA
    4 days ago
  • $142.3k - $263.3k

     ...from you!The Apple Services Engineering (ASE) organization is responsible for building powerful platforms that enable engineers to deliver...  ...experiences to customers.Our compute team is responsible for...  ...teams across Apple to build reliable, high-performance compute infrastructure... 
    Relocation

    Jobleads-US

    Seattle, WA
    4 days ago
  •  ...Senior Site Reliability Engineer (SRE) Location: Seattle, hybrid - 2 times a week in the office Job Type: Full-time, direct hire Industry...  ...with service mesh technologies (Istio, Linkerd). Background in internal platform/developer experience (DevEx) teams.... 
    Full time
    Work at office

    TalentDome Staffing

    Seattle, WA
    12 hours ago
  • $127k - $249k

    Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that...  ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper).... 
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    Seattle, WA
    3 days ago
  • $151.2k - $204.6k

    Would you like to be an engineer who builds the systems that power advertising at scale,...  ...advertising queries every day, where latency, reliability, and quality translate directly into...  ...Development Engineer, operating as a Site Reliability Engineer, to raise the reliability... 
    Flexible hours

    Amazon

    Seattle, WA
    4 days ago
  •  ...relevant scripting or programming languages (Ruby, Perl, Python, Shell, PowerShell, etc.) • Experience with Configuration Management platforms (Chef, Ansible, CFEngine, Puppet, etc.) • Database Administration - setup, configuration and basic database troubleshooting skills... 

    Comtech

    Seattle, WA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer - Compute Platform. Be the first to apply!