Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Principal Site Reliability Engineer

Bybit

Senior Principal Site Reliability Engineer

Hong Kong SAR

About Us

Established in 2018, Bybit is one of the world's leading cryptocurrency exchanges and digital financial platforms, serving over 80 million users across more than 200 countries and regions. Powered by world-class technology and a user-first mindset, Bybit delivers a seamless ecosystem across trading, payments, wealth management, custody, institutional services, and Web3 — connecting users to the future of digital finance. Our core values define how we build. We listen, care and improve to create products and experiences that put users first. Backed by a global team of ambitious builders, problem-solvers, and innovators, we foster a high-performance and fast-moving environment where talent is empowered to drive real impact at the global scale. Supported by 24/7 multilingual customer service and a strong commitment to innovation, we are shaping the future of finance through technology, collaboration, and bold execution. Today, Bybit is recognized as one of the most trusted and transparent platforms in the digital asset industry, continuing to expand its global presence while building the infrastructure for the next generation of financial services.

Core Responsibilities

Chaos Engineering Platform Architecture & Development (50%)

  • Design and build an enterprise-grade chaos engineering platform supporting multi-cluster (K8s + EC2 hybrid), multi-region, and multi-environment (testnet/mainnet) deployments
  • Core capability development:
    • Fault Injection Engine: Pod-level / Node-level / AZ-level fault simulation, network latency / packet loss / partition, dependency timeout / error injection
    • Production Safety Assurance: Blast radius control, one-click Kill Switch, automatic rollback, real-time impact monitoring
    • Traffic Isolation: Experiment traffic tagging and isolation to ensure fault injection does not impact real users
    • Fault Isolation: Precise impact scoping at service / cluster / AZ granularity
    • Design experiment orchestration capabilities supporting complex fault scenario composition (e.g., simultaneous network latency + downstream timeout + cache invalidation)
    • Deep integration with existing monitoring, alerting, and SLO systems to achieve an automated closed loop: inject fault → observe impact → determine pass/fail

Production Resilience Validation Framework (30%)

  • Define safety standards and approval workflows for mainnet fault injection
  • Design and drive routine chaos experiments:
    • Daily patrol-level experiments: Low-risk experiments executed automatically on a daily/weekly basis
    • Periodic validation experiments: Monthly/quarterly resilience verification of critical paths
    • Large-scale drills: Cross-AZ / cross-region disaster recovery failover validation
    • Establish a resilience scoring system to quantify system health based on experiment results
    • Deliver improvement recommendations and drive business teams to remediate identified weaknesses

Technology Selection & Team Enablement (20%)

  • Evaluate and select the technology foundation (Chaos Mesh / Litmus / custom components — hybrid strategy)
  • Develop chaos engineering best practices and playbooks to enable SRE teams and application developers
  • Mentor and grow the team (2–3 engineers) in chaos engineering capabilities
  • Stay current with industry developments and introduce cutting-edge practices (e.g., AI-driven fault scenario discovery)
Requirements Must-Have:
  • 8+ years of backend / infrastructure engineering experience, with 3+ years dedicated to chaos engineering or stability engineering
  • Hands-on experience with large-scale fault injection in production environments (not just test environments), with deep understanding of production safety constraints
  • Expert-level proficiency in Kubernetes fault injection (Chaos Mesh / Litmus / custom solutions), familiar with CRD / Operator development
  • Proficient in at least one backend language (Go preferred), with platform-level system architecture design capability
  • Deep understanding of distributed system failure modes (network partitions, split-brain, cascading failures, data inconsistency, etc.)
  • Familiarity with observability tech stack (Prometheus / Grafana / Thanos / OpenTelemetry)
  • Excellent technical documentation and solution design skills
Nice-to-Have:
  • Experience in financial / trading system stability (understanding of transaction consistency and fund safety constraints)
  • Experience building SLO / Error Budget frameworks
  • Experience building automated fault recovery (self-healing) systems
  • Familiarity with AWS infrastructure (EC2 / EKS / Multi-AZ / Multi-Region)
  • Knowledge of Netflix Chaos Engineering / AWS Fault Injection Simulator / Gremlin
  • Open-source community contributions (Chaos Mesh / Litmus or similar projects)
Soft Skills:
  • Ability to balance "safety" and "validation depth" — not afraid of production injection, while maintaining strict risk control
  • Strong cross-team collaboration and influence — chaos engineering requires buy-in from business teams; this role demands persuasion skills
  • Self-driven, capable of independently planning and executing in ambiguous situations
Why Join Us

At Bybit, we are committed to fostering a supportive and enriching work environment. Our benefits include: - Study Growth Fund: We support your professional development and continuous learning. - Internal Events: Participate in regular team-building activities, workshops, and events designed to promote collaboration and innovation. - Global Collaboration: Be part of a diverse, international team, working alongside colleagues from around the world. - Career Advancement: Access opportunities for growth and advancement within a rapidly expanding global company. - Internal Mobility: Grow with us- Your long-term development is important to us. We offer internal job opportunities to help build your career path.

Vacancy posted 5 days ago
Similar jobs that could be interesting for youBased on the Senior Principal Site Reliability Engineer in United States vacancy
  •  ...pages for the websites you use every day. Our team of software engineers & web developers create OneLink services and tools that provide...  ...smarts and customized attention that makes our clients’ translated sites fly!   Job Description Build responsive, web-based user... 
    Principal
    Senior
    Full time

    As Translations

    Remote
    1 day ago
  • $196k - $269.5k

    Senior Principal AI Agent EngineerThe Software Engineering team delivers next-generation software application enhancements and new products for a changing world. Working at the cutting edge, we design and develop software for platforms, peripherals, applications and diagnostics... 
    Principal
    Senior

    Dell Technologies

    Austin, TX
    5 days ago
  • DescriptionSobre nosotros Worley es una empresa global de expertos en energía, químicos y recursos naturales, con sede en Australia. Trabajamos en asociación con nuestros clientes para desarrollar proyectos y generar valor a lo largo del ciclo de vida de sus activos. Nos...
    Principal
    Senior

    Worley

    Houston, TX
    4 days ago
  • $139.7k - $232.9k

     ...implementing, and continuously improving highly reliable, scalable, and resilient platform...  ...as a subject matter expert (SME) in Site Reliability Engineering, driving reliability engineering...  ...standards, and partners with senior stakeholders to improve system stability... 
    Principal
    Full time
    Work experience placement

    M&T Bank

    Buffalo, NY
    1 day ago
  • $130k - $200k

     ...Senior / Principal Flight Software Engineer – Space Systems *REMOTE* Working with a leading U.S. aerospace company looking to add Senior and Principal-level Flight Software Engineers to their satellite team. This is a hands-on role focused on building and testing... 
    Principal
    Senior
    Remote job
    Contract work

    Red Canyon Engineering & Software

    Merritt Island, FL
    1 day ago
  • DescriptionSAIC is hiring a Senior Principal Software Systems Engineer to join the Army UAS Training Test team located in Huntsville, Alabama (Redstone Arsenal...  ...productsDesired Skills:Familiarity with providing on site engineering, sustainment, and training... 
    Principal
    Senior
    Interim role
    Work at office

    Science Applications International Corporation

    Huntsville, AL
    4 days ago
  • $150k - $190k

    DescriptionKforce has a client that is seeking a Senior Principal Software Engineer (Delivery & Architecture) in New York, NY.Overview:We are seeking...  ...in ambiguous environments and prioritizes high-quality, reliable delivery.Key Responsibilities:* Oversee the end-to-end... 
    Principal
    Senior

    KForce

    New York, NY
    3 days ago
  • $142.8k - $274.8k

     ...yearEmployment type: Full-TimeWork site: 0 days / week in-office -...  ...: Site Reliability EngineeringCompany: MicrosoftOverviewMicrosoft...  ...demanding workloads. As a Principal Site Reliability Engineer, you will set technical...  ...as an actively engaged senior on-call engineer (OCE),... 
    Principal
    Ongoing contract
    Work at office
    Local area

    Microsoft

    Redmond, WA
    3 days ago
  •  ...sponsorship.Maintain and enhance the reliability, availability, and...  ...page of the Navy Federal Career Site.Protect Yourself from Job Scams...  ...degree in computer science, engineering, or the equivalent...  ...experts, and leaders; work with senior management on complex issuesLead... 
    Principal
    Internship
    Monday to Friday

    Navy Federal Credit Union

    Vienna, VA
    1 day ago
  •  ...Evaluate applications, platforms, and vendors to assess resiliency, reliability, and operational risk.Design and implement processes that...  ...and reliability tooling.Actively participate in reliability engineering and resilience communities of practice, contributing to... 
    Senior
    Full time

    Vanguard

    Wayne, PA
    2 days ago
  •  ...us to start Caring. Connecting. Growing together.We are seeking a Principal Site Reliability Engineer (SRE) to define and scale reliability practices across large-scale cloud platforms.This is a senior individual contributor role focused on setting SRE standards, influencing... 
    Principal
    Minimum wage
    Full time
    Work experience placement
    Work at office
    Local area
    Remote work

    UnitedHealth Group

    Minnetonka, MN
    1 day ago
  •  ...instructions on how to do this,please click this link or view the document - "How to copy from a Word Document" located on the Taleo Support site under the section titled Quick Reference Guides.QualificationsNote: Recruiter to paste Qualifications/Requirements of the job here.... 
    Principal
    Senior

    Worley

    Baton Rouge, LA
    1 day ago
  •  ...Infrastructure Code. Builds reliability into the ecosystem by...  ...in resiliency engineering and observability by developing...  ...techniques with site reliability engineering...  ...processes.Advises senior management on technical...  ...years of experience as a Principal Site Reliability... 
    Principal
    Full time

    Fidelity Investments

    Texas
    5 days ago
  •  ...TechMContact: Meghana GorusuCompany: SRI Tech SolutionsJob Title: Senior Site Reliability EngineerLocation: Plano , TX (remote)Years of Experience: 8...  ...are seeking a highly skilled Senior Site Reliability Engineer (SRE) to join our dynamic team. The ideal candidate will... 
    Senior
    Remote work

    SRI Tech

    Plano, TX
    4 days ago
  • $170k - $220k

    Who We're Looking ForWe’re looking for a hands-on, high-agency Site Reliability Engineer to help shape and scale the reliability layer of our stack. You'll own the release pipeline end-to-end — managing daily releases, weekly deploys, and hotfixes — while also automating... 
    Senior

    Supio

    Seattle, WA
    4 days ago
  • About the RoleWe’re looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You’ll partner with engineers and data scientists to build, automate, and maintain the infrastructure... 
    Senior

    Alembic

    San Francisco, CA
    4 days ago
  •  ...Working remotely within the United States, the full-time Principal Site Reliability Engineer will lead project work to enhance platform reliability, mentor junior engineers, and engage in incident response while collaborating closely with product stakeholders and architects... 
    Principal
    Full time
    Remote work

    Virtual Vocations Inc

    United States
    2 days ago
  • $65 - $75 per hour

    DescriptionKforce has a client seeking a remote Senior Site Reliability Engineer to be a l be a leading member of the team working with a diverse range of technologies. You will enjoy working in a friendly environment and benefit from our investment in staff. The role... 
    Senior
    Remote work

    KForce

    Boca Raton, FL
    3 days ago
  • IXL Learning, developer of personalized learning products used by millions of people globally, is seeking a Senior Site Reliability Engineer to join our team, and help maintain the reliability and optimal performance of our products. We are seeking engineers with a passion... 
    Senior
    Work at office
    Immediate start

    IXL Learning

    Raleigh, NC
    1 day ago
  • $175.5k - $235.4k

     ...MyDisneyExperience and Hey, Disney!This role sits in the Commerce Site Reliability Engineering (SRE) specifically supporting Ecommerce , Consumer...  ...Products Technology teams from across the company. The Principal of DXT SRE will report to the Director of DXT Commerce SREAbout... 
    Principal
    Worldwide

    Disney Interactive

    Orlando, FL
    2 days ago
  • Senior Site Reliability Engineer ILocationSan Jose, Costa Rica - RemoteSummary of roleOwn availability, the most important product feature, by continually striving for sustained operational excellence of Sumo’s planet-scale observability and security products. Work with... 
    Senior
    Flexible hours

    Sumo Logic

    San Jose, CA
    2 days ago
  • $152.6k - $191.5k

     ...responsible for partnering with leaders across engineering and technology to define objective reliability goals for services. Key responsibilities include...  ...and continuous improvement.Position Summary:The Senior GCP Site Reliability Engineer acts as an advanced senior... 
    Senior
    Full time
    Work at office
    Day shift

    Bank of America

    Plano, TX
    2 days ago
  •  ...work from home day is currently Tuesday.Engineering at Lambda is responsible for building and...  ...and networking teams to improve service reliability and deployment workflowsDeploy and...  ...rotationYouHave 5+ years of experience in Site Reliability Engineering, Production Engineering... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    1 day ago
  • $104.9k - $174.7k

     ...Data Management. You can learn more about LexisNexis Risk at the link below, About the Role:We are hiring a hands-on Senior Site Reliability Engineer (SRE) to actively build, operate, and improve the reliability of our production systems. This is not a purely advisory... 
    Senior
    Full time
    Work at office
    Local area
    Remote work
    Work from home

    RELX Group

    Buford, GA
    3 days ago
  • $150k - $195k

    DescriptionKforce has a client in Orem, UT that is seeking a Senior or Principal Radar Systems Engineer. We are working directly with the hiring manager on this exclusive search assignment. This role will be onsite. The client offers a competitive compensation package... 
    Principal
    Senior

    KForce

    Orem, UT
    3 days ago
  • $168k - $270.25k

    NVIDIA is looking for a Senior Site Reliability Engineer (SRE) to join its GeForce Now (GFN) team. SRE at NVIDIA ensures that our internal and external-facing GPU cloud gaming services have reliability and uptime as promised to the users and at the same time enables developers... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    5 days ago
  • $210k - $230k

    GovCIO is currently hiring for a Senior Site Reliability Engineer (SRE) to design, implement, and maintain highly available, scalable, and resilient infrastructure systems. The ideal candidate will bridge the gap between development and operations, focusing on automation... 
    Senior
    Currently hiring
    Remote work

    Govcio

    Arlington, VA
    5 days ago
  •  ...Lambda’s designated work from home day is currently Tuesday.Engineering at Lambda is responsible for building and scaling our cloud offering...  ...and SLIs for Kubernetes services, workloads, and platform reliability.You6+ years of experience in a SRE, operations engineer, or... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    1 day ago
  • LeanData helps the world’s fastest-growing companies automate, simplify, and accelerate revenue.We are looking for a Senior Site Reliability Engineer to lead the strategic evolution of our cloud infrastructure. Reporting directly to the SVP of Engineering, this role is... 
    Senior
    Full time
    Work at office
    2 days per week

    LeanData

    Santa Clara, CA
    4 days ago
  • Site Reliability Engineer - Equity Trading PlatformLocation: New York | Practice Area: Capital Markets - Technology & Engineering | Type: PermanentKeep critical equity trading platforms resilient, reliable, and ready for the markets.The RoleWe are seeking a highly motivated... 
    Principal
    Permanent employment
    Work at office
    Weekend work
    Afternoon shift

    Capco

    New York, NY
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Principal Site Reliability Engineer. Be the first to apply!