Senior Principal Site Reliability Engineer
Bybit
Senior Principal Site Reliability Engineer
Hong Kong SAR
About Us
Established in 2018, Bybit is one of the world's leading cryptocurrency exchanges and digital financial platforms, serving over 80 million users across more than 200 countries and regions. Powered by world-class technology and a user-first mindset, Bybit delivers a seamless ecosystem across trading, payments, wealth management, custody, institutional services, and Web3 — connecting users to the future of digital finance. Our core values define how we build. We listen, care and improve to create products and experiences that put users first. Backed by a global team of ambitious builders, problem-solvers, and innovators, we foster a high-performance and fast-moving environment where talent is empowered to drive real impact at the global scale. Supported by 24/7 multilingual customer service and a strong commitment to innovation, we are shaping the future of finance through technology, collaboration, and bold execution. Today, Bybit is recognized as one of the most trusted and transparent platforms in the digital asset industry, continuing to expand its global presence while building the infrastructure for the next generation of financial services.
Core Responsibilities
Chaos Engineering Platform Architecture & Development (50%)
- Design and build an enterprise-grade chaos engineering platform supporting multi-cluster (K8s + EC2 hybrid), multi-region, and multi-environment (testnet/mainnet) deployments
- Core capability development:
- Fault Injection Engine: Pod-level / Node-level / AZ-level fault simulation, network latency / packet loss / partition, dependency timeout / error injection
- Production Safety Assurance: Blast radius control, one-click Kill Switch, automatic rollback, real-time impact monitoring
- Traffic Isolation: Experiment traffic tagging and isolation to ensure fault injection does not impact real users
- Fault Isolation: Precise impact scoping at service / cluster / AZ granularity
- Design experiment orchestration capabilities supporting complex fault scenario composition (e.g., simultaneous network latency + downstream timeout + cache invalidation)
- Deep integration with existing monitoring, alerting, and SLO systems to achieve an automated closed loop: inject fault → observe impact → determine pass/fail
Production Resilience Validation Framework (30%)
- Define safety standards and approval workflows for mainnet fault injection
- Design and drive routine chaos experiments:
- Daily patrol-level experiments: Low-risk experiments executed automatically on a daily/weekly basis
- Periodic validation experiments: Monthly/quarterly resilience verification of critical paths
- Large-scale drills: Cross-AZ / cross-region disaster recovery failover validation
- Establish a resilience scoring system to quantify system health based on experiment results
- Deliver improvement recommendations and drive business teams to remediate identified weaknesses
Technology Selection & Team Enablement (20%)
- Evaluate and select the technology foundation (Chaos Mesh / Litmus / custom components — hybrid strategy)
- Develop chaos engineering best practices and playbooks to enable SRE teams and application developers
- Mentor and grow the team (2–3 engineers) in chaos engineering capabilities
- Stay current with industry developments and introduce cutting-edge practices (e.g., AI-driven fault scenario discovery)
Requirements Must-Have:
- 8+ years of backend / infrastructure engineering experience, with 3+ years dedicated to chaos engineering or stability engineering
- Hands-on experience with large-scale fault injection in production environments (not just test environments), with deep understanding of production safety constraints
- Expert-level proficiency in Kubernetes fault injection (Chaos Mesh / Litmus / custom solutions), familiar with CRD / Operator development
- Proficient in at least one backend language (Go preferred), with platform-level system architecture design capability
- Deep understanding of distributed system failure modes (network partitions, split-brain, cascading failures, data inconsistency, etc.)
- Familiarity with observability tech stack (Prometheus / Grafana / Thanos / OpenTelemetry)
- Excellent technical documentation and solution design skills
Nice-to-Have:
- Experience in financial / trading system stability (understanding of transaction consistency and fund safety constraints)
- Experience building SLO / Error Budget frameworks
- Experience building automated fault recovery (self-healing) systems
- Familiarity with AWS infrastructure (EC2 / EKS / Multi-AZ / Multi-Region)
- Knowledge of Netflix Chaos Engineering / AWS Fault Injection Simulator / Gremlin
- Open-source community contributions (Chaos Mesh / Litmus or similar projects)
Soft Skills:
- Ability to balance "safety" and "validation depth" — not afraid of production injection, while maintaining strict risk control
- Strong cross-team collaboration and influence — chaos engineering requires buy-in from business teams; this role demands persuasion skills
- Self-driven, capable of independently planning and executing in ambiguous situations
Why Join Us
At Bybit, we are committed to fostering a supportive and enriching work environment. Our benefits include: - Study Growth Fund: We support your professional development and continuous learning. - Internal Events: Participate in regular team-building activities, workshops, and events designed to promote collaboration and innovation. - Global Collaboration: Be part of a diverse, international team, working alongside colleagues from around the world. - Career Advancement: Access opportunities for growth and advancement within a rapidly expanding global company. - Internal Mobility: Grow with us- Your long-term development is important to us. We offer internal job opportunities to help build your career path.
- ...pages for the websites you use every day. Our team of software engineers & web developers create OneLink services and tools that provide... ...smarts and customized attention that makes our clients’ translated sites fly! Job Description Build responsive, web-based user...PrincipalSeniorFull time
$196k - $269.5k
Senior Principal AI Agent EngineerThe Software Engineering team delivers next-generation software application enhancements and new products for a changing world. Working at the cutting edge, we design and develop software for platforms, peripherals, applications and diagnostics...PrincipalSenior- DescriptionSobre nosotros Worley es una empresa global de expertos en energía, químicos y recursos naturales, con sede en Australia. Trabajamos en asociación con nuestros clientes para desarrollar proyectos y generar valor a lo largo del ciclo de vida de sus activos. Nos...PrincipalSenior
$139.7k - $232.9k
...implementing, and continuously improving highly reliable, scalable, and resilient platform... ...as a subject matter expert (SME) in Site Reliability Engineering, driving reliability engineering... ...standards, and partners with senior stakeholders to improve system stability...PrincipalFull timeWork experience placement$130k - $200k
...Senior / Principal Flight Software Engineer – Space Systems *REMOTE* Working with a leading U.S. aerospace company looking to add Senior and Principal-level Flight Software Engineers to their satellite team. This is a hands-on role focused on building and testing...PrincipalSeniorRemote jobContract work- DescriptionSAIC is hiring a Senior Principal Software Systems Engineer to join the Army UAS Training Test team located in Huntsville, Alabama (Redstone Arsenal... ...productsDesired Skills:Familiarity with providing on site engineering, sustainment, and training...PrincipalSeniorInterim roleWork at office
$150k - $190k
DescriptionKforce has a client that is seeking a Senior Principal Software Engineer (Delivery & Architecture) in New York, NY.Overview:We are seeking... ...in ambiguous environments and prioritizes high-quality, reliable delivery.Key Responsibilities:* Oversee the end-to-end...PrincipalSenior$142.8k - $274.8k
...yearEmployment type: Full-TimeWork site: 0 days / week in-office -... ...: Site Reliability EngineeringCompany: MicrosoftOverviewMicrosoft... ...demanding workloads. As a Principal Site Reliability Engineer, you will set technical... ...as an actively engaged senior on-call engineer (OCE),...PrincipalOngoing contractWork at officeLocal area- ...sponsorship.Maintain and enhance the reliability, availability, and... ...page of the Navy Federal Career Site.Protect Yourself from Job Scams... ...degree in computer science, engineering, or the equivalent... ...experts, and leaders; work with senior management on complex issuesLead...PrincipalInternshipMonday to Friday
- ...Evaluate applications, platforms, and vendors to assess resiliency, reliability, and operational risk.Design and implement processes that... ...and reliability tooling.Actively participate in reliability engineering and resilience communities of practice, contributing to...SeniorFull time
- ...us to start Caring. Connecting. Growing together.We are seeking a Principal Site Reliability Engineer (SRE) to define and scale reliability practices across large-scale cloud platforms.This is a senior individual contributor role focused on setting SRE standards, influencing...PrincipalMinimum wageFull timeWork experience placementWork at officeLocal areaRemote work
- ...instructions on how to do this,please click this link or view the document - "How to copy from a Word Document" located on the Taleo Support site under the section titled Quick Reference Guides.QualificationsNote: Recruiter to paste Qualifications/Requirements of the job here....PrincipalSenior
- ...Infrastructure Code. Builds reliability into the ecosystem by... ...in resiliency engineering and observability by developing... ...techniques with site reliability engineering... ...processes.Advises senior management on technical... ...years of experience as a Principal Site Reliability...PrincipalFull time
- ...TechMContact: Meghana GorusuCompany: SRI Tech SolutionsJob Title: Senior Site Reliability EngineerLocation: Plano , TX (remote)Years of Experience: 8... ...are seeking a highly skilled Senior Site Reliability Engineer (SRE) to join our dynamic team. The ideal candidate will...SeniorRemote work
$170k - $220k
Who We're Looking ForWe’re looking for a hands-on, high-agency Site Reliability Engineer to help shape and scale the reliability layer of our stack. You'll own the release pipeline end-to-end — managing daily releases, weekly deploys, and hotfixes — while also automating...Senior- About the RoleWe’re looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You’ll partner with engineers and data scientists to build, automate, and maintain the infrastructure...Senior
- ...Working remotely within the United States, the full-time Principal Site Reliability Engineer will lead project work to enhance platform reliability, mentor junior engineers, and engage in incident response while collaborating closely with product stakeholders and architects...PrincipalFull timeRemote work
$65 - $75 per hour
DescriptionKforce has a client seeking a remote Senior Site Reliability Engineer to be a l be a leading member of the team working with a diverse range of technologies. You will enjoy working in a friendly environment and benefit from our investment in staff. The role...SeniorRemote work- IXL Learning, developer of personalized learning products used by millions of people globally, is seeking a Senior Site Reliability Engineer to join our team, and help maintain the reliability and optimal performance of our products. We are seeking engineers with a passion...SeniorWork at officeImmediate start
$175.5k - $235.4k
...MyDisneyExperience and Hey, Disney!This role sits in the Commerce Site Reliability Engineering (SRE) specifically supporting Ecommerce , Consumer... ...Products Technology teams from across the company. The Principal of DXT SRE will report to the Director of DXT Commerce SREAbout...PrincipalWorldwide- Senior Site Reliability Engineer ILocationSan Jose, Costa Rica - RemoteSummary of roleOwn availability, the most important product feature, by continually striving for sustained operational excellence of Sumo’s planet-scale observability and security products. Work with...SeniorFlexible hours
$152.6k - $191.5k
...responsible for partnering with leaders across engineering and technology to define objective reliability goals for services. Key responsibilities include... ...and continuous improvement.Position Summary:The Senior GCP Site Reliability Engineer acts as an advanced senior...SeniorFull timeWork at officeDay shift- ...work from home day is currently Tuesday.Engineering at Lambda is responsible for building and... ...and networking teams to improve service reliability and deployment workflowsDeploy and... ...rotationYouHave 5+ years of experience in Site Reliability Engineering, Production Engineering...SeniorWork at officeLocal areaWork from homeFlexible hours
$104.9k - $174.7k
...Data Management. You can learn more about LexisNexis Risk at the link below, About the Role:We are hiring a hands-on Senior Site Reliability Engineer (SRE) to actively build, operate, and improve the reliability of our production systems. This is not a purely advisory...SeniorFull timeWork at officeLocal areaRemote workWork from home$150k - $195k
DescriptionKforce has a client in Orem, UT that is seeking a Senior or Principal Radar Systems Engineer. We are working directly with the hiring manager on this exclusive search assignment. This role will be onsite. The client offers a competitive compensation package...PrincipalSenior$168k - $270.25k
NVIDIA is looking for a Senior Site Reliability Engineer (SRE) to join its GeForce Now (GFN) team. SRE at NVIDIA ensures that our internal and external-facing GPU cloud gaming services have reliability and uptime as promised to the users and at the same time enables developers...SeniorFull time$210k - $230k
GovCIO is currently hiring for a Senior Site Reliability Engineer (SRE) to design, implement, and maintain highly available, scalable, and resilient infrastructure systems. The ideal candidate will bridge the gap between development and operations, focusing on automation...SeniorCurrently hiringRemote work- ...Lambda’s designated work from home day is currently Tuesday.Engineering at Lambda is responsible for building and scaling our cloud offering... ...and SLIs for Kubernetes services, workloads, and platform reliability.You6+ years of experience in a SRE, operations engineer, or...SeniorWork at officeLocal areaWork from homeFlexible hours
- LeanData helps the world’s fastest-growing companies automate, simplify, and accelerate revenue.We are looking for a Senior Site Reliability Engineer to lead the strategic evolution of our cloud infrastructure. Reporting directly to the SVP of Engineering, this role is...SeniorFull timeWork at office2 days per week
- Site Reliability Engineer - Equity Trading PlatformLocation: New York | Practice Area: Capital Markets - Technology & Engineering | Type: PermanentKeep critical equity trading platforms resilient, reliable, and ready for the markets.The RoleWe are seeking a highly motivated...PrincipalPermanent employmentWork at officeWeekend workAfternoon shift
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Principal Site Reliability Engineer. Be the first to apply!
- associate director engineering United States
- principal network engineer United States
- senior director engineering United States
- director mechanical engineering United States
- civil engineer project manager United States
- aerospace engineering director United States
- principal developer United States
- clinical engineering director United States
- chief design engineer United States
- principal test engineer United States


