Senior Principal Site Reliability Engineer
Bybit
Established in 2018, Bybit is one of the world’s leading cryptocurrency exchanges and digital financial platforms, serving over 80 million users across more than 200 countries and regions. Powered by world-class technology and a user-first mindset, Bybit delivers a seamless ecosystem across trading, payments, wealth management, custody, institutional services, and Web3 — connecting users to the future of digital finance.
Our core values define how we build. We listen, care and improve to create products and experiences that put users first. Backed by a global team of ambitious builders, problem-solvers, and innovators, we foster a high-performance and fast-moving environment where talent is empowered to drive real impact at the global scale. Supported by 24/7 multilingual customer service and a strong commitment to innovation, we are shaping the future of finance through technology, collaboration, and bold execution.
Today, Bybit is recognized as one of the most trusted and transparent platforms in the digital asset industry, continuing to expand its global presence while building the infrastructure for the next generation of financial services.
Core Responsibilities
Chaos Engineering Platform Architecture & Development (50%)- Design and build an enterprise-grade chaos engineering platform supporting multi-cluster (K8s + EC2 hybrid), multi-region, and multi-environment (testnet/mainnet) deployments
- Core capability development:
- Fault Injection Engine: Pod-level / Node-level / AZ-level fault simulation, network latency / packet loss / partition, dependency timeout / error injection
- Production Safety Assurance: Blast radius control, one-click Kill Switch, automatic rollback, real-time impact monitoring
- Traffic Isolation: Experiment traffic tagging and isolation to ensure fault injection does not impact real users
- Fault Isolation: Precise impact scoping at service / cluster / AZ granularity
- Design experiment orchestration capabilities supporting complex fault scenario composition (e.g., simultaneous network latency + downstream timeout + cache invalidation)
- Deep integration with existing monitoring, alerting, and SLO systems to achieve an automated closed loop: inject fault → observe impact → determine pass/fail
- Define safety standards and approval workflows for mainnet fault injection
- Design and drive routine chaos experiments:
- Daily patrol-level experiments: Low-risk experiments executed automatically on a daily/weekly basis
- Periodic validation experiments: Monthly/quarterly resilience verification of critical paths
- Large-scale drills: Cross-AZ / cross-region disaster recovery failover validation
- Establish a resilience scoring system to quantify system health based on experiment results
- Deliver improvement recommendations and drive business teams to remediate identified weaknesses
- Evaluate and select the technology foundation (Chaos Mesh / Litmus / custom components — hybrid strategy)
- Develop chaos engineering best practices and playbooks to enable SRE teams and application developers
- Mentor and grow the team (2–3 engineers) in chaos engineering capabilities
- Stay current with industry developments and introduce cutting-edge practices (e.g., AI-driven fault scenario discovery)
Requirements
Must-Have:
- 8+ years of backend / infrastructure engineering experience, with 3+ years dedicated to chaos engineering or stability engineering
- Hands-on experience with large-scale fault injection in production environments (not just test environments), with deep understanding of production safety constraints
- Expert-level proficiency in Kubernetes fault injection (Chaos Mesh / Litmus / custom solutions), familiar with CRD / Operator development
- Proficient in at least one backend language (Go preferred), with platform-level system architecture design capability
- Deep understanding of distributed system failure modes (network partitions, split-brain, cascading failures, data inconsistency, etc.)
- Familiarity with observability tech stack (Prometheus / Grafana / Thanos / OpenTelemetry)
- Excellent technical documentation and solution design skills
- Experience in financial / trading system stability (understanding of transaction consistency and fund safety constraints)
- Experience building SLO / Error Budget frameworks
- Experience building automated fault recovery (self-healing) systems
- Familiarity with AWS infrastructure (EC2 / EKS / Multi-AZ / Multi-Region)
- Knowledge of Netflix Chaos Engineering / AWS Fault Injection Simulator / Gremlin
- Open-source community contributions (Chaos Mesh / Litmus or similar projects)
- Ability to balance "safety" and "validation depth" — not afraid of production injection, while maintaining strict risk control
- Strong cross-team collaboration and influence — chaos engineering requires buy-in from business teams; this role demands persuasion skills
- Self-driven, capable of independently planning and executing in ambiguous situations
At Bybit, we are committed to fostering a supportive and enriching work environment.
Our benefits include:
- Study Growth Fund: We support your professional development and continuous learning.
- Internal Events: Participate in regular team-building activities, workshops, and events designed to promote collaboration and innovation.
- Global Collaboration: Be part of a diverse, international team, working alongside colleagues from around the world.
- Career Advancement: Access opportunities for growth and advancement within a rapidly expanding global company.
- Internal Mobility: Grow with us- Your long-term development is important to us. We offer internal job opportunities to help build your career path.
- ...studied excellent open source component source code and experience is preferred 1. Proficient in NoSQL cache, message queue, search engines, such as Redis, Kafka, Elasticsearch, etc 1. Skilled in system analysis and design, code refactoring, with experience in large-...PrincipalSeniorFull time
- ...Next.js framework principles and core features (SSR/SSG); able to build scalable architectures on top of it. 4. Familiar with engineering tools such as Webpack and CI/CD pipelines; possesses product thinking and cross-team collaboration skills. 5. Committed to clean...PrincipalSeniorFull time
- ...complex ideas with key stakeholders across Engineering, Product, Data, Marketing, Sales, etc.... ...continuous improvement of platform reliability, availability and manageability Develop... ...offs to non-technical stakeholders and senior leadership. ~ Strong analytical and...PrincipalSeniorRemote jobFull timeWork at officeFlexible hours
- ...complex ideas with key stakeholders across Engineering, Product, Data, Marketing, Sales, etc.... ...continuous improvement of platform reliability, availability and manageability Develop... ...Extensive experience managing multiple senior stakeholders with differing priorities,...PrincipalSeniorRemote jobFull timeWork at officeFlexible hours
- ...Responsibilities Own end-to-end technical solution design and delivery for company-wide technical initiatives Lead the engineering implementation of AI capabilities, including but not limited to AI-assisted development, AI code review, AI observability analysis...PrincipalSeniorFull time
- ...financial services across 50+ global sites, each subject to different compliance... ...versioned, reproducible units. As a Senior DevOps Engineer for the Financial Cloud Toolchain, you... ...based on dependency graph analysis Reliability & Compliance Automation (10%) Automate...PrincipalFull time
- About Us Established in 2018, Bybit is one of the world’s leading cryptocurrency exchanges and digital financial platforms, serving over 80 million users across more than 200 countries and regions. Powered by world-class technology and a user-first mindset, Bybit delivers...PrincipalFull time
- ...problems within large-scale backend infrastructure Provide reliable and reactive trading information streams to our customers.... ...and Infra team to improve the technology stack for long-term engineering initiatives. Manage documentation for all implemented code...Senior
$10k - $12k
• For production support the consultants are responsible for carrying out the defect rectifications to solve BI problems in the production system • Incidence Analysis • Defect Rectifications • Process Chain Monitoring I. Monitor daily, weekly and monthly process...SeniorContract workShift work- ...requirements 3. Participate in business requirements analysis, system design, development, testing, and deployment; deliver high-quality engineering work 4. Write technical and system documentation to ensure code maintainability and team collaboration efficiency Requirements...PrincipalFull time
- ...solutions Own the development, refactoring, upgrade, and maintenance of CRM systems — including ticketing, agent workspace, rule engine, workflow orchestration, IM, and AI-assisted modules Tackle complex technical challenges, lead cross-team project delivery,...PrincipalFull time
- ...the next generation of financial services. Key Responsibilities 1. Agent SDK & Runtime — Design and implement the orchestration engine, runtime, and core primitives; ensure the agentic loop is both steerable and verifiable while remaining conversational and adaptive...PrincipalFull timeContract work
$197k
...deployments and ensure optimal performance, cost efficiency, and reliability of data processing systems. Data Quality Management:... ...Contribute to architectural decisions and drive adoption of data engineering best practices across the organization. Documentation &...SeniorRemote jobWork at officeFlexible hours- ...As an early adopter of React Native, we currently maintain one of the richest React Native applications on the app stores. As a RN engineer, you’ll come to us having experienced many aspects of RN development, from state management to creating lovely interfaces. The...SeniorRemote jobFull timeWork at officeFlexible hours
- About Us Established in 2018, Bybit is one of the world’s leading cryptocurrency exchanges and digital financial platforms, serving over 80 million users across more than 200 countries and regions. Powered by world-class technology and a user-first mindset, Bybit delivers...SeniorFull time
- ...etc.) in data streaming architectures ~ Proficient with real-time data transport protocols such as WebSocket and gRPC ~ Strong engineering discipline — clean code structure and a habit of conducting code reviews Why Join Us At Bybit, we are committed to...PrincipalFull time
- About Us Established in 2018, Bybit is one of the world’s leading cryptocurrency exchanges and digital financial platforms, serving over 80 million users across more than 200 countries and regions. Powered by world-class technology and a user-first mindset, Bybit delivers...PrincipalFull time
- ...financial services. Key Responsibilities 1. Recommendation Engine & Full-Pipeline Development Multi-stage Engine Development:... .... Global Distributed Tracing: In complex multi-region/multi-site/multi-language cross-border IDC environments, build and maintain...SeniorFull time
- ...4x7 (rotation basis). To assist the system integration deployment project. To obtain job related Industry Certification for principal partnership requirement. EXPERIENCE / SKILLS REQUIRED: Candidate must possess at least a Bachelor's Degree in relevant IT related...SeniorWork at office
- ...ABOUT YOU We are looking for a Mid/Senior Developer who is highly skilled, reliable, and excited about building payment solutions to join our payments engineering team in Kuala Lumpur. The best candidate will be someone who thrives in a fast-paced, highly collaborative...SeniorFull timeWork at officeLocal areaWorldwideFlexible hours
- ...Kuala Lumpur, Malaysia, with a requirement to work from the office three days per week, and partners closely with hiring managers and senior stakeholders across all three markets. This Recruiter will manage approximately 15–20 mid- to senior-level requisitions across...PrincipalRemote jobFull timeWork at officeFlexible hours3 days per week
- About Us Established in 2018, Bybit is one of the world’s leading cryptocurrency exchanges and digital financial platforms, serving over 80 million users across more than 200 countries and regions. Powered by world-class technology and a user-first mindset, Bybit delivers...PrincipalFull time
- ...to emerging technologies and new industry trends is a plus. 1. Experience in developing performance testing platforms or chaos engineering platforms is a strong plus. Why Join Us At Bybit, we are committed to fostering a supportive and enriching work environment....PrincipalFull time
- ...ledger, channel statements, and clearing accounts 2. Abstract complex, fragmented fund logic into standardised workflows, rule-based engines, and configurable reusable product capabilities 3. Design and build the company-level reconciliation product system covering...PrincipalFull time
- ...Responsibilities: - Work closely with stakeholders, fellow engineers, designer, software tester. You will be presenting and communicating... ...solutions to hard problems that consider scale, security, reliability and cost - Design and create test data and/or environment to...SeniorFlexible hours
- Company : Sime Darby Auto Engineering Sdn. Bhd. To assist Production Executive to coordinate between Supplier, Production & Logistic on parts quality issueSenior
$14k - $29k
1SENIOR PHP DEVELOPER MYR 6,000 - MYR 12,000 Requirement : ● Degree or you just passionate in programming. ● We don't mind your qualifications, we do mind your work... let the work speak for you! ● At least 4 years of experience in web development using LAMP stack...Senior- ...ABOUT YOU We are looking for a Senior Frontend Engineer who is creative, detail-oriented, and passionate about crafting exceptional user experiences to join our frontend engineering team in Kuala Lumpur. The best candidate will be someone who thrives in a fast-paced...SeniorFull timeWork at officeLocal areaWorldwideFlexible hours
- ...disrupting traditional finance. Embrace the dynamism and challenges within our diverse collective of pioneers and disruptors. As a Senior Backend Engineer at MoneyLion, your expertise will be pivotal in architecting the backbone of our fintech solutions. Dive into our forward-...SeniorRemote jobFull timeFlexible hours
- ...release traffic allocation, A/B experiment design, and effectiveness evaluation. Drive strategy integration into the real-time decision engine for production launch. Continuously counter black market actors, forming a closed loop of risk identification → intervention → post...PrincipalFull time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Principal Site Reliability Engineer. Be the first to apply!
- chief engineer Malaysia
- engineering director Malaysia
- hotel chief engineer Malaysia
- principal developer Malaysia
- principal engineer Malaysia
- data center chief engineer Malaysia
- general engineer Malaysia
- senior international accountant Malaysia
- senior performance engineer Malaysia
- senior implementation project manager Malaysia
