Site Reliability Engineer - Telemetry
Kraken
Building the Future of Open Finance
Payward - the parent company behind Kraken, NinjaTrader, Breakout, xStocks, Payward Services and CF Benchmarks - has spent the last 15 years building one of the most modern and globally accessible financial infrastructure platforms in the industry, built to advance an open, global financial system.
Before you apply, we encourage you to explore our [culture page]( to understand what drives us and how we work.
The team
Founded in 2011, Kraken is one of the world's longest-standing crypto platforms, trusted by over 10 million individuals and institutions across the globe. It offers spot trading, margin, futures, staking, and OTC services, with products built for both individual investors and institutional clients.
Join our engineering team as a Site Reliability Engineer focused on the shared telemetry platform that helps teams understand, operate, and improve production services. You will work across metrics, logs, traces, alerting, dashboards, and profiling systems that make our platform observable at scale.
As a member of the Telemetry team, you will help keep the systems that collect, store, query, and route operational data reliable and scalable. This is a hands-on role for an engineer who enjoys distributed systems, automation, observability, and solving production problems.
You will work with Software Engineers, Platform teams, Security Engineers, and SREs to improve operational practices and support mission-critical services. The role includes incident response, on-call, platform improvements, and helping teams use telemetry effectively.
The opportunity
- Operate and improve the shared platform for metrics, logs, traces, alerting, dashboards, and profiling.
- Maintain metrics collection, long-term storage, querying, dashboards, and alerting using Prometheus-compatible systems, VictoriaMetrics, Grafana, and modern alerting tools.
- Operate log pipelines using Vector, Splunk, and Loki, including reliability, throughput, and troubleshooting.
- Operate distributed tracing and profiling capabilities using Grafana Alloy, Tempo, OpenTelemetry, and Pyroscope.
- Deploy and manage telemetry services using Terraform, Terragrunt, and container orchestration across multiple environments.
- Troubleshoot missing data, slow queries, broken alerts, pipeline backpressure, and capacity issues.
- Build reusable configuration and automation that helps teams manage dashboards, alerts, and telemetry integrations safely.
- Participate in incident response and on-call, write runbooks, and improve the platform using what we learn from incidents.
What you bring
- 3+ years of experience as a Site Reliability Engineer, Platform/Infrastructure Engineer, Observability Engineer, or similar production engineering role.
- Comfortable managing production systems at scale that collect, process, store, and serve telemetry such as metrics, logs, traces, or profiles.
- Experience with Prometheus or a Prometheus-compatible monitoring stack, including metrics collection, querying, and alerting.
- Experience troubleshooting distributed production systems, including availability, latency, data flow, and capacity issues.
- Experience with Infrastructure as Code, particularly Terraform, and CI/CD.
- Experience operating containerised workloads with Nomad, Kubernetes, or similar platforms.
- Solid scripting/programming ability and comfort using AI tools and agents (e.g., Claude) to accelerate delivery.
- Strong incident response, documentation, and collaboration skills.
Nice to haves
- Experience with VictoriaMetrics, Grafana, Tempo, Loki, Vector, Splunk, Alertmanager, or OpenTelemetry.
- Experience with PromQL or LogQL, dashboards or alerts as code.
- Experience with maintaining operators and their related CRDs in Kubernetes.
- Experience with Consul, Vault, AWS, and on-premises or datacentre infrastructure.
- Experience operating high-volume logging, streaming, or data pipelines.
- Experience making practical trade-offs between observability data volume, performance, and cost.
- Background in highly regulated or financial services environments where change management and audit trails are critical.
Our commitment
Payward is powered by people from around the world and we celebrate the diverse talents, backgrounds, contributions, and unique perspectives that everyone brings to the table. We hire based on merit, seeking out people with the right abilities, knowledge, and skills for the job. We encourage you to apply for roles where you don't fully meet the listed requirements, especially if you're passionate or knowledgeable about crypto.
We may ask candidates to complete job-related skills or work-style assessments as part of our hiring process. These assessments evaluate competencies relevant to the role and are applied consistently across candidates for similar positions. Results are considered alongside experience and interviews, and are not the sole basis for any employment decision.
As an equal opportunity employer, we don't tolerate discrimination or harassment of any kind, whether based on race, ethnicity, age, gender identity, citizenship, religion, sexual orientation, disability, pregnancy, veteran status, or any other protected characteristic as outlined by federal, state, or local laws.
Stay connected[Follow us on Twitter](
[Learn on the Kraken Blog](
[Connect on LinkedIn](
[Candidate Privacy Notice](
$141.8k - $195k
...Join the company that's building the telemetry infrastructure for the AI era. At Cribl, we partner with IT and Security teams... ...You'll Love This Role Cribl Inc is seeking a Senior Site Reliability Engineer to join our mission where you will unlock the value of all...SuggestedTemporary workRemote work$207k - $300k
...pushing for changes that improve reliability and velocity.Practice... ...degree in Computer Science or Engineering, or a related field.Experience... ...complex client-to-cloud telemetry systems, CUJ monitoring frameworks... ...engineering organizations.Site Reliability Engineering (SRE...Suggested$76k - $127k
...products and services that help people, businesses and governments realize their greatest potential. Title and Summary Site Reliability Engineer II Who is Mastercard? At Mastercard technology, we work to connect and power an inclusive, digital economy that benefits...SuggestedFull timePart timeWorldwideFlexible hours- ...services that help people, businesses and governments realize their greatest potential. Title and Summary Senior Site Reliability Engineer Overview-The ProCOM team is looking for a Site Reliability Engineering (SRE) who can help us solve problems, build our...SuggestedFull timePart timeImmediate startWorldwideFlexible hours
$76k - $127k
...products and services that help people, businesses and governments realize their greatest potential. Title and Summary Site Reliability Engineer II The BizOps team at Mastercard is looking for a Site Reliability Engineer (SRE) who thrives on solving complex...SuggestedFull timePart timeWorldwideFlexible hoursEarly shift$96k - $163k
...products and services that help people, businesses and governments realize their greatest potential. Title and Summary Senior Site Reliability Engineer, Performance Engineering Senior Site Reliability Engineer, Performance Engineering Payment Optimization unifies...Full timePart timeWorldwideFlexible hours$96k - $163k
...products and services that help people, businesses and governments realize their greatest potential. Title and Summary Senior Site Reliability Engineer Who is Mastercard? At Mastercard technology, we work to connect and power an inclusive, digital economy that...Full timePart timeWorldwideFlexible hours- ...products and services that help people, businesses and governments realize their greatest potential. Title and Summary Lead Site Reliability Engineer Job Description Summary Overview: Who is Mastercard? At Mastercard technology, we work to connect and power an...Full timePart timeWorldwideFlexible hours
$122k - $207k
...and services that help people, businesses and governments realize their greatest potential. Title and Summary Manager, Site Reliability Engineering Who is Mastercard? At Mastercard technology, we work to connect and power an inclusive, digital economy that...Full timePart timeWorldwideFlexible hours$99k - $225k
...automation pipelines.Conduct traffic flow analysis using Illumio VEN telemetry and assist in building policy recommendations to reduce attack... ...total benefits by visiting the Resource page on our Careers site and reviewing Our Employee Benefits page.Salary at Booz Allen...Full timeContract workPart timeWork at officeLocal areaRemote work- A Sr Platform Engineer will assist in all aspects of lifecycle management of enterprise virtualization platforms and... ...through to retirement), leveraging automation and systems telemetry to maintain platform reliability, optimize capacity, and support enterprise technology...Local areaRemote workFlexible hours
$146k - $194k
...focused on positioning Anduril as a lead provider of specialized engineering and products for Intelligence Community (IC) customers. We... ...pressing national security requirements.ABOUT THE JOBAs a Site Reliability Engineer, your primary mission is to ensure the health,...Full timeWork experience placementImmediate startRemote work- ...and responsible for ensuring the availability, scalability, and reliability of systems and applications.What will be your responsibilities... ...using tools like Terraform or CloudFormation.Mentor junior engineers and provide technical guidance.Stay up-to-date with industry trends...Work at officeRemote work
$62k - $141k
Site Reliability EngineerThe Opportunity: Engineering to make a system more resilient and efficient frees up time and money to build more capabilities. Whether you come from a background in network engineering, systems administration, or software development, if you have...Full timeContract workPart timeWork at officeLocal areaRemote work$127k - $249k
The TeamPlatform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions... ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper)....Work at officeLocal areaRemote workWorldwideFlexible hours$158.5k - $172k
...exceptional value they deserve.About The OpportunityAs a Senior Engineer on the Runtime Automation team, you will design, automate, and... .... This is a high-impact position driving continuous reliability, deep system optimization, and automation across our entire technology...Full timeTemporary workWork at officeFlexible hours3 days per week$110k - $120k
...expertise, scale, and technology.Job DescriptionJob Title: Site Reliability Engineer (SRE) / L3 Support EngineerGetting to know us:As a leading... ...services.Proactively identify emerging issues using telemetry, logs, metrics, and distributed tracing.Investigate, troubleshoot...Ongoing contractFull timeCasual workRemote workFlexible hours$140k - $150k
WORK OPTION: Remote_________________The NBA is hiring a Senior Site Reliability Engineer (SRE) - Messaging & Collaboration to ensure the availability, performance, and reliability of enterprise messaging and collaboration platforms, including Microsoft Exchange Online (...Full timeTemporary workLocal areaRemote workWeekend work- Reliability Engineering Design, implement, and operate scalable, resilient, and highly available systems on Google Cloud Platform. Improve service... ...Skills, and Abilities Three or more years of experience in Site Reliability Engineering, platform engineering, DevOps, cloud...Remote work
$134.25k - $214.8k
...matters at a company where you matter.Your ImpactAre you an engineer who gets excited about the challenge of making complex distributed... ...it.You will be part of the Observability team within Axon's Site Reliability organization — a focused team responsible for Axon's metrics,...Work experience placementWork at officeRemote work$150k - $180k
...Umbra.About the JobWe are seeking an experienced SeniorSite Reliability Engineer to help design, build, operate, and scale the mission- and business... ...impact across the organization.This position is based on-site in either our Arlington, VA office, Reston, VA office or...Permanent employmentFull timeWork at officeLocal areaRemote workWorldwide- ...generative AI and cloud-native platforms to advanced release engineering practices, our teams are redefining how financial technology... ...AI-driven solutions that accelerate development and improve reliability. Your work will directly influence how GM Financial leverages...H1bWork at officeRemote workVisa sponsorshipFlexible hours2 days per week
$87.12k - $151.25k
...to be part of an inclusive, adaptable, and forward-thinking organization, apply now.We are currently seeking a Digital Site Reliability Sr Engineer - Remote to join our team in Memphis, Tennessee (US-TN), United States (US).Digital Site Reliability Senior EngineerWe are...Temporary workWork at officeRemote workFlexible hours$72.8k - $130k
...modernization program aimed at updating and enhancing enterprise technology systems in accordance with modern design standardsThe Site Reliability Engineer will architect, develop, and maintain Optum Serve's cloud environment in both the commercial and government clouds. The...Minimum wageFull timeWork experience placementWork at officeLocal areaRemote work$87.1k - $157.45k
...throughout the entire USG arsenal. Our team of hackers, engineers, makers, and shakers brings deep experience across... ...to come in and help us build systems that stay reliable when things get complicated.We need a Site Reliability Engineer who has experience building, deploying...Full timeWork from homeFlexible hours- ...an SRE to join our infrastructure team. This role will be responsible for building software to ensure the reliability of our back-end systems, working with engineers who develop them, and planning for our future growth. You will work with our existing production...WorldwideHome officeFlexible hours
- Edmond, OKYouVersion - YouVersion Engineering /Full-Time/ Salary /On-siteThe YouVersion Senior Site Reliability Engineer is responsible for ensuring the integrity, performance, reliability, and cost-effectiveness of the cloud-based infrastructure and related systems supporting...Full timeContract workTemporary workWork experience placementCasual workInternshipLocal areaWorldwide
$165k - $190k
...DevOps / SRE TeamThe DevOps/SRE team at Obsidian ensures that engineering excellence translates into stable, scalable, and high-... ...security platformAddress complex challenges around scalability, reliability, observability, and cost efficiencyCollaborate with Engineering...Work from home$80k - $133k
...degree, Four (4) years additional experience will be needed.Minimum Four (4) years of experience in IT administration, software engineering, or platform engineering, with a focus on AWS cloud infrastructure and enterprise systems.One(1)+ years of experience deploying and...Permanent employmentFull timeContract workRemote workFlexible hours- ...The selected colleague will work at an MUFG office or client sites four days per week and work remotely one day. A member of... ...MUFG is seeking a highly motivated Certified Sr. Cloud Site Reliability Engineer to build a robust, scalable, and reliable web application environment...Full timeWork at officeLocal areaRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer - Telemetry. Be the first to apply!
- site reliability engineer Remote
- site reliability engineer sre Remote
- site reliability engineer remote Remote
- website coordinator Remote
- on-site clinical research associate (traveling/remote) Remote
- site safety Remote
- junior website developer Remote
- construction site safety Remote
- IT site lead Remote
- website content developer Remote


