Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer - Telemetry

Full-time

Kraken

Building the Future of Open Finance

Payward - the parent company behind Kraken, NinjaTrader, Breakout, xStocks, Payward Services and CF Benchmarks - has spent the last 15 years building one of the most modern and globally accessible financial infrastructure platforms in the industry, built to advance an open, global financial system.

Before you apply, we encourage you to explore our [culture page]( to understand what drives us and how we work.

The team

Founded in 2011, Kraken is one of the world's longest-standing crypto platforms, trusted by over 10 million individuals and institutions across the globe. It offers spot trading, margin, futures, staking, and OTC services, with products built for both individual investors and institutional clients.

Join our engineering team as a Site Reliability Engineer focused on the shared telemetry platform that helps teams understand, operate, and improve production services. You will work across metrics, logs, traces, alerting, dashboards, and profiling systems that make our platform observable at scale.

As a member of the Telemetry team, you will help keep the systems that collect, store, query, and route operational data reliable and scalable. This is a hands-on role for an engineer who enjoys distributed systems, automation, observability, and solving production problems.

You will work with Software Engineers, Platform teams, Security Engineers, and SREs to improve operational practices and support mission-critical services. The role includes incident response, on-call, platform improvements, and helping teams use telemetry effectively.

The opportunity

  • Operate and improve the shared platform for metrics, logs, traces, alerting, dashboards, and profiling.
  • Maintain metrics collection, long-term storage, querying, dashboards, and alerting using Prometheus-compatible systems, VictoriaMetrics, Grafana, and modern alerting tools.
  • Operate log pipelines using Vector, Splunk, and Loki, including reliability, throughput, and troubleshooting.
  • Operate distributed tracing and profiling capabilities using Grafana Alloy, Tempo, OpenTelemetry, and Pyroscope.
  • Deploy and manage telemetry services using Terraform, Terragrunt, and container orchestration across multiple environments.
  • Troubleshoot missing data, slow queries, broken alerts, pipeline backpressure, and capacity issues.
  • Build reusable configuration and automation that helps teams manage dashboards, alerts, and telemetry integrations safely.
  • Participate in incident response and on-call, write runbooks, and improve the platform using what we learn from incidents.

What you bring

  • 3+ years of experience as a Site Reliability Engineer, Platform/Infrastructure Engineer, Observability Engineer, or similar production engineering role.
  • Comfortable managing production systems at scale that collect, process, store, and serve telemetry such as metrics, logs, traces, or profiles.
  • Experience with Prometheus or a Prometheus-compatible monitoring stack, including metrics collection, querying, and alerting.
  • Experience troubleshooting distributed production systems, including availability, latency, data flow, and capacity issues.
  • Experience with Infrastructure as Code, particularly Terraform, and CI/CD.
  • Experience operating containerised workloads with Nomad, Kubernetes, or similar platforms.
  • Solid scripting/programming ability and comfort using AI tools and agents (e.g., Claude) to accelerate delivery.
  • Strong incident response, documentation, and collaboration skills.

Nice to haves

  • Experience with VictoriaMetrics, Grafana, Tempo, Loki, Vector, Splunk, Alertmanager, or OpenTelemetry.
  • Experience with PromQL or LogQL, dashboards or alerts as code.
  • Experience with maintaining operators and their related CRDs in Kubernetes.
  • Experience with Consul, Vault, AWS, and on-premises or datacentre infrastructure.
  • Experience operating high-volume logging, streaming, or data pipelines.
  • Experience making practical trade-offs between observability data volume, performance, and cost.
  • Background in highly regulated or financial services environments where change management and audit trails are critical.
Unless a specific application deadline is stated in the job posting, applications are accepted on an ongoing basis. Please note, applicants are permitted to redact or remove information on their resume that identifies age, date of birth, or dates of attendance at or graduation from an educational institution. We consider qualified applicants with criminal histories for employment on our team, assessing candidates in a manner consistent with the requirements of the San Francisco Fair Chance Ordinance.

Our commitment

Payward is powered by people from around the world and we celebrate the diverse talents, backgrounds, contributions, and unique perspectives that everyone brings to the table. We hire based on merit, seeking out people with the right abilities, knowledge, and skills for the job. We encourage you to apply for roles where you don't fully meet the listed requirements, especially if you're passionate or knowledgeable about crypto.

We may ask candidates to complete job-related skills or work-style assessments as part of our hiring process. These assessments evaluate competencies relevant to the role and are applied consistently across candidates for similar positions. Results are considered alongside experience and interviews, and are not the sole basis for any employment decision.

As an equal opportunity employer, we don't tolerate discrimination or harassment of any kind, whether based on race, ethnicity, age, gender identity, citizenship, religion, sexual orientation, disability, pregnancy, veteran status, or any other protected characteristic as outlined by federal, state, or local laws.

Stay connected

[Follow us on Twitter](

[Learn on the Kraken Blog](

[Connect on LinkedIn](

[Candidate Privacy Notice](

Vacancy posted 22 days ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer - Telemetry in Remote vacancy
  • $141.8k - $195k

     ...Join the company that's building the telemetry infrastructure for the AI era. At Cribl, we partner with IT and Security teams...  ...You'll Love This Role Cribl Inc is seeking a Senior Site Reliability Engineer to join our mission where you will unlock the value of all... 
    Suggested
    Temporary work
    Remote work

    Cribl

    United States
    5 days ago
  • $207k - $300k

     ...pushing for changes that improve reliability and velocity.Practice...  ...degree in Computer Science or Engineering, or a related field.Experience...  ...complex client-to-cloud telemetry systems, CUJ monitoring frameworks...  ...engineering organizations.Site Reliability Engineering (SRE... 
    Suggested

    Google

    San Francisco, CA
    1 day ago
  • $76k - $127k

     ...products and services that help people, businesses and governments realize their greatest potential. Title and Summary Site Reliability Engineer II Who is Mastercard? At Mastercard technology, we work to connect and power an inclusive, digital economy that benefits... 
    Suggested
    Full time
    Part time
    Worldwide
    Flexible hours

    Mastercard

    O Fallon, MO
    36 minutes ago
  •  ...services that help people, businesses and governments realize their greatest potential. Title and Summary Senior Site Reliability Engineer Overview-The ProCOM team is looking for a Site Reliability Engineering (SRE) who can help us solve problems, build our... 
    Suggested
    Full time
    Part time
    Immediate start
    Worldwide
    Flexible hours

    Mastercard

    O Fallon, MO
    36 minutes ago
  • $76k - $127k

     ...products and services that help people, businesses and governments realize their greatest potential. Title and Summary Site Reliability Engineer II The BizOps team at Mastercard is looking for a Site Reliability Engineer (SRE) who thrives on solving complex... 
    Suggested
    Full time
    Part time
    Worldwide
    Flexible hours
    Early shift

    Mastercard

    O Fallon, MO
    36 minutes ago
  • $96k - $163k

     ...products and services that help people, businesses and governments realize their greatest potential. Title and Summary Senior Site Reliability Engineer, Performance Engineering Senior Site Reliability Engineer, Performance Engineering Payment Optimization unifies... 
    Full time
    Part time
    Worldwide
    Flexible hours

    Mastercard

    O Fallon, MO
    36 minutes ago
  • $96k - $163k

     ...products and services that help people, businesses and governments realize their greatest potential. Title and Summary Senior Site Reliability Engineer Who is Mastercard? At Mastercard technology, we work to connect and power an inclusive, digital economy that... 
    Full time
    Part time
    Worldwide
    Flexible hours

    Mastercard

    O Fallon, MO
    36 minutes ago
  •  ...products and services that help people, businesses and governments realize their greatest potential. Title and Summary Lead Site Reliability Engineer Job Description Summary Overview: Who is Mastercard? At Mastercard technology, we work to connect and power an... 
    Full time
    Part time
    Worldwide
    Flexible hours

    Mastercard

    O Fallon, MO
    36 minutes ago
  • $122k - $207k

     ...and services that help people, businesses and governments realize their greatest potential. Title and Summary Manager, Site Reliability Engineering Who is Mastercard? At Mastercard technology, we work to connect and power an inclusive, digital economy that... 
    Full time
    Part time
    Worldwide
    Flexible hours

    Mastercard

    O Fallon, MO
    36 minutes ago
  • $99k - $225k

     ...automation pipelines.Conduct traffic flow analysis using Illumio VEN telemetry and assist in building policy recommendations to reduce attack...  ...total benefits by visiting the Resource page on our Careers site and reviewing Our Employee Benefits page.Salary at Booz Allen... 
    Full time
    Contract work
    Part time
    Work at office
    Local area
    Remote work

    Booz Allen Hamilton

    Norfolk, VA
    1 day ago
  • A Sr Platform Engineer will assist in all aspects of lifecycle management of enterprise virtualization platforms and...  ...through to retirement), leveraging automation and systems telemetry to maintain platform reliability, optimize capacity, and support enterprise technology... 
    Local area
    Remote work
    Flexible hours

    O'Reilly Auto Parts

    Springfield, MO
    3 days ago
  • $146k - $194k

     ...focused on positioning Anduril as a lead provider of specialized engineering and products for Intelligence Community (IC) customers. We...  ...pressing national security requirements.ABOUT THE JOBAs a Site Reliability Engineer, your primary mission is to ensure the health,... 
    Full time
    Work experience placement
    Immediate start
    Remote work

    Anduril Industries

    Reston, VA
    4 days ago
  •  ...and responsible for ensuring the availability, scalability, and reliability of systems and applications.What will be your responsibilities...  ...using tools like Terraform or CloudFormation.Mentor junior engineers and provide technical guidance.Stay up-to-date with industry trends... 
    Work at office
    Remote work

    Interactive Brokers

    Greenwich, CT
    5 days ago
  • $62k - $141k

    Site Reliability EngineerThe Opportunity: Engineering to make a system more resilient and efficient frees up time and money to build more capabilities. Whether you come from a background in network engineering, systems administration, or software development, if you have... 
    Full time
    Contract work
    Part time
    Work at office
    Local area
    Remote work

    Booz Allen Hamilton

    Chantilly, Loudoun County, VA
    5 days ago
  • $127k - $249k

    The TeamPlatform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions...  ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper).... 
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    Austin, TX
    3 days ago
  • $158.5k - $172k

     ...exceptional value they deserve.About The OpportunityAs a Senior Engineer on the Runtime Automation team, you will design, automate, and...  .... This is a high-impact position driving continuous reliability, deep system optimization, and automation across our entire technology... 
    Full time
    Temporary work
    Work at office
    Flexible hours
    3 days per week

    GrubHub

    New York, NY
    1 day ago
  • $110k - $120k

     ...expertise, scale, and technology.Job DescriptionJob Title: Site Reliability Engineer (SRE) / L3 Support EngineerGetting to know us:As a leading...  ...services.Proactively identify emerging issues using telemetry, logs, metrics, and distributed tracing.Investigate, troubleshoot... 
    Ongoing contract
    Full time
    Casual work
    Remote work
    Flexible hours

    SS&C Technologies

    Pennsylvania
    3 days ago
  • $140k - $150k

    WORK OPTION: Remote_________________The NBA is hiring a Senior Site Reliability Engineer (SRE) - Messaging & Collaboration to ensure the availability, performance, and reliability of enterprise messaging and collaboration platforms, including Microsoft Exchange Online (... 
    Full time
    Temporary work
    Local area
    Remote work
    Weekend work

    National Basketball Association

    Secaucus, NJ
    2 days ago
  • Reliability Engineering Design, implement, and operate scalable, resilient, and highly available systems on Google Cloud Platform. Improve service...  ...Skills, and Abilities Three or more years of experience in Site Reliability Engineering, platform engineering, DevOps, cloud... 
    Remote work

    Patterson-UTI

    Houston, TX
    4 days ago
  • $134.25k - $214.8k

     ...matters at a company where you matter.Your ImpactAre you an engineer who gets excited about the challenge of making complex distributed...  ...it.You will be part of the Observability team within Axon's Site Reliability organization — a focused team responsible for Axon's metrics,... 
    Work experience placement
    Work at office
    Remote work

    Axon

    Seattle, WA
    2 days ago
  • $150k - $180k

     ...Umbra.About the JobWe are seeking an experienced SeniorSite Reliability Engineer to help design, build, operate, and scale the mission- and business...  ...impact across the organization.This position is based on-site in either our Arlington, VA office, Reston, VA office or... 
    Permanent employment
    Full time
    Work at office
    Local area
    Remote work
    Worldwide

    Umbra

    Arlington, VA
    1 day ago
  •  ...generative AI and cloud-native platforms to advanced release engineering practices, our teams are redefining how financial technology...  ...AI-driven solutions that accelerate development and improve reliability. Your work will directly influence how GM Financial leverages... 
    H1b
    Work at office
    Remote work
    Visa sponsorship
    Flexible hours
    2 days per week

    GM Financial

    Arlington, TX
    1 day ago
  • $87.12k - $151.25k

     ...to be part of an inclusive, adaptable, and forward-thinking organization, apply now.We are currently seeking a Digital Site Reliability Sr Engineer - Remote to join our team in Memphis, Tennessee (US-TN), United States (US).Digital Site Reliability Senior EngineerWe are... 
    Temporary work
    Work at office
    Remote work
    Flexible hours

    NTT DATA

    Memphis, TN
    3 days ago
  • $72.8k - $130k

     ...modernization program aimed at updating and enhancing enterprise technology systems in accordance with modern design standardsThe Site Reliability Engineer will architect, develop, and maintain Optum Serve's cloud environment in both the commercial and government clouds. The... 
    Minimum wage
    Full time
    Work experience placement
    Work at office
    Local area
    Remote work

    UnitedHealth Group

    Eden Prairie, MN
    1 day ago
  • $87.1k - $157.45k

     ...throughout the entire USG arsenal. Our team of hackers, engineers, makers, and shakers brings deep experience across...  ...to come in and help us build systems that stay reliable when things get complicated.We need a Site Reliability Engineer who has experience building, deploying... 
    Full time
    Work from home
    Flexible hours

    Leidos

    Chantilly, Loudoun County, VA
    1 day ago
  •  ...an SRE to join our infrastructure team. This role will be responsible for building software to ensure the reliability of our back-end systems, working with engineers who develop them, and planning for our future growth. You will work with our existing production... 
    Worldwide
    Home office
    Flexible hours

    Superhuman

    San Francisco, CA
    1 day ago
  • Edmond, OKYouVersion - YouVersion Engineering /Full-Time/ Salary /On-siteThe YouVersion Senior Site Reliability Engineer is responsible for ensuring the integrity, performance, reliability, and cost-effectiveness of the cloud-based infrastructure and related systems supporting... 
    Full time
    Contract work
    Temporary work
    Work experience placement
    Casual work
    Internship
    Local area
    Worldwide

    Life.Church

    Edmond, OK
    1 day ago
  • $165k - $190k

     ...DevOps / SRE TeamThe DevOps/SRE team at Obsidian ensures that engineering excellence translates into stable, scalable, and high-...  ...security platformAddress complex challenges around scalability, reliability, observability, and cost efficiencyCollaborate with Engineering... 
    Work from home

    Obsidian Security

    Palo Alto, CA
    3 days ago
  • $80k - $133k

     ...degree, Four (4) years additional experience will be needed.Minimum Four (4) years of experience in IT administration, software engineering, or platform engineering, with a focus on AWS cloud infrastructure and enterprise systems.One(1)+ years of experience deploying and... 
    Permanent employment
    Full time
    Contract work
    Remote work
    Flexible hours

    Guidehouse

    McLean, VA
    3 days ago
  •  ...The selected colleague will work at an MUFG office or client sites four days per week and work remotely one day. A member of...  ...MUFG is seeking a highly motivated Certified Sr. Cloud Site Reliability Engineer to build a robust, scalable, and reliable web application environment... 
    Full time
    Work at office
    Local area
    Remote work

    MUFG

    Tampa, FL
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer - Telemetry. Be the first to apply!