Site Reliability Engineer - AI Agents [Remote]
Kraken
- Remote job
Building the Future of Open Finance
Payward - the parent company behind Kraken, NinjaTrader, Breakout, xStocks, Payward Services and CF Benchmarks - has spent the last 15 years building one of the most modern and globally accessible financial infrastructure platforms in the industry, built to advance an open, global financial system.
Before you apply, we encourage you to explore our [culture page]( to understand what drives us and how we work.
The team
Founded in 2011, Kraken is one of the world's longest-standing crypto platforms, trusted by over 10 million individuals and institutions across the globe. It offers spot trading, margin, futures, staking, and OTC services, with products built for both individual investors and institutional clients.
The AI Infrastructure team sits within the Data organization and is responsible for building, operating, and scaling the systems that power AI agents in production — both internal tools and external-facing products. Working closely with the AI and Agent Systems teams, this group ensures that the orchestration, execution, and model-serving layers underpinning agentic workflows are reliable, observable, and built to scale.
This team operates at the intersection of data infrastructure and applied AI — a space that moves fast and demands engineers who can bring production discipline to emerging technology. You'll partner across Data Engineering, ML, and product-facing teams to harden agent infrastructure and keep it running at the standards our users expect.
Importantly, this is a platform engineering team. Beyond operating infrastructure, the team is responsible for building the APIs, SDKs, and platform capabilities that enable AI, Data, and Engineering teams to safely and efficiently consume agent infrastructure as a service. Success in this role requires thinking beyond infrastructure operations and toward developer experience, platform adoption, and long-term scalability.
The opportunity
- Design, build, and operate the infrastructure layer supporting AI agent workflows in production
- Ensure reliability, scalability, and observability of agentic systems across internal and external products
- Design and develop platform services, APIs, SDKs, and self-service capabilities that allow engineering teams to easily consume AI infrastructure and agent platform services
- Manage and maintain the compute, orchestration, and serving infrastructure powering model inference and agent execution
- Implement robust monitoring, alerting, and incident response procedures tailored to AI/ML workloads
- Utilize Infrastructure as Code (IaC) tools such as Terraform to provision and manage cloud (AWS) infrastructure components
- Build and maintain CI/CD pipelines that support rapid, reliable deployment of AI services and agent workflows
- Define and implement guardrails, failure handling, and recovery patterns specific to agentic and LLM-powered systems
- Collaborate with AI and Data Engineering teams to translate experimental agent prototypes into hardened production systems
- Manage containerized workloads using Kubernetes, ensuring efficient deployment, scaling, and orchestration of AI services
- Implement access controls and security best practices across AI infrastructure environments
- Document architecture, runbooks, and best practices to support knowledge sharing across the team
What You Bring
- 5+ years of experience as a Site Reliability Engineer, Infrastructure Engineer, Platform Engineer, or similar role in a production environment
- Hands-on experience supporting ML infrastructure, model serving, or MLOps workflows in production
- Experience building developer platforms, internal tooling, APIs, or SDKs consumed by engineering teams at scale
- Strong understanding of platform engineering principles, including developer experience, self-service infrastructure, and API-driven platform design
- Proficiency with Infrastructure as Code tools, particularly Terraform
- Experience with containerization and orchestration, particularly Kubernetes and Docker
- Solid understanding of cloud infrastructure, preferably AWS
- Strong scripting skills (bash/shell) and proficiency in at least one programming language (Python preferred)
- Experience designing and operating observability, monitoring, and alerting systems
- Experience implementing incident response procedures and participating in on-call rotations
- Strong collaboration skills working across data, AI, and engineering teams
- High ownership mindset in a fast-moving, high-stakes production environment
Nice to haves
- Experience building or operating infrastructure for agent-based or LLM-powered systems
- Familiarity with agent orchestration frameworks (e.g., LangGraph, CrewAI, or similar)
- Background in data infrastructure, including familiarity with Airflow, Kafka, Spark, or data lake tooling
- Experience with CI/CD pipelines and deployment automation for AI/ML workloads
- Exposure to evaluation frameworks and model performance monitoring at scale
- Experience working in fast-moving 0→1 environments or platform-building teams
- Experience building SDKs, developer tooling, or internal platform products with a strong focus on usability and adoption
- Experience with Cloudflare's cloud platform and product ecosystem, including networking, security, performance, and Zero Trust solutions
Payward is powered by people from around the world and we celebrate the diverse talents, backgrounds, contributions, and unique perspectives that everyone brings to the table. We hire based on merit, seeking out people with the right abilities, knowledge, and skills for the job. We encourage you to apply for roles where you don't fully meet the listed requirements, especially if you're passionate or knowledgeable about crypto.
We may ask candidates to complete job-related skills or work-style assessments as part of our hiring process. These assessments evaluate competencies relevant to the role and are applied consistently across candidates for similar positions. Results are considered alongside experience and interviews, and are not the sole basis for any employment decision.
As an equal opportunity employer, we don't tolerate discrimination or harassment of any kind, whether based on race, ethnicity, age, gender identity, citizenship, religion, sexual orientation, disability, pregnancy, veteran status, or any other protected characteristic as outlined by federal, state, or local laws.
Stay in the know[Follow us on Twitter](
[Learn on the Kraken Blog](
[Connect on LinkedIn](
[Candidate Privacy Notice](
$100k - $130k
...individual investors and institutional clients. Join our engineering team as a Senior Site Reliability Engineer focused on the edge and ingress layer that... ...least one of Python or Golang) and comfort using AI tools and agents (e.g., Claude) to accelerate delivery. ~...SuggestedRemote jobFull timeWork at officeLocal area- ...place for you? Join us! Role Overview The Sr. Manager, Site Reliability Engineering (London) is an experienced leader responsible for overseeing... ...development, testing, and maintenance of Python tools. Leverage AI to maximize efficiency. On-Call & Weekend Testing: Lead...SuggestedFull timeWork at officeLocal areaWeekend work
- ...individual investors and institutional clients. The Agent Systems team is a 0→1 engineering group building internal AI-powered agents that interact directly with... ...are expected, but the bar for production reliability remains high. Engineers on this team move fluidly...SuggestedRemote jobFull timeLocal area
- ...Associates (CA) Working closely with the other Engineering teams to ensure the systems that are... ...Hybrid working: remote + on client site (Sheffield) Benefits Our people are... ...innovation. We are one of the world's leading AI and digital infrastructure providers,...SuggestedWork at officeRemote workFlexible hours
- ...products and services that help people, businesses and governments realize their greatest potential. Title and Summary Senior Site Reliability Engineer - Connex About Mastercard Mastercard is a global technology company in the payments industry, dedicated to enabling...SuggestedFull timeWorldwideFlexible hoursRotating shift
- ...Senior Customer Solutions Engineer and empower developers and agents to own and ensure application performance and reliability - from pull request to post... ...into a single AI-native workflow, so engineering... ...weeks of paid parental leave ~ Annual company on-site...Full timeRemote workFlexible hours
- ...access management. This role sits in our Platform Engineering team, which delivers features and design work that spans... ...The team is focused on building Ping’s identity for AI offerings to secure and govern AI agents. We are looking for a Senior Software Engineer who...Full timeLocal areaWorldwideFlexible hours
- ...Description We're seeking a talented and motivated full-time Software Engineer, AI Enablement to help Tailscale's engineering organization get... ..., internal tooling) that help engineers use coding agents like Claude Code , Codex , OpenCode , and Pi more effectively...Full time
- ...think and move like a start-up - and we're at the forefront of AI-driven development in games. We're looking for curious, ambitious... ...As a Junior Programmer, you'll work alongside experienced engineers to help build, improve and support Avakin Life, one of the world...Full timeWork from home
- ...collaborate and effectively build scalable, reliable, and secure backend systems that support... ...: ~5+ years experience in software engineering ~ Professional understanding of agile concepts... ..., or ELK Stack Experience with AI/ML integration in backend systems. Benefits...Full timeRemote workWork from homeWorldwideFlexible hours
$10k
...banking rails — giving every other team a reliable foundation to build on. What You'll Do... ...Need ~5+ years of backend or platform engineering experience in high-scale production environments... ...(Global) • Flexible PTO • Unlimited AI token usage • Centralized home-office...Remote jobFull timeWork at officeHome officeRelocation packageFlexible hours- ...consumer discussion boards on the modern internet. Position Overview We are seeking a highly technical, high-agency Staff Site Reliability Engineer to lead reliability engineering initiatives for our critical user-facing systems at absolute internet scale. Sitting at...Full timeRemote workFlexible hours
- ...Requirements EDUCATION AND EXPERIENCE: ~ BS in Computer Science, Engineering, or related field. ~10+ years’ software development... ...and DICOM is preferred. Early adopter mindset for utilizing AI coding assistants and emerging technologies is preferred....Full time
$500 per month
...market in real-time analytics, data warehousing, observability, and AI workloads. The company’s sustained, accelerating momentum... ...-source database for real-time apps and analytics. Our Core Engineering teams own the heart of our ClickHouse Open Source project. We are...Remote jobFull timeLocal areaHome officeFlexible hours- ...About AI Acquisition: AI Acquisition (aiacquisition.com) is a premier, internationally... ...frontier models, manipulate parallel multi-agent harnesses, and deploy highly resilient... ...connection-obsessed, and systems-minded AI Systems Engineer to join our decentralized core Operations...Full timeLocal areaRemote workWork from homeShift work
£85k - £110k per year
...micro-savings portfolios, Monzo continuously engineers human-centric financial products that... ...microservices, and deploy highly secure, low-latency AI solutions transforming financial lives.... ...serve machine learning models safely and reliably cleanly natively utilizing Machine...Full timeWork at officeRemote workWork from homeVisa sponsorshipRelocation packageFlexible hoursShift work- ...deploys ‘Marie’, a highly advanced AI phone assistant purpose-built... ...audio-proficient Senior AI Engineer | Voice to take comprehensive... ...maintaining, and scaling phone agents that process live... ...intricate audio-model concepts into reliable product features. Hands-on...Permanent employmentFull timeWork at officeLocal areaRemote workShift work
£1,500 per month
...clients, ranging from content management systems, to complex rules engines, finance calculators and tools. At ITG, Lead Java Developers... ...technologies with which they may be unfamiliar. We integrate AI tooling across the full development lifecycle — from business...Full timeFlexible hours- ...using the power of generative AI. Our core platform introduces... .... By combining practical engineering with advanced language models... ...CTO, and Product teams to ship reliable, user-facing features that extract... ...sophisticated multi-step AI agents and tool-use systems that...Shorter hoursPermanent employmentFull timeContract workLocal areaImmediate startRemote workShift work
- ...content is as discoverable as possible across both traditional search engines and AI-powered search. Be the technical SEO lead for a team of seven writers. Conduct technical SEO audits (site architecture, crawling, indexation, Core Web Vitals, schema, internal linking...Full timeFreelance
$500 per month
...market in real-time analytics, data warehousing, observability, and AI workloads. The company’s sustained, accelerating momentum... .... Come be a part of our journey! About The Team The Core Engineering team is responsible for working on the beloved ClickHouse open-...Remote jobFull timeLocal areaWorldwideHome officeFlexible hours- ...Founded in 2017, Wayve is the leading developer of Embodied AI technology. Our advanced AI software and foundation models enable... ...performance function. We are specifically looking for product-oriented engineers who will own the autonomy performance bar that determines...Full timeWork at officeWork from home
- ..., we're scaling rapidly and investing in AI to transform how we operate and how we serve... ...looking for an experienced Azure Cloud Engineer with deep Terraform and DevOps expertise... ...troubleshoot and optimise the performance, reliability and cost of Azure workloads. As the...Full timeRemote workFlexible hoursRotating shift
- ...Wayve is the leading developer of Embodied AI technology. Our advanced AI software and... ...operator support, balancing latency, reliability, and safety constraints. As a foundational... ...looking for product-oriented engineers who are adept at converting deep technical...Odd jobFull timeContract workWork at officeRemote workWork from home
£97k - £105k per year
...build an adaptable and scalable operation, increasingly powered by AI, to deliver the crucial insights necessary to confidently deploy... ...the entire Waymo Driver. A successful flywheel will be the core engine for scaling our technology, enabling faster international expansion...Full timeWork experience placementLocal area- ...accurate, trusted, and acted upon - not just looked at. Keep the AI in our BI sharp Continuously train and improve our BI tool’s... ...steward of how well our BI tool thinks. Partner with data engineering to optimise performance, and build what's missing Own and...Full time
- ...About aion Aion is the enterprise AI platform, a full-stack solution for building,... ...architectures for AI applications, intelligent agents, automation workflows, and enterprise... ...strategies. Collaborate closely with engineering teams to ensure successful implementation...Full time
- ...work The team Payward has world-class engineering, products, infrastructure, and regulatory... ...capabilities, developer tooling, and agent-native interfaces that can power the next... ...Payward the default choice for developers and AI agents building in financial services....Remote jobFull timeLocal area
- ...protect, optimize, and transform biomaterials engineering for industrial manufacturing at scale.... ..., and automated test suits to guarantee reliable platform scaling as user adoption grows.... ...the core application interface scaling our AI-powered biomaterial innovations....Full timeWork at officeRemote workWork from homeHome officeShift work
- ...data governance and validation workflows that maximize model reliability for top AI organizations at scale. Operating as an impact-focused... ...seeking a highly analytical, systems-minded Senior Backend Engineer (Go) to join an elite group of core technologists building...Hourly payDaily paidFull timeContract workFor contractorsFreelanceRemote workWork from home
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer - AI Agents [Remote]. Be the first to apply!
- site safety United Kingdom
- junior website developer United Kingdom
- construction site safety United Kingdom
- on-site clinical research associate (traveling/remote) United Kingdom
- remote data entry no experience United Kingdom
- remote from anywhere United Kingdom
- tableau developer remote United Kingdom
- remote health information technology United Kingdom
- part time remote tech United Kingdom
- chief marketing officer remote United Kingdom
















