Site Reliability Engineer
Yuno Inc
Site Reliability Engineer
Yuno is the AI-native operating system of global commerce, powering the financial infrastructure of enterprise merchants, banks, and wallets. Through a single API, Yuno connects them to pay-ins, payouts, fraud prevention, KYC/KYB, and stablecoins globally, so they can operate everywhere. Agnostic by design and connected to 1,000+ payment methods and 460+ integrations in 190+ countries, Yuno optimizes acceptance rates, reduces costs, and strengthens security through specialized AI agents that learn from every transaction. Global brands including McDonald's, NetEase Games, GoFundMe, and Rappi run their payments on Yuno.
About The Role
Yuno is looking for a Staff Site Reliability Engineer to set the technical direction for reliability across our infrastructure — starting with the platform that provisions, deploys, and manages AI agents at scale on AWS, the system powering payments across 190+ countries. The platform is in production and growing, and we need the most senior reliability voice in the room to evolve the architecture and make sure it stays reliable, observable, and ready to scale.
This is not a "maintain what exists" role, and it's not a single-system role. You'll own the reliability strategy — driving architectural decisions, designing event-driven communication, defining how we measure and defend reliability, and setting the standards other engineering teams build on.
How AI Shows Up In This Role
The platform you own is Yuno's AI agent infrastructure — provisioning and deploying AI agents at scale, plus the agents that route payments and prevent fraud. Keeping the AI-native layer reliable is the core of the role
AI is our default execution layer: you're encouraged to use AI-assisted tooling across automation, runbooks, incident analysis, and root-cause investigations, and to help define how the wider engineering org adopts it. We care how you use it, not whether you do
Your Contribution Will Be
Reliability strategy and standards — define the SLO culture, error-budget policy, and incident practices that scale across engineering teams, turning reliability from firefighting into a measurable, org-wide discipline
Platform architecture and evolution — drive architectural decisions as the platform matures; the deciding voice on choosing technologies, designing systems, and when to evolve the infrastructure
Messaging and event-driven architecture — design and own the messaging layer for inter-service communication, replacing synchronous patterns with durable, reliable async messaging
Infrastructure and deployment — own the cloud infrastructure, automate provisioning with IaC, and ensure the platform scales reliably as transaction volume grows
Observability — build the monitoring, tracing, and alerting that keeps the platform healthy; when something breaks at 3am, your dashboards and alerts should explain why before anyone has to dig
Incident leadership and mentorship — act as the senior escalation point for the hardest production problems, run blameless postmortems and root-cause analyses that turn into permanent fixes, and raise the reliability bar by mentoring senior and mid-level engineers
Chaos engineering mindset — continuous fault injection and resilience experiments that surface weaknesses before they turn into incidents, plus identifying and proposing resilience patterns to prevent those failures from reaching production.
What Success Looks Like
Within your first 6–12 months, you've set the reliability strategy for the platform, driven at least one major architectural evolution (event-driven messaging, streaming reliability, or observability), and engineering teams have adopted the SLO and error-budget framework you defined. You're the person Yuno trusts with the hardest reliability calls.
Skills You Need
Minimum Qualifications
Event-driven architecture and messaging systems — you've designed and owned systems around message queues (Kafka, NATS, RabbitMQ) and understand at-least-once delivery, consumer groups, dead letters, and backpressure; you've migrated a system from synchronous to async
Deep AWS — EC2, VPC, IAM, S3, and RDS — with strong networking fundamentals, since inter-service communication runs over the internal VPC
Infrastructure as Code — Terraform or Pulumi, reviewed in PRs rather than clicked in consoles
Kubernetes and Docker in production — container lifecycle, resource limits, health checks, and orchestration at scale
Observability and SLOs — Datadog fluency or equivalent (dashboards, monitors, APM, distributed tracing), and a track record defining and operating SLOs, SLIs, and error budgets across services
Chaos engineering and resilience testing — hands-on experience with fault injection, game days, or chaos experiments (Gremlin, Chaos Mesh, AWS FIS, or similar) to harden production systems
Distributed systems debugging — you've diagnosed async flows and cascading failures in production and can explain what broke and how you fixed it; comfortable coding for automation and tooling (Go, Python, or similar)
Databases — solid SQL (PostgreSQL) and NoSQL (MongoDB, Redis): when to use each, indexing, replication, and performance tuning
Proven technical leadership — you've set reliability standards, influenced architecture across teams, and mentored engineers, not just owned your own scope
English — advanced proficiency, written and spoken
Preferred Qualifications
AI / MLOps infrastructure — running AI workloads in production (model serving, LLM inference, GPU/resource management, and agent evaluation/observability tools like LangFuse, LangSmith, Braintrust, or MLflow)
Multi-tenant container platforms — running customer or user workloads in containers (Replit, Railway, Fly.io, or internal PaaS)
Data pipelines and orchestration — Airflow, Prefect, or similar; data warehouses like Databricks, Snowflake, or BigQuery a plus
Incident management and on-call tooling — PagerDuty, Opsgenie, or incident.io
Experience in the payments industry
Nice to Have
ECS experience
s6-overlay for container process supervision
Experience with AI agent framework ecosystems
Spanish proficiency
What We Offer At Yuno
Competitive Compensation
Remote Work — you can work from everywhere
Home Office Bonus — a one-time allowance to set up your ideal home office
Work Equipment
Stock Options
Health Plan wherever you are
Flexible Days Off
Language, Professional, and Personal Growth courses
Disclaimer
We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed or wish to exercise your data protection rights, please contact us at View email address on click.appcast.io.
$138.4k - $173k
...infrastructure as well as help improve the reliability, quality of services and overall... ...recovery. You’ll collaborate or embed with engineering teams, helping them to improve the reliability... ...about our locations by visiting our site.Compensation & BenefitsThe base salary that...SuggestedFull timeFlexible hours$143k - $191k
...mission critical capabilities to our customers. System Deployment Engineers work in complex environments with shared environmental... ...customersParticipate in customer demonstrations and exercisesWork with site reliability engineers to provide and refine requirements for tooling and...SuggestedFull timeTemporary workWork experience placementImmediate start- ...thousands of companies. Join us as we help people all over the world thrive at work.Location: Salt Lake City, UTAs a Senior Site Reliability Engineer, you will help define the future of reliability for our world-class employee recognition platform. You'll leverage...SuggestedFull timeShift work
$166k - $220k
...globally. We work with mission partners and operators to deploy reliable and robust capabilities on operationally-relevant fielding... ...scalable deployment solution must be reached. As a Senior Software Engineer, you will re-imagine the infrastructure pipeline required to convert...SuggestedFull timeWork experience placementImmediate start$35 - $45 per hour
DescriptionKforce has a client that is seeking a remote Site Reliability Engineer to join their team.Summary:The team consists of systems that can track lead management, job management and sales management. It is built on Salesforce but underpinned by a lot of Java/API'...SuggestedRemote work$170k - $220k
Who We're Looking ForWe’re looking for a hands-on, high-agency Site Reliability Engineer to help shape and scale the reliability layer of our stack. You'll own the release pipeline end-to-end — managing daily releases, weekly deploys, and hotfixes — while also automating...$152.5k - $205k
...flexible work environment where new ideas are encouraged and everyone is a stakeholder.What you’ll be responsible forThe Site Reliability Engineer builds and maintains shared platform capabilities, common libraries, and infrastructure that help Circle teams ship secure...Flexible hours$104.9k - $174.7k
...Data Management. You can learn more about LexisNexis Risk at the link below, About the Role:We are hiring a hands-on Senior Site Reliability Engineer (SRE) to actively build, operate, and improve the reliability of our production systems. This is not a purely advisory...Full timeWork at officeLocal areaRemote workWork from home$166k - $220k
...requirements and customer expectations. Our systems integration engineers internalize the nuances of each deployment, ensuring the... ...-to-end solutions we ship.ABOUT THE JOBWe are looking for a Site Reliability Engineer (SRE) to join AGD, our rapidly growing team in Irvine...Full timeWork experience placementImmediate start$86.6k - $144.4k
...platforms and using automation to solve complex security and reliability challenges?Do you enjoy shaping the future of security... ...You can learn more about LexisNexis Risk at our TeamOur Site Reliability Engineering (SRE) team plays a critical role in ensuring the...Full timeLocal area$125k - $145k
...SpaceX is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SITE RELIABILITY ENGINEER, GNCSpaceX’s mission is to make humanity multiplanetary by developing fully and rapidly reusable launch systems capable of...Permanent employmentTemporary workFlexible hoursWeekend work- Company DescriptionComtech LLC is a woman-owned small business focused on delivering end-to-end solutions and products. Since 1998, we have successfully serviced enterprises across the public and private sectors, and the Department of Defense. Our services span all aspects...
$81.1k - $187k
...architect infrastructure and service to ensure reliability and functionality. Forecasts demands and... ...impact and develops knowledge of site reliability trends.Only Oracle brings together... ...guidance and mentorship to junior engineers. Communicate status, risks, blockers, and...Temporary workFlexible hours- LeanData helps the world’s fastest-growing companies automate, simplify, and accelerate revenue.We are looking for a Senior Site Reliability Engineer to lead the strategic evolution of our cloud infrastructure. Reporting directly to the SVP of Engineering, this role is...Full timeWork at office2 days per week
$148k - $235.75k
...see how you can make a lasting impact on the world.Join our team of innovative engineers who are building an AI Data Center AIOps platform that turns raw, high-volume telemetry into reliable, job-centric insights and automation for GPU fleets. We’re hiring a DevOps Engineer...Full time- ...selected candidate for this role to work on site in the specified location(s).Schwab... ...their money by delivering innovative and reliable technology solutions that support investing... ...Within the Bank Platform Operations and Engineering organization, you will help ensure the...Full timeWork at office
$130k - $140k
...mid-market firms, rely on SS&C for expertise, scale, and technology.Job DescriptionSite Reliability EngineerLocation(s): Waltham, MA | HybridAbout the RoleSr Site Reliability Engineer- Guardian of the products to ensuring systems are reliable, scalable, and efficient...Ongoing contractFull timeTemporary workWork experience placement$98.58k - $138.02k
...Northern California / Silicon Valley Region / Denver, COProduct Engineering - DevOps /Full Time /HybridRestaurant365 is a SaaS company... ...office locations: Austin, TX; Irvine, CA; or Akron, OH. The Site Reliability Engineer II will be responsible for supporting, enhancing,...Full timeWork at office- ...communities.This is a Lead Software Production Management & Reliability Engineering position at Director level which is part of the job family responsible... ...across the business.Job SummaryWe are looking for a Site Reliability Engineer with a minimum of 5 years of industry...Flexible hoursWeekend work
- ...Georgia, and serves customers in more than 35 countries worldwide.Position OverviewWe are seeking a highly experienced Senior Site Reliability Engineer (Unified Observability) to lead the design, implementation, and operational maturity of the F1 Next Generation Customer...Full timeWorldwideFlexible hours
$145.7k - $218.5k
...synonymous with entertainment excellence and creativity.Service Reliability EngineerDo you want to use transformative technologies to... ...scalability and efficiency? Do you want a career that combines your engineering skills and your passion for video gaming? Are you fascinated...Work experience placementShift work$117k - $209.33k
Job Requisition ID #26WD99276Position OverviewWant to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable, secure, and scalable cloud services for Autodesk GovCloud products.As part of a new SRE team supporting...Full timeFor contractorsRemote work- Job ID: 18719802Reference Number: 23-00164Title: site reliability engineerLocation: Iselin, NJ, 08830Posted Date: 2023-01-20Contact: Shyam MaramContact Email: ****@*****.*** Phone: (***) ***-****Company: HAN Staffing Devops/SRE/Python Role Malvern PA - hybrid...
$112k - $179k
...delivery of system, network, software, and security solutions.About The RolePeraton is seeking a self-driven and resourceful Site Reliability Engineer to join our dynamic of Network and UC engineers in Washington, DC. This position combines software engineering and systems...Contract workWorldwideShift work- ...English (Required)Work Shift:1st Shift (United States of America)Please review the following job description:Lead Site Reliability & Environment Monitoring Engineer (Azure / Dynatrace / ServiceNow)We are seeking a Lead Site Reliability & Environment Monitoring Engineer to...Full timeTemporary workShift workDay shift
- ...candidates that are particularly strong in a few areas, and have some interest and capabilities in others.About the Role:As a Site Reliability Engineer, you’ll join the global Platform SRE team responsible for building, operating, and scaling Kong’s multi-region SaaS...Temporary work
- ...and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Commercial and Investment Bank, Healthcare Payments team, you will solve complex and broad...
- ...Evaluate applications, platforms, and vendors to assess resiliency, reliability, and operational risk.Design and implement processes that... ...and reliability tooling.Actively participate in reliability engineering and resilience communities of practice, contributing to...Full time
$243.29k - $295.25k
...challenges at scale, and helping to create safer, more civil shared experiences for everyone.The Infrastructure Compute Site Reliability Engineering mission is to own and manage the successful operation of our underlying cell infrastructure system, along with elements...Full timeWork experience placementH1bWork at officeLocal areaVisa sponsorshipMonday to Friday$125k - $145k
...including 90% of the Fortune 500, choose DigiCert to stop today’s threats and prepare for a quantum-safe future at Job summaryThe Site Reliability Engineer (SRE) collaborates with development teams to embed reliability, scalability, and performance best practices throughout the...Flexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
- site reliability engineer remote United States
- lead site reliability engineer United States
- site reliability engineer United States
- site reliability engineer sre United States
- site reliability engineering manager United States
- junior website developer United States
- website content developer United States
- on site coordinator United States
- after school site coordinator United States
- website coordinator United States
