Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior CloudOps & Site Reliability Engineer

Full-time

MangoApps

MangoApps runs an enterprise SaaS platform that thousands of customers depend on every day. We're hiring a senior, hands-on engineer to own the reliability, availability, security, and performance of that platform in production, primarily on AWS.

This is a deep individual-contributor role, not a management track. You'll spend your time in the systems: tuning infrastructure, building observability, automating away toil, and leading the technical response when production is on the line. You'll influence how the rest of engineering builds and operates reliable services through your work and your judgment, not through a reporting line.

If you get satisfaction from understanding a production system end to end, finding the real root cause instead of the convenient one, and making the next incident less likely, this role is built for you.

Who you are:

  • 5+ years of hands-on Cloud Operations and Site Reliability Engineering, operating production-scale SaaS (not pre-production or internal-only systems).
  • You operate AWS at production scale today and can speak in specifics about AWS compute, networking, IAM, EKS/Kubernetes, and the operational realities of running real workloads there. This is a hard requirement.
  • A second cloud (Google Cloud or Azure) is a plus, not a substitute. We value it, but AWS depth is what the role is focused on.
  • You debug Linux at the level of "why is this latency spike happening," not just "restart the service."
  • You reach for automation by reflex. Manual operational work bothers you, and you've built the tooling to remove it.

What will you own:

  • Reliability & incident response - Keep the platform available and performant. Define and continuously sharpen monitoring, alerting, and observability. Lead production troubleshooting, drive root-cause analysis, and run post-incident reviews that actually change the system afterward. Participate in the on-call rotation for the services you own.
  • AWS infrastructure & operations - Design, deploy, and optimize our AWS infrastructure - compute, storage, networking, DNS, load balancing, and security services. Drive architectural improvements to enhance reliability, scalability, performance, and cost. Own disaster recovery and business-continuity processes, and prove they work before you need them.
  • Containers & orchestration - Build and operate containerized workloads on Docker with a focus on security, performance, and predictable scaling across environments.
  • Automation & Infrastructure as Code Provision - Build the scripts and tooling that make operations boring and repeatable.
  • CI/CD & release engineering - Maintain CI/CD pipelines and deployment automation. Partner with engineering to make releases safer, faster, and easier to roll back.
  • Security & compliance - Apply cloud security practices across IAM, network security, secrets management, and vulnerability remediation. Keep infrastructure aligned to our internal security standards and compliance obligations.

Must have

  • Production-scale experience operating cloud infrastructure on AWS.
  • Deep Linux systems administration, troubleshooting, and performance tuning.
  • Hands-on Docker in production
  • Solid networking fundamentals: VPCs, routing, load balancing, DNS, VPNs, and security controls.
  • Monitoring and observability with tools such as Prometheus, Grafana, ELK/OpenSearch, Datadog, or equivalents.
  • Scripting and automation in Bash, Python, or similar.
  • Git-based workflows and CI/CD pipelines; config management with Ansible or Puppet.
  • Strong incident management and root-cause analysis instincts, with a bias toward fixing the system, not the symptom.

Nice to have

  • Production experience on a second cloud (Google Cloud or Azure).
  • IAM / SSO experience (SAML, OAuth, Okta, or similar).
  • Multi-region or multi-cloud operations at scale.
  • Background in cloud security, compliance, and governance practices.

What success looks like

  • Sustained, high platform uptime against clear SLOs.
  • Faster incident detection and resolution, with recurring failure classes systematically driven down.
  • More automation and meaningfully less manual operational toil quarter over quarter.
  • Observability is good because it lets the team see problems before customers do.
  • Engineering teams ship reliably because the operational foundation is solid.

What We're Looking For in You

  • Ownership: You take accountability for outcomes, not just tasks.
  • Problem solver: You enjoy diagnosing and resolving complex infrastructure and production challenges.
  • Continuous learner: You stay current with evolving cloud, automation, and reliability practices.
  • Collaborative: You work effectively across teams and communicate clearly, both in routine operations and in the middle of a critical incident.
  • Customer-focused: You understand that infrastructure reliability directly shapes customer experience and business success.

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Senior CloudOps & Site Reliability Engineer in Seattle, WA vacancy
  • $127k - $249k

    Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions...  ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper). As... 
    Senior
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    Seattle, WA
    1 day ago
  • $134.25k - $214.8k

     ...matters at a company where you matter.Your ImpactAre you an engineer who gets excited about the challenge of making complex distributed...  ...it.You will be part of the Observability team within Axon's Site Reliability organization — a focused team responsible for Axon's metrics,... 
    Senior
    Work experience placement
    Work at office
    Remote work

    Axon

    Seattle, WA
    10 hours ago
  • $55k - $151.47k

     ...ApplicableSpecialismIFS - Internal Firm Services - OtherManagement LevelSenior AssociateJob Description & SummaryThe OpportunityAs a Site Reliability Engineer - Senior Associate, you will play a pivotal role in enhancing the reliability, scalability, and performance of our... 
    Senior
    Full time
    H1b

    PwC

    Seattle, WA
    3 days ago
  •  ...team is responsible for the reliability, scalability, and efficiency...  ...building features, but about engineering the resilience and performance...  ...maintain system stability.As a Site Reliability Engineer, you...  ...practices while working alongside senior engineers to solve... 
    Senior

    TikTok

    Seattle, WA
    1 day ago
  • $232k - $319k

     ...to help us continue to scale the service with great people and reliable, cost-effective, and efficient infrastructure, processes, and...  ...enabled with self-serviceAccelerate the velocity of SRE and product engineering by developing robust platforms, powerful tooling, and... 
    Senior
    Permanent employment
    Local area
    Worldwide
    Flexible hours

    Okta

    Bellevue, WA
    2 days ago
  • $160k - $180k

     ...Socure is seeking a Site Reliability Engineer to own AWS infrastructure and improve Kubernetes platforms. Ideal candidates will have deep AWS expertise, strong Kubernetes fundamentals, and the ability to write production-quality code in Go or Python. The position emphasizes... 
    Senior

    Socure Inc

    Seattle, WA
    3 days ago
  • $134.25k - $214.8k

     ...real change. Constantly grow as you work hard for a mission that matters at a company where you matter.Your ImpactAs a Senior Site Reliability Engineer within the APX SRE organization, you’ll focus on delivering practical, scalable solutions to support the reliability... 
    Senior
    Work experience placement
    Work at office
    Remote work
    Flexible hours

    Axon

    Seattle, WA
    4 days ago
  •  ...This is an engineering-first Senior SRE role. We’re looking for senior engineers who have: Built and shipped significant backend systems...  ...end-to-end in production (design → launch → on-call → reliability improvements) Led incident response and driven durable follow... 
    Senior

    Practice by Numbers

    Bellevue, WA
    3 days ago
  •  ...certification), ISO 27001:2005 Information Security Management System (ISMS), and CMMI-DEV Level 3. Job Description Sr. Site Reliability Engineer Location – Seattle, WA Duration – 12 months Interview – in-person if local or Phone + Skype Minimum... 
    Senior
    Local area
    Worldwide

    Comtech LLC

    Seattle, WA
    3 days ago
  • $151.2k - $204.6k

    Would you like to be an engineer who builds the systems that power advertising...  ...every day, where latency, reliability, and quality translate directly...  ...customer trust.We are looking for a Senior System Development Engineer, operating as a Site Reliability Engineer, to raise... 
    Flexible hours

    Amazon

    Seattle, WA
    2 days ago
  • $143k - $194k

     ...mission critical capabilities to our customers. System Deployment Engineers work in complex environments with shared environmental...  ...customersParticipate in customer demonstrations and exercisesWork with site reliability engineers to provide and refine requirements for tooling and... 
    Full time
    Temporary work
    Work experience placement
    Immediate start

    Anduril Industries

    Seattle, WA
    4 days ago
  • Company DescriptionComtech LLC is a woman-owned small business focused on delivering end-to-end solutions and products. Since 1998, we have successfully serviced enterprises across the public and private sectors, and the Department of Defense. Our services span all aspects...

    Comtech

    Seattle, WA
    1 day ago
  •  ...Security team. The successful candidate will apply an engineering-based approach to solving complex security and reliability challenges, leveraging machine data analytics...  ...and product development. You will partner with senior technical leaders and cross-functional teams to... 
    Full time
    Internship
    Summer internship
    Work at office
    Local area
    Remote work
    Flexible hours

    F5 Networks

    Seattle, WA
    4 days ago
  •  ...Senior Site Reliability Engineer (SRE) Location: Seattle, hybrid - 2 times a week in the office Job Type: Full-time, direct hire Industry: High-Growth Technology / SaaS About the Role We are seeking a highly skilled Senior Site Reliability Engineer... 
    Full time
    Work at office

    TalentDome Staffing

    Seattle, WA
    3 days ago
  • $120k - $170k

    Sr. Manager/Manager Site Reliability Engineering Join to apply for the Sr. Manager/Manager Site Reliability Engineering role at Aritzia Sr. Manager...  ...from the office or from a remote space of your choosing. Seniority level Seniority level Mid-Senior level Employment type... 
    Senior
    Full time
    Work at office
    Remote work
    Flexible hours

    Aritzia

    Seattle, WA
    4 days ago
  • $152k - $241.5k

     ...the world.Join the Simulation Software team at NVIDIA as a Senior System Software Engineer! This role offers an outstanding opportunity to work on...  ...development by enabling Chips Simulation as a trusted and reliable virtual platform.What you will be doing:Drive early... 
    Senior
    Full time

    Nvidia

    Seattle, WA
    2 days ago
  • $194k - $267k

     ...do something more than once, automate it” and who can rapidly self-educate on new concepts and tools.Position Overview:The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta

    Bellevue, WA
    10 hours ago
  • $194k - $267k

     ...all in on this mission. If you are too, let's talk.Position Overview:We are seeking a highly technical StaffObservabilitySite Reliability Engineer with a specialty in Splunk to own and evolve our Splunk ecosystem. In this role, you will move beyond simple monitoring to... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta

    Bellevue, WA
    4 days ago
  • $204k - $306k

     ...all in on this mission. If you are too, let's talk.Manager, Site Reliability EngineeringSan Francisco, CaliforniaSecure Every Identity, from...  ...in our San Francisco Office. The IDaaS Site Reliability Engineering GroupOkta authenticates, authorizes and provisions millions... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours
    2 days per week

    Okta

    Bellevue, WA
    4 days ago
  •  ...where you can push the limits of what's possible.As a Lead Software Engineer at JPMorganChase within the Enterprise Technology,...  .... These benefits include comprehensive health care coverage, on-site health and wellness centers, a retirement savings plan, backup childcare... 

    JP Morgan Chase

    Seattle, WA
    2 days ago
  •  ...Engineering, Product, Design, and Marketing Engineering Compensation ~ Zone 1 Base Pay: $214K – $260K Superhuman offers...  ...role will be responsible for building software to ensure the reliability of our back-end systems, working with engineers who develop them... 
    Worldwide
    Home office
    Flexible hours

    Superhuman

    Seattle, WA
    4 days ago
  •  ...We're seeking an SRE to ensure the reliability and performance of our clients' critical systems. You'll work on observability, incident...  ...management Nice to have Experience with chaos engineering Knowledge of distributed systems Background in high-scale... 
    Remote work
    Flexible hours

    ACI Infotech

    Seattle, WA
    2 days ago
  •  ...future together. We are responsible for the reliability of all the company's major data warehouse products, services, and query engines. We serve business needs across domains...  ...practices, and emerging technologies related to site reliability and infrastructure engineering.... 

    TikTok

    Seattle, WA
    1 day ago
  •  ...A Site Reliability Engineer (SRE) is responsible for ensuring the reliability, scalability, and performance of an organization's software systems and cloud infrastructure. The role combines software engineering with IT operations to automate processes, monitor system health... 

    Mybridge

    Seattle, WA
    3 days ago
  •  ...coordinate support and resolve platform issues across CPU/radio SoCs, MCU/PIC, NPU/GPU, and peripheral devices.Support hardware engineering teams with deep technical debugging and contribute to OS/platform modernization efforts.What You’ll NeedBasic Qualifications:Bachelor... 
    Senior
    Full time
    Temporary work
    Work at office
    Immediate start
    Work visa

    Sonos

    Seattle, WA
    10 hours ago
  • $95k - $134k

     ...helped build. For more information, visit Job Application Deadline: 10/31/2026 The Opportunity DAT is looking for a Site Reliability Engineer to join our SRE platform team. This position will work hybrid in Seattle, WA or Portland, OR Candidate profile DAT... 
    Temporary work
    For contractors
    Work experience placement
    Work at office
    Local area
    Immediate start
    Flexible hours

    DAT Freight Solutions

    Seattle, WA
    3 days ago
  • $94k - $142.3k

     ...on Salesforce, customer and partner enablement, applications engineering, infrastructure, collaboration, enterprise operations,...  ...Salesforce products delivered globally, at scale, sustainably.As a Site Reliability Operations Engineer you'll be part of our internal DET Site... 
    Full time
    Shift work

    Salesforce

    Seattle, WA
    10 hours ago
  • $145k - $160k

     ...We are seeking a specialized Observability & Infrastructure Engineer to lead the monitoring, telemetry, and platform-as-code initiatives critical to our multi-region disaster recovery roadmap. You will architect and implement robust observability pipelines, ensure deep... 
    Temporary work
    Remote work
    Flexible hours

    EPAM Systems Inc

    Seattle, WA
    2 days ago
  •  ...you can push the limits of what's possible. As a Lead Software Engineer at JPMorganChase within the Enterprise Technology,...  .... These benefits include comprehensive health care coverage, on-site health and wellness centers, a retirement savings plan, backup childcare... 

    Fairygodboss

    Seattle, WA
    3 days ago
  • $264k - $379.5k

     ...the ecosystem that powers our AI-pilled engineers and our agentic organizations and reinvents...  ...a bias for measurable impact through reliability, performance, and ease of use.Bonus points...  ...job posting on the Snowflake Careers Site for salary and benefits information: careers... 
    Senior

    Snowflake

    Bellevue, WA
    10 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior CloudOps & Site Reliability Engineer. Be the first to apply!