Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior CloudOps & Site Reliability Engineer

Full-time

MangoApps

MangoApps runs an enterprise SaaS platform that thousands of customers depend on every day. We're hiring a senior, hands-on engineer to own the reliability, availability, security, and performance of that platform in production, primarily on AWS.

This is a deep individual-contributor role, not a management track. You'll spend your time in the systems: tuning infrastructure, building observability, automating away toil, and leading the technical response when production is on the line. You'll influence how the rest of engineering builds and operates reliable services through your work and your judgment, not through a reporting line.

If you get satisfaction from understanding a production system end to end, finding the real root cause instead of the convenient one, and making the next incident less likely, this role is built for you.

Who you are:

  • 5+ years of hands-on Cloud Operations and Site Reliability Engineering, operating production-scale SaaS (not pre-production or internal-only systems).
  • You operate AWS at production scale today and can speak in specifics about AWS compute, networking, IAM, EKS/Kubernetes, and the operational realities of running real workloads there. This is a hard requirement.
  • A second cloud (Google Cloud or Azure) is a plus, not a substitute. We value it, but AWS depth is what the role is focused on.
  • You debug Linux at the level of "why is this latency spike happening," not just "restart the service."
  • You reach for automation by reflex. Manual operational work bothers you, and you've built the tooling to remove it.

What will you own:

  • Reliability & incident response - Keep the platform available and performant. Define and continuously sharpen monitoring, alerting, and observability. Lead production troubleshooting, drive root-cause analysis, and run post-incident reviews that actually change the system afterward. Participate in the on-call rotation for the services you own.
  • AWS infrastructure & operations - Design, deploy, and optimize our AWS infrastructure - compute, storage, networking, DNS, load balancing, and security services. Drive architectural improvements to enhance reliability, scalability, performance, and cost. Own disaster recovery and business-continuity processes, and prove they work before you need them.
  • Containers & orchestration - Build and operate containerized workloads on Docker with a focus on security, performance, and predictable scaling across environments.
  • Automation & Infrastructure as Code Provision - Build the scripts and tooling that make operations boring and repeatable.
  • CI/CD & release engineering - Maintain CI/CD pipelines and deployment automation. Partner with engineering to make releases safer, faster, and easier to roll back.
  • Security & compliance - Apply cloud security practices across IAM, network security, secrets management, and vulnerability remediation. Keep infrastructure aligned to our internal security standards and compliance obligations.

Must have

  • Production-scale experience operating cloud infrastructure on AWS.
  • Deep Linux systems administration, troubleshooting, and performance tuning.
  • Hands-on Docker in production
  • Solid networking fundamentals: VPCs, routing, load balancing, DNS, VPNs, and security controls.
  • Monitoring and observability with tools such as Prometheus, Grafana, ELK/OpenSearch, Datadog, or equivalents.
  • Scripting and automation in Bash, Python, or similar.
  • Git-based workflows and CI/CD pipelines; config management with Ansible or Puppet.
  • Strong incident management and root-cause analysis instincts, with a bias toward fixing the system, not the symptom.

Nice to have

  • Production experience on a second cloud (Google Cloud or Azure).
  • IAM / SSO experience (SAML, OAuth, Okta, or similar).
  • Multi-region or multi-cloud operations at scale.
  • Background in cloud security, compliance, and governance practices.

What success looks like

  • Sustained, high platform uptime against clear SLOs.
  • Faster incident detection and resolution, with recurring failure classes systematically driven down.
  • More automation and meaningfully less manual operational toil quarter over quarter.
  • Observability is good because it lets the team see problems before customers do.
  • Engineering teams ship reliably because the operational foundation is solid.

What We're Looking For in You

  • Ownership: You take accountability for outcomes, not just tasks.
  • Problem solver: You enjoy diagnosing and resolving complex infrastructure and production challenges.
  • Continuous learner: You stay current with evolving cloud, automation, and reliability practices.
  • Collaborative: You work effectively across teams and communicate clearly, both in routine operations and in the middle of a critical incident.
  • Customer-focused: You understand that infrastructure reliability directly shapes customer experience and business success.

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Senior CloudOps & Site Reliability Engineer in Washington DC vacancy
  • $210k - $230k

     ...GovCIO is currently hiring for a Senior Site Reliability Engineer (SRE) to design, implement, and maintain highly available, scalable, and resilient infrastructure systems. The ideal candidate will bridge the gap between development and operations, focusing on automation... 
    Senior
    Currently hiring
    Remote work

    Govcio

    Arlington, VA
    1 day ago
  • $166k - $220k

     ...requirements and customer expectations. Our systems integration engineers internalize the nuances of each deployment, ensuring the...  ...-to-end solutions we ship.ABOUT THE JOBWe are looking for a Site Reliability Engineer (SRE) to join AGD, our rapidly growing team in Irvine... 
    Senior
    Full time
    Work experience placement
    Immediate start

    Anduril Industries

    Washington DC
    3 days ago
  • $175k - $250k

     ...Senior Cloud Infrastructure Engineer Location: San Francisco, CA. Remote unavailable. Modality: On‑Site only. Must live within commuting distance of San Francisco or be willing to...  ...ensuring scalability, performance, and reliability across environments. What You’ll Do Design... 
    Senior
    Full time
    Remote work
    Relocation
    Relocation package

    The Recruiting Guy

    Washington DC
    3 days ago
  • $207k - $284.9k

     ...We're all in on this mission. If you are too, let's talk.Senior Manager, Site Reliability EngineeringSecure Every Identity, from AI to...  ...mission. If you are too, let's talk.The Federal Operations Engineering GroupOkta's Federal Operations team supports government... 
    Senior
    Permanent employment
    Local area
    Worldwide
    Flexible hours
    Day shift

    Okta

    Washington DC
    1 day ago
  •  ..., Oracle, Cassandra, SQL Server, My SQL and Mongo DB Seniority level ~ Seniority level Mid-Senior level Employment...  ...job is posted. Sign in to set job alerts for “Senior Site Reliability Engineer” roles. Bellevue, WA $204,000.00-$259,000.00 1 day ago... 
    Senior
    Contract work
    Remote work

    Signature IT World Inc

    Washington DC
    3 days ago
  • $149.4k - $202k

     ...Senior Software Engineer- Site Reliability Engineering (SRE) DC, MD, VA, CA The Site Reliability Engineering discipline at Noctua Technology, LLC is a strategic force driving digital transformation. We treat operations as a software engineering challenge, focusing... 
    Senior
    Remote work

    Noctua Technology

    Washington DC
    5 days ago
  • $141.8k - $195k

     ...best work, grow fast, and bring their full selves to the herd. Why You'll Love This Role Cribl Inc is seeking a Senior Site Reliability Engineer to join our mission where you will unlock the value of all observability data, as we expand our team in the U.S. Cribl... 
    Senior
    Temporary work
    Remote work

    Cribl

    Washington DC
    5 days ago
  •  ...technology solutions using a tailored Agile methodology. We are seeking a highly motivated and intellectually curious Senior Site Reliability Engineer to join our team working with a Federal client. The position will be a remote role open to US citizens residing in the... 
    Senior
    Remote work

    VALID8 Financial

    Washington DC
    2 days ago
  • Role Summary The Senior Site Reliability Engineer (SRE) is a hands-on role responsible for the availability, performance, and end-to-end observability of QSR digital platforms across Mobile (iOS/Android), Web, and POS systems. This role is part of the Observability... 
    Senior
    Flexible hours

    Donato Technologies, Inc

    Washington DC
    2 days ago
  • $165k - $270k

     ...is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SR. SITE RELIABILITY ENGINEER (STARSHIELD)Starshield leverages SpaceX’s Starlink technology and launch capability to support national security efforts.... 
    Senior
    Permanent employment
    Temporary work
    Immediate start
    Weekend work

    SpaceX

    Washington DC
    1 day ago
  •  ...Washington D.C., District of Columbia, United States About the job Sr. Site Reliability Engineer Our Client is currently hiring a full-time Sr. Site Reliability Engineer (SRE), who will play a vital role in continuously driving improvements in observability, performance... 
    Senior
    Full time
    Currently hiring
    3 days per week

    CruitZi

    Washington DC
    4 days ago
  • $128.5k - $190k

     ...Bring your whole self. The Role and Team The Site Reliability Engineering organization at Medallia brings together the infrastructure...  ...that power a highly reliable global SaaS platform. As a Senior Site Reliability Engineer, you will play a key role in designing... 
    Senior
    Temporary work
    Work experience placement
    Local area

    Medallia

    McLean, VA
    2 days ago
  •  ...Job Description Job Description Role Overview We are seeking a high-caliber Site Reliability Engineer (SRE) to join our Forward Engineering team. You will be the guardian of our production ecosystems, ensuring that our complex, data-driven AI platforms remain resilient... 
    Senior
    Local area

    Tiger Analytics Inc.

    Washington DC
    a month ago
  •  ...Job Description Job Description Description: Onsite in Washington, DC   our client seeks a Sr. Site Reliability Engineer III to design, automate, and operate mission-critical systems for federal environments. The role focuses on Kubernetes or VMWare platforms,... 
    Senior
    Hourly pay
    Permanent employment
    Full time
    Local area
    Immediate start

    Eliassen Group

    Washington DC
    a month ago
  •  ...candidates that are particularly strong in a few areas, and have some interest and capabilities in others.About the Role:As a Site Reliability Engineer, you’ll join the global Platform SRE team responsible for building, operating, and scaling Kong’s multi-region SaaS... 

    Kong

    Washington DC
    5 days ago
  • $125k - $185k

    Washington, D.C.Engineering /Full-time /HybridA World-Changing CompanyPalantir builds the world’s leading software for data-driven decisions...  ...locate missing children, and more.The RoleWe’re looking for Site Reliability Engineers who can help us build, operate, and maintain high-... 
    Full time
    Work experience placement
    Work at office
    Remote work
    Work from home
    Relocation package

    Palantir Technologies

    Washington DC
    1 day ago
  • $112k - $179k

     ...delivery of system, network, software, and security solutions.About The RolePeraton is seeking a self-driven and resourceful Site Reliability Engineer to join our dynamic of Network and UC engineers in Washington, DC. This position combines software engineering and systems... 
    Contract work
    Worldwide
    Shift work

    Peraton Corporation

    Washington DC
    4 days ago
  • $135k - $154k

     ...where you matter.Your ImpactAs a contributor in the APX platform engineering organization on the CloudNet team, you are passionate about...  .... You are also obsessed about achieving the high quality and reliability our customers demand. You will work closely with sovereign... 
    Work experience placement
    Work at office
    Remote work

    Axon

    Washington DC
    1 day ago
  • $230k - $250k

     ...GovCIO is hiring a Site Reliability Engineer with an active Secret clearance to ensure reliability, scalability, performance, and availability of mission-critical systems by combining software engineering practices with infrastructure operations expertise. This role is... 
    Remote work

    Govcio

    Arlington, VA
    4 days ago
  • $125k - $185k

     ...lifesaving drugs, forecast supply chain disruptions, locate missing children, and more.The RoleWe’re looking for Forward Deployed Site Reliability Engineers who can help us build, operate, and maintain high-performance, scalable, and reliable services for our production... 
    Full time
    Work experience placement
    Work at office
    Remote work
    Work from home
    Relocation package

    Palantir Technologies

    Washington DC
    4 days ago
  •  ...Lead Site Reliability Engineer Defense Tech / National Security US Defense Tech Startup The Company Early-stage defense technology leader building modern software for air-gapped, high-side, and accredited environments. Backed by a $99M defense contract to... 
    Full time
    Contract work

    Attis

    Washington DC
    3 days ago
  •  ...Senior Site Reliability Engineer (SRE) Location: Seattle, hybrid - 2 times a week in the office Job Type: Full-time, direct hire Industry: High-Growth Technology / SaaS About the Role We are seeking a highly skilled Senior Site Reliability Engineer... 
    Full time
    Work at office

    TalentDome Staffing

    Washington DC
    3 days ago
  • $150k - $180k

     ...redefining what’s possible in remote sensing, you belong here at Umbra. About the Job We are seeking an experienced Senior Site Reliability Engineer to help design, build, operate, and scale the mission- and business-critical infrastructure that powers Umbra's... 
    Senior
    Permanent employment
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    Umbra

    Arlington, VA
    5 days ago
  • $174k - $238k

     ...work. We're all in on this mission. If you are too, let's talk.The Federal SRE TeamWe are looking for an experienced Staff Site Reliability Engineer to join Okta's Federal SRE team for the Emerging Products Group (EPG). Our mission is to build highly reliable, scalable,... 
    Local area
    Worldwide
    Flexible hours

    Okta

    Washington DC
    2 days ago
  • $182k - $250.8k

     ...the backbone of our platform's reliability and operational excellence....  ...a forward-thinking group of engineers and leaders who believe that...  ...users worldwide. As a Manager, Site Reliability Engineer, you'll...  ...learningRepresent reliability as a senior technical leader in... 
    Permanent employment
    Local area
    Remote work
    Worldwide
    Flexible hours
    Weekend work
    Weekday work

    Okta

    Washington DC
    5 days ago
  • $174k - $239k

     ...From core infrastructure to enterprise platforms, we partner across functions to drive scale, reliability, and innovation through technology.The Staff Site Reliability Engineer OpportunityOkta Federal, Inc. is looking for an experienced Staff TDI Site Reliability... 
    Work experience placement
    Local area
    Worldwide
    Flexible hours

    Okta

    Washington DC
    4 days ago
  •  ...ears, and hands on the ground at a government customer site, ensuring the reliability and performance of Twenty's mission-critical platform running...  ...of deep technical ownership and customer-facing engineering: you'll define how we measure reliability, lead incident... 
    Full time
    Contract work
    Remote work
    Flexible hours

    Twenty Inc.

    Arlington, VA
    1 day ago
  •  ...Site Reliability Engineer III (AI Platform) Location: Mount Laurel, NJ (Onsite) Duration: Contract Experience: 4+ years About the Role We are seeking a Site Reliability Engineer (SRE) III to support a cutting-edge AI Platform Engineering team responsible... 
    Contract work

    GCS Recruitment

    Laurel, MD
    2 days ago
  • $170k - $230k

     ...of clearance.*** Are you a Senior Full Stack Software Engineerwho...  ...the mothership again? Our engineers were certainly tired of the...  ...excel at delivering stable and reliable software solutions using...  ...Python/Flask. AWS Certified CloudOps Engineer, or general experience... 
    Senior
    Full time
    Contract work
    For contractors
    For subcontractor
    Work at office
    Remote work
    Work from home
    Relocation package

    GliaCell Technologies LLC

    Laurel, MD
    4 days ago
  •  ...Washington, District of Columbia, United States Contractor | On-site Job Description We are seeking an experienced Site Reliability Engineer (SRE) to help build and maintain highly reliable, scalable, and secure technology platforms. The SRE will combine software... 
    For contractors

    Mybridge

    Washington DC
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior CloudOps & Site Reliability Engineer. Be the first to apply!