Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Site Reliability Engineer

HavocAI

About Us:

Collaborative autonomy is how self-tasking teams of machines will solve hard human problems, and HavocAI is an unquestioned leader in collaborative autonomy. We set the standard for autonomous surface vessels for a wide range of defense and commercial maritime missions. Success requires us to grow quickly, and we’re looking for teammates who are passionate about solving hard problems, about pushing the envelope, and about preventing conflict and saving lives. Ambition is welcome to apply within.

About the Role

We are seeking a Senior Site Reliability Engineer (SRE) with 7+ years of experience designing, operating, and scaling highly reliable distributed systems. In this role, you will be a key technical leader within the Cloud Platform team, responsible for ensuring the availability, performance, and resilience of mission-critical services supporting autonomy, simulation, and data-intensive workloads.

You will work closely with Cloud Platform, DevOps, Data Engineering, and Autonomy teams to establish reliability standards, improve operational maturity, and build systems that scale safely under real-world conditions. The ideal candidate is deeply technical, calm under pressure, and experienced in owning reliability outcomes end-to-end.

Responsibilities
Reliability Engineering & Architecture
  • Design and evolve reliability architecture for distributed and cloud-hosted systems.

  • Define and implement SRE best practices, including SLIs, SLOs, error budgets, and capacity planning.

  • Partner with platform and application teams to design systems for reliability, scalability, and operability.

  • Identify and mitigate systemic reliability risks across infrastructure and services.

Operations & Incident Management
  • Lead incident response processes including on-call rotations, escalation, and post-incident reviews.

  • Conduct root cause analysis for complex production incidents and drive long-term improvements.

  • Improve operational readiness through runbooks, automation, and resilience testing.

  • Reduce operational toil through tooling, automation, and process improvements.

Observability & Performance
  • Design and maintain observability systems for metrics, logging, tracing, and alerting.

  • Ensure services and data pipelines are observable, debuggable, and performant in production.

  • Drive performance analysis and tuning across infrastructure and service layers.

Automation & Platform Collaboration
  • Build automation to improve system reliability, deployment safety, and recovery processes.

  • Partner with DevOps and Cloud Platform teams on CI/CD reliability, rollout strategies, and safe deployment patterns.

  • Support and improve Kubernetes-based environments and containerized workloads.

Security & Resilience
  • Collaborate with security teams to ensure secure and resilient system design.

  • Participate in disaster recovery planning and testing.

  • Maintain strong operational practices around access control, secrets management, and change management.

Requirements
  • 7+ years of experience in SRE, infrastructure, or systems engineering roles.

  • Strong experience operating large-scale distributed production systems.

  • Deep understanding of Linux systems, networking, and distributed systems fundamentals.

  • Hands-on experience with Kubernetes and container orchestration.

  • Programming or scripting experience in Go, Python, or similar languages.

  • Experience designing and operating observability systems for production environments.

  • Proven ability to lead incident response and reliability improvements.

  • Strong communication skills and ability to collaborate across engineering teams.

  • Must be a US Citizen.

  • Must be Eligible to obtain a Government Clearance - if required.

Nice to Have
  • Experience supporting autonomy, robotics, simulation, or real-time systems.

  • Familiarity with AWS and large-scale cloud infrastructure.

  • Experience with chaos engineering, fault injection, or resilience testing.

  • Knowledge of CI/CD systems and progressive delivery practices.

  • Experience working in high-reliability or safety-critical environments.

Benefits:
  • 100% Employer paid Health, Dental and Vision Insurance for you and your families

  • Life Insurance (Employer Paid)

  • Ability to participate in the companies 401k program (Matching)

  • Unlimited PTO policy with an enforced 2 week minimum

  • Equity Package

  • Work / Home Office Stipend

  • Global Entry

  • 16 Week Paid Parental Leave

  • Monthly Health and Wellness Stipend


Our Values:
  • Innovation: We are driven to break new ground. Every day presents an opportunity to challenge the status quo, think boldly, and deliver advanced solutions that transform the future of defense technology.

  • Integrity: We hold ourselves to the highest ethical standards, ensuring transparency, accountability, and trust in all our actions and partnerships.

  • Mission-Driven: We are focused on achieving impactful outcomes that align with our core mission—protecting lives through innovation.

  • Forward-Leaning: We continuously seek out new opportunities and remain at the forefront of technological advancements. We embrace change and anticipate the challenges of tomorrow with confidence and creativity.

  • Ownership of All Tasks: At HavocAI, no problem is too complex or too trivial. We believe that greatness comes from tackling the hardest challenges, but also in handling the smallest, sometimes thankless, tasks with the same level of commitment and care.

  • Servant Leadership: We lead by serving others, whether it’s supporting our employees, partners, or the broader community. Empowering those around us is key to achieving long-term success and making a lasting impact.

HavocAI is an Equal Opportunity Employer and is committed to creating an inclusive and diverse workplace. We welcome applicants from all backgrounds and do not discriminate based on race, color, religion, gender, sexual orientation, age, national origin, disability, veteran status, or any other legally protected status.

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Senior Site Reliability Engineer in United States vacancy
  •  ...Sight Machine is seeking a senior Cloud Infrastructure IC to lead reliability, automation, and scale across our platform. You will drive IaC, CI/CD, observability and operate agentic AI systems, mentoring engineers and guiding architectural decisions while staying hands... 
    Senior

    Jobless

    Ann Arbor, MI
    2 days ago
  •  ...Senior Site Reliability Engineer Company: CyberArk Work Type: Remote Employment: Full Time Location: US Seniority: Mid Level Technologies: AWS, Kubernetes, Terraform, CloudFormation, Ansible, CloudWatch, Grafana, Datadog, OpenSearch, PagerDuty Requirements: Senior SRE... 
    Senior
    Full time
    Remote work

    CyberArk

    United States
    3 days ago
  •  ...Apple Service Engineering (ASE) seeks a senior SRE software engineer to own the architectural direction of Kubernetes internals powering Apple services...  ...will define controllers and namespace management, raise reliability, and contribute to upstream Kubernetes. The role... 
    Senior

    Socket

    San Francisco, CA
    3 days ago
  •  ...data and AI infrastructure provider is seeking a Senior Staff Technical Program Manager for Reliability to enhance the reliability and performance of their...  ...involves leading programs in partnership with senior engineering leaders, requiring over 10 years of experience in... 
    Senior

    Menlo Ventures

    Bellevue, WA
    1 day ago
  •  ...Lambda Inc. in San Francisco is seeking a Storage Engineer to own the reliability, performance, and capacity of our production storage fleet across multiple data centers, using a software-defined data plane. You will build monitoring, dashboards, and alerting for storage... 
    Senior

    Lambda

    San Francisco, CA
    2 days ago
  •  ...Oracle Cloud Infrastructure (OCI) seeks a Senior Principal Engineer to lead the design and implementation of reliability validation for OCI control plane services, focusing on a high-performance, low-level systems approach. You will mentor engineers, define validation... 
    Senior

    Oracle

    Santa Clara, CA
    1 day ago
  •  ...Senior Site Reliability Engineer Company: Sphera Work Type: Remote Employment: Full Time Location: US Seniority: Senior Level Technologies: Terraform, ARM templates, Kubernetes, Azure, SonarCloud, CheckPoint, Hadoop, Kafka, Presto, NewRelic, CI/CD, Linux, Windows, Redis... 
    Senior
    Full time
    Remote work

    Sphera

    United States
    3 days ago
  • $166k - $244k

    # Senior Software Engineer, Site Reliability EngineeringGoogle • onsite • 601 N 34th St, Seattle, WA 98103, USA • full\_timePay: USD 166000.00 - USD 244000.00 / unspecifiedBusinesses of all shapes and sizes rely on Google’s unparalleled advertising solutions to help them... 
    Senior
    Temporary work

    Epic Games

    Seattle, WA
    3 days ago
  •  ...infra has to match. The role We're looking for a Senior SRE to own the reliability, scalability, and operational posture of Satsuma's multi...  ...-assisted development workflows Partner closely with engineering on reliability reviews and architecture decisions... 
    Senior

    Satsuma

    United States
    3 days ago
  • $182.8k - $247.3k

     ...mission to develop education for our half a billion (and growing!) learners around the world. About the role... As a Senior Site Reliability Engineer, you will work closely with both product and platform engineering teams to ensure Duolingo’s sophisticated distributed... 
    Senior
    Work experience placement

    Socket

    Eastern, KY
    2 days ago
  •  ...automated detection, drain/cordon/taint, workload rescheduling. Feed the AIOps substrate The remediation-actuator and workflow engine land here — you make the control plane safe for automated action. Your CRDs are the schema the platform's predictors and... 
    Senior
    Local area

    Bitdeer Technologies Group

    San Jose, CA
    21 hours ago
  • $160k - $180k

     ...have a big impact. See Arkestro in action at arkestro.com. About the Role Arkestro is hiring for a Senior SRE Engineer to manage our performance and reliability for our software platform and infrastructure. The right candidate will own and develop our... 
    Senior
    Local area
    Remote work

    Arkestro

    United States
    3 days ago
  •  ...TechMContact: Meghana GorusuCompany: SRI Tech SolutionsJob Title: Senior Site Reliability EngineerLocation: Plano , TX (remote)Years of Experience:8...  ...are seeking a highly skilled Senior Site Reliability Engineer (SRE) to join our dynamic team. The ideal candidate will have... 
    Senior
    Remote work

    Sri Tech

    Plano, TX
    3 hours ago
  • $152k - $195k

     ...investors including Silver Lake Waterman, Moody’s, Sequoia Capital, GV and Riverwood Capital. About the Team: As a Senior Site Reliability Engineer, you will be a key technical leader driving the design and optimization of our Kubernetes-based infrastructure and CI/... 
    Senior
    Remote work

    SecurityScorecard

    United States
    4 days ago
  •  ...Evaluate applications, platforms, and vendors to assess resiliency, reliability, and operational risk.Design and implement processes that...  ...and reliability tooling.Actively participate in reliability engineering and resilience communities of practice, contributing to... 
    Senior
    Full time

    Vanguard Services Inc

    Dallas, TX
    3 hours ago
  • $170k - $220k

    Who We're Looking ForWe’re looking for a hands-on, high-agency Site Reliability Engineer to help shape and scale the reliability layer of our stack. You'll own the release pipeline end-to-end — managing daily releases, weekly deploys, and hotfixes — while also automating... 
    Senior

    Supio

    Seattle, WA
    3 hours ago
  •  ...Senior Site Reliability Engineer Remote – Home Based Job Summary We’re partnering with a company in the SaaS space to find a Senior Site Reliability Engineer . In this role, you’ll be part of the IT Operations group responsible for maintaining all environments... 
    Senior
    Temporary work
    Remote work
    Work from home
    Flexible hours

    SourceDirect Talent

    United States
    21 hours ago
  • $96k - $163k

     ...and services that help people, businesses and governments realize their greatest potential. Title and Summary Senior Site Reliability Engineer, Performance Engineering Senior Site Reliability Engineer, Performance Engineering Payment Optimization... 
    Senior
    Full time
    Part time
    Worldwide
    Flexible hours

    Mastercard

    O Fallon, MO
    1 day ago
  • $145k - $193k

     ...entertainment, we want to talk to you. About the Role & Team The SRE team at PENN Entertainment is looking for a Senior Site Reliability Engineer to help build and operate the infrastructure behind a large-scale sports betting and media platform. You'll own critical... 
    Senior
    Remote work

    Penn Interactive

    United States
    4 days ago
  •  ...role, we encourage you to apply. The Role  As a Senior Platform Engineer, you are a champion for DevOps and SRE culture and industry...  .... \n What You Will Be Doing Improving production reliability and system resilience within an SRE scoped team Championing... 
    Senior
    Remote work
    Flexible hours

    Megaport

    United States
    2 days ago
  •  ...and performance our customers have come to expect, and help raise the reliability bar as we grow. What you would do: Design, build, and operate the shared platform foundations engineers ship on every day: GCP infrastructure, Kubernetes, networking, routing,... 
    Senior
    Remote work
    Worldwide
    Flexible hours

    Sanity

    United States
    3 days ago
  •  ...and services that help people, businesses and governments realize their greatest potential. Title and Summary Senior Site Reliability Engineer Overview-The ProCOM team is looking for a Site Reliability Engineering (SRE) who can help us solve problems, build... 
    Senior
    Full time
    Part time
    Immediate start
    Worldwide
    Flexible hours

    Mastercard

    O Fallon, MO
    1 day ago
  • $141.8k - $195k

     ...best work, grow fast, and bring their full selves to the herd. Why You’ll Love This Role Cribl Inc is seeking a Senior Site Reliability Engineer to join our mission where you will unlock the value of all observability data, as we expand our team in the U.S. Cribl... 
    Senior
    Temporary work
    Remote work

    Cribl

    United States
    3 days ago
  •  ...We are seeking a Senior Database Reliability Engineer (DBRE) to design, operate, and improve reliable, scalable, secure, and highly available database...  ...data platforms. The role combines database engineering, site reliability engineering, Linux systems administration, and... 
    Senior

    Neshent Technologies

    Los Gatos, CA
    21 hours ago
  •  ...About The Role: We're looking for a Senior Site Reliability Engineer to help us mature and scale the infrastructure behind our multi-cloud SaaS platform. Most of our footprint runs on Microsoft Azure, built from the ground up around cloud architecture principles:... 
    Senior
    Remote work
    Flexible hours

    Dental Intelligence

    United States
    2 days ago
  •  ...healthcare organizations maintain accurate, compliant, and reliable provider networks at scale. Our vision is simple: One...  ...of patients. About the Role We're looking for a Senior Site Reliability Engineer who takes ownership seriously - someone who designs for... 
    Senior
    Remote work

    CertifyOS

    United States
    3 days ago
  • $200k - $260k

     ...the future of professional services is being written today — and we're just getting started.Role OverviewAs a Software Engineer on the Site Reliability team at Harvey, you will ensure the reliability, scalability, and performance of our legal AI platform. You'll join a... 
    Senior
    Relocation package

    Harvey

    San Francisco, CA
    2 days ago
  • $152.6k - $191.5k

     ...is responsible for partnering with leaders across engineering and technology to define objective reliability goals for services. Key responsibilities include...  ...and continuous improvement.Position Summary: The Senior Site Reliability Engineer acts as an advanced senior individual... 
    Senior
    Full time
    Work at office
    Day shift

    Bank of America

    Charlotte, MI
    3 hours ago
  •  ...Leading initiatives to transform the IT Compute Core Team, the full-time Senior Staff Site Reliability Engineer will design, scale, and deploy core infrastructure services while optimizing performance and reliability across on-prem and cloud environments. Key responsibilities... 
    Senior
    Full time
    Remote work

    Virtual Vocations Inc

    United States
    1 day ago
  • $96k - $163k

     ...and services that help people, businesses and governments realize their greatest potential. Title and Summary Senior Site Reliability Engineer Who is Mastercard? At Mastercard technology, we work to connect and power an inclusive, digital economy that... 
    Senior
    Full time
    Part time
    Worldwide
    Flexible hours

    Mastercard

    O Fallon, MO
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!