Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer

On-Board Services

Cloud Operations Engineer III

Function: Engineering 

Reports to: Manager, Cloud Operations 

Location: Remote - United States 

Position Summary

The Cloud Operations Engineer III is a senior member of the cloud operations team, responsible for the reliability, observability, performance, and operational security of our multi-product SaaS platform. This role owns our Datadog observability practice — instrumentation standards, dashboards, SLOs, monitors, and alert routing — and leads the migration off our legacy monitoring stack. 

It is an engineering role, not a ticket-queue role: the expectation is that recurring operational work gets replaced with code. The Cloud Operations Engineer III participates in on-call, incident response and is measured on fewer customer-impacting incidents, faster detection and recovery, and less manual work year over year. 

 
The ideal candidate is a proactive problem-solver who thrives in dynamic, evolving environments and works effectively across departments to address complex challenges. They have experience partnering with cross-functional teams to understand and document requirements, then translating those needs into meaningful dashboards that improve service visibility (Observability) and support informed decision-making. They are passionate about automation, process improvement, and eliminating unnecessary manual effort. They confidently propose better approaches when opportunities for improvement arise. 

Key Responsibilities

Observability and Datadog Ownership

  • Own the Datadog platform across all products and environments, including agent lifecycle, instrumentation standards, unified service tagging, and per-cluster configuration. 
  • Instrument services for APM and distributed tracing, log collection, and synthetic monitoring; partner with engineering teams to close instrumentation gaps in both legacy and modern codebases. 
  • Build and maintain the dashboard, monitor, and SLO catalog; define SLIs and error budgets for critical user journeys and use them to drive prioritization with engineering and product. 
  • Design high-signal alerting: reduce noise and duplicate alerts, tune thresholds, and ensure every alert has an owner and a runbook. 

Automation and Toil Elimination

  • Develop and maintain automation in PowerShell, Python, and Bash for provisioning, configuration, diagnostics, remediation, and reporting. 
  • Extend our infrastructure-as-code estate — Bicep modules, Kubernetes manifests, Helm releases, and Azure DevOps pipeline templates — so environments and regions are reproducible and drift-free. 
  • Convert manual runbooks into automated or self-service workflows: cluster upgrades, secret and certificate rotation, tenant provisioning, data retention purges, and access provisioning. 

Security, Documentation, and Mentorship

  • Implement and maintain platform security controls and audit-ready operational evidence: managed identities, secret and key rotation, least-privilege access, and image and dependency scanning. 
  • Author and maintain runbooks, on-call guides, and architecture documentation, and provide technical leadership and mentorship to junior engineers on observability, automation, and incident response. 

Skills and Experience Needed

  • Bachelor's degree in Computer Science, Information Technology, or a related field, or equivalent practical experience. 
  • 5-7 years of professional experience in cloud operations, site reliability, platform, or DevOps engineering for production SaaS systems. 
  • Demonstrated hands-on depth with a modern observability platform — Datadog strongly preferred — including APM and distributed tracing, log pipelines and indexing controls, dashboards, monitors, and SLOs. 
  • Strong scripting and automation ability in PowerShell, with the judgment to write tooling that other engineers can safely operate. 
  • Strong knowledge of containers, container orchestration, and the Kubernetes ecosystem, including autoscaling, cluster upgrades, and diagnosing pod-level failures. 
  • Production experience with Azure — Kubernetes Service, Azure SQL, Cosmos DB, Redis, Service Bus, Key Vault, and Entra ID — or equivalent depth in another major cloud. 
  • Experience with infrastructure-as-code and CI/CD pipeline authoring (Bicep or Terraform; Helm or Kustomize; Azure DevOps preferred). 
  • Proven incident response experience in a customer-facing production environment, including on-call participation and leading post-incident reviews. 
  • Experience operating multi-region, multi-tenant systems. 
  • Strong knowledge of platform security and operational best practices: secret and key rotation, least-privilege access, and vulnerability remediation. 
  • Excellent problem-solving and analytical abilities, with strong written communication for runbooks, incident updates, and technical proposals. 
  • Strong communication, and teamwork skills, including the ability to work effectively with legacy systems and their constraints. 
  • Nice to have: experience migrating from a legacy monitoring stack to a consolidated observability platform; relevant Azure, Kubernetes, or Datadog certifications. 

Competencies

Accountability 

Adaptability 

AI Curiosity/Innovation 

Applied Learning 

Business Acumen 

Collaboration 

Customer Focus 

Dealing w/Ambiguity 

Decision Making 

Driving for Results 

Initiating Action 

Planning and Organizing 

Technical/Professional Knowledge 

About the Company:

Boards set the standard for what organizations can achieve. At OnBoard, our board management software helps boards function at a higher level so every organization can make a bigger difference in the world.

Launched in 2011, today, OnBoard serves as the board intelligence platform for more than 5,000 organizations and their 12,000 boards and committees in 60 countries worldwide. With customers in higher education, nonprofit, healthcare systems, government, and enterprise business, OnBoard is the leading board management provider.

OnBoard has grown from a class project at Purdue University in West Lafayette, Indiana in 2003 into the world’s leading board management software platform today. Backed by JMI Equity and the acquisitions of eScribe and Govenda, OnBoard is positioned to become the industry leader in Board Management and Meeting Solutions for private and public sector entities.

Benefits and Perks:

  • Fully remote work with company provided equipment (laptop, software, etc.) 
  • Employment with a growing, casual, fun, philanthropic minded company
  • US Based Employees
    • Comprehensive, high-quality medical/prescription drug plan options, as well as dental and vision plan offerings.   
    • An employer contribution to your Health Savings Account (HSA) if you participate in a High Deductible Healthcare Plan.  
    • Medical Flexible Spending Accounts available.   
    • Dependent Care Flexible Spending Accounts available.  
    • Basic life insurance in the amount of $50,000 or 1 X’s your salary (whichever is higher).
    • Short and long-term disability and Accidental Death and Dismemberment benefits at no cost to you.  
    • 401K Retirement Savings Plan with automatic enrollment at the first of the month following 60 days of employment at 5% to help you secure your financial freedom. We offer a generous company match that starts on the first of the month following 60 days of employment. The company match is dollar for dollar on the first 3% of your pay that you contribute and $0.50 on the dollar on the next 2%, for a total match of 4%. 
    • Paid Time Off (PTO)/Holiday 
  • CAN Based Employees

    • Employer paid Life and Accidental Death Insurance
    • Contribution to Health Care Spending Account
    • Dependent Life Insurance
    • Optional Life Insurance
    • LTD Insurance
    • Drug and Paramedical Coverage
    • Dental Insurance
    • Vision Insurance
    • EAP
  • AUS Based employees
    • Superannuation rate of 12% 
    • Monthly stipend of $400 AUD to purchase private medical insurance 
  • UK Based Employees (via EPG)
    • Pension - Aegon
      • Passageways/OnBoard contributes 8% of the employee's basic salary
      • Employees can contribute up to 100% of salary subject to max limits
      • Enrolled from Day 1 of employment
    • Private Medical Insurance
    • Life Assurance
    • Income Protection
    • Critical Illness
    • Employee Assistance Programme
    • Serious Illness Benefit
    • View email address on click.appcast.io
    • Cashplan

Diversity Statement - Culture of Togetherness: 

At OnBoard, our mission is to encourage and celebrate a culture of togetherness. We acknowledge that uniqueness is powerful, and we welcome, foster, and appreciate all. Diversity, Equity, and Inclusiveness fuel the Pathfinder atmosphere and all our efforts. Our power is in our people and we Pledge 1% to give back to our communities and across the globe.

OnBoard is an equal opportunity employer and committed to a diverse and inclusive working environment. We  do not discriminate based on race, national origin, gender, gender identity, sexual orientation, protected veteran status, disability, age, or other legally protected status.

Interview Transparency & Technology Disclosure
We use video/audio recordings and artificial intelligence (AI) tools during our interview process to transcribe responses, evaluate skills, and streamline evaluations. Your data is processed securely and handled in line with our Privacy Policy and local data protection laws

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in United States vacancy
  •  ...Fluency: English (Required)Work Shift:1st shift (United States of America)Please review the following job description:The Site Reliability Engineer role focuses on enhancing the reliability and operational excellence of enterprise platforms across hybrid cloud and on-premises... 
    Suggested
    Permanent employment
    Full time
    Part time
    H1b
    Work at office
    Local area
    Immediate start
    Work visa
    Monday to Friday
    Shift work
    Day shift

    Truist

    Raleigh, NC
    5 days ago
  •  ...and responsible for ensuring the availability, scalability, and reliability of systems and applications.What will be your responsibilities...  ...using tools like Terraform or CloudFormation.Mentor junior engineers and provide technical guidance.Stay up-to-date with industry trends... 
    Suggested
    Work at office
    Remote work

    Interactive Brokers

    Greenwich, CT
    4 days ago
  • $148.5k - $223.9k

     ...right place! Agentforce is the future of AI, and you are the future of Salesforce.Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in San Francisco. Working closely with counterparts in the Infrastructure and R&D organizations,... 
    Suggested
    Full time
    Worldwide
    Weekend work

    Salesforce

    San Francisco, CA
    4 days ago
  • $104.9k - $174.7k

     ...Data Management. You can learn more about LexisNexis Risk at the link below, About the Role:We are hiring a hands-on Senior Site Reliability Engineer (SRE) to actively build, operate, and improve the reliability of our production systems. This is not a purely advisory... 
    Suggested
    Full time
    Work at office
    Local area
    Remote work
    Work from home

    RELX Group

    Atlanta, GA
    1 day ago
  • Job ID: 28091697Reference Number: 26-00413Title: Site Reliability EngineerLocation: Iselin, NJ, 08830Posted Date: 2026-04-27Contact: Deepak...  ...Phone: (***) ***-****Company: HAN Staffing As a Site Reliability Engineer at JPMorgan Chase within the Commercial & Investment Banking,... 
    Suggested

    HAN Staffing

    Iselin, NJ
    1 day ago
  • $45 - $85 per hour

    DescriptionThe Resy Site Reliability Engineering groups goal is to ensure Resy Customers can always use the service reliably. We're looking for engineers to be part of an empowered, self-organizing group, with the opportunity to use modern languages and tools and to operate... 
    Contract work
    Temporary work

    TEKsystems

    Phoenix, AZ
    5 days ago
  •  ...importance of in-office collaboration and fully intend for the selected candidate for this role to work on site in the specified location(s).As a Site Reliability Engineer supporting the Cashiering organization, you will play a critical role in ensuring the stability,... 
    Full time
    Work at office

    The Charles Schwab Corporation

    Southlake, TX
    20 hours ago
  • $160k - $180k

     ...expertise, and world-class customer satisfaction. The Platform Engineering group at CentralReach builds the underlying technologies that...  ...in Software Engineering to drive adoption of modern reliability practices like SLOs, error budget policies, actionable alerts... 
    Full time
    Worldwide

    CentralReach

    Holmdel, NJ
    20 hours ago
  •  ...and foster a dynamic work environment where new ideas thrive. Are you ready to join our team and make an impact?As a Senior Site Reliability Engineer at TeamViewer, you’ll be a key player in ensuring the reliability, scalability, and performance of our Azure-based SaaS... 
    Temporary work
    Casual work
    Worldwide

    TeamViewer

    Austin, TX
    1 day ago
  •  ...and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Corporate and Investment Banking team, you will solve complex and broad business problems... 

    JP Morgan Chase

    Orem, UT
    5 days ago
  • $182.8k - $247.3k

     ...mission to develop education for our half a billion (and growing!) learners around the world.About the role...As a Senior Site Reliability Engineer, you will work closely with both product and platform engineering teams to ensure Duolingo’s sophisticated distributed systems... 
    Work experience placement

    Duolingo

    Pittsburgh, PA
    2 days ago
  •  ...and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems. As a Site Reliability Engineer III at JPMorgan Chase within the IP, you will solve complex and broad business problems with simple and straightforward solutions... 

    JP Morgan Chase

    Plano, TX
    4 days ago
  • $141k - $216.6k

     ...—it means helping shape the future of emergency response and building a safer, more connected world.Position OverviewAs a Site Reliability Engineer, you'll own the reliability, observability, and operational excellence of our Unified Call (UC) platform—the mission-critical... 
    Work experience placement
    Work at office

    Axon

    New York, NY
    3 days ago
  • About the RoleWe’re looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You’ll partner with engineers and data scientists to build, automate, and maintain the infrastructure... 

    Alembic

    San Francisco, CA
    5 days ago
  • The Senior Site Reliability Engineer is responsible for improving the reliability, availability, scalability, and operational excellence of our critical infrastructure platforms and services. This role partners closely with Engineering, Security, and Infrastructure teams... 
    Full time
    Work at office
    Local area

    Castleton Commodities International

    Stamford, CT
    5 days ago
  •  ...selected candidate for this role to work on site in the specified location(s).The Client...  ...team is responsible for ensuring the reliability, scalability, and operational excellence...  ...around the clock. As a Site Reliability Engineer, you will partner across application engineering... 
    Full time
    Work at office

    The Charles Schwab Corporation

    Austin, TX
    20 hours ago
  • Qualifications: 8+ years of Software Engineering experience, or equivalent demonstrated through...  ...implement and maintain scalable and reliable infrastructure on Google Cloud Platform...  ...vendor resources Willingness to work on-site at stated location in the job openingDepartment... 
    Contract work
    For contractors
    Work experience placement

    Cedent Consulting

    Dallas, TX
    4 hours ago
  • $230k - $250k

     ...network. It's the foundation for autonomous networking, giving engineers and AI agents the ability to know the impact of every change...  ...how things have always been done.Forward is looking for a Site Reliability EngineerAbout the Role This is not a "keep the lights on"... 
    Night shift

    Forward Networks

    Santa Clara, CA
    4 hours ago
  • $167.7k - $245.2k

     ...very effective.We’re looking for talented engineers with a software or operations background...  ...development teams to ensure the reliability, performance and security of our infrastructure...  ...insurance. Please see the Cisco careers site to discover more benefits and perks.... 
    Full time
    Temporary work
    Work at office
    Local area
    Flexible hours
    1 day per week

    CISCO Systems

    New York, NY
    4 days ago
  • $143k - $191k

     ...mission critical capabilities to our customers. System Deployment Engineers work in complex environments with shared environmental...  ...customersParticipate in customer demonstrations and exercisesWork with site reliability engineers to provide and refine requirements for tooling and... 
    Full time
    Temporary work
    Work experience placement
    Immediate start

    Anduril Industries

    Seattle, WA
    4 hours ago
  •  ...importance of in-office collaboration and fully intend for the selected candidate for this role to work on site in the specified location(s). As a Senior Site Reliability Engineer within the CET SAvE organization, you will play a critical leadership role advancing the... 
    Full time
    Work at office

    The Charles Schwab Corporation

    Southlake, TX
    1 hour ago
  • Senior Site Reliability EngineerLocation: Exton or Philadelphia, PA (Hybrid - 3 times a week in-office)Position SummaryAre you ready to start...  ...looking for you!We are looking for a Senior Site Reliability Engineer to take on the responsibility of automating cloud-based... 
    Casual work
    Work at office
    Worldwide

    Bentley Systems

    Philadelphia, PA
    3 days ago
  • We are looking for an experienced Site Reliability Engineer (SRE) to strengthen observability and operational resilience across a Microsoft Azure environment. This long-term Contract role will work closely with DevOps and engineering teams to establish monitoring standards... 
    Long term contract

    Robert Half

    Maumee, OH
    1 day ago
  • $158.5k - $172k

     ...exceptional value they deserve.About The OpportunityAs a Senior Engineer on the Runtime Automation team, you will design, automate, and...  .... This is a high-impact position driving continuous reliability, deep system optimization, and automation across our entire technology... 
    Full time
    Temporary work
    Work at office
    Flexible hours
    3 days per week

    GrubHub

    New York, NY
    4 hours ago
  •  ...some of the most important challenges in global education. Client is currently seeking a talented Software Engineer who is able to work into the Site Reliability Engineer role. This candidate is expected to work towards becoming a Subject Matter Expert in the cloud space... 
    Remote work

    Intelliswift

    Durham, NC
    5 days ago
  • $165k - $225.6k

     ...From core infrastructure to enterprise platforms, we partner across functions to drive scale, reliability, and innovation through technology.The Senior Site Reliability Engineer OpportunityReporting to the Manager, Site Reliability Engineering, this role will help build,... 
    Permanent employment
    Local area
    Worldwide
    Flexible hours

    Okta

    Bellevue, WA
    5 days ago
  • Inspire Brands is hiring two Senior Site Reliability Engineers to help build and scale reliable, resilient, and observable systems supporting high-traffic, customer-facing digital platforms. These role blends software engineering, systems thinking, and operational excellence... 
    Worldwide

    Inspire Brands

    Atlanta, GA
    5 days ago
  • Senior Site Reliability Engineer ILocationSan Jose, Costa Rica - RemoteSummary of roleOwn availability, the most important product feature, by continually striving for sustained operational excellence of Sumo’s planet-scale observability and security products. Work with... 
    Flexible hours

    Sumo Logic

    San Jose, CA
    3 days ago
  • $166k - $220k

     ...requirements and customer expectations. Our systems integration engineers internalize the nuances of each deployment, ensuring the...  ...-to-end solutions we ship.ABOUT THE JOBWe are looking for a Site Reliability Engineer (SRE) to join AGD, our rapidly growing team in Irvine... 
    Full time
    Work experience placement
    Immediate start

    Anduril Industries

    Costa Mesa, CA
    4 days ago
  • $102.1k - $202.2k

     ...per yearEmployment type: Full-TimeWork site: Fully on-siteRole type: Individual ContributorTravel...  ...: Software EngineeringDiscipline: Site Reliability EngineeringCompany:...  ...at the intersection of large-scale cloud engineering, service reliability, and operational excellence... 
    Ongoing contract
    Local area
    Worldwide

    Microsoft

    Reston, VA
    5 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!