Site Reliability Engineer
OnBoard Group
Reports to: Manager,Cloud Operations
Location: Remote - United States
Position Summary
The Cloud Operations Engineer III is a senior member of thecloud operationsteam, responsible for the reliability, observability, performance, and operational security of our multi-product SaaS platform. This role owns our Datadog observability practice — instrumentation standards, dashboards, SLOs, monitors, and alert routing — and leads the migration off our legacy monitoring stack. The idealcandidateis aproactiveproblem-solverwhothrives indynamic,evolvingenvironmentsand workseffectivelyacrossdepartmentsto address complexchallenges.They haveexperiencepartneringwith cross-functionalteams tounderstandanddocumentrequirements,thentranslatingthose needsintomeaningfuldashboardsthatimproveservicevisibility(Observability) and support informeddecision-making.They are passionate about automation,process improvement,andeliminatingunnecessarymanualeffort.Theyconfidentlyproposebetterapproacheswhenopportunitiesfor improvementarise.
Key Responsibilities
Observability and Datadog Ownership
- Own the Datadog platform across all products and environments, including agent lifecycle, instrumentation standards, unified service tagging, and per-cluster configuration.
- Instrument services for APM and distributed tracing, log collection, and synthetic monitoring; partner withengineeringteams to close instrumentation gaps in both legacy and modern codebases.
- Build and maintain the dashboard, monitor, and SLO catalog; define SLIs and error budgets for critical user journeys and use them to drive prioritization with engineering and product.
- Design high-signal alerting: reduce noise and duplicate alerts, tune thresholds, and ensure every alert has an owner and a runbook.
Automation and Toil Elimination
- Develop andmaintainautomation in PowerShell, Python, and Bash for provisioning, configuration, diagnostics, remediation, and reporting.
- Extend our infrastructure-as-code estate — Bicep modules, Kubernetes manifests, Helm releases, and Azure DevOps pipeline templates — so environments and regions are reproducible and drift-free.
- Convert manual runbooks into automated or self-service workflows: cluster upgrades, secret and certificate rotation, tenant provisioning, data retention purges, and access provisioning.
Security, Documentation, and Mentorship
- Implement andmaintainplatform security controls and audit-ready operational evidence: managed identities, secret and key rotation, least-privilege access, and image and dependency scanning.
- Author andmaintainrunbooks, on-call guides, and architecture documentation, and provide technical leadership and mentorship to junior engineers on observability, automation, and incident response.
Skills and Experience Needed
- Bachelor's degree in Computer Science, Information Technology, or a related field, or equivalent practical experience.
- 5-7 years of professional experience in cloud operations, site reliability, platform, or DevOps engineering forproductionSaaS systems.
- Demonstrated hands-on depth with a modern observability platform — Datadog strongly preferred — including APM and distributed tracing, log pipelines and indexing controls, dashboards, monitors, and SLOs.
- Strong scripting and automation ability in PowerShell, with the judgment to write tooling that other engineers can safelyoperate.
- Strong knowledge of containers, container orchestration, and the Kubernetes ecosystem, including autoscaling, cluster upgrades, and diagnosing pod-level failures.
- Production experience with Azure — Kubernetes Service, Azure SQL, Cosmos DB, Redis, Service Bus, Key Vault, and Entra ID — or equivalent depth in another major cloud.
- Experience with infrastructure-as-code and CI/CD pipeline authoring (Bicep or Terraform; Helm orKustomize; Azure DevOps preferred).
- Proven incident response experience in a customer-facing production environment, including on-call participation and leading post-incident reviews.
- Strong knowledge of platform security and operational best practices: secret and key rotation, least-privilege access, and vulnerability remediation.
- Excellent problem-solving and analytical abilities, with strong written communication for runbooks, incident updates, and technical proposals.
- Strongcommunication,and teamwork skills, including the ability to work effectively with legacy systems and their constraints.
- Nice tohave:experience migrating from a legacy monitoring stack to aconsolidatedobservability platform; relevant Azure, Kubernetes, or Datadog certifications.
Accountability
AI Curiosity & Innovation
Business Acumen
Customer Focus
Dealing with Ambiguity
Decision Making
Driving for Results
Initiating Action
Technical/Professional Knowledge
About the Company:
Boards set the standard for what organizations can achieve. At OnBoard, our board management software helps boards function at a higher level so every organization can make a bigger difference in the world.
Launched in 2011, today, OnBoard serves as the board intelligence platform for more than 5,000 organizations and their 12,000 boards and committees in 60 countries worldwide. With customers in higher education, nonprofit, healthcare systems, government, and enterprise business, OnBoard is the leading board management provider.
OnBoard has grown from a class project at Purdue University in West Lafayette, Indiana in 2003 into the world’s leading board management software platform today. Backed by JMI Equity and the acquisitions of eScribe and Govenda, OnBoard is positioned to become the industry leader in Board Management and Meeting Solutions for private and public sector entities.
Benefits and Perks:
- Fully remote work with company provided equipment (laptop, software, etc.)
- Employment with a growing, casual, fun, philanthropic minded company
- US Based Employees
- Comprehensive, high-quality medical/prescription drug plan options, as well as dental and vision plan offerings.
- An employer contribution to your Health Savings Account (HSA) if you participate in a High Deductible Healthcare Plan.
- Medical Flexible Spending Accounts available.
- Dependent Care Flexible Spending Accounts available.
- Basic life insurance in the amount of $50,000 or 1 X's your salary (whichever is higher) .
- Short and long-term disability and Accidental Death and Dismemberment benefits at no cost to you.
- 401K Retirement Savings Plan with automatic enrollment at the first of the month following 60 days of employment at 5% to help you secure your financial freedom. We offer a generous company match that starts on the first of the month following 60 days of employment. The company match is dollar for dollar on the first 3% of your pay that you contribute and $0.50 on the dollar on the next 2%, for a total match of 4%.
- Paid Time Off (PTO)/Holiday
- CAN Based Employees
- Employer paid Life and Accidental Death Insurance
- Contribution to Health Care Spending Account
- Dependent Life Insurance
- Optional Life Insurance
- LTD Insurance
- Drug and Paramedical Coverage
- Dental Insurance
- Vision Insurance
- EAP
- AUS Based employees
- Superannuation rate of 12%
- Monthly stipend of $400 AUD to purchase private medical insurance
- UK Based Employees (via EPG)
- Pension - Aegon
- Passageways/OnBoard contributes 8% of the employee's basic salary
- Employees can contribute up to 100% of salary subject to max limits
- Enrolled from Day 1 of employment
- Private Medical Insurance
- Life Assurance
- Income Protection
- Critical Illness
- Employee Assistance Programme
- Serious Illness Benefit
- View email address on click.appcast.io
- Cashplan
- Pension - Aegon
Diversity Statement - Culture of Togetherness:
AtOnBoard, our mission is to encourage and celebrate a culture of togetherness. We acknowledge that uniqueness is powerful, and we welcome, foster, and appreciate all. Diversity, Equity, and Inclusiveness fuelthe Pathfinder atmosphere and all our efforts. Our power is in our people and we Pledge 1% to give back to our communities and across the globe.
OnBoardis an equal opportunity employer and committed to a diverse and inclusive working environment.We do not discriminate based on race, national origin, gender, gender identity, sexual orientation, protected veteran status, disability, age, or other legally protected status.
Interview Transparency & Technology Disclosure We use video/audio recordings and artificial intelligence (AI) tools during our interview process to transcribe responses, evaluate skills, and streamline evaluations. Your data is processed securely and handled in line with our Privacy Policy and local data protection laws
#J-18808-Ljbffr- ...Engineering, Product, Design, and Marketing Engineering Compensation ~ Zone 1 Base Pay: $214K – $260K Superhuman offers... ...role will be responsible for building software to ensure the reliability of our back-end systems, working with engineers who develop them...SuggestedWorldwideHome officeFlexible hours
- ...Cloudflare, GitHub Actions, PostgreSQL, Redis/BullMQ, Node.js/NestJS, Datadog, TypeScript, React, SQL Position: Senior Site Reliability Engineer Engagement period: Ongoing Interview timeline: ASAP Interview process: 1) CV review 2) Interview with our CTO 3)...SuggestedContract workImmediate start
$115.5k - $164.8k
...matters at a company where you matter. Your Impact As an engineer on the APX SRE CloudOps team, you will spend a significant portion... ...that replace what previously required human intervention with reliable, tested automation. You will also participate in on‑call...SuggestedWork experience placementWork at officeRemote work$140k - $195k
...Improve reliability, observability, service health, incident response, and operational readiness. CodeVertex works across data... ..., secure systems, and operational clarity matter. The Site Reliability Engineer role helps turn business needs into reliable execution, whether...SuggestedRemote work- ...This is an engineering-first Senior SRE role. We’re looking for senior engineers who have: Built and shipped significant backend... ...services end-to-end in production (design → launch → on-call → reliability improvements) Led incident response and driven durable...Suggested
- ...profitable developer-tooling company whose product is used by engineering teams at thousands of software companies for application... ...well-resourced group of nine. As Senior SRE you will lead reliability initiatives across the platform — from defining and driving SLOs...
- ...personalized care faster. We are building AI agents to support the full arc of the patient journey. The Opportunity: Machine Learning Engineer Patients count on our platform 24/7. You'll build and maintain the tooling, alerts and incident-response playbooks that keep...
$182.8k - $247.3k
...to develop education for our half a billion (and growing!) learners around the world. About the role... As a Senior Site Reliability Engineer, you will work closely with both product and platform engineering teams to ensure Duolingo’s sophisticated distributed systems...Work experience placement$180k - $230k
...Acceleration Job Description We're looking for a Senior SRE to own the reliability, scalability, and observability of our production systems. You'll work closely with platform and data engineering to keep high-throughput, data-intensive services running at the...Work at officeLocal areaImmediate startRemote work3 days per week- ...A senior Site Reliability Engineer will join an established infrastructure function responsible for highly available, security-conscious cloud systems supporting complex business-critical workloads. You’ll take significant ownership of reliability, scalability, and...Full timeRemote work
$90k - $100k
...Site Reliability Engineer The Opportunity We are looking for a highly capable engineer to join our Platform and Site Reliability engineering team. You will be responsible for building, maintaining and operating the infrastructure platform on which all Ookla services...Flexible hours$110k - $145k
...operations. You will liaise with product and engineering teams to ensure applications and... ...feedback loop for platform and product reliability. The ideal candidate is a solutions-oriented... ...experience as a platform engineer, site reliability engineer, systems engineer...Work experience placement- ...As a Senior Site Reliability Engineer on our cloud engineering team, you'll keep our production environment healthy, secure, and running smoothly. This is an operations-focused role: you'll own the day-to-day administration of our AWS accounts and databases, backup posture...Work experience placement
$107.9k - $195.05k
...The Digital Sector at Leidos currently has an opening for a Site Reliability Engineer (SRE) / Senior Cloud Engineer to work in our Baltimore, Maryland office. This is an exciting opportunity to use your experience helping the Center for Medicare and Medicaid Services (...Contract workWork at office- ...% uptime. You'll own SLOs, incident response, and production reliability for a system that processes millions of identity verifications... ...Sentry error tracking, structured logging Implement chaos engineering practices to proactively identify failure modes Optimize...Remote work
$148.5k - $223.9k
...Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in San Francisco. Working closely with counterparts in the Infrastructure and R&D organizations, this organization provides a global team of engineers monitoring cloud service...WorldwideWeekend work- ...Discover exciting DevOps job opportunities and connect with 28,396 DevOps professionals. The Senior Site Reliability Engineer role at Jobicy is designed for experienced professionals who are passionate about enhancing system reliability and operational efficiency. The...Remote workFlexible hours
$114k - $148k
...Total compensation is based on experience, skills, and location using objective, job-related criteria. Summary As a Site Reliability Engineer, you will focus on ensuring the platform and services customers rely on are reliable, performant, and highly available. If...Work experience placement- ...democratizing software development by removing traditional barriers to application creation. About the role: Join our Site Reliability Engineering team and help ensure the reliability, scalability, and performance of Replit's infrastructure that serves millions of...Full timeTemporary workWork at officeWorldwideFlexible hours
- ...Quarterhill is seeking a Senior Site Reliability Engineer (SRE) to join our growing team. This role is an exciting opportunity to contribute to the reliability and performance of smart transportation systems, including a next-generation, cloud-native tolling platform that...Local area
- ...knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who... ...knowledge with their teammates. ABOUT THE ROLE: As a Site Reliability Engineer focused on campus reliability, you will design what...Night shift
$104.9k - $174.7k
...Technology Senior Site Reliability Engineer II The SRE role is responsible for improving the reliability, availability, performance, and operational quality of production systems. This role provides technical input into project plans, schedules, methodologies, and...Temporary workLocal area$110k - $145k
...global, with headquarters in Denver, Colorado, and offices across the U.S., Canada, and India. We are seeking a Senior Site Reliability Engineer to own the reliability, scalability, performance, and operational integrity of critical production services. This role is...Flexible hours- ...We're seeking an SRE to ensure the reliability and performance of our clients' critical systems. You'll work on observability, incident... ...management Nice to have Experience with chaos engineering Knowledge of distributed systems Background in high-scale...Remote workFlexible hours
$185 per hour
...important clinical and business workflows, so they must be available, secure, and easy to operate. We are hiring a Senior Site Reliability Engineer to join our Security and Site Reliability team. You will focus on a stable, scalable AWS and Kubernetes platform, reliable...Work at office- ...match. The role We're looking for a Senior SRE to own the reliability, scalability, and operational posture of Satsuma's multi-... ...using AI-assisted development workflows Partner closely with engineering on reliability reviews and architecture decisions ~5-8...
$135k - $160k
...About the Role We are looking for a Site Reliability Engineer to help us evolve and safeguard the infrastructure powering healthcare experiences for millions of patients, all while keeping operational toil to a minimum for the Fabric tech community. What You’ll Do...$150k - $220k
...Senior Site Reliability EngineerJob detailsDepartment / EngineeringRemoteFull-time$150,000 USD - $220,000 USD## About UsMetaRouter is a customer... ...architecture.## About The RoleAs a Senior Site Reliability Engineer, you own significant pieces of our infrastructure and...Full timeRemote work$135.2k - $181.2k
...enhance electrical, mechanical, and sensor-based systems to ensure reliability and performance. Configure, calibrate, and validate... ...professional development, including an interest in emerging data engineering tools and methodologies. Preferred Qualifications: ~8+...Worldwide- ...Washington, District of Columbia, United States Contractor | On-site Job Description We are seeking an experienced Site Reliability Engineer (SRE) to help build and maintain highly reliable, scalable, and secure technology platforms. The SRE will combine software...For contractors
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
- site reliability engineer Eastern, KY
- site reliability engineer remote Eastern, KY
- site reliability engineer sre Eastern, KY
- site safety Eastern, KY
- on-site clinical research associate (traveling/remote) Eastern, KY
- construction site safety Eastern, KY
- junior website developer Eastern, KY
- historic site Eastern, KY
- IT site lead Eastern, KY
- site leader Eastern, KY

