Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer

Greenhouse Software

Sand Technologies is a global Physical AI company using data and AI to make critical industries work better. We partner with governments, cities and enterprises to improve how essential systems operate across healthcare, water, energy, telecommunications and infrastructure.

Our work delivers proven real-world impact. We have built AI systems that help manage London’s water supply, supported telecom network planning across hundreds of cities, and developed digital healthcare platforms serving tens of millions of people across Africa. From intelligent command centers to AI-powered infrastructure platforms, we help organizations sense, analyze and act in complex environments.

Our people are ambitious, curious and relentlessly practical. Our teams work alongside clients in the field, solving hard problems and deploying solutions that last. With colleagues across Africa, Europe, the UK and the US, we operate across the full stack - from research and engineering to deployment and capability building.

Our mission is simple: to harness AI to solve humanity’s most pressing challenges.

Our platform

Sand builds the intelligent platforms behind critical infrastructure: water, healthcare, telecommunications, energy, for enterprises and governments worldwide. Our AI and data intelligence systems help manage London's water supply and power digital health platforms serving tens of millions of people. We have delivered successfully into over fifty client environments globally and are rapidly growing our footprint in the US. The successful applicant for this role will have an opportunity to grow with these expansion efforts, taking on significant technical and operational leadership responsibilities as we scale.

We are investing in a platform, Symmetri, which underpins and accelerates our client impact. It integrates heterogeneous operational data, models the entities and relationships of an entire domain, and lets operators of critical infrastructure understand the state of their world, reason about it, and act. It runs in multi-tenant cloud, in single-tenant customer environments, on premises, and is already deployed and in use in multiple fully air-gapped installations at ministry and utility level.

Our US customers are city governments, water utilities and public agencies who will run their operations on these systems and audit how we deploy and manage them.

About the role

Sand operates in partnership with senior executives at our customers. Your primary client within these relationships will be the customer's enterprise IT organization. You will spend as much time in front of a municipal CIO's security team, an architecture review board and a cloud governance committee as you will in a terminal. You will be asked where the data sits, who has access to it, how it is encrypted, and what happens during an incident, and you will need to answer those questions accurately and calmly without escalating every one of them.

Symmetri is designed to be highly modular and extensible, supporting rapid development and the mechanisms to support reusable intelligence capabilities across clients without sharing their data. As it continues to become the standard way we deliver, the line between deploying the platform and building it converges. You will be expected to see that coming and drive the strategy and execution that helps us get there.

Your role will be to lead and deliver the deployment and the operation of cloud infrastructure, including Symmetri environments, for US customers, end to end.

What you’ll do
  • Provisioning: Standing up infrastructure for a new environment, where the shape of the build depends on the environment type and variant, from managed cloud through single-tenant to on premises based on established patterns and practices.
  • Configuration: Tuning environments to a customer's needs as we land and expand on the use cases they are consuming: performance, scale, resource ceilings, network boundaries, identity, applying their security model rather than ours.
  • Lifecycle: Running the upgrade and maintenance cycles, and coordinating releases across live customer production estates without breaking the operations that depend on them.
  • Operation: Monitoring, alerting, and being first responder for critical production outages at any time of the day or night - which we see as an exceptional occurrence to be solved for, not the norm.
  • Collaboration & Growth: You will be the first US member of our Delivery Infra & Enablement function, but you will not need to work in isolation. You will collaborate directly with an established and experienced remote enablement team and regional operations counterparts in other geographies. This will assist you to establish a full US-based operations team and set the standard of how US deployment and operations actually works in practice.
Your first 6 months
  • You have taken a US customer environment from provisioning to production and you are the named owner of it.
  • Runbooks for that environment exist, are accurate, and someone other than you (or their agent) could follow them safely.
  • Monitoring and alerting are in place and tuned, so that a page means something and silence means something too.
  • You have been through at least one customer security review as our technical voice in the room.
  • You can say, with evidence, which parts of our global deployment approach need to work differently in the United States.
Who you are
  • 5 to 8 years in platform engineering, cloud infrastructure, DevOps or site reliability, operating production systems that mattered to someone.
  • Cloud infrastructure, hands on. Azure and AWS. You have provisioned and run production data infrastructure rather than only managing it, with a strong understanding of networking, IAM and secure data handling, and how to design effectively within a client’s existing landscape and governance constraints.
  • Networking you can reason about under pressure. VPCs and VNets, subnets, security groups, DNS, load balancing, private connectivity, and the patterns for getting services to talk to each other across a customer's internal and external boundaries.
  • Kubernetes and containers. Pods, deployments, services, namespaces. Comfortable in kubectl or any kube-api interface for inspection and troubleshooting. You do not need to have written an operator.
  • Infrastructure as code. You are comfortable with using IaC by default, from the start, with more than one toolchain. Demonstrable experience across one or more of Pulumi, Terragrunt, Terraform, OpenTofu, CDK or equivalents is required. Pulumi is used for Symmetri and experience with it is a strong plus.
  • Python for automation. As an organisation with strong data science roots we are strongly biased to python for many coding use cases, so familiarity is important and experience is a plus.
  • Change management and branching discipline. Standard git workflows, and the willingness to follow and improve a team's existing strategy rather than your own.
  • Resource planning. Sizing compute, memory and storage against what a customer actually needs and what their environment can actually give you, including when that environment is on premises and finite.
  • Incident response. You have been on call. You know the difference between fixing an outage and fixing the cause of one, you are able to run effective root cause analysis processes that result in long term improvements.
  • Client-facing competence with enterprise IT. You can hold a technical conversation with a customer's infrastructure and security people, be trusted by them, and represent a commitment without over-promising.

You will not be equally strong across all of these aspects. Strong in most, familiar with the rest, unafraid of learning fast while being aware of and drawing in support for your current limitations.

  • Public sector, utility or other regulated operational environments, and the security review processes that come with them
  • Air-gapped, sovereign or disconnected deployments of data-intensive systems
  • Observability and monitoring stacks, and building the alerting rather than only responding to it
  • Serving ML models and LLM systems in production, including offline
Location

We are looking to hire on the East Coast in the United States, as this is where the majority of our clients are based.

How we work

This role has an on-call component and accountability. Real production systems, real operational consequences.

It also has less scaffolding than you might expect. We optimize first to fall in love with our clients' challenges and to deliver impactful solutions. This means you will not always receive neatly scoped work, and part of the job is creating that clarity yourself and sharing it. We are not looking for someone to invent everything from scratch or replicate exactly what they have done in the past either.

We believe in strong opinions, loosely held. We have established patterns and approaches that work, built by teams that have run these environments for years, and the person who succeeds here is one who learns them properly first and then improves them while participating in raising our global standard.

Due to the highly collaborative and internationally distributed nature of our work, successful candidates must be comfortable operating in small teams while contributing to larger, globally coordinated efforts. A strong sense of ownership, self-motivation and discipline in maintaining clear and consistent communication through virtual collaboration tools and video conferencing is essential.

#J-18808-Ljbffr
Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in Eastern, KY vacancy
  • $115.5k - $164.8k

     ...matters at a company where you matter. Your Impact As an engineer on the APX SRE CloudOps team, you will spend a significant portion...  ...that replace what previously required human intervention with reliable, tested automation. You will also participate in on‑call... 
    Suggested
    Work experience placement
    Work at office
    Remote work

    Koitecc Solutions

    Eastern, KY
    1 day ago
  •  ...This is an engineering-first Senior SRE role. We’re looking for senior engineers who have: Built and shipped significant backend...  ...services end-to-end in production (design → launch → on-call → reliability improvements) Led incident response and driven durable... 
    Suggested

    Practice by Numbers

    Eastern, KY
    3 days ago
  •  ...Cloudflare, GitHub Actions, PostgreSQL, Redis/BullMQ, Node.js/NestJS, Datadog, TypeScript, React, SQL Position: Senior Site Reliability Engineer Engagement period: Ongoing Interview timeline: ASAP Interview process: 1) CV review 2) Interview with our CTO 3)... 
    Suggested
    Contract work
    Immediate start

    Devspace

    Eastern, KY
    1 day ago
  •  ...intuitive solutions that provide the power of computing without the complexity of programming. InRule Technology is seeking a Site Reliability Engineer to join our growing team of experts operating our decision intelligence platform. You'll design, implement, and support... 
    Suggested
    Work at office
    Remote work
    Worldwide
    Flexible hours
    Shift work

    Rippling

    Eastern, KY
    2 days ago
  • $140k - $195k

     ...Improve reliability, observability, service health, incident response, and operational readiness. CodeVertex works across data...  ..., secure systems, and operational clarity matter. The Site Reliability Engineer role helps turn business needs into reliable execution, whether... 
    Suggested
    Remote work

    Codevertex Innovations

    Eastern, KY
    1 day ago
  •  ...democratizing software development by removing traditional barriers to application creation. About the role: Join our Site Reliability Engineering team and help ensure the reliability, scalability, and performance of Replit's infrastructure that serves millions of... 
    Full time
    Temporary work
    Work at office
    Worldwide
    Flexible hours

    Replit

    Eastern, KY
    1 day ago
  •  ...remoteseniorfull-time# Senior Site Reliability EngineerLaravelRemote / Argentina, Brazil, Denmark, Portugal, United Kingdom Salary not disclosedlaravelphpdockerawskubernetesdescription##...  ...their dreams. We are looking for a Senior Site Reliability Engineer to help us scale that mission by ensuring our global... 
    Remote work

    LaraBench

    Eastern, KY
    1 day ago
  •  ...profitable developer-tooling company whose product is used by engineering teams at thousands of software companies for application...  ...well-resourced group of nine. As Senior SRE you will lead reliability initiatives across the platform — from defining and driving SLOs... 

    Kovoro

    Eastern, KY
    1 day ago
  •  ...world, Blackpoint Cyber is in hyper-growth mode, fueled by a recent $190m series C round. SUMMARY We're hiring a Senior Site Reliability Engineer to design, implement, and maintain our cloud and on-premise infrastructure and CI/CD pipelines, with a focus on automation... 
    Local area

    Blackpoint Cyber

    Eastern, KY
    1 day ago
  •  ...Engineering, Product, Design, and Marketing Engineering Compensation ~ Zone 1 Base Pay: $214K – $260K Superhuman offers...  ...role will be responsible for building software to ensure the reliability of our back-end systems, working with engineers who develop them... 
    Worldwide
    Home office
    Flexible hours

    Superhuman

    Eastern, KY
    3 days ago
  • $110k - $145k

     ...global, with headquarters in Denver, Colorado, and offices across the U.S., Canada, and India. We are seeking a Senior Site Reliability Engineer to own the reliability, scalability, performance, and operational integrity of critical production services. This role is... 
    Flexible hours

    Hirebridge

    Eastern, KY
    1 day ago
  •  ...Participated in an on-call rotation, responding to incidents, troubleshooting under pressure, and driving postmortems to improve system reliability over time Design systems with security in mind, applying principles like least privilege and threat modeling Bring a strong... 
    Work at office
    Remote work
    Home office

    Arbitrum Inc

    Eastern, KY
    3 days ago
  • $107.9k - $195.05k

     ...The Digital Sector at Leidos currently has an opening for a Site Reliability Engineer (SRE) / Senior Cloud Engineer to work in our Baltimore, Maryland office. This is an exciting opportunity to use your experience helping the Center for Medicare and Medicaid Services (... 
    Contract work
    Work at office

    Koitecc Solutions

    Eastern, KY
    1 day ago
  • $148.5k - $223.9k

     ...Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in San Francisco. Working closely with counterparts in the Infrastructure and R&D organizations, this organization provides a global team of engineers monitoring cloud service... 
    Worldwide
    Weekend work

    Salesforce.Com Inc

    Eastern, KY
    1 day ago
  • $110k - $145k

     ...operations. You will liaise with product and engineering teams to ensure applications and...  ...feedback loop for platform and product reliability. The ideal candidate is a solutions-oriented...  ...experience as a platform engineer, site reliability engineer, systems engineer... 
    Work experience placement

    CoSM

    Eastern, KY
    1 day ago
  •  ...Discover exciting DevOps job opportunities and connect with 28,396 DevOps professionals. The Senior Site Reliability Engineer role at Jobicy is designed for experienced professionals who are passionate about enhancing system reliability and operational efficiency. The... 
    Remote work
    Flexible hours

    DevOpsChat

    Eastern, KY
    1 day ago
  •  ...% uptime. You'll own SLOs, incident response, and production reliability for a system that processes millions of identity verifications...  ...Sentry error tracking, structured logging Implement chaos engineering practices to proactively identify failure modes Optimize... 
    Remote work

    Xident B.V.

    Eastern, KY
    1 day ago
  •  ...As a Senior Site Reliability Engineer on our cloud engineering team, you'll keep our production environment healthy, secure, and running smoothly. This is an operations-focused role: you'll own the day-to-day administration of our AWS accounts and databases, backup posture... 
    Work experience placement

    MeridianLink

    Eastern, KY
    1 day ago
  • $114.4k - $171.6k

    ## Site Reliability Engineer IIApply: Hybrid: Reston: Full time: Posted Today: RP1038765At F5, we strive to bring a better digital world to life. Our teams empower organizations across the globe to create, secure, and run applications that enhance how we experience our... 
    Full time

    F5 Networks

    Eastern, KY
    1 day ago
  • $180k - $230k

     ...Acceleration Job Description We're looking for a Senior SRE to own the reliability, scalability, and observability of our production systems. You'll work closely with platform and data engineering to keep high-throughput, data-intensive services running at the... 
    Work at office
    Local area
    Immediate start
    Remote work
    3 days per week

    GridCARE

    Eastern, KY
    1 day ago
  •  ...A senior Site Reliability Engineer will join an established infrastructure function responsible for highly available, security-conscious cloud systems supporting complex business-critical workloads. You’ll take significant ownership of reliability, scalability, and... 
    Full time
    Remote work

    Motion Recruitment Partners LLC

    Eastern, KY
    1 day ago
  • $90k - $100k

     ...Site Reliability Engineer The Opportunity We are looking for a highly capable engineer to join our Platform and Site Reliability engineering team. You will be responsible for building, maintaining and operating the infrastructure platform on which all Ookla services... 
    Flexible hours

    Ookla

    Eastern, KY
    3 days ago
  • $104.9k - $174.7k

     ...Technology Senior Site Reliability Engineer II The SRE role is responsible for improving the reliability, availability, performance, and operational quality of production systems. This role provides technical input into project plans, schedules, methodologies, and... 
    Temporary work
    Local area

    LexisNexis Risk Solutions

    Eastern, KY
    4 days ago
  •  ...Quarterhill is seeking a Senior Site Reliability Engineer (SRE) to join our growing team. This role is an exciting opportunity to contribute to the reliability and performance of smart transportation systems, including a next-generation, cloud-native tolling platform that... 
    Local area

    Electronic Transaction Consultants

    Eastern, KY
    1 day ago
  •  ...knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who...  ...knowledge with their teammates. ABOUT THE ROLE: As a Site Reliability Engineer focused on campus reliability, you will design what... 
    Night shift

    SpaceX

    Eastern, KY
    3 days ago
  •  ...Own reliability for client platforms: SLOs, observability and incident practice. You will join a delivery team working directly with client engineers. We do not run a bench — everyone is on real work from week one. What you will do Define and track SLOs... 

    RTS RTech Solutions

    Eastern, KY
    2 days ago
  • $135k - $160k

     ...About the Role We are looking for a Site Reliability Engineer to help us evolve and safeguard the infrastructure powering healthcare experiences for millions of patients, all while keeping operational toil to a minimum for the Fabric tech community. What You’ll Do... 

    Fabric Labs

    Eastern, KY
    3 days ago
  •  ...practical to use as money at a global scale. Our team includes engineers and researchers from organizations such as Blockstream,...  ...and blockchain projects. Role Overview As a Senior Site Reliability Engineer (SRE), you will lead the design, scalability, and reliability... 

    Alpen Labs Inc.

    Eastern, KY
    1 day ago
  • $150k - $220k

     ...Senior Site Reliability EngineerJob detailsDepartment / EngineeringRemoteFull-time$150,000 USD - $220,000 USD## About UsMetaRouter is a customer...  ...architecture.## About The RoleAs a Senior Site Reliability Engineer, you own significant pieces of our infrastructure and... 
    Full time
    Remote work

    Deel

    Eastern, KY
    2 days ago
  •  ...match. The role We're looking for a Senior SRE to own the reliability, scalability, and operational posture of Satsuma's multi-...  ...using AI-assisted development workflows Partner closely with engineering on reliability reviews and architecture decisions ~5-8... 

    Satsuma AI, Inc.

    Eastern, KY
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!