Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Site Reliability Engineer

Full-time

jobgether

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Site Reliability Engineer based in United States.

This is an opportunity to join a critical AI Hardware SRE team responsible for the reliability of next-generation dedicated AI infrastructure.
You will help scale and optimize high-density hardware and software environments across regional data centers.
The role combines automation, observability, infrastructure engineering, networking, and real-time incident response.
You will build Python-based tooling, infrastructure-as-code utilities, telemetry pipelines, and intelligent monitoring solutions.
Your work will directly improve uptime, performance, scalability, and operational efficiency for business-critical systems.
You will collaborate with engineering teams, infrastructure vendors, and field technicians to solve complex reliability challenges.
This role is ideal for an experienced SRE who thrives on ownership, ambiguity, automation, and production-scale infrastructure.

Accountabilities

  • Develop and scale robust Python-based tooling, infrastructure-as-code utilities, and automation frameworks to eliminate operational toil and streamline fleet-wide provisioning.
  • Build automated workflows and API integrations across corporate ticketing systems to accelerate resolution of hardware and network incidents.
  • Apply modern AI and LLM-based development tools to improve technical execution, automate scripting, and evaluate complex infrastructure systems.
  • Work with advanced private cloud and compute technologies to improve availability, latency, scalability, and overall health across high-density hardware environments.
  • Design and implement telemetry pipelines, Prometheus and Grafana dashboards, and AI-driven anomaly detection for bare-metal and virtualized infrastructure.
  • Define operational KPIs, monitoring standards, telemetry baselines, alerting thresholds, and operational readiness criteria for new services and infrastructure deployments.
  • Participate in a 24x7x365 on-call rotation, leading real-time incident response and managing high-severity service disruptions through automated PagerDuty and Slack workflows.
  • Develop detailed technical runbooks, lead incident response bridges, and drive blameless post-mortems that identify systemic improvements and prevent recurring issues.
  • Partner with infrastructure vendors and coordinate on-site field technicians to support hardware reliability, break-fix activities, and uptime objectives.
  • Collaborate across engineering and infrastructure teams to identify reliability gaps, establish best practices, and deliver production-grade solutions to ambiguous technical challenges.

Requirements

  • 5+ years of relevant Site Reliability Engineering, infrastructure engineering, systems engineering, or related experience, along with a Bachelor’s degree in Computer Science or a related technical field.
  • Exceptional proficiency in Python and experience developing scalable operational tooling, API integrations, automation frameworks, and infrastructure utilities.
  • Hands-on experience with modern observability technologies such as Prometheus, Grafana, OpenTelemetry, and Loki, as well as familiarity with time-series monitoring and telemetry systems.
  • Strong understanding of advanced networking concepts, including high-bandwidth routing and switching, BGP, and dual-stack IPv4/IPv6 environments.
  • Experience designing and launching new services with clear operational readiness requirements, telemetry baselines, monitoring strategies, and alerting thresholds.
  • Extensive experience creating technical runbooks, leading complex incident response processes, and conducting comprehensive, blameless post-mortems.
  • Strong understanding of distributed infrastructure, high-density compute environments, private cloud technologies, and large-scale content or infrastructure delivery challenges.
  • Ability to leverage AI-assisted development tools and LLM-based approaches to accelerate engineering workflows and solve technical problems effectively.
  • Proven ability to take ownership of ambiguous and complex technical challenges, coordinate cross-functional teams, and drive solutions through to production.
  • Strong communication and collaboration skills, with the ability to work effectively with engineering teams, vendors, and field operations.
  • Willingness to participate in a 24x7x365 on-call rotation and respond effectively to high-severity production incidents.

Benefits

  • Comprehensive benefits designed to support employee health, well-being, financial security, and life beyond work.
  • Flexible working options that allow employees to work from home, in an office, or through a combination of both, depending on role and business needs.
  • Opportunity to work on cutting-edge AI hardware, private cloud, distributed infrastructure, and edge technologies.
  • Exposure to large-scale, business-critical systems serving global digital experiences.
  • Collaborative environment with opportunities to work alongside experienced infrastructure, engineering, and technology professionals.
  • Opportunities to develop expertise in SRE, observability, automation, AI-assisted engineering, networking, and high-density compute.
  • Support for professional growth and continued development within a technology-focused environment.

How Jobgether works:

We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.

We appreciate your interest and wish you the best!

Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.

#LI-CL1

We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.

Vacancy posted 7 days ago
Similar jobs that could be interesting for youBased on the Senior Site Reliability Engineer in United States vacancy
  •  ...Evaluate applications, platforms, and vendors to assess resiliency, reliability, and operational risk.Design and implement processes that...  ...and reliability tooling.Actively participate in reliability engineering and resilience communities of practice, contributing to... 
    Senior
    Full time

    Vanguard

    Charlotte, NC
    1 day ago
  •  ...TechMContact: Meghana GorusuCompany: SRI Tech SolutionsJob Title: Senior Site Reliability EngineerLocation: Plano , TX (remote)Years of Experience: 8...  ...are seeking a highly skilled Senior Site Reliability Engineer (SRE) to join our dynamic team. The ideal candidate will... 
    Senior
    Remote work

    SRI Tech

    Plano, TX
    3 days ago
  • About the RoleWe’re looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You’ll partner with engineers and data scientists to build, automate, and maintain the infrastructure... 
    Senior

    Alembic

    San Francisco, CA
    3 days ago
  • $170k - $220k

    Who We're Looking ForWe’re looking for a hands-on, high-agency Site Reliability Engineer to help shape and scale the reliability layer of our stack. You'll own the release pipeline end-to-end — managing daily releases, weekly deploys, and hotfixes — while also automating... 
    Senior

    Supio

    Seattle, WA
    3 days ago
  • Inspire Brands is hiring two Senior Site Reliability Engineers to help build and scale reliable, resilient, and observable systems supporting high-traffic, customer-facing digital platforms. These role blends software engineering, systems thinking, and operational excellence... 
    Senior
    Worldwide

    Inspire Brands

    Atlanta, GA
    2 days ago
  • IXL Learning, developer of personalized learning products used by millions of people globally, is seeking a Senior Site Reliability Engineer to join our team, and help maintain the reliability and optimal performance of our products. We are seeking engineers with a passion... 
    Senior
    Work at office
    Immediate start

    IXL Learning

    Raleigh, NC
    12 hours ago
  • Senior Site Reliability Engineer ILocationSan Jose, Costa Rica - RemoteSummary of roleOwn availability, the most important product feature, by continually striving for sustained operational excellence of Sumo’s planet-scale observability and security products. Work with... 
    Senior
    Flexible hours

    Sumo Logic

    San Jose, CA
    1 day ago
  • $104.9k - $174.7k

    About the role:A FinOps Site Reliability Engineer (SRE) bridges the gap between engineering, operations, and financial governance by embedding cost optimization into infrastructure design, automation, monitoring, and operational processes. A FinOps SRE proactively identifies... 
    Senior
    Full time
    Local area

    RELX Group

    Boca Raton, FL
    2 days ago
  •  ...work from home day is currently Tuesday.Engineering at Lambda is responsible for building and...  ...and networking teams to improve service reliability and deployment workflowsDeploy and...  ...rotationYouHave 5+ years of experience in Site Reliability Engineering, Production Engineering... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    12 hours ago
  • $152.6k - $191.5k

     ...responsible for partnering with leaders across engineering and technology to define objective reliability goals for services. Key responsibilities include...  ...and continuous improvement.Position Summary:The Senior GCP Site Reliability Engineer acts as an advanced senior... 
    Senior
    Full time
    Work at office
    Day shift

    Bank of America

    Charlotte, NC
    1 day ago
  •  ...Lambda’s designated work from home day is currently Tuesday.Engineering at Lambda is responsible for building and scaling our cloud offering...  ...and SLIs for Kubernetes services, workloads, and platform reliability.You6+ years of experience in a SRE, operations engineer, or... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    2 days ago
  • $168k - $270.25k

     ...phenomenal people like you to help us accelerate the next wave of artificial intelligence.Join our team at NVIDIA as a Senior Site reliability engineer focused on HPC storage and play a crucial role in designing, implementing, and optimizing on-prem High-Performance... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $152.5k - $205k

     ...flexible work environment where new ideas are encouraged and everyone is a stakeholder.What you’ll be responsible for:As a Senior Site Reliability Engineer on Circle’s platform team, you’ll design, build, and operate the secure, scalable platform infrastructure behind... 
    Senior
    Flexible hours

    Circle

    San Francisco, CA
    2 days ago
  • $210k - $230k

    GovCIO is currently hiring for a Senior Site Reliability Engineer (SRE) to design, implement, and maintain highly available, scalable, and resilient infrastructure systems. The ideal candidate will bridge the gap between development and operations, focusing on automation... 
    Senior
    Currently hiring
    Remote work

    Govcio

    Arlington, VA
    12 hours ago
  • $160k - $240k

     ...millions of times a day - quickly, reliably, and securely. Any time you...  ...at Fiserv.Job TitleSenior Site Reliability EngineerWhat does a successful Site Reliability Engineer do at Fiserv?You will join our...  ...operations or DevOps at a mid-to-senior level.Strong shell scripting... 
    Senior
    Full time

    Fiserv

    Sunnyvale, CA
    2 days ago
  • $174k - $252k

     ...systems by pushing for changes that improve reliability and velocity.Practice sustainable...  ...:Bachelor’s degree in Computer Science, Engineering, a related field, or equivalent practical...  ...degree in Computer Science or Engineering.Site Reliability Engineering (SRE) is what you... 
    Senior

    Google

    Cambridge, MA
    3 days ago
  • $15k

     ...benefits packages, technology talks by our experts, a beautiful modern office, daily catered lunches, and more.As a Senior Cluster Site Reliability Engineer (SRE), you will help scale our research compute cluster to meet our growing needs, and you will leverage... 
    Senior
    Work at office
    Local area
    Remote work

    The Voleon Group

    Berkeley, CA
    1 day ago
  • LeanData helps the world’s fastest-growing companies automate, simplify, and accelerate revenue.We are looking for a Senior Site Reliability Engineer to lead the strategic evolution of our cloud infrastructure. Reporting directly to the SVP of Engineering, this role is... 
    Senior
    Full time
    Work at office
    2 days per week

    LeanData

    Santa Clara, CA
    3 days ago
  •  ...professionalism. We are seeking an experienced AWS solution design engineer/architect to join our infrastructure cloud team. The...  ...product features efficiently and confidently them into production.As Senior SRE, you will be responsible for providing leadership, design and... 
    Senior

    Black Knight Financial Services

    Jacksonville, FL
    2 days ago
  • We are looking for a Senior or Staff level Site Reliability Engineer to strengthen the reliability, scalability, and operational maturity of our platform in San Francisco, California. This role will focus on improving service health, refining observability, and partnering... 
    Senior

    Robert Half

    San Francisco, CA
    3 days ago
  • $104.9k - $174.7k

    Are you passionate about improving reliability, scalability, and resilience in complex database...  ....Own prioritization of reliability engineering tasks within team backlogs.Lead incident...  ...a Service (IaaS).Background in DevOps, site reliability engineering practices, or related... 
    Senior
    Full time
    Local area

    RELX Group

    Texas
    12 hours ago
  • $160k - $200k

     ...data, ideally using promQLKey Responsibilities:Mentor and evangelize on observability best practices, SLIs/SLOs, and reliability culture across engineering teams. Contributing to and maintaining Tulip's triage & remediation processes as a player / coachPerform incident... 
    Senior
    Temporary work
    Work at office
    Local area
    Flexible hours
    3 days per week

    Tulip Interface

    Somerville, MA
    4 days ago
  •  ...and foster a dynamic work environment where new ideas thrive. Are you ready to join our team and make an impact?As a Senior Site Reliability Engineer at TeamViewer, you’ll be a key player in ensuring the reliability, scalability, and performance of our Azure-based SaaS... 
    Senior
    Temporary work
    Casual work
    Worldwide

    TeamViewer

    Austin, TX
    4 days ago
  • $267k - $356k

     ...day is currently Tuesday.Lambda's Storage Engineering team is the backbone behind our world-...  ...workloads in the industry, which means reliability and performance aren't just goals—they're...  ...defined storage across new and existing sites using tools such as Ansible, Jenkins etc... 
    Senior
    Work experience placement
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    1 day ago
  •  ...candidates that are particularly strong in a few areas, and have some interest and capabilities in others.About the Role:As a Site Reliability Engineer, you’ll join the global Platform SRE team responsible for building, operating, and scaling Kong’s multi-region SaaS... 
    Senior
    Temporary work

    Kong

    Washington DC
    3 days ago
  • $80k - $140k

    Job DescriptionRBC Wealth Management Technology is seeking a Senior Site Reliability Engineer to join its Wealth Management SRE Team. This team is responsible for ensuring the performance, availability, resilience, and operational excellence of critical applications and... 
    Senior
    Full time
    Flexible hours
    Shift work

    Royal Bank of Canada

    Minneapolis, MN
    4 days ago
  • $119.8k - $234.7k

     ...yearEmployment type: Full-TimeWork site: 3 days / week in-officeRole type: Individual...  ...EngineeringDiscipline: Site Reliability EngineeringCompany: MicrosoftOverviewMicrosoft...  ...’s most demanding workloads. As a Senior Site Reliability Engineer, you will lead reliability... 
    Senior
    Ongoing contract
    Local area
    3 days per week

    Microsoft

    Redmond, WA
    4 days ago
  •  ...Georgia, and serves customers in more than 35 countries worldwide.Position OverviewWe are seeking a highly experienced Senior Site Reliability Engineer (Unified Observability) to lead the design, implementation, and operational maturity of the F1 Next Generation... 
    Senior
    Full time
    Worldwide
    Flexible hours

    NCR

    Atlanta, GA
    2 days ago
  • $101k - $161k

     ...several prestigious awards, such as Best Engineering Team, Best Company for Diversity,...  ...DescriptionWho You'll Work WithWe’re looking for Site Reliability Engineers to join our growing Arista’s...  ...: EngineeringExperience level: Mid-Senior LevelIndustry: Computer Networking
    Senior

    Arista Networks

    Santa Clara, CA
    1 day ago
  • The Senior Site Reliability Engineer is responsible for improving the reliability, availability, scalability, and operational excellence of our critical infrastructure platforms and services. This role partners closely with Engineering, Security, and Infrastructure teams... 
    Senior
    Full time
    Work at office
    Local area

    Castleton Commodities International

    Stamford, CT
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!