Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Site Reliability Engineer

AgileEngine

Senior Site Reliability Engineer

AgileEngine is an Inc. 5000 company that creates award-winning software for Fortune 500 brands and trailblazing startups across 17+ industries. We rank among the leaders in areas like application development and AI/ML, and our people-first culture has earned us multiple Best Place to Work awards. If you're looking for a place to grow, make an impact, and work with people who care, we'd love to meet you!

We are looking for a Senior Site Reliability Engineer to support platform reliability, monitoring, and modernization across Kubernetes-based microservices environments with a strong observability focus. You will build and maintain Datadog solutions including dashboards, alerts, APM, metrics, logging, and tracing, integrate observability tooling into AWS and CI/CD pipelines, and automate monitoring and operational tasks using Python.

The role blends software engineering (60–70%) with site reliability engineering (30–40%) and requires JST timezone overlap.

What You Will Do:

  • Support platform reliability, monitoring, and continuous improvement across internal systems.
  • Work in Kubernetes-based environments.
  • Build and maintain observability solutions, with a focus on Datadog.
  • Configure dashboards, alerts, APM, metrics, logging, and tracing.
  • Monitor containerized and microservices-based applications.
  • Integrate observability tools into AWS environments.
  • Integrate observability into CI/CD pipelines.
  • Automate monitoring and operational tasks using scripting (Python preferred).
  • Install and configure Datadog agents and integrations.
  • Manage API keys and secure configurations.
  • Manage user roles and access controls within observability platforms.
  • Lead maintenance efforts and platform improvements while driving reliability, scalability, and performance.

Must Haves:

  • Strong proficiency in Python, JavaScript (Node.js), or Java.
  • Hands-on experience with API integrations (designing, consuming, and integrating).
  • Strong experience working in Kubernetes environments (deployment, operations, monitoring).
  • Experience with Datadog (preferred) or similar tools (Prometheus, Grafana).
  • Ability to configure dashboards, alerts, and APM (tracing, metrics, logging).
  • Experience monitoring containerized/microservices architectures.
  • Hands-on experience with AWS.
  • Experience integrating observability tools into cloud environments.
  • Experience integrating observability into CI/CD pipelines.
  • Ability to automate monitoring and operational tasks using scripting (Python preferred).
  • Upper-intermediate English level.

Nice To Haves:

  • Experience owning and operating an internal engineering platform.
  • Demonstrated ownership of reliability, scalability, and performance.
  • Proven ability to proactively lead maintenance efforts and platform improvements (not just reactive support).
  • Familiarity with Golang.
  • Experience with additional observability tools such as New Relic, Dynatrace, Elastic, or Splunk Observability.

Perks and Benefits:

  • Growth without limits: build your skills through mentorship, internal TechTalks, challenging projects, and a dedicated annual learning budget.
  • Competitive compensation: get recognition that reflects your skills and impact, with regular performance and compensation reviews.
  • Flexibility: work 100% remotely with flexible hours that support focus, autonomy, and a healthy work rhythm.
  • Meaningful, modern projects: build impactful products using modern technologies alongside global teams and leading brands.
  • Collaborative culture: join a supportive environment with zero micromanagement where ideas are welcomed and contributions are recognized.
  • Well-being & support: access local well-being programs and people-focused support tailored to your location.
Vacancy posted 17 hours ago
Similar jobs that could be interesting for youBased on the Senior Site Reliability Engineer in United States vacancy
  •  ...Evaluate applications, platforms, and vendors to assess resiliency, reliability, and operational risk.Design and implement processes that...  ...and reliability tooling.Actively participate in reliability engineering and resilience communities of practice, contributing to... 
    Senior
    Full time

    Vanguard

    Wayne, PA
    2 days ago
  •  ...TechMContact: Meghana GorusuCompany: SRI Tech SolutionsJob Title: Senior Site Reliability EngineerLocation: Plano , TX (remote)Years of Experience: 8...  ...are seeking a highly skilled Senior Site Reliability Engineer (SRE) to join our dynamic team. The ideal candidate will... 
    Senior
    Remote work

    SRI Tech

    Plano, TX
    4 days ago
  • About the RoleWe’re looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You’ll partner with engineers and data scientists to build, automate, and maintain the infrastructure... 
    Senior

    Alembic

    San Francisco, CA
    4 days ago
  • Inspire Brands is hiring two Senior Site Reliability Engineers to help build and scale reliable, resilient, and observable systems supporting high-traffic, customer-facing digital platforms. These role blends software engineering, systems thinking, and operational excellence... 
    Senior
    Worldwide

    Inspire Brands

    Atlanta, GA
    3 days ago
  • $170k - $220k

    Who We're Looking ForWe’re looking for a hands-on, high-agency Site Reliability Engineer to help shape and scale the reliability layer of our stack. You'll own the release pipeline end-to-end — managing daily releases, weekly deploys, and hotfixes — while also automating... 
    Senior

    Supio

    Seattle, WA
    4 days ago
  • Senior Site Reliability Engineer ILocationSan Jose, Costa Rica - RemoteSummary of roleOwn availability, the most important product feature, by continually striving for sustained operational excellence of Sumo’s planet-scale observability and security products. Work with... 
    Senior
    Flexible hours

    Sumo Logic

    San Jose, CA
    2 days ago
  • IXL Learning, developer of personalized learning products used by millions of people globally, is seeking a Senior Site Reliability Engineer to join our team, and help maintain the reliability and optimal performance of our products. We are seeking engineers with a passion... 
    Senior
    Work at office
    Immediate start

    IXL Learning

    Raleigh, NC
    1 day ago
  • $104.9k - $174.7k

    About the role:A FinOps Site Reliability Engineer (SRE) bridges the gap between engineering, operations, and financial governance by embedding cost optimization into infrastructure design, automation, monitoring, and operational processes. A FinOps SRE proactively identifies... 
    Senior
    Full time
    Local area

    LexisNexis Risk Solutions Group

    Boca Raton, FL
    3 days ago
  • $174k - $252k

     ...systems by pushing for changes that improve reliability and velocity.Practice sustainable...  ...:Bachelor’s degree in Computer Science, Engineering, a related field, or equivalent practical...  ...degree in Computer Science or Engineering.Site Reliability Engineering (SRE) is what you... 
    Senior

    Google

    Sunnyvale, CA
    4 days ago
  • $152.6k - $191.5k

     ...responsible for partnering with leaders across engineering and technology to define objective reliability goals for services. Key responsibilities include...  ...and continuous improvement.Position Summary:The Senior GCP Site Reliability Engineer acts as an advanced senior... 
    Senior
    Full time
    Work at office
    Day shift

    Bank of America

    Plano, TX
    2 days ago
  • $104.9k - $174.7k

     ...Data Management. You can learn more about LexisNexis Risk at the link below, About the Role:We are hiring a hands-on Senior Site Reliability Engineer (SRE) to actively build, operate, and improve the reliability of our production systems. This is not a purely advisory... 
    Senior
    Full time
    Work at office
    Local area
    Remote work
    Work from home

    RELX Group

    Buford, GA
    3 days ago
  •  ...work from home day is currently Tuesday.Engineering at Lambda is responsible for building and...  ...and networking teams to improve service reliability and deployment workflowsDeploy and...  ...rotationYouHave 5+ years of experience in Site Reliability Engineering, Production Engineering... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    1 day ago
  • $168k - $270.25k

    NVIDIA is looking for a Senior Site Reliability Engineer (SRE) to join its GeForce Now (GFN) team. SRE at NVIDIA ensures that our internal and external-facing GPU cloud gaming services have reliability and uptime as promised to the users and at the same time enables developers... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    5 days ago
  • LeanData helps the world’s fastest-growing companies automate, simplify, and accelerate revenue.We are looking for a Senior Site Reliability Engineer to lead the strategic evolution of our cloud infrastructure. Reporting directly to the SVP of Engineering, this role is... 
    Senior
    Full time
    Work at office
    2 days per week

    LeanData

    Santa Clara, CA
    4 days ago
  •  ...Lambda’s designated work from home day is currently Tuesday.Engineering at Lambda is responsible for building and scaling our cloud offering...  ...and SLIs for Kubernetes services, workloads, and platform reliability.You6+ years of experience in a SRE, operations engineer, or... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    2 days ago
  • $210k - $230k

    GovCIO is currently hiring for a Senior Site Reliability Engineer (SRE) to design, implement, and maintain highly available, scalable, and resilient infrastructure systems. The ideal candidate will bridge the gap between development and operations, focusing on automation... 
    Senior
    Currently hiring
    Remote work

    Govcio

    Arlington, VA
    5 days ago
  • $168k - $270.25k

     ...phenomenal people like you to help us accelerate the next wave of artificial intelligence.Join our team at NVIDIA as a Senior Site reliability engineer focused on HPC storage and play a crucial role in designing, implementing, and optimizing on-prem High-Performance... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $152.5k - $205k

     ...flexible work environment where new ideas are encouraged and everyone is a stakeholder.What you’ll be responsible for:As a Senior Site Reliability Engineer on Circle’s platform team, you’ll design, build, and operate the secure, scalable platform infrastructure behind... 
    Senior
    Flexible hours

    Circle

    San Francisco, CA
    3 days ago
  •  ...candidates that are particularly strong in a few areas, and have some interest and capabilities in others.About the Role:As a Site Reliability Engineer, you’ll join the global Platform SRE team responsible for building, operating, and scaling Kong’s multi-region SaaS... 
    Senior
    Temporary work

    Kong

    Washington DC
    4 days ago
  • $104.9k - $174.7k

    Are you passionate about improving reliability, scalability, and resilience in complex database...  ....Own prioritization of reliability engineering tasks within team backlogs.Lead incident...  ...a Service (IaaS).Background in DevOps, site reliability engineering practices, or related... 
    Senior
    Full time
    Local area

    LexisNexis Risk Solutions Group

    Texas
    4 days ago
  • $160k - $240k

     ...millions of times a day - quickly, reliably, and securely. Any time you...  ...at Fiserv.Job TitleSenior Site Reliability EngineerWhat does a successful Site Reliability Engineer do at Fiserv?You will join our...  ...operations or DevOps at a mid-to-senior level.Strong shell scripting... 
    Senior
    Full time

    Fiserv

    Sunnyvale, CA
    3 days ago
  • Job Description:Note: Fidelity will not provide immigration sponsorship for this positionThe RoleOur Site Reliability Engineering group within Enterprise Infrastructure combines Operations Excellence with the Development Experience to deliver services at high scale, high... 
    Senior
    Full time

    Fidelity Investments

    Durham, NC
    3 days ago
  • $80k - $140k

    Job DescriptionRBC Wealth Management Technology is seeking a Senior Site Reliability Engineer to join its Wealth Management SRE Team. This team is responsible for ensuring the performance, availability, resilience, and operational excellence of critical applications and... 
    Senior
    Full time
    Flexible hours
    Shift work

    Royal Bank of Canada

    Minneapolis, MN
    17 hours ago
  •  ...Georgia, and serves customers in more than 35 countries worldwide.Position OverviewWe are seeking a highly experienced Senior Site Reliability Engineer (Unified Observability) to lead the design, implementation, and operational maturity of the F1 Next Generation... 
    Senior
    Full time
    Worldwide
    Flexible hours

    NCR

    Atlanta, GA
    3 days ago
  • $166k - $258k

     ...Seattle office a minimum of 4 days/week in order to be considered for this position.Nordstrom is looking for a Senior Engineer 2 to join our Site Reliability Engineering (SRE) team — and we think that person could be you.You'll help build the scalable, reliable, and resilient... 
    Senior
    Full time
    Work at office

    Nordstrom

    Seattle, WA
    3 days ago
  • $148k - $235.75k

     ...see how you can make a lasting impact on the world.Join our team of innovative engineers who are building an AI Data Center AIOps platform that turns raw, high-volume telemetry into reliable, job-centric insights and automation for GPU fleets. We’re hiring a DevOps Engineer... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $90k - $180k

     ...nutritionals and branded generic medicines. Our 115,000 colleagues serve people in more than 160 countries.About the RoleThis Senior Site Reliability Engineer position works on-site out of our Sylmar, CA or Sunnyvale, CA location in the Cardiac Rhythm Management Division.We... 
    Senior
    Remote work

    Abbott

    Sunnyvale, CA
    1 day ago
  •  ...and foster a dynamic work environment where new ideas thrive. Are you ready to join our team and make an impact?As a Senior Site Reliability Engineer at TeamViewer, you’ll be a key player in ensuring the reliability, scalability, and performance of our Azure-based SaaS... 
    Senior
    Temporary work
    Casual work
    Worldwide

    TeamViewer

    Austin, TX
    17 hours ago
  •  ...importance of in-office collaboration and fully intend for the selected candidate for this role to work on site in the specified location(s). As a Senior Site Reliability Engineer within the CET SAvE organization, you will play a critical leadership role advancing the... 
    Senior
    Full time
    Work at office

    The Charles Schwab Corporation

    Southlake, TX
    4 days ago
  • $160k - $200k

     ...data, ideally using promQLKey Responsibilities:Mentor and evangelize on observability best practices, SLIs/SLOs, and reliability culture across engineering teams. Contributing to and maintaining Tulip's triage & remediation processes as a player / coachPerform incident... 
    Senior
    Temporary work
    Work at office
    Local area
    Flexible hours
    3 days per week

    Tulip Interface

    Somerville, MA
    17 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!