Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Site Reliability Engineer

$20k

ServiceTitan

Ready to be a Titan?

We're looking for a Senior Site Reliability Engineer to join our Site Reliability & Infrastructure Engineering team. We run entirely on the cloud, and this team owns the reliability and health of the applications running on top of it - designing the signals that tell us when something's wrong, and building the systems that keep ServiceTitan running better, faster, and cheaper as we scale.

We make a huge impact on thousands of companies in the U.S. and abroad by enabling them to be more efficient and effective at running their business. Our Site Reliability and Infrastructure Engineering team centralizes the concerns of measurement and guidance so every engineer can improve availability and efficiency in their own area of the ServiceTitan cloud. We have a cultural foundation built on diversity, inclusion, and innovation, and we want you and your ideas to thrive at ServiceTitan. Come join us.

What You'll Do
  • Participate in an on-call rotation, using runbooks and playbooks to diagnose and resolve production issues (e.g., adjusting Horizontal Pod Autoscaler rules in response to load).
  • Design, build, and maintain observability dashboards and alerting grounded in Service Level Indicators (SLIs) and Service Level Objectives (SLOs).
  • Operate and improve our Kubernetes-based compute platform, which runs the large majority of our infrastructure.
  • Work across cloud networking and infrastructure (Azure/AWS) to support reliable, scalable systems.
  • Investigate and resolve production incidents, including root-cause analysis and follow-up remediation work.
  • Partner with product engineering teams to review architecture and infrastructure decisions before they ship.
  • Build and maintain automation that reduces manual, repetitive operational work across the team.
  • Write and maintain runbooks and documentation so on-call knowledge is shared across the team, not siloed with one person.
  • Help define non-functional requirements - scalability, availability, performance - for new systems as they're designed.
  • Collaborate across engineering teams to adopt best practices in reliability and observability.
  • Contribute to CI/CD pipelines and help teams ship changes safely and quickly.
What You'll Bring
  • Kubernetes (must-have): strong, hands-on understanding of Kubernetes as a system.
  • SRE principles: practical experience with SLIs, SLOs, and error budgets - able to speak to how you've defined and monitored these on real systems, not just definitions.
  • Cloud engineering & networking: solid grounding in AWS or Azure, including networking fundamentals (subnetting, IP addressing).
  • Observability: deep experience with at least one modern observability stack (OpenTelemetry, Prometheus, Grafana, Datadog, or Elasticsearch) and the ability to translate that understanding across tools.
  • CI/CD: strong understanding of a CI/CD system - GitHub Actions preferred, but TeamCity, Azure DevOps, or GitLab CI experience is acceptable.
  • Strong programming skills with the ability to build web applications - ideally with solid working knowledge of .NET and ASP.NET. We're also open to strong Python (Flask, FastAPI) or Java (Spring) backgrounds. The coding assessment will be tailored to whichever language/framework you're most comfortable in.
  • Experience with distributed systems and their common failure modes (retries, timeouts, cascading failures).
  • Strong production troubleshooting skills - comfortable diagnosing issues under pressure.
  • 8-10+ years of relevant hands-on experience.
  • Nice-to-have: database experience (not mandatory - databases are monitored by the same team, not owned individually).
About You

You're someone who enjoys being directly accountable for the reliability of a business-critical, large-scale enterprise system. You're comfortable guiding and making decisions with limited information, and capable of operating within the trade-offs between solving for immediate needs versus bigger-scale solutions. You feel rewarded by developing an operability culture in a quickly growing and changing environment, and you're comfortable owning a wide and diverse set of problem areas.

Be Human With Us:
Being human isn't about checking every box on a list. It's about the experiences we have, people we meet, and the perspectives we share. So, if you have the skills but are hesitant to apply because of your background, apply anyway. We need amazing people like you to help us challenge the conventional and think differently about the problems that we're solving. We're in this together. Come be human, with us.


Use of AI Technology:

We use technology, including automated and AI-assisted tools, to support certain aspects of our recruitment process. These tools are designed to improve efficiency and enhance the candidate experience. AI tools are not used to make hiring decisions; all hiring decisions are made by our hiring teams.

What We Offer:

When you join our team, you're not just accepting a job. You're making a career move. Here's how we'll support you in doing some of the most impactful work of your career:
  • Flextime, recognition, and support for autonomous work : Flexible time off with ample learning and development opportunities to continue growing your career. We offer a comprehensive onboarding program, leadership training for Titans at all levels, and other programs and events. Great work is rewarded through Bonusly, peer-nominated awards, and more.
  • Holistic health and wellness benefits: Company-paid medical, dental, and vision (with 100% employer paid options and 90% coverage for dependents), FSA and HSA, 401k match, and telehealth options including memberships to One Medical.
  • Support for Titans at all stages of life: Parental leave and support, up to $20k in fertility services (i.e. IUI and IVF), surrogacy, and adoption reimbursement, on demand maternity support through Maven Maternity, free breast milk shipping through Maven Milk, pet insurance, legal advisory services, financial planning tools, and more.

At ServiceTitan, we celebrate individuality and uniqueness. We believe that the convergence of fresh perspectives and experiences from all walks of life is what makes our product and culture so great. We strongly encourage people from underrepresented groups to apply. We do not discriminate against employees based on race, color, religion, sex, national origin, gender identity or expression, age, disability, pregnancy (including childbirth, breastfeeding, or related medical condition), genetic information, protected military or veteran status, sexual orientation, or any other characteristic protected by applicable federal, state or local laws.
ServiceTitan is committed to fair and equitable compensation for all of our employees. We thoughtfully consider a wide range of factors when determining individual compensation, which may change over time. We comply with all applicable minimum wage laws. For candidates in the United States, the good faith salary ranges estimate for this role isZone 1: $147,600 USD - $221,400 USD Applicable for: CA, CT, DC, MD, MA, NJ, NY, VA, and WAZone 2: $137,900 USD - $206,900 USD Applicable for: All other US locations.International Compensation for candidates residing outside the United States will vary by location and will be discussed during the hiring process. Actual compensation within a range is determined by factors including relevant experience, skill set, qualifications, and performance. In addition to base salary, our total compensation package includes an annual bonus, equity, and a holistic suite of benefits.
Vacancy posted 1 hour ago
Similar jobs that could be interesting for youBased on the Senior Site Reliability Engineer in United States vacancy
  • We are looking for an experienced Senior Site Reliability Engineer to join our team and drive the reliability and functionality of our critical infrastructure and applications. In this role, you will be a key contributor, taking ownership of designing and architecting... 
    Senior
    Full time
    Flexible hours

    Oracle

    Nashville, TN
    1 day ago
  •  ...dynamic team with a broad knowledge of how Oracle’s cloud platform works. You’ll partner with customer support, service owners, and engineering teams around the globe to ensure high-quality service for customers. Note - this role is not a Monday to Friday core hours... 
    Senior
    Full time
    Monday to Friday
    Flexible hours
    Shift work
    Night shift

    Oracle

    Reston, VA
    2 days ago
  • $160k - $180k

     ...can have a big impact. See Arkestro in action at arkestro.com. About the Role Arkestro is hiring for a Senior SRE Engineer to manage our performance and reliability for our software platform and infrastructure. The right candidate will own and develop our... 
    Senior
    Full time
    Local area
    Remote work

    Arkestro

    United States
    3 days ago
  • $158.5k - $172k

     ...exceptional value they deserve. About The Opportunity As a Senior Engineer on the Runtime Automation team, you will design, automate,...  ...environment. This is a high-impact position driving continuous reliability, deep system optimization, and automation across our entire... 
    Senior
    Full time
    Work at office
    3 days per week

    Wonder

    Chicago, IL
    1 day ago
  • $127k - $249k

    THE TEAM Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational...  ..., alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper).... 
    Senior
    Full time
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    United States
    3 days ago
  • $90k - $180k

     ...medicines. Our 115,000 colleagues serve people in more than 160 countries. JOB DESCRIPTION: About the Role This Senior Site Reliability Engineer position works on-site out of our Sylmar, CA or Sunnyvale, CA location in the Cardiac Rhythm Management Division.... 
    Senior
    Full time
    Remote work
    Shift work

    Abbott

    Sunnyvale, CA
    3 days ago
  •  ...our Series B and have grown 800% over the last 12 months. Engineering at Ivo Engineers at Ivo are inventors. Ivo was first-to-...  ...expect us to hit our SLAs. What? We’re looking for an Senior Site level Reliability Engineer as part of Infrastructure team to: Own uptime,... 
    Senior
    Contract work
    Work at office
    Remote work
    Visa sponsorship
    Relocation package
    Flexible hours

    Ivo Inc.

    San Francisco, CA
    6 hours ago
  • $232k - $319k

     ...to help us continue to scale the service with great people and reliable, cost-effective, and efficient infrastructure, processes, and...  ...with self-service Accelerate the velocity of SRE and product engineering by developing robust platforms, powerful tooling, and... 
    Senior
    Permanent employment
    Full time
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    1 day ago
  • $160k - $180k

     ...Socure is seeking a Site Reliability Engineer in New York to enhance our identity trust infrastructure. In this role, you will take full ownership of AWS and Kubernetes platforms, ensuring high reliability and operability. The ideal candidate will possess extensive experience... 
    Senior

    Socure Inc

    New York, NY
    19 hours ago
  •  ...A leading technology firm is in search of a Senior Wireless Network Site Reliability Engineer to manage and enhance their wireless network infrastructure. The ideal candidate has over 8 years of experience in wireless network operations and a strong background in wireless... 
    Senior

    TechDigital Group

    Santa Clara, CA
    19 hours ago
  •  ...The Depository Trust & Clearing Corporation (DTCC) is seeking a Senior Application Support Engineer (SRE) to enhance reliability, scalability, and performance of mission-critical applications. You will apply SRE principles across engineering, infrastructure, and operations... 
    Senior

    The Depository Trust & Clearing Corporation

    Dallas, TX
    20 hours ago
  •  ...Karsun Solutions, LLC is seeking a Site Reliability Manager to lead a multi-disciplinary team responsible for reliability, security, and platform lifecycle across AWS-based services. The role emphasizes collaboration, observability, and continuous improvement in a client... 
    Senior

    Karsun Solutions

    New York, NY
    20 hours ago
  •  ...The Consulting Solutions is seeking an experienced Senior / Staff Engineer for our SRE, InfraSec team in Seattle. The role involves leading the security of cloud-based infrastructure, mentoring a team of SREs, and collaborating with other engineering teams to ensure high... 
    Senior
    Remote work

    The Consulting Solutions

    New York, NY
    19 hours ago
  •  ...optimize production infrastructure across CI/CD, cloud deployments, and security. You will collaborate with our internal product and engineering teams to keep services scalable, secure, and highly available. The role emphasizes GitHub Actions, Terraform, Vercel, AWS core... 
    Senior
    Remote work

    Remote Leverage

    New York, NY
    20 hours ago
  •  ...PVH (Tommy Hilfiger/Calvin Klein) seeks a Senior Software Engineer to own the reliability and performance of our Kubernetes-based data platform across multi-region deployments. You will design scalable infrastructure, optimize deployment pipelines, and strengthen security... 
    Senior

    PVH (Tommy Hilfiger/Calvin Klein)

    Livingston, NJ
    19 hours ago
  •  ...A global leader in fast food is seeking a Senior Manager for Edge Operations/SRE in Chicago. This pivotal role involves leading edge...  ..., collaborating across teams to ensure high availability and reliability of the platform. Candidates should have 10+ years in infrastructure... 
    Senior

    McDonald's Corporation

    Chicago, IL
    19 hours ago
  • $145k - $165k

     ...A technology solutions firm in Sunnyvale, CA is looking for a highly experienced Site Reliability Engineer (SRE). This role involves maintaining uptime and performance across systems. Exceptional Linux expertise and automation skills in Bash and Python are crucial. Key... 
    Senior

    Bolt Graphics, Inc.

    Sunnyvale, CA
    20 hours ago
  •  ...Karsun Solutions in the DMV area is seeking a Site Reliability Manager to ensure reliability, scalability, and performance of our systems. You will lead a team focusing on Application Reliability, DevSecOps, and Platform Lifecycle Management. The ideal candidate has 1... 
    Senior

    Karsun Solutions

    New York, NY
    19 hours ago
  •  ...The Home Depot is seeking a Senior Software Reliability Engineer to join the Platform Reliability Engineering team, ensuring the resilience, performance, and security of our enterprise Cloud Platform. You will mentor junior engineers, lead incident triage, root cause... 
    Senior

    Home Depot

    Atlanta, GA
    20 hours ago
  • $149.4k - $202k

     ...Noctua Technology is seeking a Senior Software Engineer specializing in Site Reliability Engineering to join their team. This role focuses on the reliability and performance of cloud-native applications, emphasizing Infrastructure as Code and automation. The ideal candidate... 
    Senior
    Remote work

    Noctua Technology

    Virginia, MN
    20 hours ago
  •  ...A technology solutions company is seeking a Senior Databricks Administrator located in Georgia. This role involves designing and optimizing a scalable Databricks platform for AI and ML workloads. The ideal candidate has proven experience with Terraform, strong programming... 
    Senior
    Contract work

    New Era Technology Europe

    Alpharetta, GA
    20 hours ago
  •  ...Motion Recruitment Partners LLC is seeking a Senior Java Applications Administrator / SRE in Dallas to own production Java environments...  ..., and cloud automation. You will optimize performance, drive reliability, and mentor teammates while aligning with enterprise security... 
    Senior

    Motion Recruitment Partners LLC

    Dallas, TX
    20 hours ago
  •  ...Karsun Solutions, LLC seeks a Site Reliability Manager to lead a team focused on application reliability, DevSecOps, and platform lifecycle management. You will drive production reliability, standardize observability with Datadog, and manage incident post-mortems and... 
    Senior

    Karsun Solutions

    Charleston, WV
    20 hours ago
  • $135k - $145k

     ...A medical equipment manufacturing company is hiring a Senior Site Reliability Engineer in Carlsbad, CA. The role involves ensuring system reliability, automating operational tasks, and managing incident response. Candidates should have extensive experience in SRE, strong... 
    Senior

    ATEC Spine

    Carlsbad, CA
    20 hours ago
  •  ...A leading livestream shopping platform is seeking a Senior Software Engineer for the Logistics Platform team. This role focuses on improving logistical data systems, enhancing buyer and seller experience, and fostering collaboration across departments. Ideal candidates... 
    Senior
    Remote work

    Whatnot

    Los Angeles, CA
    1 day ago
  •  ...Passionate about designing and automating cloud platforms, the full-time Senior Azure Site Reliability Engineer will focus on improving reliability through Infrastructure as Code, automation, and observability while managing Azure cloud infrastructure and collaborating... 
    Senior
    Full time
    Remote work

    Virtual Vocations Inc

    United States
    20 hours ago
  •  ...Senior Site Reliability Engineer (SRE) Our client is a global technology consulting and digital solutions company that enables enterprises across industries to reimagine business models, accelerate innovation, and maximize growth by harnessing digital technologies.... 
    Senior
    Local area

    E-Solutions

    Los Angeles, CA
    4 days ago
  • $210k - $240k

     ...Join to apply for the Senior Site Reliability Engineer role at Alembic Technologies This range is provided by Alembic Technologies. Your actual pay will be based on your skills and experience — talk with your recruiter to learn more. Base pay range $210,000.00/yr - $2... 
    Senior
    Full time

    Alembic Technologies

    San Francisco, CA
    1 day ago
  •  ...Overview: Senior Site Reliability Engineer (SRE) Location: Chicago, IL (Onsite) Type: Contract Role Overview: We are seeking a Senior Site Reliability Engineer (SRE) with strong expertise in AWS infrastructure, automation, observability, and production... 
    Senior
    Contract work

    Purple Drive

    Chicago, IL
    2 days ago
  •  ...Resiliency And Reliability Engineer Core Responsibilities: Evaluate applications, platforms, and vendors to assess resiliency, reliability, and operational risk. Design and implement processes that enforce enterprise resiliency and reliability standards. Lead... 
    Senior

    Vanguard

    Charlotte, NC
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!