Senior Site Reliability Engineer
$20kServiceTitan
Senior Site Reliability Engineer
We're looking for a Senior Site Reliability Engineer to join our Site Reliability & Infrastructure Engineering team. We run entirely on the cloud, and this team owns the reliability and health of the applications running on top of it — designing the signals that tell us when something's wrong, and building the systems that keep ServiceTitan running better, faster, and cheaper as we scale.
We make a huge impact on thousands of companies in the U.S. and abroad by enabling them to be more efficient and effective at running their business. Our Site Reliability and Infrastructure Engineering team centralizes the concerns of measurement and guidance so every engineer can improve availability and efficiency in their own area of the ServiceTitan cloud. We have a cultural foundation built on diversity, inclusion, and innovation, and we want you and your ideas to thrive at ServiceTitan. Come join us.
What You'll Do
- Participate in an on-call rotation, using runbooks and playbooks to diagnose and resolve production issues (e.g., adjusting Horizontal Pod Autoscaler rules in response to load).
- Design, build, and maintain observability dashboards and alerting grounded in Service Level Indicators (SLIs) and Service Level Objectives (SLOs).
- Operate and improve our Kubernetes-based compute platform, which runs the large majority of our infrastructure.
- Work across cloud networking and infrastructure (Azure/AWS) to support reliable, scalable systems.
- Investigate and resolve production incidents, including root-cause analysis and follow-up remediation work.
- Build and operate AI agents and automation that take on manual, repetitive SRE work directly — not just tools that assist a human doing the work.
- Partner with product engineering teams to review architecture and infrastructure decisions before they ship.
- Write and maintain runbooks and documentation so on-call knowledge is shared across the team, not siloed with one person.
- Help define non-functional requirements — scalability, availability, performance — for new systems as they're designed.
- Collaborate across engineering teams to adopt best practices in reliability and observability.
- Contribute to CI/CD pipelines and help teams ship changes safely and quickly.
What You'll Bring
Kubernetes (must-have): strong, hands-on understanding of Kubernetes as a system.
AI-native SRE practice (must-have): hands-on, personal experience using AI tools and agents (e.g., Claude Code, GitHub Copilot Workspace, Cursor, or custom agents built on Claude/MCP servers) to diagnose, automate, and resolve infrastructure and reliability work.
Agentic depth: You've built or operated agents that autonomously monitor infrastructure and take action (e.g., an agent that watches system load and scales, remediates, or escalates without a human in the loop) — not just general AI coding assistance. You can speak concretely to how these agents are actually built and operated (MCP protocol, what a harness is, how to differentiate agent-design strategies rather than just naming tools), to context management (e.g., progressive-disclosure strategies for surfacing the right information without dumping everything into context), and to securing agent actions (scoping authorization, guardrails, what's available in the AI infra ecosystem to enforce it).
SRE application: That agentic work is pointed at infrastructure and reliability problems specifically — you use AI to move at a materially faster pace in the SRE domain, not as a general-purpose coding aid.
SRE principles: practical experience with SLIs, SLOs, and error budgets — able to speak to how you've defined and monitored these on real systems, not just definitions.
Cloud engineering & networking: solid grounding in AWS or Azure, including networking fundamentals (subnetting, IP addressing).
Observability: deep experience with at least one modern observability stack (OpenTelemetry, Prometheus, Grafana, Datadog, or Elasticsearch) and the ability to translate that understanding across tools.
CI/CD: strong understanding of a CI/CD system — GitHub Actions preferred, but TeamCity, Azure DevOps, or GitLab CI experience is acceptable.
Programming: strong programming skills with the ability to build web applications — ideally with solid working knowledge of .NET and ASP.NET. We're also open to strong Python (Flask, FastAPI) or Java (Spring) backgrounds. The coding assessment will be tailored to whichever language/framework you're most comfortable in.
Experience with distributed systems and their common failure modes (retries, timeouts, cascading failures).
Strong production troubleshooting skills — comfortable diagnosing issues under pressure.
8-10+ years of relevant hands-on experience.
Nice-to-have: database experience (not mandatory — databases are monitored by the same team, not owned individually).
About You
You're someone who enjoys being directly accountable for the reliability of a business-critical, large-scale enterprise system. You're comfortable guiding and making decisions with limited information, and capable of operating within the trade-offs between solving for immediate needs versus bigger-scale solutions. You feel rewarded by developing an operability culture in a quickly growing and changing environment, and you're comfortable owning a wide and diverse set of problem areas.
Be Human With Us: Being human isn't about checking every box on a list. It's about the experiences we have, people we meet, and the perspectives we share. So, if you have the skills but are hesitant to apply because of your background, apply anyway. We need amazing people like you to help us challenge the conventional and think differently about the problems that we're solving. We're in this together. Come be human, with us.
Use of AI Technology:
We use technology, including automated and AI-assisted tools, to support certain aspects of our recruitment process. These tools are designed to improve efficiency and enhance the candidate experience. AI tools are not used to make hiring decisions; all hiring decisions are made by our hiring teams.
What We Offer:
- Flextime, recognition, and support for autonomous work: Flexible time off with ample learning and development opportunities to continue growing your career. We offer a comprehensive onboarding program, leadership training for Titans at all levels, and other programs and events. Great work is rewarded through Bonusly, peer-nominated awards, and more.
- Holistic health and wellness benefits: Company-paid medical, dental, and vision (with 100% employer paid options and 90% coverage for dependents), FSA and HSA, 401k match, and telehealth options including memberships to One Medical.
- Support for Titans at all stages of life: Parental leave and support, up to $20k in fertility services (i.e. IUI and IVF), surrogacy, and adoption reimbursement, on demand maternity support through Maven Maternity, free breast milk shipping through Maven Milk, pet insurance, legal advisory services, financial planning tools, and more.
At ServiceTitan, we celebrate individuality and uniqueness. We believe that the convergence of fresh perspectives and experiences from all walks of life is what makes our product and culture so great. We strongly encourage people from underrepresented groups to apply. We do not discriminate against employees based on race, color, religion, sex, national origin, gender identity or expression, age, disability, pregnancy (including childbirth, breastfeeding, or related medical condition), genetic information, protected military or veteran status, sexual orientation, or any other characteristic protected by applicable federal, state or local laws.
ServiceTitan is committed to fair and equitable compensation for all of our employees. We thoughtfully consider a wide range of factors when determining individual compensation, which may change over time. We comply with all applicable minimum wage laws. For candidates in the United States, the good faith salary ranges estimate for this role is Zone 1: $147,600 USD - $221,400 USD Applicable for: CA, CT, DC, MD, MA, NJ, NY, VA, and WA Zone 2: $137,900 USD - $206,900 USD Applicable for: All other US locations. International Compensation for candidates residing outside the United States will vary by location and will be discussed during the hiring process. Actual compensation within a range is determined by factors including relevant experience, skill set, qualifications, and performance. In
- ...We are seeking a Staff Site Reliability Engineer to serve as the foundational Technical Lead for our Platform Engineering SRE organization. In this role, you will be the primary architect and visionary for the core technology foundations. As the technical lead for all...SeniorFull time
- ...Position Summary We are seeking a Lead Site Reliability & Environment Monitoring Engineer to establish and evolve our enterprise observability and monitoring strategy across cloud and application platforms. This is a full-time leadership role responsible for owning...SeniorFull timeShift work
- ...thousands of customers depend on every day. We're hiring a senior, hands-on engineer to own the reliability, availability, security, and performance of that... ...you are: ~5+ years of hands-on Cloud Operations and Site Reliability Engineering, operating production-scale SaaS...SeniorFull time
- ...A leading High-Frequency Trading firm is seeking an experienced Senior Site Reliability Engineer. This is a critical role where reliability, performance, automation and operational excellence are essential. You will work closely with software engineers, infrastructure...SeniorFull time
- ...TechMContact: Meghana GorusuCompany: SRI Tech SolutionsJob Title: Senior Site Reliability EngineerLocation: Plano , TX (remote)Years of Experience: 8... ...are seeking a highly skilled Senior Site Reliability Engineer (SRE) to join our dynamic team. The ideal candidate will...SeniorRemote work
- ...expertise, scale, and agility needed to move forward with confidence. Visit to know more about us. Role : Senior Python Site Reliability Engineer Location : Pennington, NJ/ Jersey City, NJ Mode : Hybrid ( (Hybrid, minimum 3 days onsite per week) Job Description...SeniorFull time3 days per week
- About the RoleWe’re looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You’ll partner with engineers and data scientists to build, automate, and maintain the infrastructure...Senior
$104.9k - $174.7k
About the role:A FinOps Site Reliability Engineer (SRE) bridges the gap between engineering, operations, and financial governance by embedding cost optimization into infrastructure design, automation, monitoring, and operational processes. A FinOps SRE proactively identifies...SeniorFull timeLocal area- ...Sinchan Chakraborty at email address Sinchan Chakraborty can be reached on # (***) ***-****. We have Contract role Senior Site Reliability Engineer-Hybrid for client at Louisville, KY. Please let me know if you or any of your friends would be interested in this...SeniorPermanent employmentContract workWork experience placementWork at officeLocal areaShift work
- IXL Learning, developer of personalized learning products used by millions of people globally, is seeking a Senior Site Reliability Engineer to join our team, and help maintain the reliability and optimal performance of our products. We are seeking engineers with a passion...SeniorWork at officeImmediate start
$210k - $230k
...GovCIO is currently hiring for a Senior Site Reliability Engineer (SRE) to design, implement, and maintain highly available, scalable, and resilient infrastructure systems. The ideal candidate will bridge the gap between development and operations, focusing on automation...SeniorCurrently hiringRemote work$174k - $252k
...systems by pushing for changes that improve reliability and velocity.Practice sustainable... ...:Bachelor’s degree in Computer Science, Engineering, a related field, or equivalent practical... ...degree in Computer Science or Engineering.Site Reliability Engineering (SRE) is what you...Senior- Senior Site Reliability Engineer (SRE)Salt Lake City, UTAre you passionate about building highly reliable, scalable cloud platforms that power mission-critical applications? We're partnering with an innovative technology company that's investing heavily in platform reliability...SeniorWork at officeRemote work1 day per week
- ...and best in class outcomesVisionary in future focused problem-solvingExceptional in execution and impactThe RoleAs a Senior Site Reliability Engineer at Proofpoint you will develop a deep understanding of the various services and applications that come together to deliver...SeniorFull timeFlexible hours
$182.8k - $247.3k
...changing mission to develop education for our half a billion (and growing!) learners around the world.About the role...As a Senior Site Reliability Engineer, you will work closely with both product and platform engineering teams to ensure Duolingo’s sophisticated distributed...SeniorWork experience placement- ...work from home day is currently Tuesday.Engineering at Lambda is responsible for building and... ...and networking teams to improve service reliability and deployment workflowsDeploy and... ...rotationYouHave 5+ years of experience in Site Reliability Engineering, Production Engineering...SeniorWork at officeLocal areaWork from homeFlexible hours
$168k - $270.25k
...phenomenal people like you to help us accelerate the next wave of artificial intelligence.Join our team at NVIDIA as a Senior Site reliability engineer focused on HPC storage and play a crucial role in designing, implementing, and optimizing on-prem High-Performance...SeniorFull time- ...CyLogic is seeking a highly experienced Senior SRE / DevOps Engineer to build, operate, and improve scalable, reliable, and observable platforms supporting mission-critical applications such as Omnissa Workspace ONE. This role combines deep expertise in automation, CI...SeniorFull time
- ...professionalism. We are seeking an experienced AWS solution design engineer/architect to join our infrastructure cloud team. The... ...product features efficiently and confidently them into production.As Senior SRE, you will be responsible for providing leadership, design and...Senior
$152.5k - $205k
...flexible work environment where new ideas are encouraged and everyone is a stakeholder.What you’ll be responsible for:As a Senior Site Reliability Engineer on Circle’s platform team, you’ll design, build, and operate the secure, scalable platform infrastructure behind...SeniorFlexible hours- ...importance of in-office collaboration and fully intend for the selected candidate for this role to work on site in the specified location(s).As a Site Reliability Engineer supporting the Cashiering organization, you will play a critical role in ensuring the stability,...SeniorFull timeWork at office
- ...Lambda’s designated work from home day is currently Tuesday.Engineering at Lambda is responsible for building and scaling our cloud offering... ...and SLIs for Kubernetes services, workloads, and platform reliability.You6+ years of experience in a SRE, operations engineer, or...SeniorWork at officeLocal areaWork from homeFlexible hours
$160k - $240k
...millions of times a day - quickly, reliably, and securely. Any time you... ...at Fiserv.Job TitleSenior Site Reliability EngineerWhat does a successful Site Reliability Engineer do at Fiserv?You will join our... ...operations or DevOps at a mid-to-senior level.Strong shell scripting...SeniorFull time$190.8k - $267.1k
...helping Reddit grow its business. The reliability of our Ads systems directly impacts advertiser... ...team partners closely with Ads Engineering to improve reliability, scalability, operational... ...advertiser trust. We’re looking for a Senior Site Reliability Engineer to build, operate,...SeniorFor contractorsWork experience placement$148k - $235.75k
...see how you can make a lasting impact on the world.Join our team of innovative engineers who are building an AI Data Center AIOps platform that turns raw, high-volume telemetry into reliable, job-centric insights and automation for GPU fleets. We’re hiring a DevOps Engineer...SeniorFull time- Job-ID29423084Reference26-27966Remote50% Remote Overview Client US is seeking an experienced and pragmatic Senior Site Reliability Engineer to own the reliability, design, implementation, and continuous improvement of the infrastructure that powers restaurant technology...SeniorPermanent employmentWork experience placementRemote workShift work
$91.7k - $163.7k
...potential to change lives. Ready to build the next breakthrough? Join us to start Caring. Connecting. Growing together. The Site Reliability Engineer will architect, develop, and maintain Optum Serve's cloud environment in both the commercial and government cloud. The...SeniorMinimum wageFull timeWork experience placementWork at officeLocal areaRemote work$267k - $356k
...day is currently Tuesday.Lambda's Storage Engineering team is the backbone behind our world-... ...workloads in the industry, which means reliability and performance aren't just goals—they're... ...defined storage across new and existing sites using tools such as Ansible, Jenkins etc...SeniorWork experience placementWork at officeLocal areaWork from homeFlexible hours$80k - $140k
Job DescriptionRBC Wealth Management Technology is seeking a Senior Site Reliability Engineer to join its Wealth Management SRE Team. This team is responsible for ensuring the performance, availability, resilience, and operational excellence of critical applications and...SeniorFull timeFlexible hoursShift work$160k - $200k
...data, ideally using promQLKey Responsibilities:Mentor and evangelize on observability best practices, SLIs/SLOs, and reliability culture across engineering teams. Contributing to and maintaining Tulip's triage & remediation processes as a player / coachPerform incident...SeniorTemporary workWork at officeLocal areaFlexible hours3 days per week
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!
- site reliability engineering manager United States
- site reliability engineer sre United States
- site reliability engineer United States
- site reliability engineer remote United States
- senior business controller United States
- senior service associate United States
- senior safety specialist United States
- civitas senior living United States
- senior learning manager United States
- senior merchandising manager United States


