Senior Site Reliability Engineer
$175k - $200kOrder
Order.co is the System of Action for the Office of the CFO, transforming the way businesses purchase and pay into an intuitive, B2C-like shopping experience. Order.co leverages embedded AI agents and embedded financial products to reinvent the way businesses connect with their vendors.
End users enjoy a seamless, zero-training buying experience, while finance and procurement leaders gain a single platform to orchestrate how the business "should operate". The result is an all-in-one solution that serves as a gravitational pull for spend and data, automating and eliminating procurement and finance workflows from requisition to reconciliation along the way.
Founded in 2016 and headquartered in New York City, Order.co oversees nearly half a billion in annualized spend across hundreds of customers like WeWork, SoulCycle, Lume, and [solidcore]. Order.co has raised $75M in funding from industry-leading investors like MIT, Stage 2 Capital, Rally Ventures, 645 Ventures, and more. Order.co has been proudly named a 50 to Watch by Spend Matters and a Best Place to Work by BuiltIn and Inc. Magazine. The Role As a Senior Site Reliability Engineer on the Platform team, you will ensure that software systems are reliable, scalable, performant, and operationally efficient. You blend software engineering skills with infrastructure and operations expertise to keep critical systems running smoothly while enabling rapid product development.
Responsibilities Reliability Engineering & Infrastructure Ownership
- Design, build, and operate highly available, scalable, and fault-tolerant infrastructure and platform services
- Own reliability, availability, latency, and operational excellence for critical production systems and services
- Define and maintain service level objectives (SLOs), service level indicators (SLIs), and error budgets across platform systems
- Lead incident response efforts for complex production outages; drive root-cause analysis and long-term remediation actions
- Build resilient systems that gracefully handle failures, traffic spikes, dependency degradation, and regional outages
- Continuously improve system reliability through automation, observability, performance tuning, and capacity planning
- Develop infrastructure automation and self-service tooling to reduce operational toil and improve engineering velocity
- Build and maintain CI/CD pipelines, deployment automation, and release engineering workflows
- Implement infrastructure as code (IaC) practices using tools such as Terraform, CloudFormation, and container orchestration
- Improve developer experience by building reliable internal platforms, operational tooling, and standardized deployment patterns
- Drive adoption of GitOps, immutable infrastructure, and automated remediation patterns
- Design and maintain comprehensive monitoring, logging, tracing, and alerting systems for distributed services
- Establish actionable alerting standards that reduce noise while improving incident detection and response times
- Analyze production trends, system bottlenecks, and failure patterns to proactively prevent incidents
- Lead operational readiness reviews, disaster recovery planning, and game-day exercises
- Improve mean time to detect (MTTD) and mean time to recovery (MTTR) through tooling, automation, and process refinement
- Participate actively in architecture and infrastructure design reviews
- Propose scalable and reliable platform designs that account for multi-region deployment, redundancy, failover, and security considerations
- Evaluate trade-offs between reliability, scalability, operational complexity, and engineering velocity
- Identify systemic risks and operational gaps before they become production incidents
- Partner with engineering teams to ensure services are designed with operability, observability, and resilience in mind from day one
- Approach infrastructure and operational practices with a strong security mindset
- Implement and maintain secure cloud networking, secrets management, IAM policies, and infrastructure hardening standards
- Partner with Security and Compliance teams to ensure systems meet organizational and regulatory requirements
- Drive operational best practices around vulnerability management, patching, and production access controls
- Scope and estimate infrastructure and reliability initiatives accurately
- Coordinate production rollouts, maintenance events, and reliability improvements across teams
- Communicate operational risks, dependencies, and incident impacts clearly to technical and non-technical stakeholders
- Collaborate closely with Software Engineering, Security, Product, and Operations teams to improve platform reliability and scalability
- Serve as a trusted escalation point during critical production incidents
- Mentor junior and mid-level engineers on reliability engineering principles, operational excellence, and infrastructure best practices
- Raise the operational maturity of the engineering organization through documentation, reviews, and technical guidance
- Drive improvements in team standards around observability, incident management, automation, and infrastructure design
- Influence technical decisions through credibility, operational expertise, and strong engineering judgment
- You are motivated by accountability - you own outcomes, not just tasks
- You are results-oriented and measure success by shipped, working software
- You are motivated by correctness in code that touches money - the consequences of a bug land on real customer balances, and you take that seriously
- You love helping people on your team grow and improve
- Writing tests is an integral part of your development process, not an afterthought
- You know how to design and build software incrementally - you don't need a complete spec to make progress
- Collaborating with the people around you to achieve a goal motivates you
- You are collaborative, open-minded, and actively developing your craft
- You are curious and pragmatic about AI-driven solutions - you apply them where they add real value and stay skeptical where they don't
- Familiarity with AI-assisted development tools - you understand how they work, where they help, and where they fail. Prior hands-on use is a plus; intellectual curiosity and the instinct to evaluate AI output critically are what matter
- Strong foundation in computer science fundamentals: data structures, algorithms, and system design
- Familiarity with building production-grade applications and services using Ruby and Ruby on Rails
- Deep expertise with Linux systems administration and production troubleshooting
- Strong experience operating cloud infrastructure at scale, particularly within AWS environments
- Experience with Kubernetes, container orchestration, and cloud-native infrastructure patterns
- Proficiency with infrastructure as code tools such as Terraform or CloudFormation
- Expertise designing and operating CI/CD pipelines and deployment automation systems
- Deep understanding of observability tooling including Datadog, OpenTelemetry, or similar platforms
- Strong knowledge of distributed systems reliability patterns including redundancy, failover, autoscaling, rate limiting, and graceful degradation
- Experience building automation and operational tooling using languages such as Python, Go, Bash, or Ruby
- Strong understanding of networking fundamentals including DNS, load balancing, TLS, VPNs, firewalls, and service discovery
- Hands-on experience with incident response, root-cause analysis, and production operations in high-availability environments
- Familiarity with SRE methodologies including SLOs, SLIs, error budgets, capacity planning, and operational maturity modeling
- Experience implementing secure infrastructure and cloud security best practices including IAM, secrets management, and vulnerability remediation
- Proven ability to design scalable, resilient, and maintainable platform systems and APIs
- Experience supporting distributed microservices architectures and event-driven systems
- Strong understanding of operational excellence principles including automation-first engineering and toil reduction
- Experience using AI-assisted engineering tools (e.g., Claude, GitHub Copilot) as force multipliers while applying sound operational and engineering judgment
- Excellent debugging and systems thinking skills across infrastructure, networking, application, and platform layers
- Reliable delivery of complex work - consistently ships multi-part solutions on time with low defect rates
- Low defects in owned areas - proactively monitors and improves the quality of the systems they own; that means incident-free quarters in code paths that move funds and clean reconciliation against vendor reports
- Measurable mentorship impact - engineers around you write better code because of your reviews and guidance
- Competitive compensation including base salary, bonus, and equity
- Employer-sponsored 401(k) with match
- Comprehensive medical, dental, and vision coverage
- Flexible time off and hybrid work environment
The anticipated annual salary range for this role is $175,000 - $200,000 . Actual compensation and title will be commensurate with experience, qualifications, knowledge, and skills.
- ...Evaluate applications, platforms, and vendors to assess resiliency, reliability, and operational risk.Design and implement processes that... ...and reliability tooling.Actively participate in reliability engineering and resilience communities of practice, contributing to...SeniorFull time
$170k - $220k
Who We're Looking ForWe’re looking for a hands-on, high-agency Site Reliability Engineer to help shape and scale the reliability layer of our stack. You'll own the release pipeline end-to-end — managing daily releases, weekly deploys, and hotfixes — while also automating...Senior$65 - $75 per hour
DescriptionKforce has a client seeking a remote Senior Site Reliability Engineer to be a l be a leading member of the team working with a diverse range of technologies. You will enjoy working in a friendly environment and benefit from our investment in staff. The role...SeniorRemote work- About the RoleWe’re looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You’ll partner with engineers and data scientists to build, automate, and maintain the infrastructure...Senior
- ...TechMContact: Meghana GorusuCompany: SRI Tech SolutionsJob Title: Senior Site Reliability EngineerLocation: Plano , TX (remote)Years of Experience: 8... ...are seeking a highly skilled Senior Site Reliability Engineer (SRE) to join our dynamic team. The ideal candidate will...SeniorRemote work
$130k - $200k
IXL Learning, developer of personalized learning products used by millions of people globally, is seeking a Senior Site Reliability Engineer to join our team, and help maintain the reliability and optimal performance of our products. We are seeking engineers with a passion...SeniorFull timeWork at officeImmediate start- Senior Site Reliability Engineer ILocationSan Jose, Costa Rica - RemoteSummary of roleOwn availability, the most important product feature, by continually striving for sustained operational excellence of Sumo’s planet-scale observability and security products. Work with...SeniorFlexible hours
$168k - $270.25k
NVIDIA is looking for a Senior Site Reliability Engineer (SRE) to join its GeForce Now (GFN) team. SRE at NVIDIA ensures that our internal and external-facing GPU cloud gaming services have reliability and uptime as promised to the users and at the same time enables developers...SeniorFull time$152.6k - $191.5k
...responsible for partnering with leaders across engineering and technology to define objective reliability goals for services. Key responsibilities include... ...and continuous improvement.Position Summary:The Senior GCP Site Reliability Engineer acts as an advanced senior...SeniorFull timeWork at officeDay shift$104.9k - $174.7k
...Data Management. You can learn more about LexisNexis Risk at the link below, About the Role:We are hiring a hands-on Senior Site Reliability Engineer (SRE) to actively build, operate, and improve the reliability of our production systems. This is not a purely advisory...SeniorFull timeWork at officeLocal areaRemote workWork from home$210k - $230k
GovCIO is currently hiring for a Senior Site Reliability Engineer (SRE) to design, implement, and maintain highly available, scalable, and resilient infrastructure systems. The ideal candidate will bridge the gap between development and operations, focusing on automation...SeniorCurrently hiringRemote work- Job Description:About the Role: We are looking for a Senior SRE to join our Platform Engineering team as the operations owner of our observability platforms. You’ll be responsible for the reliability, scalability, and continued evolution of the tools that give our engineering...SeniorFull time
- ...professionalism. We are seeking an experienced AWS solution design engineer/architect to join our infrastructure cloud team. The... ...product features efficiently and confidently them into production.As Senior SRE, you will be responsible for providing leadership, design and...Senior
$168k - $270.25k
...phenomenal people like you to help us accelerate the next wave of artificial intelligence.Join our team at NVIDIA as a Senior Site reliability engineer focused on HPC storage and play a crucial role in designing, implementing, and optimizing on-prem High-Performance...SeniorFull time- ...work from home day is currently Tuesday.Engineering at Lambda is responsible for building and... ...and networking teams to improve service reliability and deployment workflowsDeploy and... ...rotationYouHave 5+ years of experience in Site Reliability Engineering, Production Engineering...SeniorWork at officeLocal areaWork from homeFlexible hours
- ...Lambda’s designated work from home day is currently Tuesday.Engineering at Lambda is responsible for building and scaling our cloud offering... ...and SLIs for Kubernetes services, workloads, and platform reliability.You6+ years of experience in a SRE, operations engineer, or...SeniorWork at officeLocal areaWork from homeFlexible hours
- LeanData helps the world’s fastest-growing companies automate, simplify, and accelerate revenue.We are looking for a Senior Site Reliability Engineer to lead the strategic evolution of our cloud infrastructure. Reporting directly to the SVP of Engineering, this role is...SeniorFull timeWork at office2 days per week
- ...candidates that are particularly strong in a few areas, and have some interest and capabilities in others.About the Role:As a Site Reliability Engineer, you’ll join the global Platform SRE team responsible for building, operating, and scaling Kong’s multi-region SaaS...SeniorTemporary work
$80k - $140k
Job DescriptionRBC Wealth Management Technology is seeking a Senior Site Reliability Engineer to join its Wealth Management SRE Team. This team is responsible for ensuring the performance, availability, resilience, and operational excellence of critical applications and...SeniorFull timeFlexible hoursShift work- Job Description:Note: Fidelity will not provide immigration sponsorship for this positionThe RoleOur Site Reliability Engineering group within Enterprise Infrastructure combines Operations Excellence with the Development Experience to deliver services at high scale, high...SeniorFull time
$119.8k - $234.7k
...yearEmployment type: Full-TimeWork site: 3 days / week in-officeRole type: Individual... ...EngineeringDiscipline: Site Reliability EngineeringCompany: MicrosoftOverviewMicrosoft... ...’s most demanding workloads. As a Senior Site Reliability Engineer, you will lead reliability...SeniorOngoing contractLocal area3 days per week- ...the team takes that seriously!The RoleThe Senior SRE at 2K is a hands-on technical leader... ...regions while partnering with network engineers, systems architects, and game studio developers... ...technical direction, influencing reliability from architecture review through production...Senior
$104.9k - $174.7k
Are you passionate about improving reliability, scalability, and resilience in complex database... ....Own prioritization of reliability engineering tasks within team backlogs.Lead incident... ...a Service (IaaS).Background in DevOps, site reliability engineering practices, or related...SeniorFull timeLocal area$98k - $176k
...joy of everyday life. We bring that vision to life through our values and culture. Learn more about Target here. As a Senior Site Reliability Engineer within Digital Enablement, you specialize in building and supporting the platforms and tools that enable teams to deliver...SeniorFull timeTemporary workWork experience placementFlexible hours- ...and foster a dynamic work environment where new ideas thrive. Are you ready to join our team and make an impact?As a Senior Site Reliability Engineer at TeamViewer, you’ll be a key player in ensuring the reliability, scalability, and performance of our Azure-based SaaS...SeniorTemporary workCasual workWorldwide
$158.5k - $172k
...the exceptional value they deserve.About The OpportunityAs a Senior Engineer on the Runtime Automation team, you will design, automate,... ...environment. This is a high-impact position driving continuous reliability, deep system optimization, and automation across our entire...SeniorFull timeTemporary workWork at officeFlexible hours3 days per week- Reliability Engineering Design, implement, and operate scalable, resilient, and highly available systems on Google Cloud Platform. Improve service... ...Skills, and Abilities Three or more years of experience in Site Reliability Engineering, platform engineering, DevOps, cloud...SeniorRemote work
$267k - $356k
...day is currently Tuesday.Lambda's Storage Engineering team is the backbone behind our world-... ...workloads in the industry, which means reliability and performance aren't just goals—they're... ...defined storage across new and existing sites using tools such as Ansible, Jenkins etc...SeniorWork experience placementWork at officeLocal areaWork from homeFlexible hours- ...Georgia, and serves customers in more than 35 countries worldwide.Position OverviewWe are seeking a highly experienced Senior Site Reliability Engineer (Unified Observability) to lead the design, implementation, and operational maturity of the F1 Next Generation...SeniorFull timeWorldwideFlexible hours
- ...the rapidly evolving state of the art to engineer scalable, innovative, and research... ...business. About the Role:We are looking for a Senior SRE to serve as the operations owner for... ...tooling ecosystemsOwn the operational reliability of developer tooling ecosystems, including...SeniorFull timeLocal area
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!
- site reliability engineer remote United States
- site reliability engineer sre United States
- site reliability engineering manager United States
- site reliability engineer United States
- lead site reliability engineer United States
- senior groundskeeper United States
- senior maintenance supervisor United States
- senior operations associate United States
- senior safety specialist United States
- lcb senior living United States
