Senior Site Reliability Engineer
OutSystems
There are NO limits to your career: come shape the future and be part of a truly unique global culture at OutSystems! Hybrid Onsite in Menlo Park, CA Site Reliability Engineering (SRE) is a discipline that incorporates aspects of software engineering and applies them to infrastructure and operations problems. The main goals of SRE are to create scalable and highly reliable systems. Our SREs ensure our production systems' reliability, performance, and scalability while enabling rapid development and deployment of new features and services. SREs at OutSystems work closely with development teams, acting as an extension of the team, in adopting the reliability tenets with the shared goal of meeting Service Level Objectives (SLOs) and thus delivering a smooth and frictionless Customer Experience. Site Reliability Engineer Role
As an SRE at OutSystems here are your key responsibilities and duties: Lead and onboard services and teams to the reliability tenets; Establish and maintain Service Level Objectives (SLOs) and Service Level Agreements (SLAs); Design and implement scalable, reliable, and secure infrastructure, while ensuring cloud-native best practices; Collaborate with software development teams to ensure systems are resilient (observable, fault-tolerant, recoverable, scalable) and performant; Implement monitoring, alerting, logging, and tracing solutions to detect and respond to incidents; Lead incident response efforts, ensuring quick resolution and minimal downtime, and conduct RCA/post-mortems; Automate every operational task, with a special focus on fast incident detection & recovery; Programming in Python supported by Gen AI tooling to accelerate development of mission critical automation and tools. Foster a culture of continuous improvement and knowledge sharing; Communicate effectively with stakeholders, providing updates on system reliability and performance; Participate in on-call rotation to provide 24/7 support for production systems. Site Reliability Engineering Performance Indicators
The main KPIs that aid in understanding the impact and success of the SRE function at OutSystems are: SLA and Service Level Objectives (SLO) compliance; SLO Coverage and Detection Ratio; MTTA - Mean time to acknowledge; MTTR - Mean time to resolve. Qualifications and Skills
To illustrate the desired profile for a Site Reliability Engineer. Nevertheless, the selection of candidates will always vary depending on specific knowledge of the field and prior experience. Qualifications
BS/MS in Computer Science or Equivalent 6+ years of experience in Site Reliability Engineering, managing infrastructure and services at scale History of end-to-end project delivery Experience managing Hadoop and Kubernetes infrastructure and related services, or equivalent experience Advanced knowledge of Linux, Networking, and Containers Proficiency in at least one high-level programming language (Python, GoLang etc.). Strong troubleshooting and debugging skills. Fluency in English and excellent communication skills. An understanding or hands-on experience with Prompt engineering in software development; Familiarity with AI Native IDEs or AI Assistants such as Cursor, GitHub CoPilot, and Claude. Soft Skills
Communication - able to communicate effectively (in English) both orally and written showing empathy for the other person; Collaboration - Proactive collaboration and presentation skills to effectively communicate ideas and represent the deliverables and needs of the SRE team with leadership. Humbleness - accepts mistakes and acts accordingly, with a humble attitude, apologizing for them and mitigating them ASAP to avoid higher impact. Accountability - takes ownership of problems and makes sure to see them through. Even if he does not have all the necessary knowledge to move on alone, can involve the right people to reach closure. Negotiation Skills - has tough and politically complex conversations with colleagues and customers, defusing disagreements and leading towards a mutual agreement and understanding of all parties involved. Process Oriented - is organized and able to properly follow defined processes, whilst being able to properly challenge inefficient processes and suggest improvements. Problem-solving - Has a top-down approach to problems, breaking them into smaller pieces and solving them by starting with a wider scope and narrowing it down as the analysis progresses. Has critical thinking, so can analyze information objectively and make a reasoned judgment. Technical Skills
Experience in any of the following is valued, but not fully required: Ability to establish, monitor, and improve Service Level Objectives (SLOs), Indicators (SLIs), and Agreements (SLAs) in line with business needs. Containerization technologies and orchestration platforms, mainly Kubernetes and EKS
(CKA, CKAD, CKS certifications are valued); Experience with automation and Infrastructure as Code (IaC) tools, such as AWS CloudFormation, Terraform, Puppet, Chef, Spacelift, etc; Experience with Python, Go, Bash/Shell scripting, or other automation tools/languages; Familiarity with AWS services like EC2, RDS, ELB, CloudFront, Lambda, etc; Proficiency in monitoring and troubleshooting complex distributed systems; Experience with Grafana, ELK stack, Prometheus, or others; Strong understanding of designing resilient and fault-tolerant systems; Expertise in debugging complex distributed systems. More about OutSystems OutSystems is a leading AI Development Platform built for the enterprise. Global organizations trust OutSystems to rapidly build mission-critical apps and agents, modernize legacy processes with agentic systems, and govern their entire AI portfolio across complex regulatory environments, all on one unified platform. As the future becomes agentic, our customers need us now more than ever. While AI has opened the door to extraordinary possibilities, most large organizations find themselves stuck on one side of the "enterprise gap" because AI by itself doesn't solve their complex use cases and business challenges. OutSystems bridges the "enterprise gap" by combining the speed of generative AI with a deterministic, enterprise-grade framework. We provide the tools for teams of any size to deliver high-quality, reliable AI solutions that drive real business impact. We are looking for passionate, talented, and motivated people to join us as we empower organizations to build, deploy, and scale the next generation of enterprise software. While we are leading the charge into the agentic era, our mission is broader: we are the platform enterprise leaders trust to evolve their entire business, accelerating innovation through secure, governed human-AI collaboration.
OutSystems is a global company, with more than 900k developer community members, 1,700 employees, more than 600 partners, and thousands of active customers in over 75 countries and across 21 industries. Founded in 2001, OutSystems now has offices in the United States, United Kingdom, the Netherlands, Portugal, Germany, the UAE, Japan, Hong Kong, Malaysia, Australia, India, and Singapore, and includes a thriving, worldwide community of remote employees. Our customers are some of the world's most recognizable brands across diverse industries- such as Toyota, Heineken, Bosch, KeyBank, and UCLA-who trust OutSystems to deliver ROI and transformational impact.
Consistently recognized as a leader by top analyst firms Gartner, IDC and Forrester, OutSystems continues to shape the future of enterprise software development in the agentic era. We are proud to be named a leader in more than 100 categories on G2, including #1 in Customer Satisfaction in Enterprise Low Code Development, and most recently as a leader in AI Agent Building in the G2 Spring 2026 Reports.
Working at OutSystems Our culture is built on our core values of Trust, Customer Success, Innovation, and Alignment. We operate as one global OutSystems team, taking ownership to pursue our vision of being the AI platform enterprise leaders trust to build, secure, and evolve their most critical applications and systems. What do we have to offer you?
As an SRE at OutSystems here are your key responsibilities and duties: Lead and onboard services and teams to the reliability tenets; Establish and maintain Service Level Objectives (SLOs) and Service Level Agreements (SLAs); Design and implement scalable, reliable, and secure infrastructure, while ensuring cloud-native best practices; Collaborate with software development teams to ensure systems are resilient (observable, fault-tolerant, recoverable, scalable) and performant; Implement monitoring, alerting, logging, and tracing solutions to detect and respond to incidents; Lead incident response efforts, ensuring quick resolution and minimal downtime, and conduct RCA/post-mortems; Automate every operational task, with a special focus on fast incident detection & recovery; Programming in Python supported by Gen AI tooling to accelerate development of mission critical automation and tools. Foster a culture of continuous improvement and knowledge sharing; Communicate effectively with stakeholders, providing updates on system reliability and performance; Participate in on-call rotation to provide 24/7 support for production systems. Site Reliability Engineering Performance Indicators
The main KPIs that aid in understanding the impact and success of the SRE function at OutSystems are: SLA and Service Level Objectives (SLO) compliance; SLO Coverage and Detection Ratio; MTTA - Mean time to acknowledge; MTTR - Mean time to resolve. Qualifications and Skills
To illustrate the desired profile for a Site Reliability Engineer. Nevertheless, the selection of candidates will always vary depending on specific knowledge of the field and prior experience. Qualifications
BS/MS in Computer Science or Equivalent 6+ years of experience in Site Reliability Engineering, managing infrastructure and services at scale History of end-to-end project delivery Experience managing Hadoop and Kubernetes infrastructure and related services, or equivalent experience Advanced knowledge of Linux, Networking, and Containers Proficiency in at least one high-level programming language (Python, GoLang etc.). Strong troubleshooting and debugging skills. Fluency in English and excellent communication skills. An understanding or hands-on experience with Prompt engineering in software development; Familiarity with AI Native IDEs or AI Assistants such as Cursor, GitHub CoPilot, and Claude. Soft Skills
Communication - able to communicate effectively (in English) both orally and written showing empathy for the other person; Collaboration - Proactive collaboration and presentation skills to effectively communicate ideas and represent the deliverables and needs of the SRE team with leadership. Humbleness - accepts mistakes and acts accordingly, with a humble attitude, apologizing for them and mitigating them ASAP to avoid higher impact. Accountability - takes ownership of problems and makes sure to see them through. Even if he does not have all the necessary knowledge to move on alone, can involve the right people to reach closure. Negotiation Skills - has tough and politically complex conversations with colleagues and customers, defusing disagreements and leading towards a mutual agreement and understanding of all parties involved. Process Oriented - is organized and able to properly follow defined processes, whilst being able to properly challenge inefficient processes and suggest improvements. Problem-solving - Has a top-down approach to problems, breaking them into smaller pieces and solving them by starting with a wider scope and narrowing it down as the analysis progresses. Has critical thinking, so can analyze information objectively and make a reasoned judgment. Technical Skills
Experience in any of the following is valued, but not fully required: Ability to establish, monitor, and improve Service Level Objectives (SLOs), Indicators (SLIs), and Agreements (SLAs) in line with business needs. Containerization technologies and orchestration platforms, mainly Kubernetes and EKS
(CKA, CKAD, CKS certifications are valued); Experience with automation and Infrastructure as Code (IaC) tools, such as AWS CloudFormation, Terraform, Puppet, Chef, Spacelift, etc; Experience with Python, Go, Bash/Shell scripting, or other automation tools/languages; Familiarity with AWS services like EC2, RDS, ELB, CloudFront, Lambda, etc; Proficiency in monitoring and troubleshooting complex distributed systems; Experience with Grafana, ELK stack, Prometheus, or others; Strong understanding of designing resilient and fault-tolerant systems; Expertise in debugging complex distributed systems. More about OutSystems OutSystems is a leading AI Development Platform built for the enterprise. Global organizations trust OutSystems to rapidly build mission-critical apps and agents, modernize legacy processes with agentic systems, and govern their entire AI portfolio across complex regulatory environments, all on one unified platform. As the future becomes agentic, our customers need us now more than ever. While AI has opened the door to extraordinary possibilities, most large organizations find themselves stuck on one side of the "enterprise gap" because AI by itself doesn't solve their complex use cases and business challenges. OutSystems bridges the "enterprise gap" by combining the speed of generative AI with a deterministic, enterprise-grade framework. We provide the tools for teams of any size to deliver high-quality, reliable AI solutions that drive real business impact. We are looking for passionate, talented, and motivated people to join us as we empower organizations to build, deploy, and scale the next generation of enterprise software. While we are leading the charge into the agentic era, our mission is broader: we are the platform enterprise leaders trust to evolve their entire business, accelerating innovation through secure, governed human-AI collaboration.
OutSystems is a global company, with more than 900k developer community members, 1,700 employees, more than 600 partners, and thousands of active customers in over 75 countries and across 21 industries. Founded in 2001, OutSystems now has offices in the United States, United Kingdom, the Netherlands, Portugal, Germany, the UAE, Japan, Hong Kong, Malaysia, Australia, India, and Singapore, and includes a thriving, worldwide community of remote employees. Our customers are some of the world's most recognizable brands across diverse industries- such as Toyota, Heineken, Bosch, KeyBank, and UCLA-who trust OutSystems to deliver ROI and transformational impact.
Consistently recognized as a leader by top analyst firms Gartner, IDC and Forrester, OutSystems continues to shape the future of enterprise software development in the agentic era. We are proud to be named a leader in more than 100 categories on G2, including #1 in Customer Satisfaction in Enterprise Low Code Development, and most recently as a leader in AI Agent Building in the G2 Spring 2026 Reports.
Working at OutSystems Our culture is built on our core values of Trust, Customer Success, Innovation, and Alignment. We operate as one global OutSystems team, taking ownership to pursue our vision of being the AI platform enterprise leaders trust to build, secure, and evolve their most critical applications and systems. What do we have to offer you?
- A company at the vanguard of the agentic revolution, where we don't just react to AI innovation-we architect it. Joining OutSystems means stepping onto a high-growth rocket ship that combines the fearless agility of a startup with the sophisticated, global foundation of an enterprise powerhouse.
- Real growth opportunities. We don't just talk about development; we invest in it through structured programs designed to scale your expertise. Whether you are aiming for vertical progression, exploring lateral moves into new domains, or mastering specialized AI skills through our Professional Development Fund and Internal Mobility Program, we provide the resources to get you there.
- A global collective of world-class talent, where you'll collaborate with enterprise software legends and sought-after thought leaders. At OutSystems, our industry experts aren't just visionaries-they are accessible, approachable mentors who are deeply invested in your growth as we architect the agentic future together.
Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Senior Site Reliability Engineer in San Francisco, CA vacancy
- About the RoleWe’re looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You’ll partner with engineers and data scientists to build, automate, and maintain the infrastructure...Senior
$152.5k - $205k
...flexible work environment where new ideas are encouraged and everyone is a stakeholder.What you’ll be responsible for:As a Senior Site Reliability Engineer on Circle’s platform team, you’ll design, build, and operate the secure, scalable platform infrastructure behind...SeniorFlexible hours$127k - $249k
The TeamPlatform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions... ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper)....SeniorWork at officeLocal areaRemote workWorldwideFlexible hours$117k - $209.33k
Job Requisition ID #26WD99273Position OverviewWant to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable, secure, and scalable cloud services for Autodesk GovCloud products.As part of a new SRE team supporting...SeniorFull timeFor contractors- ...’s build what’s next.About the teamThe Engineering team at Airwallex is a diverse group of... ...ownership, working together to build scalable, reliable, and secure products that empower... ...our Global services.What you’ll doAs a Senior Site Reliability Engineer, you’ll work...SeniorTemporary workLocal areaWorldwide
$165k - $225.6k
...From core infrastructure to enterprise platforms, we partner across functions to drive scale, reliability, and innovation through technology.The Senior Site Reliability Engineer OpportunityReporting to the Manager, Site Reliability Engineering, this role will help build,...SeniorPermanent employmentLocal areaWorldwideFlexible hours$167.7k - $245.2k
...requiring approximately 2 days per week on-site at Cisco offices in either San... ...behave as intended, improving reliability and reducing risks. This unified approach... ...enhanced observability and control.As a Senior Site Reliability Engineer (SRE), you will build, operate, and...SeniorFull timeTemporary workLocal areaFlexible hours2 days per week$165k - $241.4k
...very effective.We’re looking for talented engineers with a software or operations background... ...development teams to ensure the reliability, performance and security of our infrastructure... ...insurance. Please see the Cisco careers site to discover more benefits and perks....SeniorFull timeTemporary workWork at officeLocal areaFlexible hours1 day per week$148.5k - $223.9k
...right place! Agentforce is the future of AI, and you are the future of Salesforce.Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in San Francisco. Working closely with counterparts in the Infrastructure and R&D organizations...SeniorFull timeWorldwideWeekend work$220k - $235k
...are seeking a strategic, high-output Staff/Senior Staff SRE to define the future of our cloud platform and champion engineering excellence across Ironclad. In this role,... ...leadership and strategic direction for the Site Reliability Engineering team and our broader Cloud...SeniorFull timeContract workWork at office- ...getting here.)About the RoleWe're building infrastructure that has to perform under real-world scale, reliability, and security demands — and we're looking for an engineer who wants to own the foundation it runs on. This isn't a traditional "keep the lights on" role.You'...Senior
$215k - $275k
...by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date.About the role:Anyscale is looking for a Senior Site Reliability Engineer to join the Infrastructure team. Anyscale aims to provide the next generation of tools and infrastructure to make developing...SeniorWork at office$127k - $249k
We are looking for an experienced Senior or Staff Engineer for our SRE, InfraSec team, to guide the security of our cloud-based infrastructure. As a Staff SRE, you will be very hands-on technically while also mentoring a small team of SREs.The InfraSec team collaborates...SeniorLocal areaRemote workWorldwideFlexible hours$165k - $241.4k
...within Cisco’s Networking, Security, Collaboration, and Observability portfolios.Your ImpactWe are seeking a skilled Senior Site Reliability Engineer (SRE) in Production Engineering with a strong background in SaaS and operations. You will design and manage large-scale...SeniorFull timeTemporary workWork at officeLocal areaFlexible hours1 day per week- ...founders with PhDs in AI, Math, and Computer Science - is poised to redefine computing. About the Role We're seeking a Site Reliability Engineer to ensure Hyperbolic's GPU marketplace and AI infrastructure operate with exceptional reliability, performance, and...Senior
- ...About the job Senior Site Reliability Engineer About the Company Stellar is a decentralized, public blockchain that gives developers the tools to create experiences that are more like cash than crypto. The network is faster, cheaper, and far more energy-efficient...Senior
$160k - $250k
...public clouds when the right fit. As we continue to commercialize our machine learning models, we also need to grow our DevOps and Site Reliability team to maintain the reliability of our enterprise SaaS offering for our customers. Our ideal candidate is someone who is able...Senior- ...Engineering Hiring Sprint We're growing our engineering team and are accelerating hiring through a focused Engineering Hiring Sprint... ...: Platform Engineers Database Engineers Site Reliability Engineers Extensibility API Engineers AI Agents Engineers...SeniorWork at officeLocal areaFlexible hours
- ...Udaip Cloud-Based Data And Ai Platform Engineer At U.S. Bank, we're on a journey to do our best. Helping the customers and businesses we serve to make better and smarter financial decisions and enabling the communities we support to grow and succeed. We believe it...SeniorTemporary workWork experience placement
$81.1k - $187k
...Site Reliability Engineer 3 We are looking for a Site Reliability Engineer 3 to support mission-critical cloud services and production operations. The role focuses on improving service reliability, reducing operational risk, automating repetitive tasks, and driving...SeniorTemporary workImmediate startFlexible hoursShift work- ...Series B and have grown 800% over the last 12 months. Engineering at Ivo Engineers at Ivo are inventors. Ivo was first-to... ...expect us to hit our SLAs. What? We're looking for a Senior Site level Reliability Engineer as part of Infrastructure team to: Own uptime...SeniorContract workWork at officeRemote workVisa sponsorshipRelocation packageFlexible hours
$166.9k - $225.9k
...Summary: Drata's SRE team operates as both a central engineering function and an embedded reliability practice. You'll be part of a close-knit SRE team... ...What you'll bring: ~6+ years of experience in Site Reliability Engineering, Cloud Engineering, or building...SeniorWork at officeImmediate startWorldwideMonday to FridayFlexible hours$210.8k - $272.8k
About Thumbtack Thumbtack helps millions of people confidently care for their homes. About the Site Reliability Engineering Team The Site Reliability Engineering team focuses on creating and maintaining a reliable, secure, and scalable platform vital for a seamless user...SeniorLocal area$175k - $250k
...00.00/yr - $250,000.00/yr Job Title: Senior Cloud Infrastructure Engineer Location: San Francisco, CA. Remote unavailable. Modality: On-Site only. Must live within commuting distance... ...scalability, performance, and reliability across environments. What You’ll Do...SeniorFull timeRemote workRelocationRelocation package$232k - $319k
...to help us continue to scale the service with great people and reliable, cost-effective, and efficient infrastructure, processes, and... ...with self-service Accelerate the velocity of SRE and product engineering by developing robust platforms, powerful tooling, and...SeniorPermanent employmentLocal areaWorldwideFlexible hours$15k
...benefits packages, technology talks by our experts, a beautiful modern office, daily catered lunches, and more.As a Senior Cluster Site Reliability Engineer (SRE), you will help scale our research compute cluster to meet our growing needs, and you will leverage...SeniorWork at officeLocal areaRemote work$139.76k - $287.75k
...to grow their business.We are seeking a Senior Site ReliabilityEngineer to help operate,... ...will be instrumental in advancing the reliability, scalability, automation, observability... ...The ideal candidate is a highly hands-on engineer with strong production experience and a...SeniorWork at officeLocal areaRelocationRelocation package$174.92k - $209.91k
...High-Performance Engineer For Site Reliability Engineering Team Fivetran is building data pipelines to power the modern data stack for thousands of companies. Fivetran is looking for a high-performance, experienced engineer to be a part of a team of Site Reliability...SeniorFull timeWork at officeRemote work$174.92k - $209.91k
...same: to make access to data as simple and reliable as electricity. With Fivetran, customer... ..., canonical and ready to query, with no engineering or maintenance required. We're proud... ...integrate our teams, systems, and career sites. About the Role Fivetran is building...SeniorFull timeWork at officeRemote work- ...Job Description Job Description Senior Site Reliability Engineer (Payments Infrastructure) Kody is seeking a Senior Site Reliability Engineer to ensure the reliability, availability, scalability, and operational excellence of our global payment platform. You will...Senior
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!
Related searches
- site reliability engineer San Francisco, CA
- site reliability engineer sre San Francisco, CA
- site reliability engineer remote San Francisco, CA
- senior business analyst San Francisco, CA
- senior risk manager San Francisco, CA
- senior cost estimator San Francisco, CA
- senior manager tax San Francisco, CA
- senior automation engineer San Francisco, CA
- senior devops San Francisco, CA
- senior recruiter San Francisco, CA

