Senior Site Reliability Engineer
Lodgify
Senior Site Reliability Engineer
Lodgify is a fast-growing scale-up company leading the vacation rental industry. Backed by $30M in funding, our platform empowers property owners and managers worldwide to efficiently manage and grow their business through technology.
Headquartered in sunny Barcelona, we're now a team of 380+ people representing over 60 nationalities, united by a passion for transforming the future of short-term rentals.
Are you a systems-minded engineer who cares deeply about reliability, scalability, and production excellence? Join Lodgify as a Senior Site Reliability Engineer and help our engineering teams build and operate services that are reliable, observable, scalable, and resilient by design. In this role, you will improve the reliability of our shared infrastructure and product services while helping teams own their systems in production. You will work in the Platform team to strengthen observability, reduce operational toil, improve incident response, define practical SRE standards, and improve the reliability of critical infrastructure and delivery workflows.
How will you make an impact?
- Define meaningful SLIs, SLOs, and reliability targets for the platform.
- Collaborate with the software engineering teams to define and achieve the best practices for software observability, SLIs, SLOs and reliability.
- Strengthen production readiness by improving service ownership, observability, alerting, runbooks, scaling assumptions, rollback paths, and failure-mode preparedness.
- Improve the reliability, scalability, and performance of cloud, Kubernetes, and shared infrastructure, including how systems scale during growth, traffic spikes, and dependency failures.
- Build actionable observability using metrics, logs, traces, and golden signals, with tools such as Datadog, Prometheus, and Grafana.
- Implement operational and security best practices through guidelines, policies and automation.
- Reduce alert noise and improve signal quality so teams can detect, understand, and resolve issues quickly.
- Automate repetitive operational work using Python or other languages, turning recurring manual work into safer automation and clearer runbooks.
- Implement self-service Internal Developer Platform features via APIs and Kubernetes operators.
- Improve deployment safety, rollbackability, and release observability.
- Improve reliability of critical stateful systems such as databases, caches, queues, and streaming platforms.
- Participate in on-call, troubleshoot, and coordinate incident response, and facilitate blameless post-incident reviews that turn into concrete improvements.
- Execute disaster recovery drills and analyse cloud/platform usage to identify cost and resource-efficiency gains without compromising reliability.
What makes you a great fit?
- You have 7+ years of production experience operating Kubernetes-based platforms and cloud infrastructure.
- You understand and apply SRE practices: SLIs, SLOs, error budgets, production readiness, incident response, post-incident learning, toil reduction, scalability, capacity planning, high availability, backups, and disaster recovery.
- You can design and improve observability and alerting for critical systems using metrics, logs, traces, and golden signals, and are comfortable troubleshooting complex distributed systems to identify systemic reliability improvements.
- You can write maintainable software to automate operational tasks and reduce manual intervention.
- You have experience with stateful production systems such as relational databases, caches, queues, or streaming platforms.
- You know how to balance reliability, performance, cost, and delivery speed pragmatically.
- You are comfortable working in a transitional environment where SRE practices are being introduced while critical infrastructure and delivery systems still need hands-on reliability support.
- You collaborate effectively with Engineering, Platform, Security, and Product stakeholders.
- You communicate clearly, document well, and enjoy coaching teams toward stronger production ownership.
- You model initiative and accountability, raising risks early and driving improvements through to completion.
What does success look like?
- Critical services have clear owners, meaningful SLIs/SLOs, actionable alerts, dashboards, runbooks, and production readiness coverage.
- Reliability targets are consistently met across critical infrastructure and services.
- Operational toil and manual intervention are measurably reduced through automation and safer workflows.
- MTTR improves through reduced alert noise, better signal quality, stronger observability, and clear incident response playbooks and escalation paths.
- Post-incident actions are tracked, completed, and used to reduce repeat incidents.
- Disaster recovery exercises validate that critical services and infrastructure can recover within agreed expectations.
- Cloud and infrastructure resources are optimised without sacrificing performance, elasticity, or resilience.
You'll be part of a growing, dynamic company with a truly international team. At Lodgify, we are full of contagious energy, hard work, and passion for what we do. We celebrate diversity and are proud to acknowledge a variety of backgrounds, perspectives and skills in our team; committed to creating a workplace where everyone is heard and feels a sense of belonging.
What's in it for you?
Remote Flexibility: The freedom to work from home any day that works for you.
Time to Recharge: 25 working days of paid vacation and Jornada Intensiva in August.
Alan Health Insurance: Premium health, dental, and mental health support via Alan. Pre-existing conditions are covered.
Meal Perk: €150/month allowance on your Alan card + 50% off Ametller Origen prepared dishes at the office.
Tax-Free Savings: Increase your take-home pay by using Flexible Remuneration for extra meal costs (up to €70/mo) and public transport (up to €136/mo).
Home Office Gear: We provide a table, ergonomic chair, and monitor for your home setup.
Language Learning: Free Spanish classes.
Referrals: Cash rewards for bringing in new talent.
Social Life: Daily office breakfast and monthly team events
Dynamic Hub: A high-energy, inclusive environment designed for collaboration and connection with a team that represents over 60 countries.
*Benefits offered may differ based on the type of contract that is issued
So, what are you waiting for? Apply now!
All applications and CVs must be submitted in English
- ...Apple Service Engineering (ASE) seeks a senior SRE software engineer to own the architectural direction of Kubernetes internals powering Apple services... ...will define controllers and namespace management, raise reliability, and contribute to upstream Kubernetes. The role...Senior
- ...Sight Machine is seeking a senior Cloud Infrastructure IC to lead reliability, automation, and scale across our platform. You will drive IaC, CI/CD, observability and operate agentic AI systems, mentoring engineers and guiding architectural decisions while staying hands...Senior
- ...Lambda Inc. in San Francisco is seeking a Storage Engineer to own the reliability, performance, and capacity of our production storage fleet across multiple data centers, using a software-defined data plane. You will build monitoring, dashboards, and alerting for storage...Senior
- ...Senior Site Reliability Engineer Company: CyberArk Work Type: Remote Employment: Full Time Location: US Seniority: Mid Level Technologies: AWS, Kubernetes, Terraform, CloudFormation, Ansible, CloudWatch, Grafana, Datadog, OpenSearch, PagerDuty Requirements: Senior SRE...SeniorFull timeRemote work
- ...data and AI infrastructure provider is seeking a Senior Staff Technical Program Manager for Reliability to enhance the reliability and performance of their... ...involves leading programs in partnership with senior engineering leaders, requiring over 10 years of experience in...Senior
- ...Oracle Cloud Infrastructure (OCI) seeks a Senior Principal Engineer to lead the design and implementation of reliability validation for OCI control plane services, focusing on a high-performance, low-level systems approach. You will mentor engineers, define validation...Senior
- ...Senior Site Reliability Engineer Company: Sphera Work Type: Remote Employment: Full Time Location: US Seniority: Senior Level Technologies: Terraform, ARM templates, Kubernetes, Azure, SonarCloud, CheckPoint, Hadoop, Kafka, Presto, NewRelic, CI/CD, Linux, Windows, Redis...SeniorFull timeRemote work
$166k - $244k
# Senior Software Engineer, Site Reliability EngineeringGoogle • onsite • 601 N 34th St, Seattle, WA 98103, USA • full\_timePay: USD 166000.00 - USD 244000.00 / unspecifiedBusinesses of all shapes and sizes rely on Google’s unparalleled advertising solutions to help them...SeniorTemporary work- ...Avalara, Inc. is seeking a senior reliability engineer to lead how reliability is engineered across Avalara's global SaaS platform as we move toward an AI-first operating model. You will build a modern, automation-first reliability ecosystem that improves stability,...Senior
- ...Senior Site Reliability Engineer We are looking for a Senior Site Reliability Engineer with Cloud platform experience. This individual will be part of a team responsible for operating and maintaining production clusters and developing our observability solutions; they...SeniorRemote work
- ...Senior Site Reliability Engineer Teikametrics is seeking a Senior Site Reliability Engineer to enhance and maintain our cloud infrastructure for hosting applications and platforms. This position involves collaborating within a DevOps model to design and deploy automation...SeniorRemote work
- ...Senior Sre We're hiring a Senior SRE based in Latin America to work alongside our US-based engineering team, building out observability, on-call coverage, and deployment automation for a client with strict compliance requirements. We're specifically looking for someone...SeniorFull timeRemote work
- ...Job Summary We are seeking a Senior SRE / DevSecOps Engineer with strong experience in Kubernetes, AWS... ...troubleshooting. The role will focus on platform reliability, incident management, SLO/SLI... ...with SLO/SLI governance and site reliability practices. ~ Strong understanding...Senior
- ...We are seeking a Senior Database Reliability Engineer (DBRE) to design, operate, and improve reliable, scalable, secure, and highly available database... ...data platforms. The role combines database engineering, site reliability engineering, Linux systems administration, and...Senior
- Job Posting Datavant recognizes the importance of information security and data privacy, including in its hiring processes and recruitment. Datavant encourages all potential job applicants to take precautions against potential phishing schemes or other scams that improperly...SeniorFixed term contractWork at officeLocal area
- ...automated detection, drain/cordon/taint, workload rescheduling. Feed the AIOps substrate The remediation-actuator and workflow engine land here — you make the control plane safe for automated action. Your CRDs are the schema the platform's predictors and...SeniorLocal area
- ...Senior Site Reliability Engineer (SRE) Our client is a global technology consulting and digital solutions company that enables enterprises across industries to reimagine business models, accelerate innovation, and maximize growth by harnessing digital technologies....SeniorLocal area
$182.8k - $247.3k
...mission to develop education for our half a billion (and growing!) learners around the world. About the role... As a Senior Site Reliability Engineer, you will work closely with both product and platform engineering teams to ensure Duolingo’s sophisticated distributed...SeniorWork experience placement$160k - $180k
...have a big impact. See Arkestro in action at arkestro.com. About the Role Arkestro is hiring for a Senior SRE Engineer to manage our performance and reliability for our software platform and infrastructure. The right candidate will own and develop our...SeniorLocal areaRemote work$152k - $195k
...investors including Silver Lake Waterman, Moody’s, Sequoia Capital, GV and Riverwood Capital. About the Team: As a Senior Site Reliability Engineer, you will be a key technical leader driving the design and optimization of our Kubernetes-based infrastructure and CI/...SeniorRemote work$140k - $180k
...Senior Site Reliability Engineer We're hiring a Senior Site Reliability Engineer to become Endear's first dedicated reliability hire. This is a hands-on builder role for someone who wants to solve the underlying systems problems that create on-call burden—not simply...SeniorRemote workWork from homeHome officeFlexible hours$55 - $60 per hour
...on W2 #LP Job Summary: This senior-level role focuses on building and maintaining... ...candidate will possess strong cloud engineering judgment and the ability to mentor... ...services. ~ Strong understanding of reliability, security, cost management, and developer...SeniorHourly payContract workRemote work- ...Senior Site Reliability Engineer As Senior Site Reliability Engineer, you will define and operate the architectural backbone of Chalice's AI platform, reporting directly to the VP of Engineering and working cross-functionally with Engineering, Data Science, Machine...Senior
- ...training, where they work alongside former pro and D1 athletes coaching them at the highest standard. We are hiring a Senior Site Reliability Engineer on a part-time consulting contract to audit our systems and make the changes required to keep them running as we scale...SeniorFull timeContract workPart timeRemote workFlexible hours
$145k - $193k
...entertainment, we want to talk to you. About the Role & Team The SRE team at PENN Entertainment is looking for a Senior Site Reliability Engineer to help build and operate the infrastructure behind a large-scale sports betting and media platform. You'll own critical...SeniorRemote work- ...and services that help people, businesses and governments realize their greatest potential. Title and Summary Senior Site Reliability Engineer Overview-The ProCOM team is looking for a Site Reliability Engineering (SRE) who can help us solve problems, build...SeniorFull timePart timeImmediate startWorldwideFlexible hours
$115k - $160k
...collaborating with Barclays to connect them with exceptional professionals for this role. Embark on a transformative journey as a Senior Site Reliability Engineer - AVP - Credit Trade Floor. At Barclays, our vision is clear –to redefine the future of banking and help craft...SeniorHourly payWork at office- ...Senior Site Reliability Engineer Remote – Home Based Job Summary We’re partnering with a company in the SaaS space to find a Senior Site Reliability Engineer . In this role, you’ll be part of the IT Operations group responsible for maintaining all environments...SeniorTemporary workRemote workWork from homeFlexible hours
- ...infra has to match. The role We're looking for a Senior SRE to own the reliability, scalability, and operational posture of Satsuma's multi... ...-assisted development workflows Partner closely with engineering on reliability reviews and architecture decisions...Senior
- ...shaping it, one bold step at a time. To those who see AI as a driver of progress, come build the future together. As a Senior Site Reliability Engineer, you'll build and scale the critical infrastructure behind every product. In this role, you'll take on complex...SeniorImmediate startRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!
- site reliability engineer sre United States
- site reliability engineering manager United States
- site reliability engineer United States
- site reliability engineer remote United States
- senior technical service engineer United States
- senior functional safety engineer United States
- senior technical consultant United States
- senior director product management United States
- senior vendor manager United States
- senior vice president human resources United States

