Site Reliability Engineer (SRE)
re-tool®
Software Reliability Engineer
Good software has to run where customers need it. For many of Retool's largest customers, that means running Retool in their own infrastructure, behind their own controls, with the reliability and operational clarity they would expect from any critical system. Retool's Core Infrastructure team owns the systems that make this possible: Retool Cloud, managed single tenant environments, BYOC (bring-your-own-cloud) environments, Kubernetes and Helm deployments, Docker Compose, and the migration paths between them. It is a broad surface area, and it is one of the biggest levers we have for making Retool work for enterprise customers.
The work is not clean-room infrastructure. Customers run different clouds, different versions, different deployment models, and different levels of operational maturity. A bad upgrade experience can leave a customer many versions behind. A manual Terraform run can become the bottleneck during a launch or incident.
We are hiring SREs who want to turn that mess into leverage. You will help us reduce customer toil, automate upgrades and infrastructure changes, build reliability tooling across Retool Cloud and customer-owned environments, and make Retool easier to deploy and operate at enterprise scale. The strongest candidates are comfortable debugging Kubernetes, Terraform, AWS, Postgres, networking, and deployment problems, then stepping back and building the automation or product surface that prevents the same problem from happening again.
What you'll do:
- Own reliability across Retool Cloud, managed single tenant, BYOC, and self-hosted deployment paths, including provisioning, upgrades, migrations, configuration changes, and production escalations.
- Build the automation that turns today's manual infrastructure work into repeatable systems: Terraform runs, customer environment updates, upgrade workflows, secret rotations, and migration steps.
- Improve observability for Retool Cloud, self-hosted customers, and internal operators. We care less about exposing every metric and more about turning health signals into clear status, likely causes, and recommended actions.
- Design safer deployment, upgrade, and rollback paths so Cloud and managed customers can stay current
- Help move customers from legacy or less-supported deployment models toward supported paths such as Retool's official deployment paths (Blueprints, Kubernetes, and Helm), with migration flows that are repeatable enough for customers, Support, and TAMs to trust.
- Partner with product engineers on infrastructure requirements for new Retool products, especially when they introduce new dependencies
- Lead through ambiguity, make careful risk calls, and communicate clearly while things are moving quickly.
- Write the docs, runbooks, design notes, and migration guides that make complex systems understandable to other engineers and to customers.
What we're looking for:
Infrastructure fundamentals
- Deep experience operating production infrastructure in AWS.
- Experience improving reliability for customer-facing SaaS systems.
- Strong Kubernetes fundamentals.
- Real Terraform or infrastructure-as-code experience.
- Good operational judgment around databases, especially Postgres.
Reliability and automation
- Experience building or operating observability systems.
- Programming ability in a language such as Go, Python, TypeScript, Java, or Ruby.
- A bias toward automation. If you find yourself doing the same operational task twice, you should start thinking about the interface, workflow, or tool that eliminates the third time.
Customer and team judgment
- Clear written communication.
- Comfort working directly with customer-facing teams and, when useful, customers themselves.
What makes SREs successful here:
You will do well here if you like infrastructure that sits close to real customer pain. Some days that means debugging a specific customer environment. Other days it means improving Retool Cloud reliability or designing the migration path so the next 25 customers do not need that same debugging session. We value SREs who are ambitious, curious, energetic, and careful with the details. Retool moves quickly, priorities can change, and the systems are not always as clean as we want them to be. The work needs SREs who can get their hands dirty, tell the truth about tradeoffs, and leave the system better than they found it.
- ...Senior Site Reliability Engineer (SRE) Our client is seeking a Senior Site Reliability Engineer (SRE) with 10–15 years of experience to support front-office trading systems in a production environment. This role focuses on troubleshooting complex trading infrastructure...Suggested
- ...Site Reliability Engineer (SRE) Job Title Site Reliability Engineer (SRE) Job Summary We are seeking a skilled Site Reliability Engineer (SRE) to build, automate, and maintain highly available, scalable, and reliable infrastructure and applications...SuggestedFlexible hours
- ...We are seeking a highly motivated Site Reliability Engineer (SRE) to join the Equity Trading Platform Engineering team, supporting critical trading applications and Fidessa flows used by internal business users and external clients. This role is responsible for ensuring...SuggestedPermanent employmentWork at officeAfternoon shift
- ...Senior Site Reliability Engineer (SRE) Our client is a global technology consulting and digital solutions company that enables enterprises across industries to reimagine business models, accelerate innovation, and maximize growth by harnessing digital technologies....SuggestedLocal area
$120k - $180k
...leaders in space and defense. You will be the first dedicated Site Reliability Engineer and own critical infrastructure end to end. This is a... ...and operate cloud and on-premises infrastructure as the sole SRE. Architect migrations from AWS into on-premises and air-gapped...SuggestedPermanent employmentFull timeRelocation package$207k - $300k
...by pushing for changes that improve reliability and velocity.Practice sustainable... ...Master's degree in Computer Science or Engineering.Experience mentoring engineers and... ...across cross-functional teams.Site Reliability Engineering (SRE) combines software and systems engineering...$80k - $95k
...join our dynamic team supporting the company’s users, applications, and web-based product offerings. In this role, the Site Reliability Engineer (SRE) will play a key role in maintaining resources at peak efficiency to guarantee staff are able to perform their...Remote workVisa sponsorshipWork visa$150k - $160k
Front-End & AdTech Site Reliability Engineer (SRE)Haymarket Media, Inc. is seeking a Front-End & AdTech Site Reliability Engineer (SRE) to join the Engineering team. This position is located in our New York, NY office; three (3) days in office depending on business needs...Work at officeLocal area- ...Site Reliability Engineer I, Abhishek, would like to share a job opportunity as Site Reliability Engineer in Jacksonville, FL, Cary, NC or New York, NY (Onsite) location for a Fulltime position. In case, if you are not comfortable with this location, please share your...Full timeWork visa
- ...SRE/DevOps Engineer Versana is an industry-backed data and technology company on a mission to... ...SLOs) and indicators. Improve system reliability and resiliency. Conduct post-... ...Have: ~5+ years of experience as a Site Reliability Engineer or similar role....Work experience placementLocal area
- ...development, cloud infrastructure, DevOps, SRE, and platform engineering. You will test AI-generated commands,... ...workflows for accuracy and reliability. Work with AWS, Azure, GCP, Kubernetes... ...DevOps Cloud Infrastructure Site Reliability Engineering (SRE) Platform...Remote jobFor contractors
$140k - $155k
...Salary: $140,000 - 155,000 per year Requirements: Over 8 years of experience in Site Reliability Engineering, DevOps, or Production Engineering Demonstrated leadership capabilities as a technical lead or team supervisor, with a focus on overseeing and guiding engineers...Full time- ...Job title : Platform Engineer / SRE-DevOps Engineer Amazon Redshift Location: Remote Duration: 3+Months Key Responsibilities Design, deploy, and maintain Amazon Redshift environments and supporting AWS infrastructure. Automate infrastructure...Remote work
- The Role:GIPHY is seeking a highly experienced Site Reliability Engineer to join our SRE team. You will help design, build, operate, and evolve the infrastructure that powers GIPHY, including our cloud environment, Kubernetes clusters, and CI/CD platforms.You will also...Full timeWork experience placementRemote work
$139k - $257.55k
The ChallengeThe Adobe Creative Community CCM organization is seeking an outstanding Senior Site Reliability Engineer (SRE) to support innovation through machine learning, autonomous AI workflows, and cloud-native infrastructure. Adobe Stock gives designers and businesses...Full timeTemporary workLocal areaRemote workWorldwide$120k - $150k
...to our teams and communities.We are currently looking for a Site Reliability Engineer to join our Platform Engineering team in New York, NY.About... ...our Platform Engineering team as a Site Reliability Engineer (SRE), where you will help operate and improve the reliability of...Full time$185.5k - $232k
...Senior Site Reliability Engineer New York, NY; Boston, MA; San Francisco, CA About Formation Bio Formation Bio is a tech and AI driven pharma... ...Share knowledge and mentor engineers on infrastructure and SRE fundamentals. About You ~5+ years of relevant...Work experience placementWork at officeLocal areaRelocation3 days per week- ...Senior Site Reliability Engineer (SRE) Plenful is hiring a Senior Site Reliability Engineer (SRE) to keep our production systems reliable, performant, and scalable as we grow. This role is centered on operating real systems at scale — not just building infrastructure...Full timeWork at officeRemote workFlexible hours2 days per week
$150k - $170k
...Senior Site Reliability Engineer – Zip Co Join to apply for the Senior Site Reliability Engineer role at Zip Co At Zip, we build cloud‑native... ...experience to spearhead our Site Reliability Engineering (SRE) initiatives and mentor our engineering team. We offer a...Casual workWork at officeRemote workFlexible hours$260k - $300k
...makers of Devin, the first AI software engineer. Our team is extremely talent-dense.... ...expects. You will own both the production reliability of our user-facing products and the... ...Strong software engineering fundamentals; SRE at Cognition means writing real code, not...$104k - $178k
...Sr. Site Reliability Engineer I You will join the Site Reliability Engineering (SRE) team within DoubleVerify's Technology organization. The team is responsible for building and maintaining the reliability, scalability, and performance of DV's digital media measurement...- ...SRE Engineer Location: New York, NY, USA Exp: 8-12 Years Client: Amex Job Description: SRE Engineer (This is not a Devops role, strictly need an SRE Engineer, who has great analytical skills and is a good incident manager as well) This is an SRE role supporting...
- ...Site Reliability Engineer Mistral provides full-stack AI solutions: from frontier models to developer tools, applications, and compute. We partner... ...are seeking highly experienced Site Reliability Engineers (SRE) to shape the reliability, scalability and performance of...Relocation package
$130k - $200k
...makes AI work. The Role This is a career-level SRE role for someone who wants to own systems, not just watch... ...real surface area: the automation and tooling other engineers depend on, and the reliability of production services running AI and GPU workloads at...Shift work- ...Site Reliability Engineer Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence... .... You'll work on projects like these as part of the SRE team: Improve Baseten SRE Practices, by instrumenting...Flexible hours
$189k - $283.6k
...The Role As a member of the SRE team, you will proactively and reactively improve the reliability of Block's platform and critical infrastructure... ...~ A strong desire to perform and grow as an engineer ~5+ years of software development experience...Full timeLocal areaRemote workRelocation packageFlexible hoursShift work$111k - $218k
...The Site Reliability Engineering team designs and builds the global infrastructure on which we deploy our services, focusing on the above mentioned... ...and comply with various data sovereignty requirements. The SRE Team\'s mission is to build this increasingly complex infrastructure...Local areaWorldwideFlexible hours- ...the world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Commercial & Investment Bank... ...speed repeatable triage.Lead L1/L2 production support using SRE practices: quickly triage incidents, investigate likely causes...Shift work
$200k - $250k
Hudson River Trading (HRT) is seeking a Senior Site Reliability Engineer to join our growing Enterprise SRE team. This team is responsible for developing and maintaining productivity service infrastructure for the entire firm, both on-prem and in the cloud. They ensure...Work at officeLocal areaImmediate start- ...significant effect in a realm tailored for top achievers in site reliability. As a Lead Site Reliability Engineer at JPMorgan Chase within the Commercial & Investment... ...systems thinking and root-cause practices, expanding SRE adoption (observability, monitoring, automation,...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer (SRE). Be the first to apply!
- site reliability engineer sre New York, NY
- site reliability engineer New York, NY
- site reliability engineer remote New York, NY
- remote website tester New York, NY
- IT site lead New York, NY
- site safety New York, NY
- site merchandiser New York, NY
- website content developer New York, NY
- site leader New York, NY
- on-site clinical research associate (traveling/remote) New York, NY


