Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer (SRE)

re-tool®

Software Reliability Engineer

Good software has to run where customers need it. For many of Retool's largest customers, that means running Retool in their own infrastructure, behind their own controls, with the reliability and operational clarity they would expect from any critical system. Retool's Core Infrastructure team owns the systems that make this possible: Retool Cloud, managed single tenant environments, BYOC (bring-your-own-cloud) environments, Kubernetes and Helm deployments, Docker Compose, and the migration paths between them. It is a broad surface area, and it is one of the biggest levers we have for making Retool work for enterprise customers.

The work is not clean-room infrastructure. Customers run different clouds, different versions, different deployment models, and different levels of operational maturity. A bad upgrade experience can leave a customer many versions behind. A manual Terraform run can become the bottleneck during a launch or incident.

We are hiring SREs who want to turn that mess into leverage. You will help us reduce customer toil, automate upgrades and infrastructure changes, build reliability tooling across Retool Cloud and customer-owned environments, and make Retool easier to deploy and operate at enterprise scale. The strongest candidates are comfortable debugging Kubernetes, Terraform, AWS, Postgres, networking, and deployment problems, then stepping back and building the automation or product surface that prevents the same problem from happening again.

What you'll do:

  • Own reliability across Retool Cloud, managed single tenant, BYOC, and self-hosted deployment paths, including provisioning, upgrades, migrations, configuration changes, and production escalations.
  • Build the automation that turns today's manual infrastructure work into repeatable systems: Terraform runs, customer environment updates, upgrade workflows, secret rotations, and migration steps.
  • Improve observability for Retool Cloud, self-hosted customers, and internal operators. We care less about exposing every metric and more about turning health signals into clear status, likely causes, and recommended actions.
  • Design safer deployment, upgrade, and rollback paths so Cloud and managed customers can stay current
  • Help move customers from legacy or less-supported deployment models toward supported paths such as Retool's official deployment paths (Blueprints, Kubernetes, and Helm), with migration flows that are repeatable enough for customers, Support, and TAMs to trust.
  • Partner with product engineers on infrastructure requirements for new Retool products, especially when they introduce new dependencies
  • Lead through ambiguity, make careful risk calls, and communicate clearly while things are moving quickly.
  • Write the docs, runbooks, design notes, and migration guides that make complex systems understandable to other engineers and to customers.

What we're looking for:

Infrastructure fundamentals

  • Deep experience operating production infrastructure in AWS.
  • Experience improving reliability for customer-facing SaaS systems.
  • Strong Kubernetes fundamentals.
  • Real Terraform or infrastructure-as-code experience.
  • Good operational judgment around databases, especially Postgres.

Reliability and automation

  • Experience building or operating observability systems.
  • Programming ability in a language such as Go, Python, TypeScript, Java, or Ruby.
  • A bias toward automation. If you find yourself doing the same operational task twice, you should start thinking about the interface, workflow, or tool that eliminates the third time.

Customer and team judgment

  • Clear written communication.
  • Comfort working directly with customer-facing teams and, when useful, customers themselves.

What makes SREs successful here:

You will do well here if you like infrastructure that sits close to real customer pain. Some days that means debugging a specific customer environment. Other days it means improving Retool Cloud reliability or designing the migration path so the next 25 customers do not need that same debugging session. We value SREs who are ambitious, curious, energetic, and careful with the details. Retool moves quickly, priorities can change, and the systems are not always as clean as we want them to be. The work needs SREs who can get their hands dirty, tell the truth about tradeoffs, and leave the system better than they found it.

Vacancy posted 5 days ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer (SRE) in New York, NY vacancy
  •  ...Senior Site Reliability Engineer (SRE) Our client is seeking a Senior Site Reliability Engineer (SRE) with 10–15 years of experience to support front-office trading systems in a production environment. This role focuses on troubleshooting complex trading infrastructure... 
    Suggested

    TheStaffed

    New York, NY
    2 days ago
  •  ...Site Reliability Engineer (SRE) Job Title Site Reliability Engineer (SRE) Job Summary We are seeking a skilled Site Reliability Engineer (SRE) to build, automate, and maintain highly available, scalable, and reliable infrastructure and applications... 
    Suggested
    Flexible hours

    Ova Technologies

    New York, NY
    3 days ago
  •  ...We are seeking a highly motivated Site Reliability Engineer (SRE) to join the Equity Trading Platform Engineering team, supporting critical trading applications and Fidessa flows used by internal business users and external clients. This role is responsible for ensuring... 
    Suggested
    Permanent employment
    Work at office
    Afternoon shift

    RIT Solutions, Inc.

    New York, NY
    3 days ago
  •  ...Senior Site Reliability Engineer (SRE) Our client is a global technology consulting and digital solutions company that enables enterprises across industries to reimagine business models, accelerate innovation, and maximize growth by harnessing digital technologies.... 
    Suggested
    Local area

    E-Solutions

    New York, NY
    2 days ago
  • $120k - $180k

     ...leaders in space and defense. You will be the first dedicated Site Reliability Engineer and own critical infrastructure end to end. This is a...  ...and operate cloud and on-premises infrastructure as the sole SRE. Architect migrations from AWS into on-premises and air-gapped... 
    Suggested
    Permanent employment
    Full time
    Relocation package

    Raydar Inc

    New York, NY
    1 day ago
  • $207k - $300k

     ...by pushing for changes that improve reliability and velocity.Practice sustainable...  ...Master's degree in Computer Science or Engineering.Experience mentoring engineers and...  ...across cross-functional teams.Site Reliability Engineering (SRE) combines software and systems engineering... 

    Google

    New York, NY
    2 days ago
  • $80k - $95k

     ...join our dynamic team supporting the company’s users, applications, and web-based product offerings. In this role, the Site Reliability Engineer (SRE) will play a key role in maintaining resources at peak efficiency to guarantee staff are able to perform their... 
    Remote work
    Visa sponsorship
    Work visa

    RANE Network

    New York, NY
    a month ago
  • $150k - $160k

    Front-End & AdTech Site Reliability Engineer (SRE)Haymarket Media, Inc. is seeking a Front-End & AdTech Site Reliability Engineer (SRE) to join the Engineering team. This position is located in our New York, NY office; three (3) days in office depending on business needs... 
    Work at office
    Local area

    Haymarket Media Group

    New York, NY
    2 days ago
  •  ...Site Reliability Engineer I, Abhishek, would like to share a job opportunity as Site Reliability Engineer in Jacksonville, FL, Cary, NC or New York, NY (Onsite) location for a Fulltime position. In case, if you are not comfortable with this location, please share your... 
    Full time
    Work visa

    Syntricate Technologies

    New York, NY
    1 day ago
  •  ...SRE/DevOps Engineer Versana is an industry-backed data and technology company on a mission to...  ...SLOs) and indicators. Improve system reliability and resiliency. Conduct post-...  ...Have: ~5+ years of experience as a Site Reliability Engineer or similar role.... 
    Work experience placement
    Local area

    Versana

    New York, NY
    3 days ago
  •  ...development, cloud infrastructure, DevOps, SRE, and platform engineering. You will test AI-generated commands,...  ...workflows for accuracy and reliability. Work with AWS, Azure, GCP, Kubernetes...  ...DevOps Cloud Infrastructure Site Reliability Engineering (SRE) Platform... 
    Remote job
    For contractors

    YO AI Labs

    New York, NY
    2 days ago
  • $140k - $155k

     ...Salary: $140,000 - 155,000 per year Requirements: Over 8 years of experience in Site Reliability Engineering, DevOps, or Production Engineering Demonstrated leadership capabilities as a technical lead or team supervisor, with a focus on overseeing and guiding engineers... 
    Full time

    EPAM Systems

    New York, NY
    11 days ago
  •  ...Job title : Platform Engineer / SRE-DevOps Engineer Amazon Redshift Location: Remote Duration: 3+Months Key Responsibilities Design, deploy, and maintain Amazon Redshift environments and supporting AWS infrastructure. Automate infrastructure... 
    Remote work

    Conch Technologies Inc

    New York, NY
    2 days ago
  • The Role:GIPHY is seeking a highly experienced Site Reliability Engineer to join our SRE team. You will help design, build, operate, and evolve the infrastructure that powers GIPHY, including our cloud environment, Kubernetes clusters, and CI/CD platforms.You will also... 
    Full time
    Work experience placement
    Remote work

    Shutterstock

    New York, NY
    3 days ago
  • $139k - $257.55k

    The ChallengeThe Adobe Creative Community CCM organization is seeking an outstanding Senior Site Reliability Engineer (SRE) to support innovation through machine learning, autonomous AI workflows, and cloud-native infrastructure. Adobe Stock gives designers and businesses... 
    Full time
    Temporary work
    Local area
    Remote work
    Worldwide

    Adobe Systems

    New York, NY
    4 days ago
  • $120k - $150k

     ...to our teams and communities.We are currently looking for a Site Reliability Engineer to join our Platform Engineering team in New York, NY.About...  ...our Platform Engineering team as a Site Reliability Engineer (SRE), where you will help operate and improve the reliability of... 
    Full time

    Piper Sandler Companies

    New York, NY
    2 days ago
  • $185.5k - $232k

     ...Senior Site Reliability Engineer New York, NY; Boston, MA; San Francisco, CA About Formation Bio Formation Bio is a tech and AI driven pharma...  ...Share knowledge and mentor engineers on infrastructure and SRE fundamentals. About You ~5+ years of relevant... 
    Work experience placement
    Work at office
    Local area
    Relocation
    3 days per week

    Formation Bio (Formerly TrailSpark)

    New York, NY
    17 hours ago
  •  ...Senior Site Reliability Engineer (SRE) Plenful is hiring a Senior Site Reliability Engineer (SRE) to keep our production systems reliable, performant, and scalable as we grow. This role is centered on operating real systems at scale — not just building infrastructure... 
    Full time
    Work at office
    Remote work
    Flexible hours
    2 days per week

    Plenful

    New York, NY
    4 days ago
  • $150k - $170k

     ...Senior Site Reliability Engineer – Zip Co Join to apply for the Senior Site Reliability Engineer role at Zip Co At Zip, we build cloud‑native...  ...experience to spearhead our Site Reliability Engineering (SRE) initiatives and mentor our engineering team. We offer a... 
    Casual work
    Work at office
    Remote work
    Flexible hours

    ZIP

    New York, NY
    7 hours ago
  • $260k - $300k

     ...makers of Devin, the first AI software engineer. Our team is extremely talent-dense....  ...expects. You will own both the production reliability of our user-facing products and the...  ...Strong software engineering fundamentals; SRE at Cognition means writing real code, not... 

    Cognition AI

    New York, NY
    4 days ago
  • $104k - $178k

     ...Sr. Site Reliability Engineer I You will join the Site Reliability Engineering (SRE) team within DoubleVerify's Technology organization. The team is responsible for building and maintaining the reliability, scalability, and performance of DV's digital media measurement... 

    DoubleVerify

    New York, NY
    1 day ago
  •  ...SRE Engineer Location: New York, NY, USA Exp: 8-12 Years Client: Amex Job Description: SRE Engineer (This is not a Devops role, strictly need an SRE Engineer, who has great analytical skills and is a good incident manager as well) This is an SRE role supporting... 

    IVID TEK INC

    New York, NY
    4 days ago
  •  ...Site Reliability Engineer Mistral provides full-stack AI solutions: from frontier models to developer tools, applications, and compute. We partner...  ...are seeking highly experienced Site Reliability Engineers (SRE) to shape the reliability, scalability and performance of... 
    Relocation package

    Mistral AI

    New York, NY
    4 days ago
  • $130k - $200k

     ...makes AI work. The Role This is a career-level SRE role for someone who wants to own systems, not just watch...  ...real surface area: the automation and tooling other engineers depend on, and the reliability of production services running AI and GPU workloads at... 
    Shift work

    nScale

    New York, NY
    1 day ago
  •  ...Site Reliability Engineer Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence...  .... You'll work on projects like these as part of the SRE team: Improve Baseten SRE Practices, by instrumenting... 
    Flexible hours

    Baseten

    New York, NY
    1 day ago
  • $189k - $283.6k

     ...The Role As a member of the SRE team, you will proactively and reactively improve the reliability of Block's platform and critical infrastructure...  ...~ A strong desire to perform and grow as an engineer ~5+ years of software development experience... 
    Full time
    Local area
    Remote work
    Relocation package
    Flexible hours
    Shift work

    Block USA

    New York, NY
    1 day ago
  • $111k - $218k

     ...The Site Reliability Engineering team designs and builds the global infrastructure on which we deploy our services, focusing on the above mentioned...  ...and comply with various data sovereignty requirements. The SRE Team\'s mission is to build this increasingly complex infrastructure... 
    Local area
    Worldwide
    Flexible hours

    MongoDB

    New York, NY
    2 days ago
  •  ...the world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Commercial & Investment Bank...  ...speed repeatable triage.Lead L1/L2 production support using SRE practices: quickly triage incidents, investigate likely causes... 
    Shift work

    JP Morgan Chase

    New York, NY
    3 days ago
  • $200k - $250k

    Hudson River Trading (HRT) is seeking a Senior Site Reliability Engineer to join our growing Enterprise SRE team. This team is responsible for developing and maintaining productivity service infrastructure for the entire firm, both on-prem and in the cloud. They ensure... 
    Work at office
    Local area
    Immediate start

    Hudson River Trading

    New York, NY
    3 days ago
  •  ...significant effect in a realm tailored for top achievers in site reliability. As a Lead Site Reliability Engineer at JPMorgan Chase within the Commercial & Investment...  ...systems thinking and root-cause practices, expanding SRE adoption (observability, monitoring, automation,... 

    JP Morgan Chase

    New York, NY
    7 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer (SRE). Be the first to apply!