Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Customer Reliability Engineer

Fluidstack

Fluidstack

We exist to make humanity more free. For most of human history, you farmed or you starved. Technology gave people more time for the things they wanted to do, instead of things they had to do. Powerful AI will be the biggest lever for human choice we've ever built - but only if models are aligned with what humanity actually wants. There are groups building AI who don't share these goals. Whoever deploys frontier compute infrastructure fastest will decide whether AI expands human freedom or shrinks it.

We're singularly focused on delivering 10 to 100s of GWs of compute faster than anyone else, rethinking every layer of the stack. We acquire power, design and build data centers, and operate them - with teams spanning hardware and software. Speed and scale are our key differentiators. Come be a part of building civilization-scale infrastructure for AI.

We hire people who care deeply about this problem space. If that is you, please apply!

How We Operate

Be a barrel. Full autonomy. Own things end to end, take on scope without being asked, no permission required to operate outside your core role.

Insane urgency. We drive everything forward as fast as possible.

Reason from first principles. Challenge every assumption. Zero analogy thinking, no egos, the best idea wins.

Love of the game. The frontier of AI is the most interesting problem of our time. We put in long hours at high intensity to push the frontier forward.

Build something that actually matters. If you're going to spend your time, spend it on something that matters to the world.

The Production Engineering Team

Examples of key problems the team is working on

Operate at the scale of a nation, not a building. The fleet you run will draw more power than some countries, on the way to 10s to 100s of GWs.

Fly the plane while it's being built. Sites come online in pieces, and you keep the live ones running flawlessly while construction continues around them.

Write the playbook, don't inherit it. No prior operations org has run at this speed and scale, so the standards you set become the standard.

Role Scope

Own reliability for named customer workloads: their clusters, their SLAs, their escalations.

Debug across the full stack, hardware to fabric to scheduler, when a training run degrades.

Run customer-facing incident communication with technical depth and no spin.

Turn recurring customer pain into engineering fixes with the production teams.

What We're Looking For

The below is a starting point. We always make space for exceptional people, so if you don't fit this role exactly, tell us where you would.

You've supported large-scale compute customers (HPC, cloud, or AI labs) at a technical level.

You debug distributed systems methodically across layers you don't own.

You've written incident updates customers trusted more after reading.

You push internal teams to fix causes, not symptoms, and follow up until they do.

Bonus: GPU training workloads. InfiniBand or RoCE. Slurm or Kubernetes. NCCL debugging.

We are committed to pay equity and transparency.

Fluidstack is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans' status, or any other characteristic protected by law. Fluidstack will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.

You will receive a confirmation email once your application has successfully been accepted. If there is an error with your submission and you did not receive a confirmation email, please email View email address on click.appcast.io with your resume/CV, the role you've applied for, and the date you submitted your application-- someone from our recruiting team will be in touch.

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Customer Reliability Engineer in New York, NY vacancy
  • $145.68k - $178.05k

    Job Description:Reliability Engineer (Tonawanda, NY)Collaborate with Innovative 3Mers Around the WorldChoosing where to start and grow your...  ...obligations to not compete against 3M or solicit its employees or customers, both during their employment, and for a period after they... 
    Customer
    Full time
    H1b
    Flexible hours

    3M

    New York, NY
    4 days ago
  • $98.18k - $115.5k

     ...a journey to do our best. Helping the customers and businesses we serve to make better...  ...andmonitoringbestpractices.ServeasatrustedadvisortoProduct,Engineering,SRE,Infrastructure,andOperationsleaders...  ...in Observability Engineering, Site Reliability Engineering (SRE), or Reliability... 
    Customer
    Full time
    Work experience placement
    Local area
    3 days per week

    US Bank

    New York, NY
    3 days ago
  • $153k - $210k

     ...Senior Software Engineer, Site Reliability Engineering Reno, NV; San Ramon, CA; NYC - Hybrid Are you passionate about building resilient...  ...constraints and reliability risks before they impact customers. Participate in an on-call rotation, triaging production... 
    Customer
    Full time

    Ridgeline

    New York, NY
    more than 2 months ago
  • The Role:GIPHY is seeking a highly experienced Site Reliability Engineer to join our SRE team. You will help design, build, operate, and evolve...  ...focused on creating an unrivaled experience for our Customers and Contributors. Our Values represent the mindset of the employee... 
    Customer
    Full time
    Work experience placement
    Remote work

    Shutterstock

    New York, NY
    3 days ago
  • $207k - $300k

     ...providing feedback to ensure best practices in reliability, security, and efficiency.Triage and...  ...development initiatives. Mentor other engineers and contribute to the engineering...  ...help developers build more sustainably. Customers in more than 200 countries and territories... 
    Customer
    Full time
    Work at office

    Google

    New York, NY
    4 days ago
  • $139k - $257.55k

     ...Community CCM organization is seeking an outstanding Senior Site Reliability Engineer (SRE) to support innovation through machine learning,...  ...tools that unleash creativity, productivity and personalized customer experiences. Adobe’s industry-leading offerings including Adobe... 
    Customer
    Full time
    Temporary work
    Local area
    Remote work
    Worldwide

    Adobe Systems

    New York, NY
    4 days ago
  •  ...Senior Site Reliability Engineer (SRE) Our client is a global technology consulting and digital solutions company that enables enterprises...  ...expertise to help drive superior competitive differentiation, customer experiences, and business outcomes in a converging world.... 
    Customer
    Local area

    E-Solutions

    New York, NY
    2 days ago
  • $207k - $300k

     ...systems by pushing for changes that improve reliability and velocity.Practice sustainable...  ...Master's degree in Computer Science or Engineering.Experience mentoring engineers and cultivating...  ...reliability, uptime appropriate to customer's needs and a fast rate of improvement.... 
    Customer

    Google

    New York, NY
    2 days ago
  •  ...Software Reliability Engineer Good software has to run where customers need it. For many of Retool's largest customers, that means running Retool in their own infrastructure, behind their own controls, with the reliability and operational clarity they would expect... 
    Customer

    re-tool®

    New York, NY
    5 days ago
  • $120k - $180k

     ...platform for aerospace, defense, and advanced-manufacturing customers. The company has raised approximately $11M, is preparing...  ...space and defense. You will be the first dedicated Site Reliability Engineer and own critical infrastructure end to end. This is a greenfield... 
    Customer
    Permanent employment
    Full time
    Relocation package

    Raydar Inc

    New York, NY
    1 day ago
  • $100k - $250k

     ...Role Roadmap As a member of Kalshi's engineering team, you'll help build the next-...  ...You'll Do Improve observability, reliability, and service availability by defining and...  ...code that supports internal and external customer needs Debug complex technical issues... 
    Customer
    Local area

    Kalshi

    New York, NY
    1 day ago
  • $150k - $170k

     ...Senior Site Reliability Engineer – Zip Co Join to apply for the Senior Site Reliability Engineer role at Zip Co At Zip, we build cloud‑native software applications that serve millions of customers and process billions of dollars in payments. We’re looking for... 
    Customer
    Casual work
    Work at office
    Remote work
    Flexible hours

    ZIP

    New York, NY
    7 hours ago
  •  ...Site Reliability Engineer Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence...  ...management. Define and instrument SLOs and SLIs across customer workloads and internal services. Navigate ambiguity, make... 
    Customer
    Flexible hours

    Baseten

    New York, NY
    1 day ago
  • $190k - $260k

     ...capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about...  ...by building high-performance, scalable and reliable machine learning systems? Do you want to help define... 
    Customer
    Full time
    Work experience placement
    Work at office
    Local area
    Remote work
    Home office

    Cohere

    New York, NY
    2 days ago
  • $150k - $220k

     ...incredible interest from investors, demand from customers, and a need to grow our team to meet...  ...in this way. The Role: As an engineering organization, we pride ourselves on...  ..., and purpose. The Manager, Site Reliability Engineering will lead Forge’s SRE team... 
    Customer
    Local area

    Forge Global

    New York, NY
    2 days ago
  • $194k - $267k

     ...strategic priorities—like reducing costs, and doing more for your customers.If you like to be challenged and have a passion for solving...  ...on new concepts and tools.Position Overview:The Site Reliability Engineer (SRE) will play a key role in building and managing... 
    Customer
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta

    New York, NY
    3 days ago
  •  ...Site Reliability Engineer Mistral provides full-stack AI solutions: from frontier models to developer tools, applications, and compute....  ...manufacturing, defense, healthcare, and the public sector—co-creating customized AI systems that they can run on their terms. We are a... 
    Customer
    Relocation package

    Mistral AI

    New York, NY
    4 days ago
  • $189k - $283.6k

     ...money. Afterpay is transforming the way customers manage their spending over time. TIDAL...  ...proactively and reactively improve the reliability of Block's platform and critical infrastructure...  ...desire to perform and grow as an engineer ~5+ years of software development... 
    Customer
    Full time
    Local area
    Remote work
    Relocation package
    Flexible hours
    Shift work

    Block USA

    New York, NY
    1 day ago
  • $111k - $218k

     ...The Site Reliability Engineering team designs and builds the global infrastructure on which we deploy our services, focusing on the above mentioned flagship MongoDB Atlas platform. As our customers grow and globalize, our services must satisfy demands for low-latency... 
    Customer
    Local area
    Worldwide
    Flexible hours

    MongoDB

    New York, NY
    2 days ago
  • $100k - $160k

     ...General Catalyst) made up of software engineers (Jane Street, Google, Stanford, Princeton...  ...ex-attorneys. To keep up with inbound customer demand, we are quickly scaling our engineering...  ...fast, always. We're hiring a Product Reliability Engineer to own the health, stability,... 
    Customer
    Permanent employment
    Work at office

    PointOne

    New York, NY
    a month ago
  • $150k - $160k

    Front-End & AdTech Site Reliability Engineer (SRE)Haymarket Media, Inc. is seeking a Front-End & AdTech Site Reliability Engineer (SRE) to join...  ..., combining the best in content, design, production and customer services. Globalization is opening up the world further and... 
    Customer
    Work at office
    Local area

    Haymarket Media Group

    New York, NY
    2 days ago
  • $195k - $275k

     ...Deployment Planning & Release Management, and the Chief Operating Office.The Reliability Operations (RO) within WMT is responsible for providing swift, courteous, and knowledgeable customer service to end users of the production systems. This position is focused on user... 
    Customer
    Temporary work
    Work at office
    Worldwide
    Night shift

    Morgan Stanley

    New York, NY
    7 hours ago
  •  ...on what matters: judgment, strategy, and outcomes. 1,000+ customers across 50+ countries trust us, including Cleary Gottlieb,...  ...the bar. This is the place. The Role As a Staff Site Reliability Engineer you'll play a lead role on the founding SRE team at our new... 
    Customer
    Work at office

    Legora

    New York, NY
    1 day ago
  •  ...future of our cloud platform and champion engineering excellence across Ironclad. In this...  ...leadership and strategic direction for the Site Reliability Engineering team and our broader Cloud...  ..., and provide support with internal or customer-facing incidents Translate the near,... 
    Customer
    Full time
    Contract work
    Work at office

    Ironclad Inc

    New York, NY
    4 days ago
  •  ...future of our cloud platform and champion engineering excellence across Ironclad. In this...  ...leadership and strategic direction for the Site Reliability Engineering team and our broader Cloud...  ..., and provide support with internal or customer-facing incidents Translate the near... 
    Customer
    Full time
    Contract work
    Work at office

    Ironclad Inc

    New York, NY
    3 days ago
  • $145k - $160k

     ...We are seeking a specialized Observability & Infrastructure Engineer to lead the monitoring, telemetry, and platform-as-code initiatives...  ...value of diversity and inclusion in creating success for our customers, business partners, shareholders, employees and communities.... 
    Customer
    Temporary work
    H1b
    Remote work
    Flexible hours

    EPAM Systems Inc

    New York, NY
    2 days ago
  • $179k - $250k

     ...fraud, and scales compliance across the customer lifecycle so financial organizations can...  ...Infrastructure Team is a small team (6 engineers) responsible for a large and growing infrastructure...  ...isn't just scale-it's making that scale reliable, secure, and operable with less manual... 
    Customer
    Work at office
    Local area
    Immediate start
    Work from home
    Home office
    Monday to Friday
    Flexible hours

    Alloy

    New York, NY
    5 days ago
  • $131k - $164k

     ...Role OverviewYou’re a seasoned Site Reliability Engineer who loves owning complex infrastructure, making things run faster, safer, and with...  ...platforms that power mission‑critical SaaS products used by customers around the world. You’ll work across Linux, Windows Server,... 
    Customer
    Work at office
    Local area
    Visa sponsorship
    Flexible hours

    Diligent

    New York, NY
    7 hours ago
  • $189.59k - $220k

     ...Director, Site Reliability Engineering NBCUniversal is one of the world’s leading media and entertainment companies. We create world‑class...  ...Software Engineering team, responsible for leading and performing custom architectural design, implementation, monitoring, and... 
    Customer
    Full time
    For contractors
    Remote work

    NBCUniversal

    New York, NY
    7 hours ago
  • Software Reliability Engineer (SRE) The Software Reliability Engineer (SRE) will play a critical role in ensuring that our Warehouse Management...  .... Code-Level Debugging Debug application code, workflows, customizations, and interfaces to identify bugs or performance... 
    Customer
    Local area
    Remote work
    Rotating shift

    Lineage Logistics

    New York, NY
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Customer Reliability Engineer. Be the first to apply!