Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Customer Reliability Engineer

Fluidstack

FluidstackWe exist to make humanity more free. For most of human history, you farmed or you starved. Technology gave people more time for the things they wanted to do, instead of things they had to do. Powerful AI will be the biggest lever for human choice we've ever built - but only if models are aligned with what humanity actually wants. There are groups building AI who don't share these goals. Whoever deploys frontier compute infrastructure fastest will decide whether AI expands human freedom or shrinks it.We're singularly focused on delivering 10 to 100s of GWs of compute faster than anyone else, rethinking every layer of the stack. We acquire power, design and build data centers, and operate them - with teams spanning hardware and software. Speed and scale are our key differentiators. Come be a part of building civilization-scale infrastructure for AI.We hire people who care deeply about this problem space. If that is you, please apply!How We OperateBe a barrel. Full autonomy. Own things end to end, take on scope without being asked, no permission required to operate outside your core role.Insane urgency. We drive everything forward as fast as possible.Reason from first principles. Challenge every assumption. Zero analogy thinking, no egos, the best idea wins.Love of the game. The frontier of AI is the most interesting problem of our time. We put in long hours at high intensity to push the frontier forward.Build something that actually matters. If you're going to spend your time, spend it on something that matters to the world.The Production Engineering TeamExamples of key problems the team is working onOperate at the scale of a nation, not a building. The fleet you run will draw more power than some countries, on the way to 10s to 100s of GWs.Fly the plane while it's being built. Sites come online in pieces, and you keep the live ones running flawlessly while construction continues around them.Write the playbook, don't inherit it. No prior operations org has run at this speed and scale, so the standards you set become the standard.Role ScopeOwn reliability for named customer workloads: their clusters, their SLAs, their escalations.Debug across the full stack, hardware to fabric to scheduler, when a training run degrades.Run customer-facing incident communication with technical depth and no spin.Turn recurring customer pain into engineering fixes with the production teams.What We're Looking ForThe below is a starting point. We always make space for exceptional people, so if you don't fit this role exactly, tell us where you would.You've supported large-scale compute customers (HPC, cloud, or AI labs) at a technical level.You debug distributed systems methodically across layers you don't own.You've written incident updates customers trusted more after reading.You push internal teams to fix causes, not symptoms, and follow up until they do.Bonus: GPU training workloads. InfiniBand or RoCE. Slurm or Kubernetes. NCCL debugging.We are committed to pay equity and transparency.Fluidstack is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans' status, or any other characteristic protected by law. Fluidstack will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.You will receive a confirmation email once your application has successfully been accepted. If there is an error with your submission and you did not receive a confirmation email, please email View email address on click.appcast.io with your resume/CV, the role you've applied for, and the date you submitted your application-- someone from our recruiting team will be in touch.

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Customer Reliability Engineer in New York, NY vacancy
  • $98.18k - $115.5k

     ...S. Bank, we’re on a journey to do our best. Helping the customers and businesses we serve to make better and smarter financial...  ..., and each other. Job DescriptionResponsibilitiesThe Reliability Observability Engineer 3 is responsible for enabling reliable, measurable, and... 
    Customer
    Full time
    Work experience placement
    Local area
    3 days per week

    US Bank

    New York, NY
    2 days ago
  •  ...About this role: As an Airflow Reliability Engineer on the Customer Reliability Engineering (CRE) team at Astronomer, you will have the opportunity to become an Apache Airflow expert, learning directly from leaders of the Airflow project. You’ll provide Apache Airflow... 
    Customer
    Full time
    Remote work
    Weekend work

    Astronomer

    New York, NY
    a month ago
  • $153k - $210k

     ...Senior Software Engineer, Site Reliability Engineering Reno, NV; San Ramon, CA; NYC - Hybrid Are you passionate about building resilient...  ...constraints and reliability risks before they impact customers. Participate in an on-call rotation, triaging production... 
    Customer
    Full time

    Ridgeline

    New York, NY
    13 hours ago
  •  ...organization is expanding, and we are seeking a Senior Site Reliability Engineer to help drive a major architectural modernization. In this...  ...learn more about our culture and commitment to our people, customers, community, environment, and shareholders. Equal... 
    Customer
    Permanent employment
    Full time
    H1b
    Local area
    Remote work
    Shift work

    Jack Henry & Associates

    New York, NY
    13 hours ago
  • $207k - $300k

     ...providing feedback to ensure best practices in reliability, security, and efficiency.Triage and...  ...development initiatives. Mentor other engineers and contribute to the engineering...  ...help developers build more sustainably. Customers in more than 200 countries and territories... 
    Customer
    Full time
    Work at office

    Google

    New York, NY
    2 days ago
  • $141k - $216.6k

     ...candor and care, seeking out diverse perspectives from our customers, communities and each other.Life at Axon is fast-paced, challenging...  ...a safer, more connected world.Position OverviewAs a Site Reliability Engineer, you'll own the reliability, observability, and operational... 
    Customer
    Work experience placement
    Work at office

    Axon

    New York, NY
    4 days ago
  • $139k - $257.55k

     ...Community CCM organization is seeking an outstanding Senior Site Reliability Engineer (SRE) to support innovation through machine learning,...  ...tools that unleash creativity, productivity and personalized customer experiences. Adobe’s industry-leading offerings including Adobe... 
    Customer
    Full time
    Temporary work
    Local area
    Remote work
    Worldwide

    Adobe Systems

    New York, NY
    3 days ago
  • $45 - $85 per hour

    DescriptionThe Site Reliability Engineering groups goal is to ensure Customers can always use the service reliably.We're looking for engineers to be part of an empowered, self-organizing group, with the opportunity to use modern languages and tools and to operate software... 
    Customer
    Contract work
    Temporary work

    TEKsystems

    New York, NY
    1 day ago
  • $167.7k - $245.2k

     ...technology portfolio and beyond, helping customers deploy at scale while also delivering...  ...very effective.We’re looking for talented engineers with a software or operations...  ...application development teams to ensure the reliability, performance and security of our infrastructure... 
    Customer
    Full time
    Temporary work
    Work at office
    Local area
    Flexible hours
    1 day per week

    CISCO Systems

    New York, NY
    13 hours ago
  • $190k - $260k

     ...capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about...  ...by building high-performance, scalable and reliable machine learning systems? Do you want to help define... 
    Customer
    Full time
    Work experience placement
    Work at office
    Local area
    Remote work
    Home office

    Cohere

    New York, NY
    1 day ago
  • $150k - $220k

     ...incredible interest from investors, demand from customers, and a need to grow our team to meet...  ...in this way. The Role: As an engineering organization, we pride ourselves on...  ..., and purpose. The Manager, Site Reliability Engineering will lead Forge’s SRE team... 
    Customer
    Local area

    Forge Global

    New York, NY
    1 day ago
  • $194k - $267k

     ...strategic priorities—like reducing costs, and doing more for your customers.If you like to be challenged and have a passion for solving...  ...on new concepts and tools.Position Overview:The Site Reliability Engineer (SRE) will play a key role in building and managing... 
    Customer
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta

    New York, NY
    2 days ago
  • $194k - $267k

     ...talk.The TeamWe are looking for an experienced Staff Site Reliability Engineer to join Okta's Emerging Products Group (EPG). Our mission is...  ...highly reliable, scalable, and secure cloud services that our customers can trust. We embrace an automation-first mindset and... 
    Customer
    Local area
    Worldwide
    Flexible hours

    Okta

    New York, NY
    4 days ago
  • $100k - $160k

     ...General Catalyst) made up of software engineers (Jane Street, Google, Stanford, Princeton...  ...ex-attorneys. To keep up with inbound customer demand, we are quickly scaling our engineering...  ...fast, always. We're hiring a Product Reliability Engineer to own the health, stability,... 
    Customer
    Permanent employment
    Work at office

    PointOne

    New York, NY
    19 days ago
  • $129k - $152k

     ...and are a growing team of world-class engineering, operations, medical affairs, marketing...  ...seeking a highly skilled, experienced Site Reliability Engineer (SRE) to join the core...  ...to Cleerly's development partners and customers.   The ideal candidate possesses strong... 
    Customer
    Full time
    Remote work

    Cleerly

    New York, NY
    13 hours ago
  • $195k - $275k

     ...Deployment Planning & Release Management, and the Chief Operating Office.The Reliability Operations (RO) within WMT is responsible for providing swift, courteous, and knowledgeable customer service to end users of the production systems. This position is focused on user... 
    Customer
    Temporary work
    Work at office
    Worldwide
    Night shift

    Morgan Stanley

    New York, NY
    4 days ago
  • $100k - $250k

     ...Role Roadmap As a member of Kalshi's engineering team, you'll help build the next-...  ...You'll Do Improve observability, reliability, and service availability by defining and...  ...code that supports internal and external customer needs Debug complex technical issues... 
    Customer
    Local area

    Kalshi Inc

    New York, NY
    5 days ago
  • $150k - $160k

    Front-End & AdTech Site Reliability Engineer (SRE)Haymarket Media, Inc. is seeking a Front-End & AdTech Site Reliability Engineer (SRE) to join...  ..., combining the best in content, design, production and customer services. Globalization is opening up the world further and... 
    Customer
    Work at office
    Local area

    Haymarket Media Group

    New York, NY
    1 day ago
  •  ...on what matters: judgment, strategy, and outcomes. 1,000+ customers across 50+ countries trust us, including Cleary Gottlieb,...  ...bar. This is the place. The role As a Senior Site Reliability Engineer you'll join the founding SRE team at our new NYC engineering... 
    Customer
    Work at office

    Legora

    New York, NY
    5 days ago
  •  ...Triomics Backend Engineer Triomics is building the agentic AI layer for oncology EHRs. Cancer hospitals spend billions on highly trained...  ...clinical documents monthly across multi-tenant deployments in customer as well as Triomics cloud environments, with GPU infrastructure... 
    Customer
    Day shift

    Triomics

    New York, NY
    4 days ago
  • $111k - $218k

     ...The Site Reliability Engineering team designs and builds the global infrastructure on which we deploy our services, focusing on the above mentioned flagship MongoDB Atlas platform. As our customers grow and globalize, our services must satisfy demands for low-latency... 
    Customer
    Local area
    Worldwide
    Flexible hours

    MongoDB

    New York, NY
    5 days ago
  • $115k - $125k

     ...Site Reliability Engineer New York City, NY Pico fuels the global capital markets community by providing exceptional market data services and customized managed infrastructure solutions. As financial industry experts at the center of markets and technology, we help... 
    Customer
    Work experience placement
    Work at office
    Work from home
    Monday to Friday
    Flexible hours
    Shift work
    Weekend work
    Afternoon shift
    Early shift

    Pico

    New York, NY
    5 days ago
  • $115k - $160k

     ...Senior Site Reliability Engineer - AVP - Credit Trade FloorEmbark on a transformative journey as a Senior Site Reliability Engineer - AVP...  ...support model and service offering to improve the service to customers and stakeholders.Execution of preventative maintenance tasks... 
    Customer
    Work at office

    hackajob

    New York, NY
    1 day ago
  •  ...generated incredible interest from investors, demand from customers, and a need to grow our team to meet the needs of more companies...  ...innovators in this way. The RoleThe Director of Platform & Reliability Engineering will lead a critical engineering organization responsible... 
    Customer
    Work at office
    Local area
    2 days per week
    3 days per week

    Forge Global

    New York, NY
    2 days ago
  • $120k - $180k

     ...platform for aerospace, defense, and advanced-manufacturing customers. The company has raised approximately $11M, is preparing...  ...space and defense. You will be the first dedicated Site Reliability Engineer and own critical infrastructure end to end. This is a greenfield... 
    Customer
    Permanent employment
    Full time
    Relocation package

    Raydar

    New York, NY
    19 days ago
  • $18 - $50 per hour

     ...focuses on growth, so our people, our business, and our customers can achieve their full potential. We're currently recruiting...  ...the program Position Overview: As an Electronics Reliability Product Engineering Intern , you will support the development of Simcenter... 
    Customer
    Remote job
    Hourly pay
    Full time
    Internship
    Work at office
    Local area

    Siemens

    New York, NY
    13 hours ago
  •  ..., more resilient organizations. The Role: As a Senior Site Reliability Engineer (SRE) at Dune Security, you will play a critical role in ensuring...  ...teams to proactively address issues before they impact customers. Key Responsibilities: Design, implement, and maintain a... 
    Customer
    Full time
    Work at office

    Dune Security

    New York, NY
    3 days ago
  • $191k - $226k

     ...else can. About the role: We are seeking a Senior Site Reliability Engineer to own the reliability, performance, and resilience of the cloud...  ...our internal platform to the same rigorous standards as our customer-facing products Enable Engineering: Build and maintain... 
    Customer
    Remote work
    Work visa
    Flexible hours

    Garner Health

    New York, NY
    10 days ago
  • $115k - $160k

     ...repeat issues. We need strong systems engineering expertise, with deep knowledge of...  ...delivering effective fixes that preserve reliability and performance. We require hands-on...  ...service offering to improve outcomes for customers and stakeholders. You will carry out... 
    Customer
    Full time
    Work at office

    Barclays

    New York, NY
    5 days ago
  • $80k - $95k

     ...and web-based product offerings. In this role, the Site Reliability Engineer (SRE) will play a key role in maintaining resources at peak...  ...to perform their functions efficiently and to ensure a good customer experience across RANE’s product offerings.The ideal candidate... 
    Customer
    Remote work
    Visa sponsorship
    Work visa

    RANE Network

    New York, NY
    a month ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Customer Reliability Engineer. Be the first to apply!