Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Site Reliability Engineer

Parasail

AI companies need inference that’s fast, reliable, and economical at scale. Parasail delivers it. We’re building an enterprise-grade inference cloud for open-weight models where customers pay for the tokens they use, and we handle everything required to serve them.

Behind one OpenAI-compatible API, we pool GPU capacity from providers around the world and continuously optimize where and how workloads run. That means turning a changing mix of hardware, networks, and infrastructure into a service customers can trust.

We’ve raised a $32 million Series A, and we’re scaling beyond trillions of tokens a day. You’ll have the ownership and reach to shape how we get there.

The Role

At Parasail, reliability is an engineering problem that spans the entire stack. A GPU fails. A provider goes down. Traffic spikes. Customers still expect their inference to work.

We’re hiring Site Reliability Engineers to build the systems that make that possible. You’ll own infrastructure across our global GPU fleet, write software that automates operations, and make the platform better at detecting, surviving, and recovering from failures.

You’ll work directly with infrastructure, platform, and inference engineers in a flat organization. We welcome SREs, software engineers, platform engineers, and systems engineers who want to build ambitious systems and take responsibility for how they perform in production.

What You’ll Do
  • Scale a global GPU fleet. Build and improve the Kubernetes infrastructure behind provisioning, networking, storage, and service deployment across providers and regions.

  • Make failure survivable. Design better isolation, failover, and recovery so hardware and infrastructure failures have less impact on customers.

  • Build software that runs infrastructure. Automate capacity expansion, deployments, and maintenance, eliminating manual work and making changes safer.

  • Make the system understandable. Develop observability and diagnostics that reveal bottlenecks, surface failures, and help engineers act quickly.

  • Own the production feedback loop. Respond to incidents, get to the root cause, and turn what you learn into stronger systems.

  • Push the platform forward. Work across the stack to improve performance, utilization, security, and reliability as inference demand grows.

What You Bring
  • Experience building and operating production infrastructure or distributed systems, with real ownership of reliability.

  • Strong Linux fundamentals and practical knowledge of networking, storage, and containers.

  • Hands-on experience running Kubernetes in production.

  • The ability to write maintainable software and automation to solve infrastructure problems.

  • A systematic approach to debugging problems that cross application, cluster, network, and hardware boundaries.

  • Good judgment about when to move quickly, when to simplify, and where reliability matters most.

  • The initiative to take a problem from investigation through implementation and work closely with teammates along the way.

Your strongest skill might be software development, distributed systems, or infrastructure operations. We’re building a team with complementary strengths; your previous job title matters less than what you can build and own.

Nice to Have
  • Experience with multi-region, multi-provider, or bare-metal infrastructure.

  • Familiarity with GPUs, model serving, or inference systems such as vLLM or SGLang.

  • Experience with infrastructure as code, CI/CD, observability, or automated recovery.

  • Experience building highly available services, multi-tenant platforms, or distributed data systems.

Why Join Parasail

The systems you build will determine how reliably and efficiently customers can run AI in production. You’ll work close to the hardware, deep in distributed systems, and alongside engineers optimizing the inference stack.

This is a small team tackling problems at substantial scale. You’ll own meaningful architecture decisions, ship improvements directly into production, and help build the foundation for the next stage of AI infrastructure.

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Senior Site Reliability Engineer in United States vacancy
  •  ...Senior Site Reliability Engineer Company: CyberArk Work Type: Remote Employment: Full Time Location: US Seniority: Mid Level Technologies: AWS, Kubernetes, Terraform, CloudFormation, Ansible, CloudWatch, Grafana, Datadog, OpenSearch, PagerDuty Requirements: Senior SRE... 
    Senior
    Full time
    Remote work

    CyberArk

    United States
    1 day ago
  •  ...Senior Site Reliability Engineer Company: Sphera Work Type: Remote Employment: Full Time Location: US Seniority: Senior Level Technologies: Terraform, ARM templates, Kubernetes, Azure, SonarCloud, CheckPoint, Hadoop, Kafka, Presto, NewRelic, CI/CD, Linux, Windows, Redis... 
    Senior
    Full time
    Remote work

    Sphera

    United States
    1 day ago
  •  ...Senior Site Reliability Engineer Company: ZetaChain Work Type: Remote Employment: Full Time Location: US Seniority: Senior Level Technologies: Go, Python, Bash, Terraform, Ansible, Kubernetes, Docker, Linux, Prometheus, Grafana, Datadog, Loki, incident.io, AWS, GCP, Bare... 
    Senior
    Full time
    Remote work

    ZetaChain

    United States
    1 day ago
  •  ...Senior Site Reliability Engineer Company: Filevine Work Type: Remote Employment: Full Time Location: US Seniority: Senior Level Technologies: Python, Bash, PowerShell, AWS, Kubernetes, EKS, CloudWatch, Lambda, S3, IAM, CI/CD, Monitoring Requirements: 8+ years in software... 
    Senior
    Full time
    Remote work

    Filevine

    United States
    1 day ago
  •  ...Job Summary We are seeking a Senior SRE / DevSecOps Engineer with strong experience in Kubernetes, AWS...  ...troubleshooting. The role will focus on platform reliability, incident management, SLO/SLI...  ...with SLO/SLI governance and site reliability practices. ~ Strong understanding... 
    Senior

    PB consulting

    Stallings, NC
    2 days ago
  •  ...Senior Site Reliability Engineer Company: Regrello Work Type: Remote Employment: Full Time Location: US Seniority: Senior Level Technologies: Go, AWS, Azure, GCP, Kubernetes, Terraform, CI/CD, GitHub Actions, GitLab CI, CircleCI, Helm, Helmfile, LaunchDarkly, Otel, Prometheus... 
    Senior
    Full time
    Remote work

    Regrello

    United States
    1 day ago
  •  .... Our infra has to match. The role We\'re looking for a Senior SRE to own the reliability, scalability, and operational posture of Satsuma\'s multi...  ...AI-assisted development workflows Partner closely with engineering on reliability reviews and architecture decisions Requirements... 
    Senior

    Satsuma

    Austin, TX
    23 hours ago
  •  ...Senior Site Reliability Engineer We are looking for a Senior Site Reliability Engineer with Cloud platform experience. This individual will be part of a team responsible for operating and maintaining production clusters and developing our observability solutions; they... 
    Senior
    Remote work

    Omilia - Conversational Intelligence

    United States
    23 hours ago
  •  ...day is an operations job. Coalfire organizes its delivery engineering into capability-focused teams, and the Run teams keep authorized...  ...provably compliant long after the build team leaves. As a Senior Site Reliability Engineer you own one operational capability, such as... 
    Senior
    Full time

    Coalfire

    Remote
    8 days ago
  •  ...Job Title Location Remote - United States Job Category Information Technology, Platform Engineering, Site Reliability Engineering Industry Computer Software, SaaS, National Security Employee Type FT Exempt Manage Others No Minimum Experience 5 Years... 
    Senior
    Remote work

    CenCore

    United States
    2 days ago
  •  ...Senior Site Reliability Engineer (SRE) Our client is a global technology consulting and digital solutions company that enables enterprises across industries to reimagine business models, accelerate innovation, and maximize growth by harnessing digital technologies.... 
    Senior
    Local area

    E-Solutions

    New York, NY
    4 days ago
  •  ...Seeking a highly skilled Senior Site Reliability Engineer, Databases to join the Platform Engineering team in a full-time remote capacity, where the engineer will manage database monitoring, backup verification, disaster recovery testing, and incident response for MySQL... 
    Senior
    Full time
    Remote work

    Virtual Vocations Inc

    United States
    3 days ago
  • $152k - $195k

     ...investors including Silver Lake Waterman, Moody’s, Sequoia Capital, GV and Riverwood Capital. About the Team: As a Senior Site Reliability Engineer, you will be a key technical leader driving the design and optimization of our Kubernetes-based infrastructure and CI/... 
    Senior
    Remote work

    SecurityScorecard

    United States
    2 days ago
  •  ...About the RoleWe’re looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You’ll partner with engineers and data scientists to build, automate, and maintain the... 
    Senior

    Alembic Limited

    San Francisco, CA
    3 days ago
  • $182.8k - $247.3k

     ...mission to develop education for our half a billion (and growing!) learners around the world. About the role... As a Senior Site Reliability Engineer, you will work closely with both product and platform engineering teams to ensure Duolingo’s sophisticated distributed... 
    Senior
    Work experience placement

    Socket

    Eastern, KY
    23 hours ago
  •  ...IXL Learning, developer of personalized learning products used by millions of people globally, is seeking a Senior Site Reliability Engineer to join our team, and help maintain the reliability and optimal performance of our products. We are seeking engineers with a passion... 
    Senior
    Work at office
    Immediate start

    IXL Learning

    Raleigh, NC
    3 days ago
  •  ...Senior Sre We're hiring a Senior SRE based in Latin America to work alongside our US-based engineering team, building out observability, on-call coverage, and deployment automation for a client with strict compliance requirements. We're specifically looking for someone... 
    Senior
    Full time
    Remote work

    zipdev

    United States
    1 day ago
  • $210k - $240k

     ...Join to apply for the Senior Site Reliability Engineer role at Alembic Technologies This range is provided by Alembic Technologies. Your actual pay will be based on your skills and experience — talk with your recruiter to learn more. Base pay range $210,000.00/yr - $2... 
    Senior
    Full time

    Alembic Technologies

    San Francisco, CA
    23 hours ago
  • Job Posting Datavant recognizes the importance of information security and data privacy, including in its hiring processes and recruitment. Datavant encourages all potential job applicants to take precautions against potential phishing schemes or other scams that improperly...
    Senior
    Fixed term contract
    Work at office
    Local area
    Remote work

    Datavant

    United States
    4 days ago
  •  ...healthcare organizations maintain accurate, compliant, and reliable provider networks at scale. Our vision is simple: One...  ...of patients. About the Role We're looking for a Senior Site Reliability Engineer who takes ownership seriously - someone who designs for... 
    Senior
    Remote work

    CertifyOS

    United States
    4 days ago
  • $175k - $250k

     ...Senior Cloud Infrastructure Engineer Location: San Francisco, CA. Remote unavailable. Modality: On‑Site only. Must live within commuting distance of San Francisco or be willing to...  ...ensuring scalability, performance, and reliability across environments. What You’ll Do Design... 
    Senior
    Full time
    Remote work
    Relocation
    Relocation package

    The Recruiting Guy

    Washington DC
    2 days ago
  • $200k - $240k

     ...problems and help health systems deliver better care, we'd love to meet you! About the role We're looking for a Senior Site Reliability Engineer to join our Infrastructure Engineering team and get their hands directly into the systems that keep our healthcare platform... 
    Senior
    Work at office
    3 days per week

    Socket

    New York, NY
    2 days ago
  •  ...Get AI-powered advice on this job and more exclusive features. Direct message the job poster from Algo Capital Group Senior Site Reliability Engineer - Observability and Automation A leading high-frequency trading firm is seeking a mid to senior-level Site Reliability... 
    Senior
    Full time
    Work at office
    Flexible hours

    Algo Capital Group

    Chicago, IL
    2 days ago
  • $168k - $200k

     ...that is passionate about creating transformative change in healthcare. What We're Looking For We're looking for a Senior Site Reliability Engineer to join our Data & ML Platform team. You'll be at the forefront of building and operating a resilient, observable,... 
    Senior
    Remote work

    Datavant

    United States
    3 days ago
  • $140k - $180k

     ...design. As Endear grows, we are investing in the reliability and platform systems that keep our product fast, resilient, and easy for engineering teams to operate. Position Overview We’re hiring a Senior Site Reliability Engineer to become Endear’s first... 
    Senior
    Remote work
    Work from home
    Home office
    Flexible hours

    Endear

    United States
    1 day ago
  • $180k - $230k

     ...Power Acceleration Job Description We're looking for a Senior SRE to own the reliability, scalability, and observability of our production systems. You'll work closely with platform and data engineering to keep high-throughput, data-intensive services running at the... 
    Senior
    Work at office
    Local area
    Immediate start
    Remote work
    3 days per week

    GridCARE, Inc.

    Eastern, KY
    2 days ago
  •  ...role, we encourage you to apply. The Role  As a Senior Platform Engineer, you are a champion for DevOps and SRE culture and industry...  .... \n What You Will Be Doing Improving production reliability and system resilience within an SRE scoped team Championing... 
    Senior
    Remote work
    Flexible hours

    Megaport

    United States
    23 hours ago
  •  ...training, where they work alongside former pro and D1 athletes coaching them at the highest standard. We are hiring a Senior Site Reliability Engineer on a part-time consulting contract to audit our systems and make the changes required to keep them running as we scale... 
    Senior
    Full time
    Contract work
    Part time
    Remote work
    Flexible hours

    Texas Sports Academy Main

    United States
    1 day ago
  •  ...Senior Site Reliability Engineer Remote – Home Based Job Summary We’re partnering with a company in the SaaS space to find a Senior Site Reliability Engineer . In this role, you’ll be part of the IT Operations group responsible for maintaining all environments... 
    Senior
    Temporary work
    Remote work
    Work from home
    Flexible hours

    SourceDirect Talent

    United States
    3 days ago
  • $118k - $177k

     ...Everforth ECS is seeking a  Senior Site Reliability Engineer  to work remotely . Everforth ECS is seeking talented professionals to join our successful and growing team in building the next-generation Continuous Diagnostics and Mitigation (CDM) Cyber data solution... 
    Senior
    Remote work

    ECS Tech Inc

    United States
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!