Senior Site Reliability Engineer
Parasail
Parasail Site Reliability Engineer Role
AI companies need inference that's fast, reliable, and economical at scale. Parasail delivers it. We're building an enterprise-grade inference cloud for open-weight models where customers pay for the tokens they use, and we handle everything required to serve them.
Behind one OpenAI-compatible API, we pool GPU capacity from providers around the world and continuously optimize where and how workloads run. That means turning a changing mix of hardware, networks, and infrastructure into a service customers can trust.
We've raised a $32 million Series A, and we're scaling beyond trillions of tokens a day. You'll have the ownership and reach to shape how we get there.
The Role
At Parasail, reliability is an engineering problem that spans the entire stack. A GPU fails. A provider goes down. Traffic spikes. Customers still expect their inference to work.
We're hiring Site Reliability Engineers to build the systems that make that possible. You'll own infrastructure across our global GPU fleet, write software that automates operations, and make the platform better at detecting, surviving, and recovering from failures.
You'll work directly with infrastructure, platform, and inference engineers in a flat organization. We welcome SREs, software engineers, platform engineers, and systems engineers who want to build ambitious systems and take responsibility for how they perform in production.
What You'll Do
Scale a global GPU fleet. Build and improve the Kubernetes infrastructure behind provisioning, networking, storage, and service deployment across providers and regions.
Make failure survivable. Design better isolation, failover, and recovery so hardware and infrastructure failures have less impact on customers.
Build software that runs infrastructure. Automate capacity expansion, deployments, and maintenance, eliminating manual work and making changes safer.
Make the system understandable. Develop observability and diagnostics that reveal bottlenecks, surface failures, and help engineers act quickly.
Own the production feedback loop. Respond to incidents, get to the root cause, and turn what you learn into stronger systems.
Push the platform forward. Work across the stack to improve performance, utilization, security, and reliability as inference demand grows.
What You Bring
Experience building and operating production infrastructure or distributed systems, with real ownership of reliability.
Strong Linux fundamentals and practical knowledge of networking, storage, and containers.
Hands-on experience running Kubernetes in production.
The ability to write maintainable software and automation to solve infrastructure problems.
A systematic approach to debugging problems that cross application, cluster, network, and hardware boundaries.
Good judgment about when to move quickly, when to simplify, and where reliability matters most.
The initiative to take a problem from investigation through implementation and work closely with teammates along the way.
Your strongest skill might be software development, distributed systems, or infrastructure operations. We're building a team with complementary strengths; your previous job title matters less than what you can build and own.
Nice to Have
Experience with multi-region, multi-provider, or bare-metal infrastructure.
Familiarity with GPUs, model serving, or inference systems such as vLLM or SGLang.
Experience with infrastructure as code, CI/CD, observability, or automated recovery.
Experience building highly available services, multi-tenant platforms, or distributed data systems.
Why Join Parasail
The systems you build will determine how reliably and efficiently customers can run AI in production. You'll work close to the hardware, deep in distributed systems, and alongside engineers optimizing the inference stack.
This is a small team tackling problems at substantial scale. You'll own meaningful architecture decisions, ship improvements directly into production, and help build the foundation for the next stage of AI infrastructure.
- ...Senior Site Reliability Engineer Company: CyberArk Work Type: Remote Employment: Full Time Location: US Seniority: Mid Level Technologies: AWS, Kubernetes, Terraform, CloudFormation, Ansible, CloudWatch, Grafana, Datadog, OpenSearch, PagerDuty Requirements: Senior SRE...SeniorFull timeRemote work
- ...Senior Site Reliability Engineer Company: Sphera Work Type: Remote Employment: Full Time Location: US Seniority: Senior Level Technologies: Terraform, ARM templates, Kubernetes, Azure, SonarCloud, CheckPoint, Hadoop, Kafka, Presto, NewRelic, CI/CD, Linux, Windows, Redis...SeniorFull timeRemote work
- ...day is an operations job. Coalfire organizes its delivery engineering into capability-focused teams, and the Run teams keep authorized... ...provably compliant long after the build team leaves. As a Senior Site Reliability Engineer you own one operational capability, such as...SeniorFull time
- ...Job Title: Senior Site Reliability Engineer (SRE) Work Location: Southlake, TX 76092 Contract duration: 12 months Interview Mode- In-person Interview Job Details: Must Have Skills:: SRE Grafana Python Nice to have skills AI Cloud...SeniorContract work
- ...Senior Site Reliability Engineer We are looking for a Senior Site Reliability Engineer with Cloud platform experience. This individual will be part of a team responsible for operating and maintaining production clusters and developing our observability solutions; they...SeniorRemote work
$191k - $226k
...and using AI to scale that impact further and faster than anyone else can. About the role: We are seeking a Senior Site Reliability Engineer to own the reliability, performance, and resilience of the cloud infrastructure powering Garner’s products and AI/ML workloads...SeniorRemote workWork visaFlexible hours- ...Job Title Location Remote - United States Job Category Information Technology, Platform Engineering, Site Reliability Engineering Industry Computer Software, SaaS, National Security Employee Type FT Exempt Manage Others No Minimum Experience 5 Years...SeniorRemote work
$152k - $195k
...investors including Silver Lake Waterman, Moody’s, Sequoia Capital, GV and Riverwood Capital. About the Team: As a Senior Site Reliability Engineer, you will be a key technical leader driving the design and optimization of our Kubernetes-based infrastructure and CI/...SeniorRemote work- ...Job Summary We are seeking a Senior SRE / DevSecOps Engineer with strong experience in Kubernetes, AWS... ...troubleshooting. The role will focus on platform reliability, incident management, SLO/SLI... ...with SLO/SLI governance and site reliability practices. ~ Strong understanding...SeniorTemporary work
$182.8k - $247.3k
...mission to develop education for our half a billion (and growing!) learners around the world. About the role... As a Senior Site Reliability Engineer, you will work closely with both product and platform engineering teams to ensure Duolingo’s sophisticated distributed...SeniorWork experience placement- ...Senior Sre We're hiring a Senior SRE based in Latin America to work alongside our US-based engineering team, building out observability, on-call coverage, and deployment automation for a client with strict compliance requirements. We're specifically looking for someone...SeniorFull timeRemote work
- ...healthcare organizations maintain accurate, compliant, and reliable provider networks at scale. Our vision is simple: One... ...of patients. About the Role We're looking for a Senior Site Reliability Engineer who takes ownership seriously - someone who designs for...SeniorRemote work
- ...Senior Site Reliability Engineer The expertise you have: Understanding of critical production applications and their infrastructure dependencies to supervise, analyze and discover anomalies. Deep knowledge of application request rates, transactions per second, client...Senior
- ...role, we encourage you to apply. The Role As a Senior Platform Engineer, you are a champion for DevOps and SRE culture and industry... .... \n What You Will Be Doing Improving production reliability and system resilience within an SRE scoped team Championing...SeniorRemote workFlexible hours
$160k - $180k
...have a big impact. See Arkestro in action at arkestro.com. About the Role Arkestro is hiring for a Senior SRE Engineer to manage our performance and reliability for our software platform and infrastructure. The right candidate will own and develop our...SeniorLocal areaRemote work- ...will be responsible for ensuring best-in-class uptime and reliability of our AI hardware infrastructure offerings. Partner... ...through efficient, data-driven solutions. As a Senior Site Reliability Engineer, you will be responsible for: Developing and scaling...SeniorWork at officeRemote work
- ...Senior Site Reliability Engineer Remote – Home Based Job Summary We’re partnering with a company in the SaaS space to find a Senior Site Reliability Engineer . In this role, you’ll be part of the IT Operations group responsible for maintaining all environments...SeniorTemporary workRemote workWork from homeFlexible hours
$141.8k - $195k
...best work, grow fast, and bring their full selves to the herd. Why You’ll Love This Role Cribl Inc is seeking a Senior Site Reliability Engineer to join our mission where you will unlock the value of all observability data, as we expand our team in the U.S. Cribl...SeniorTemporary workRemote work- ...Senior Site Reliability Engineer (SRE) Our client is a global technology consulting and digital solutions company that enables enterprises across industries to reimagine business models, accelerate innovation, and maximize growth by harnessing digital technologies....SeniorLocal area
- ...democratizing software development by removing traditional barriers to application creation. About the role: Join our Site Reliability Engineering team and help ensure the reliability, scalability, and performance of Replit's infrastructure that serves millions of...SeniorFull timeTemporary workWork at officeRemote workWorldwideFlexible hours
- ...Senior Site Reliability Engineer Company: ZetaChain Work Type: Remote Employment: Full Time Location: US Seniority: Senior Level Technologies: Go, Python, Bash, Terraform, Ansible, Kubernetes, Docker, Linux, Prometheus, Grafana, Datadog, Loki, incident.io, AWS, GCP, Bare...SeniorFull timeRemote work
$170k - $220k
Who We're Looking ForWe’re looking for a hands-on, high-agency Site Reliability Engineer to help shape and scale the reliability layer of our stack. You'll own the release pipeline end-to-end — managing daily releases, weekly deploys, and hotfixes — while also automating...Senior$145k - $193k
...entertainment, we want to talk to you. About the Role & Team The SRE team at PENN Entertainment is looking for a Senior Site Reliability Engineer to help build and operate the infrastructure behind a large-scale sports betting and media platform. You'll own critical...SeniorRemote work$129k - $161k
...Job Description Job Description Job title: Senior Site Reliability Engineer Reports to: Director, Site Reliability Engineering Department: Cloud Platforms Location: Remote Grade: 20 About Priority Commerce: Priority Commerce is a leading financial...SeniorRemote work$106.3k - $221.1k
...Senior Site Reliability Engineer At Accenture Federal Services, nothing matters more than helping the US federal government make the nation stronger and safer and life better for people. Our 13,000+ people are united in a shared purpose to pursue the limitless potential...Senior- Job Posting Datavant recognizes the importance of information security and data privacy, including in its hiring processes and recruitment. Datavant encourages all potential job applicants to take precautions against potential phishing schemes or other scams that improperly...SeniorFixed term contractWork at officeLocal areaRemote work
$140k - $180k
...decision-making, and accelerated growth in the AI-driven world. Learn more at Opportunity We’re looking for a Senior Site Reliability Engineer to help build and scale a high-impact SRE function. You’ll be a technical leader on a team responsible for improving...SeniorWork experience placementLocal areaRemote workVisa sponsorshipWork visa$180k - $230k
...Power Acceleration Job Description We're looking for a Senior SRE to own the reliability, scalability, and observability of our production systems. You'll work closely with platform and data engineering to keep high-throughput, data-intensive services running at the...SeniorWork at officeLocal areaImmediate startRemote work3 days per week- ...is responsible for partnering with leaders across engineering and technology to define objective reliability goals for services. Key responsibilities include... ...continuous improvement. Position Summary: The Senior Site Reliability Engineer acts as an advanced senior...SeniorWork at officeShift workDay shift
$400k
...in financial markets, the organization combines innovation, engineering excellence, and data-driven insights to support complex trading operations worldwide. This opportunity is for a Senior Site Reliability Engineer to join a high-performance infrastructure...SeniorFull timeWorldwide
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!
- site reliability engineer sre United States
- site reliability engineer United States
- site reliability engineering manager United States
- site reliability engineer remote United States
- lead site reliability engineer United States
- senior facilities maintenance technician United States
- senior operations technician United States
- senior operations associate United States
- senior cloud service delivery manager United States
- senior it service manager United States


