Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer

Full-time

Runware

Runware is building high-performance infrastructure and products to power the worlds intelligence. Our platform enables developers and businesses to run fast, scalable inference across image, video and emerging modalities, while our Serverless platform allows customers to deploy and scale their own AI models on production-grade GPU infrastructure.

As a Site Reliability Engineer at Runware, you will help ensure these systems remain reliable, performant and resilient as we scale. This is a highly technical, hands-on role working across software, infrastructure and production operations to improve observability, reduce incidents, eliminate operational toil and build lasting improvements across complex distributed systems.

What you’ll do

  • Own and improve the reliability, availability and performance of critical production services across the Runware platform
  • Define and evolve our reliability practices, including SLIs , SLOs , alerting, observability and production-readiness standards
  • Investigate complex production issues across distributed systems, APIs, networking, queues, databases and GPU-backed workloads, participating in our engineering on-call rotation
  • Lead and contribute to incident reviews and RCAs, turning recurring failure modes into lasting engineering improvements
  • Reduce operational toil through automation, automated remediation and improvements to deployment safety, recovery and system resilience
  • Work closely with Engineering and DevOps teams on capacity planning, performance, scaling and architectural improvements as the platform grows

Requirements

  • Have strong experience operating and troubleshooting production systems at scale in an SRE, Production Engineering, Platform Engineering or similar role
  • Have a strong understanding of distributed systems and are comfortable debugging across applications, databases, queues, containers, networking and infrastructure
  • Have experience designing and operating observability systems using metrics, logs and distributed tracing
  • Understand SRE principles including SLIs , SLOs , error budgets, capacity planning, incident management and reducing operational toil
  • Have experience with Kubernetes , containers, IaC and automated deployment practices, alongside the ability to write software and automation using languages such as Python , Go or PHP
  • Take strong ownership of production problems and are comfortable participating in an engineering on-call rotation, taking issues from initial investigation through to long-term remediation

Bonus

  • Experience operating high-throughput or low-latency APIs and distributed systems
  • Experience with bare-metal infrastructure, GPU environments or AI and ML workloads
  • Experience with RabbitMQ or other distributed messaging and queueing systems
  • Experience operating MySQL , Redis , ClickHouse or similar production data systems
  • Experience with global traffic management, load balancing, CDN platforms and hybrid infrastructure environments
  • Experience building automated scaling, capacity management or self-healing systems

Benefits

We’re a remote-first collective, meeting in person twice a year to plan, brainstorm, celebrate wins, and enjoy some face-to-face time. We have core hours for cooperative working and calls, but outside of that your calendar is yours. Work the hours that let you perform at your peak while also building a healthy life.

Our release cycles are fast and intense, but they’re followed by real downtime. After big pushes we expect the team to unplug, recharge, and come back ready & stronger than ever for the next leap.

  • Generous paid time off – vacation, sick days, public holidays
  • Meaningful stock options – share in the upside you create
  • Remote-first setup – work from home anywhere we can employ you
  • Flexible hours – own your schedule outside core collaboration blocks
  • Family leave – paid maternity, paternity, and caregiver time
  • Company retreats – twice-yearly gatherings in inspiring locations
Vacancy posted 21 days ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in United States vacancy
  • $100k - $115k

     ...Internal Developer Platform (IDP) as a product, treating engineering teams as customers and optimizing for reliability, usability, and delivery velocity.Define and...  ....4+ years of experience in Platform Engineering, Site Reliability Engineering, DevOps, or Systems Engineering... 
    Suggested
    Temporary work

    Analytic Partners

    Dallas, TX
    2 days ago
  • $62k - $141k

    Site Reliability EngineerThe Opportunity: Engineering to make a system more resilient and efficient frees up time and money to build more capabilities. Whether you come from a background in network engineering, systems administration, or software development, if you have... 
    Suggested
    Full time
    Contract work
    Part time
    Work at office
    Local area
    Remote work

    Booz Allen Hamilton

    Aurora, CO
    6 days ago
  • $98.58k - $138.02k

     ...Northern California / Silicon Valley Region / Denver, COProduct Engineering - DevOps /Full Time /HybridRestaurant365 is a SaaS company...  ...office locations: Austin, TX; Irvine, CA; or Akron, OH. The Site Reliability Engineer II will be responsible for supporting, enhancing,... 
    Suggested
    Full time
    Work at office

    Restaurant 365

    Irvine, CA
    1 day ago
  • $100k - $125k

    We are looking for a Senior Site Reliability Engineer (SRE) to join our Site Reliability Engineering Team, working closely with a dedicated product team to modernize infrastructure, strengthen system resilience, and scale our global platform, leveraging AI tools and agents... 
    Suggested
    Full time
    Local area
    Worldwide
    Flexible hours

    IDEXX Laboratories

    Westbrook, ME
    6 days ago
  • $145.7k - $218.5k

     ...synonymous with entertainment excellence and creativity.Service Reliability EngineerDo you want to use transformative technologies to...  ...scalability and efficiency? Do you want a career that combines your engineering skills and your passion for video gaming? Are you fascinated... 
    Suggested
    Work experience placement
    Shift work

    Sony Interactive Entertainment America

    Aliso Viejo, CA
    2 days ago
  • $160k - $180k

     ...solutions that advisors use in helping clients achieve wealth, independence, and purpose.The OpportunityWe are seeking a Site Reliability Engineer (SRE) to join our Charlotte-based engineering team. This role sits at the center of platform resilience — ensuring high availability... 
    Full time
    Work at office
    Flexible hours

    Asset Mark

    Charlotte, NC
    5 days ago
  •  ...and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Commercial and Investment Bank, Global Payments team, you will solve complex and broad business... 

    JP Morgan Chase

    Irvine, CA
    4 days ago
  • $152.13k - $162.13k

     ...challenge the status-quo.Unum is changing, and we’re excited about what’s next. Join us.General Summary:Unum Group seeks Site Reliability Engineers in Atlanta, GA.Applicants who are interested in this position may apply at (Ref #66753) for consideration.Design, build,... 
    Full time
    Temporary work
    Work at office
    Remote work

    Unum Group

    Atlanta, GA
    4 days ago
  • $69.8k - $148.3k

     ...of infrastructure and service to ensure reliability and functionality. Responds to infrastructure...  ...tools and develops working knowledge of site reliability trends.Only Oracle brings...  ...Science, Information Technology, Engineering, or a related field, or equivalent practical... 
    Temporary work
    Flexible hours

    Oracle Corporation

    Nashville, TN
    2 days ago
  •  ...English (Required)Work Shift:1st Shift (United States of America)Please review the following job description:Lead Site Reliability & Environment Monitoring Engineer (Azure / Dynatrace / ServiceNow)We are seeking a Lead Site Reliability & Environment Monitoring Engineer to... 
    Full time
    Temporary work
    Shift work
    Day shift

    TIH

    Charlotte, NC
    2 days ago
  •  ...Evaluate applications, platforms, and vendors to assess resiliency, reliability, and operational risk.Design and implement processes that...  ...and reliability tooling.Actively participate in reliability engineering and resilience communities of practice, contributing to... 
    Full time

    Vanguard

    Dallas, TX
    5 days ago
  • Job ID: 18719802Reference Number: 23-00164Title: site reliability engineerLocation: Iselin, NJ, 08830Posted Date: 2023-01-20Contact: Shyam MaramContact Email: ****@*****.*** Phone: (***) ***-****Company: HAN Staffing Devops/SRE/Python Role Malvern PA - hybrid... 

    HAN Staffing

    Iselin, NJ
    2 days ago
  •  ...communities.This is a Lead Software Production Management & Reliability Engineering position at Director level which is part of the job family responsible...  ...across the business.Job SummaryWe are looking for a Site Reliability Engineer with a minimum of 5 years of industry... 
    Flexible hours
    Weekend work

    Morgan Stanley

    Alpharetta, GA
    6 days ago
  • $138.4k - $173k

     ...infrastructure as well as help improve the reliability, quality of services and overall...  ...recovery. You’ll collaborate or embed with engineering teams, helping them to improve the reliability...  ...about our locations by visiting our site.Compensation & BenefitsThe base salary that... 
    Full time
    Flexible hours

    AppFolio

    San Diego, CA
    2 days ago
  •  ...across multiple clouds and regions while partnering with network engineers, systems architects, and game studio developers. This is an ownership role: driving technical direction, influencing reliability from architecture review through production operation, and closing... 

    2K Games

    Austin, TX
    4 days ago
  • $148k - $235.75k

     ...see how you can make a lasting impact on the world.Join our team of innovative engineers who are building an AI Data Center AIOps platform that turns raw, high-volume telemetry into reliable, job-centric insights and automation for GPU fleets. We’re hiring a DevOps Engineer... 
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $243.29k - $295.25k

     ...challenges at scale, and helping to create safer, more civil shared experiences for everyone.The Infrastructure Compute Site Reliability Engineering mission is to own and manage the successful operation of our underlying cell infrastructure system, along with elements... 
    Full time
    Work experience placement
    H1b
    Work at office
    Local area
    Visa sponsorship
    Monday to Friday

    Roblox

    San Mateo, CA
    5 days ago
  •  ...Georgia, and serves customers in more than 35 countries worldwide.Position OverviewWe are seeking a highly experienced Senior Site Reliability Engineer (Unified Observability) to lead the design, implementation, and operational maturity of the F1 Next Generation Customer... 
    Full time
    Worldwide
    Flexible hours

    NCR

    Atlanta, GA
    1 day ago
  • $112k - $179k

     ...delivery of system, network, software, and security solutions.About The RolePeraton is seeking a self-driven and resourceful Site Reliability Engineer to join our dynamic of Network and UC engineers in Washington, DC. This position combines software engineering and systems... 
    Contract work
    Worldwide
    Shift work

    Peraton Corporation

    Washington DC
    2 days ago
  • LeanData helps the world’s fastest-growing companies automate, simplify, and accelerate revenue.We are looking for a Senior Site Reliability Engineer to lead the strategic evolution of our cloud infrastructure. Reporting directly to the SVP of Engineering, this role is... 
    Full time
    Work at office
    2 days per week

    LeanData

    Santa Clara, CA
    2 days ago
  • $125k - $145k

     ...including 90% of the Fortune 500, choose DigiCert to stop today’s threats and prepare for a quantum-safe future at Job summaryThe Site Reliability Engineer (SRE) collaborates with development teams to embed reliability, scalability, and performance best practices throughout the... 
    Flexible hours

    DigiCert

    Lehi, UT
    4 days ago
  • $117k - $209.33k

    Job Requisition ID #26WD99276Position OverviewWant to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable, secure, and scalable cloud services for Autodesk GovCloud products.As part of a new SRE team supporting... 
    Full time
    For contractors
    Remote work

    Autodesk

    Plano, TX
    2 days ago
  •  ...and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Commercial and Investment Banking and Payment Technology Team, you will solve complex and... 
    Local area

    JP Morgan Chase

    Chicago, IL
    7 days ago
  • $134.25k - $214.8k

     ...matters at a company where you matter.Your ImpactAre you an engineer who gets excited about the challenge of making complex distributed...  ...it.You will be part of the Observability team within Axon's Site Reliability organization — a focused team responsible for Axon's metrics,... 
    Work experience placement
    Work at office
    Remote work

    Axon

    Boston, MA
    5 days ago
  • $130k - $140k

     ...mid-market firms, rely on SS&C for expertise, scale, and technology.Job DescriptionSite Reliability EngineerLocation(s): Waltham, MA | HybridAbout the RoleSr Site Reliability Engineer- Guardian of the products to ensuring systems are reliable, scalable, and efficient... 
    Ongoing contract
    Full time
    Temporary work
    Work experience placement

    SS&C Technologies

    Waltham, MA
    6 days ago
  • $165k - $230k

     ...is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SR. SITE RELIABILITY ENGINEER (STARSHIELD)Starshield leverages SpaceX’s Starlink technology and launch capability to support national security efforts.... 
    Permanent employment
    Temporary work
    Immediate start
    Weekend work

    SpaceX

    Hawthorne, CA
    4 days ago
  • As a Site Reliability Engineer, you will be responsible for: Operational Excellence & Incident Management- Maintain and monitor production systems for availability, latency, and performance.- Lead incident response efforts, including communication, resolution, and postmortem... 
    Permanent employment

    National Oilwell Varco

    Houston, TX
    2 days ago
  • $81.1k - $187k

     ...infrastructure and/or service according to terms for reliability and functionality.- Assists team members...  ...deployments.- Gains basic knowledge of site reliability trends and shares relevant...  ...are seeking a skilled Site Reliability Engineer to design, build, operate, and automate... 
    Temporary work
    Immediate start
    Flexible hours
    Shift work

    Oracle Corporation

    Reston, VA
    5 days ago
  •  ...importance of in-office collaboration and fully intend for the selected candidate for this role to work on site in the specified location(s).As a Senior Reliability Engineer, you will help shape the reliability, scalability, and operational excellence of mission-critical... 
    Full time
    Work at office

    The Charles Schwab Corporation

    Austin, TX
    4 days ago
  •  ...and foster a dynamic work environment where new ideas thrive. Are you ready to join our team and make an impact?As a Senior Site Reliability Engineer at TeamViewer, you’ll be a key player in ensuring the reliability, scalability, and performance of our Azure-based SaaS... 
    Temporary work
    Casual work
    Worldwide

    TeamViewer

    Austin, TX
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!