Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer

SpaceX

SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands‑on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.

ABOUT THE ROLE:

As a Site Reliability Engineer focused on campus reliability, you will design what the campus watches and trusts, technically command cross‑discipline SEVs, and build the guardrails that make the next incident smaller. You are the connective tissue across compute, network, storage, power, and cooling. This role demands calm incident leadership, fleet‑scale observability judgment, and the ability to drive reliability work across software and facility boundaries.

RESPONSIBILITIES:
  • Own monitoring architecture and signal quality: what we alert on, suppress, and trust. Consume NOC noise‑disposition feedback to drive suppression and redesign. Treat alert noise as a design failure, not an operator failure.
  • Provide SEV command support: technical incident leadership, bridge coordination with the NOC, and timeline and severity hygiene.
  • Run blameless postmortems and drive corrective actions to closed, not filed.
  • Lead cross‑functional reliability projects spanning compute, network, storage, and facility signal boundaries.
  • Build and maintain playbooks, run game days, and keep cross‑discipline dependency maps current. Own runbook quality jointly with the NOC (SRE designs; NOC operates and corrects).
  • Define error budgets and availability objectives at campus and service boundaries as adopted by the business.
  • Participate in on‑call rotations and incident response for SEV‑class events in the Memphis / Southaven data center campus.
BASIC QUALIFICATIONS:
  • Bachelor's degree in Systems Engineering, Computer Science, Electrical Engineering, or a related field (or equivalent experience).
  • 5+ years of experience in site reliability, systems engineering, or large‑scale production operations, preferably in high‑performance computing or data center environments.
  • Proven large‑scale incident command experience and calm technical leadership on a bridge.
  • Demonstrated monitoring and observability design at fleet or campus scale, including alert hygiene, suppression, and signal quality.
  • Experience working across at least two of: compute, network, storage, power, and cooling / facilities telemetry.
  • Experience writing and operating playbooks or runbooks with a 24/7 operations or NOC partner.
  • Proficiency in scripting (Python, Bash) for automation and analysis, plus general experience in at least one systems language (C, C++, Java, Go, Rust, or similar). Not required to be expert in all of them.
  • Excellent problem‑solving skills with a data‑driven approach to reliability engineering.
  • Ability to work collaboratively with cross‑functional teams, including NOC, data center operations, and infrastructure engineering.
PREFERRED SKILLS AND EXPERIENCE:
  • Experience in AI/ML infrastructure or supercomputing environments.
  • Hands‑on definition and use of SLOs, SLIs, and error budgets at service or campus boundaries.
  • Experience running game days, dependency mapping, and closed‑loop corrective action programs.
  • Familiarity with data center hardware and plant signals (servers, GPUs, networking, power, cooling) in addition to software telemetry.
  • Prior work in a fast‑paced startup or tech company like SpaceXAI.
#J-18808-Ljbffr
Vacancy posted 18 hours ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in Eastern, KY vacancy
  •  ...personalized care faster. We are building AI agents to support the full arc of the patient journey. The Opportunity: Machine Learning Engineer Patients count on our platform 24/7. You'll build and maintain the tooling, alerts and incident-response playbooks that keep... 
    Suggested

    Tala Health

    Eastern, KY
    4 days ago
  •  ...profitable developer-tooling company whose product is used by engineering teams at thousands of software companies for application...  ...well-resourced group of nine. As Senior SRE you will lead reliability initiatives across the platform — from defining and driving SLOs... 
    Suggested

    Kovoro

    Eastern, KY
    4 days ago
  • $182.8k - $247.3k

     ...to develop education for our half a billion (and growing!) learners around the world. About the role... As a Senior Site Reliability Engineer, you will work closely with both product and platform engineering teams to ensure Duolingo’s sophisticated distributed systems... 
    Suggested
    Work experience placement

    Socket

    Eastern, KY
    2 days ago
  • $180k - $230k

     ...Acceleration Job Description We're looking for a Senior SRE to own the reliability, scalability, and observability of our production systems. You'll work closely with platform and data engineering to keep high-throughput, data-intensive services running at the... 
    Suggested
    Work at office
    Local area
    Immediate start
    Remote work
    3 days per week

    GridCARE, Inc.

    Eastern, KY
    4 days ago
  •  ...A senior Site Reliability Engineer will join an established infrastructure function responsible for highly available, security-conscious cloud systems supporting complex business-critical workloads. You’ll take significant ownership of reliability, scalability, and... 
    Suggested
    Full time
    Remote work

    Motion Recruitment Partners LLC

    Eastern, KY
    4 days ago
  • $90k - $100k

     ...Site Reliability Engineer The Opportunity We are looking for a highly capable engineer to join our Platform and Site Reliability engineering team. You will be responsible for building, maintaining and operating the infrastructure platform on which all Ookla services... 
    Flexible hours

    Ookla

    Eastern, KY
    18 hours ago
  •  ...Engineering, Product, Design, and Marketing Engineering Compensation ~ Zone 1 Base Pay: $214K – $260K Superhuman offers...  ...role will be responsible for building software to ensure the reliability of our back-end systems, working with engineers who develop them... 
    Worldwide
    Home office
    Flexible hours

    Superhuman

    Eastern, KY
    1 day ago
  • $114k - $148k

     ...Total compensation is based on experience, skills, and location using objective, job-related criteria. Summary As a Site Reliability Engineer, you will focus on ensuring the platform and services customers rely on are reliable, performant, and highly available. If... 
    Work experience placement

    OneStream Software

    Eastern, KY
    4 days ago
  •  ...democratizing software development by removing traditional barriers to application creation. About the role: Join our Site Reliability Engineering team and help ensure the reliability, scalability, and performance of Replit's infrastructure that serves millions of... 
    Full time
    Temporary work
    Work at office
    Worldwide
    Flexible hours

    Replit

    Eastern, KY
    4 days ago
  • $107.9k - $195.05k

     ...The Digital Sector at Leidos currently has an opening for a Site Reliability Engineer (SRE) / Senior Cloud Engineer to work in our Baltimore, Maryland office. This is an exciting opportunity to use your experience helping the Center for Medicare and Medicaid Services (... 
    Contract work
    Work at office

    Koitecc Solutions

    Eastern, KY
    4 days ago
  • $148.5k - $223.9k

     ...Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in San Francisco. Working closely with counterparts in the Infrastructure and R&D organizations, this organization provides a global team of engineers monitoring cloud service... 
    Worldwide
    Weekend work

    Salesforce.Com Inc

    Eastern, KY
    4 days ago
  •  ...Discover exciting DevOps job opportunities and connect with 28,396 DevOps professionals. The Senior Site Reliability Engineer role at Jobicy is designed for experienced professionals who are passionate about enhancing system reliability and operational efficiency. The... 
    Remote work
    Flexible hours

    DevOpsChat

    Eastern, KY
    4 days ago
  •  ...% uptime. You'll own SLOs, incident response, and production reliability for a system that processes millions of identity verifications...  ...Sentry error tracking, structured logging Implement chaos engineering practices to proactively identify failure modes Optimize... 
    Remote work

    Xident B.V.

    Eastern, KY
    4 days ago
  • $110k - $145k

     ...operations. You will liaise with product and engineering teams to ensure applications and...  ...feedback loop for platform and product reliability. The ideal candidate is a solutions-oriented...  ...experience as a platform engineer, site reliability engineer, systems engineer... 
    Work experience placement

    CoSM

    Eastern, KY
    4 days ago
  •  ...As a Senior Site Reliability Engineer on our cloud engineering team, you'll keep our production environment healthy, secure, and running smoothly. This is an operations-focused role: you'll own the day-to-day administration of our AWS accounts and databases, backup posture... 
    Work experience placement

    MeridianLink

    Eastern, KY
    4 days ago
  •  ...Quarterhill is seeking a Senior Site Reliability Engineer (SRE) to join our growing team. This role is an exciting opportunity to contribute to the reliability and performance of smart transportation systems, including a next-generation, cloud-native tolling platform that... 
    Local area

    Electronic Transaction Consultants

    Eastern, KY
    4 days ago
  • $104.9k - $174.7k

     ...Technology Senior Site Reliability Engineer II The SRE role is responsible for improving the reliability, availability, performance, and operational quality of production systems. This role provides technical input into project plans, schedules, methodologies, and... 
    Temporary work
    Local area

    LexisNexis Risk Solutions

    Eastern, KY
    2 days ago
  • $110k - $145k

     ...global, with headquarters in Denver, Colorado, and offices across the U.S., Canada, and India. We are seeking a Senior Site Reliability Engineer to own the reliability, scalability, performance, and operational integrity of critical production services. This role is... 
    Flexible hours

    Hirebridge

    Eastern, KY
    4 days ago
  •  ...Cloudflare, GitHub Actions, PostgreSQL, Redis/BullMQ, Node.js/NestJS, Datadog, TypeScript, React, SQL Position: Senior Site Reliability Engineer Engagement period: Ongoing Interview timeline: ASAP Interview process: 1) CV review 2) Interview with our CTO 3)... 
    Contract work
    Immediate start

    Devspace

    Eastern, KY
    4 days ago
  • $115.5k - $164.8k

     ...matters at a company where you matter. Your Impact As an engineer on the APX SRE CloudOps team, you will spend a significant portion...  ...that replace what previously required human intervention with reliable, tested automation. You will also participate in on‑call... 
    Work experience placement
    Work at office
    Remote work

    Koitecc Solutions

    Eastern, KY
    4 days ago
  • $140k - $195k

     ...Improve reliability, observability, service health, incident response, and operational readiness. CodeVertex works across data...  ..., secure systems, and operational clarity matter. The Site Reliability Engineer role helps turn business needs into reliable execution, whether... 
    Remote work

    Codevertex Innovations

    Eastern, KY
    4 days ago
  •  ...This is an engineering-first Senior SRE role. We’re looking for senior engineers who have: Built and shipped significant backend...  ...services end-to-end in production (design → launch → on-call → reliability improvements) Led incident response and driven durable... 

    Practice by Numbers

    Eastern, KY
    17 hours ago
  •  ...intuitive solutions that provide the power of computing without the complexity of programming. InRule Technology is seeking a Site Reliability Engineer to join our growing team of experts operating our decision intelligence platform. You'll design, implement, and support... 
    Work at office
    Remote work
    Worldwide
    Flexible hours
    Shift work

    Rippling

    Eastern, KY
    13 minutes ago
  • $95k - $171k

     .... Opportunities exist to focus on GPU infrastructure, Kubernetes, and ensuring reliability for AI workloads within Akamai's serverless inference platform. As an Site Reliability Engineer II, you will be responsible for: Building and maintaining dashboards, alerts... 
    Permanent employment
    Work experience placement
    Work at office
    Remote work
    Work from home
    Worldwide
    Flexible hours

    Akamai

    Eastern, KY
    4 days ago
  •  ...join us on our mission of providing humankind access to the galaxy beyond our planet. About the Role We are seeking a Site Reliability Engineer to join our Ground Software team. As a Site Reliability Engineer, you will design, build, and operate the ground and site... 
    Full time
    Work at office

    Apex Space, Inc.

    Eastern, KY
    4 days ago
  •  ...to grow, adapt and use your skills consistently. Our customers rely on us in the moments that matter. Engineering delivers on that promise.The Senior Site Reliability Engineer is responsible for ensuring our SaaS products are fast, stable and optimized for our customers... 
    Work experience placement
    Flexible hours

    Donnelley Financial Solutions

    Eastern, KY
    4 days ago
  •  ...electronic production — all running on AWS with zero downtime tolerance for firms in active litigation. We're looking for a Site Reliability Engineer to help maintain the reliability, scalability, and security posture of that platform as we expand our AI capabilities (... 

    Nextpoint

    Eastern, KY
    4 days ago
  •  ...Washington, District of Columbia, United States Contractor | On-site Job Description We are seeking an experienced Site Reliability Engineer (SRE) to help build and maintain highly reliable, scalable, and secure technology platforms. The SRE will combine software... 
    For contractors

    Mybridge

    Eastern, KY
    18 hours ago
  •  ...Zof AI is seeking a Site Reliability Engineer to run the infrastructure that lets fleets of sandboxed agents execute customer code safely and cheaply. This role owns the execution layer of our control plane: Kubernetes and container orchestration, CI/CD pipelines, hard... 
    Full time

    Zof AI

    Eastern, KY
    4 days ago
  •  ...75+ countries, including businesses, developers, IT professionals, and individuals. About the Role We are seeking a Site Reliability Engineer II (SRE II) to help ensure the stability, scalability, and reliability of our services and infrastructure. This role focuses... 

    Webhosting.net

    Eastern, KY
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!