Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer, AI Infrastructure

$125k - $195k
Full-time

SpaceX

Responsibilities

  • Manage GPU and CPU infrastructure deployments to Top Secret data centers.
  • Provide GPU-as-a-service support for external customers using bare-metal hardware and virtualized platforms.
  • Design, validate, and productize AI cluster solutions at 100,000-plus GPU scale.
  • Develop automation for deploying and managing on-premise Kubernetes and AI clusters and operating systems.
  • Deploy and manage databases, monitoring systems, and distributed storage.
  • Collaborate with AI engineers to build scalable, operable, and maintainable products.
  • Improve services across their lifecycle, from design and deployment through operation and refinement.
  • Maintain monitoring and alerting systems and develop solutions that improve system availability.

Requirements

  • Bachelor’s degree in computer science, information systems/IT, or an engineering discipline plus at least one year of professional site reliability engineering or DevOps experience, or at least three years of such experience in lieu of a degree.
  • At least one year of professional experience with Linux operating systems.
  • Experience with Terraform, Ansible, or comparable infrastructure tools.
  • Experience with containerization technologies such as OCI containers and Kubernetes.
  • Scripting experience in Bash, Python, or similar languages.
  • Development experience in Python, C++, or Go.
  • Preferred experience includes Python-based development frameworks, managing Kubernetes clusters, Linux boot and system configuration, distributed databases and data modeling, large-scale server automation, TCP/IP networking, cloud virtualization, and NVIDIA GPU deployment stacks.
  • Preferred knowledge includes testing, continuous integration, build and deployment technologies, continuous monitoring, Bazel, Makefiles, and performance optimization.
  • An active Top Secret, Top Secret SCI, or DOE Level Q clearance is preferred; the role requires successfully obtaining and maintaining a Top Secret clearance.
  • Must be willing to work extended hours and weekends and travel domestically and globally as needed.

Benefits

  • Base salary is listed in two levels ranging from $125,000 to $195,000 annually, with a 10% clearance differential up to an additional $20,000 annually; compensation figures are excluded from benefit details.
  • Eligible employees may receive company stock or long-term cash awards, discretionary bonuses, and access to an Employee Stock Purchase Plan.
  • Benefits include medical, vision, dental, 401(k), short- and long-term disability, life insurance, paid parental leave, discounts, three weeks of paid vacation, paid holidays, and paid sick leave.
  • The role requires obtaining and maintaining a Top Secret security clearance and may require extended hours, weekend work, and future domestic or global travel.
  • ITAR eligibility requires U.S. citizenship or national status, lawful permanent residence, refugee or asylee status, or eligibility to obtain required U.S. Department of State authorizations.
Vacancy posted 13 hours ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer, AI Infrastructure in Palo Alto, CA vacancy
  • $27.12 - $51.93 per hour

     ...Team What the Role Entails Role Summary We are seeking a motivated Site Reliability Engineer (SRE) Intern to join our AI Compute team, supporting the daily operations of AI infrastructure. In this role, you will work closely with internal business teams and external... 
    Suggested
    Hourly pay
    Full time
    Internship

    Tencent

    Palo Alto, CA
    2 hours ago
  • $165k - $265k

     ...goal of enabling human life on Mars.SR. SITE RELIABILITY ENGINEER (STARSHIELD) At SpaceX we’re...  ...focused on Starshield's software and GPU infrastructure, you will design, operate and scale the...  ...validate, and productize solutions for AI clusters (100k+ GPU scale)Develop... 
    Suggested
    Permanent employment
    Temporary work
    Immediate start
    Weekend work

    SpaceX

    Palo Alto, CA
    19 hours ago
  •  ...Build the Future Workforce Wand turns AI into labor. It enables humans and AI...  ...the world’s first Agentic Labor Infrastructure enabling governments and global enterprises...  ...to establish, lead, and scale our Site Reliability Engineering function. This role combines... 
    Suggested
    Shift work

    Wand AI

    Palo Alto, CA
    10 hours ago
  • $262k - $364k

     ...within the AViD ecosystem have reliability and uptime appropriate to...  ...and performance.Build creative engineering solutions to operations and infrastructure problems, including AI-powered automation and agentic...  ...in a strategic way.Site Reliability Engineering (SRE)... 
    Suggested

    Google

    Mountain View, CA
    21 hours ago
  • $222k - $300.5k

     ....Job OverviewAbout the TeamIntuit's Infrastructure and Site Reliability organization owns the operational backbone...  .... The Fintech Platform Systems Engineering team builds and operates the AWS-...  ...defining priority for this role is AI Ops: embedding AI-driven, autonomous... 
    Suggested
    Worldwide
    Shift work

    Intuit

    Mountain View, CA
    3 days ago
  •  ...Site Reliability Engineer There are NO limits to your career: come shape the future and be part...  ...software engineering and applies them to infrastructure and operations problems. The main...  ...Programming in Python supported by Gen AI tooling to accelerate development of... 
    Immediate start
    Remote work
    Worldwide

    OutSystems

    Menlo Park, CA
    4 days ago
  •  ...Site Reliability Engineer III There's nothing more exciting than being at the center of a rapidly...  ...Chief Data & Analytics Office (CDAO) AI/ML & Data Platforms team, you will solve...  ...solutions. Through code and cloud infrastructure, you will configure, maintain, monitor... 
    Work at office

    Chase

    Palo Alto, CA
    4 days ago
  •  ...JOB DESCRIPTION Elevate your engineering prowess to unprecedented...  ...yourself among the top echelon in site reliability. As a Senior Lead Site...  ...JPMorgan Chase within the Infrastructure Platforms and Foundational...  ...Uses enterprise-authorized AI capabilities within the work... 

    J.P. Morgan

    Palo Alto, CA
    2 days ago
  •  ...the Role We're looking for a Senior Site Reliability Engineer who is equally at home writing production software and running the infrastructure it lives on — and who wants to take ownership...  ...policies appropriate to a healthcare AI platform Partner with engineers... 
    Shift work

    Hippocratic AI Inc.

    Menlo Park, CA
    3 days ago
  • $217.57k - $260k

     ...explicitly states otherwise, all roles are on-site five days per week at one of our...  ...me, we embrace the thoughtful use of AI tools in our daily work and there are...  ...Role Overview The Staff Site Reliability Engineer, Infrastructure role is building a high-scale... 
    Full time
    Temporary work
    Work at office
    Remote work
    Flexible hours
    Shift work

    ID.me

    Mountain View, CA
    4 days ago
  •  ...Lead Site Reliability Engineer Assume a critical role in defining the future of a globally recognized...  ...in their applications Implements infrastructure, configuration, and network as code...  ...Uses enterprise-authorized AI capabilities within the work environment... 

    Chase

    Palo Alto, CA
    9 hours ago
  • $168k - $270.25k

    NVIDIA DGX Cloud is developing and managing large-scale GPU infrastructure for AI research and production workloads. We are looking for Senior Reliability Engineers to help build the automation, tooling, and operational systems that make GPU clusters reliable, scalable... 
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    1 day ago
  • $168k - $270.25k

    Site Reliability Engineering (SRE) at NVIDIA is an engineering discipline to design, build and maintain...  ...that power NVIDIA’s next-generation AI-driven enterprise products and services...  ...Go, with a focus on automation and infrastructure-as-code.Experience with infrastructure... 
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $272k - $431.25k

    NVIDIA is looking for a Cloud Site Reliability Engineering Architect to work in IPP's (Infrastructure, Planning and Process) Cloud Infrastructure Team. IPP is a global...  ...bottlenecks and optimize the speed and cost efficiency of AI development and testing systems.Leading software... 
    Full time
    Work experience placement
    Worldwide

    Nvidia

    Santa Clara, CA
    3 days ago
  • Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This...  ...inference services, powered by the Wafer-Scale Engine (WSE). This team will help deliver world-class, ultra-reliable inference infrastructure for leading model builders such as OpenAI and... 
    Shift work

    Cerebras Systems

    Sunnyvale, CA
    21 hours ago
  • $168k - $270.25k

     ....Join our team at NVIDIA as a Senior Site reliability engineer focused on HPC storage and play a crucial...  ...with engineering teams to align infrastructure with their evolving needs, document best...  ...for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.... 
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $152k - $241.5k

     ...You’ll harness the power of AI to deliver groundbreaking solutions...  ...and network fabrics.Use IaC(Infrastructure‑as‑Code) and config...  ...lifecycle management, fleet reliability/auto-healing, E2E observability...  ..., or Ruby.Mentored other engineers and influenced technical direction... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $148k - $235.75k

     ...into the unlimited potential of AI to define the next era of...  ....Join our team of innovative engineers who are building an AI Data Center...  ..., high-volume telemetry into reliable, job-centric insights and...  ...automation.Manage deployment infrastructure and packaging (Helm + Terraform... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $230k - $250k

     ...foundation for autonomous networking, giving engineers and AI agents the ability to know the...  ...been done.Forward is looking for a Site Reliability EngineerAbout the Role This is not a...  ...will work closely with engineering, infrastructure, and product to ensure our platform... 
    Night shift

    Forward Networks

    Santa Clara, CA
    21 hours ago
  • $145k - $175k

     ...care. Requirements As a Senior Site Reliability Engineer at Commence, you will own the reliability...  ..., implement, and own observability infrastructure including metrics, logging, tracing,...  ..., streaming systems). Exposure to AI/ML infrastructure and the reliability... 
    Full time
    Remote work

    GrabJobs

    Santa Clara, CA
    2 days ago
  • $180k - $230k

     ...most critical constraint in AI’s growth trajectory: immediate...  ...bottleneck in the AI infrastructure race. While leading tech companies...  ...for a Senior SRE to own the reliability, scalability, and...  ...closely with platform and data engineering to keep high-throughput, data... 
    Work at office
    Local area
    Immediate start
    Remote work
    3 days per week

    GridCARE, Inc.

    Redwood City, CA
    3 days ago
  • $276.1k - $311.4k

     ...Wayve is the leading developer of Embodied AI technology. Our advanced AI software and...  ...its charter, hiring its founding engineers, establishing the operating model, and creating...  ...the technical strategy that makes reliability a first-class property of the software running... 
    Permanent employment
    Full time
    Work at office
    Work from home

    Lindus Health

    Sunnyvale, CA
    3 days ago
  •  ...Senior SRE Team: Infra Reliability · SF Bay Area / Remote (US) You'll own the GPU infrastructure Luma's research and product run on...  ...for a first-principles Linux engineer. You'll be the final escalation...  ...large-scale GPU clusters for AI/ML training or inference. Familiarity... 
    Work experience placement
    Remote work

    Luma AI

    Redwood City, CA
    3 days ago
  • $248k - $396.75k

    Site Reliability Engineering (SRE) at NVIDIA is an engineering discipline focused on designing, building...  ...systems, networking, cloud infrastructure, Kubernetes, databases, capacity management...  ...the technical direction of NVIDIA’s AI Platform Runtime and lead reliability... 
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  •  ...Job Title : Senior Site Reliability Engineer Location : Santa Clara, CA Contract ENGAGEMENT SUMMARY The Candidate will provide SRE services for AI platforms and supporting infrastructure with emphasis on reliability engineering, incident response... 
    Contract work

    VDart

    Santa Clara, CA
    3 days ago
  • $132.6k - $214.5k

     ...Integrity, and Inclusion. We weave AI into the fabric of...  ...collaborate closely with our engineering teams to develop innovative...  ...particularly GCP, to optimize our infrastructure, leveraging cloud-native...  ...the product and ensure the reliability and availability of our services... 
    Full time
    Work at office
    Visa sponsorship
    Work visa

    Palo Alto Networks

    Santa Clara, CA
    3 days ago
  •  ...Powered by the Illumio AI Security Graph, our...  ...cyber resilience for the infrastructure, systems, and...  ...running. Location: 5 on-site days a week in Sunnyvale...  ...Team's Vision: Our Engineering team is shaping the...  ...experienced Senior Site Reliability Engineer (SRE) with a... 
    Work experience placement
    Immediate start

    Illumio

    Sunnyvale, CA
    1 day ago
  •  ...Integrity, and Inclusion. We weave AI into the fabric of everything we do...  ...Lead, mentor, and develop a team of Site Reliability/Production Engineers, providing technical direction,...  ...health of critical Cortex services and infrastructure. Drive improvements in... 
    Full time
    Work at office
    Visa sponsorship
    Work visa

    Palo Alto Networks

    Santa Clara, CA
    9 hours ago
  • $155k - $260k

     ...news and information powered by advanced AI, recommendation systems, and adtech....  ...team to fulfill our mission: building the infrastructure layer for content intelligence.If you’re...  ...RequirementsRequired5+ years of backend engineering experience building and operating distributed... 
    Full time
    Local area
    Work from home

    News Break

    Mountain View, CA
    2 days ago
  • $160.36k - $240.54k

     ...driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets,...  .... Rowe Price, and other leading investors.About the RoleThe ML Infrastructure team is responsible for building & improving the core... 
    Immediate start
    Flexible hours

    Nuro

    Mountain View, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer, AI Infrastructure. Be the first to apply!