Site Reliability Engineer, AI Infrastructure
$125k - $195kFull-time
SpaceX
Responsibilities
- Manage GPU and CPU infrastructure deployments to Top Secret data centers.
- Provide GPU-as-a-service support for external customers using bare-metal hardware and virtualized platforms.
- Design, validate, and productize AI cluster solutions at 100,000-plus GPU scale.
- Develop automation for deploying and managing on-premise Kubernetes and AI clusters and operating systems.
- Deploy and manage databases, monitoring systems, and distributed storage.
- Collaborate with AI engineers to build scalable, operable, and maintainable products.
- Improve services across their lifecycle, from design and deployment through operation and refinement.
- Maintain monitoring and alerting systems and develop solutions that improve system availability.
Requirements
- Bachelor’s degree in computer science, information systems/IT, or an engineering discipline plus at least one year of professional site reliability engineering or DevOps experience, or at least three years of such experience in lieu of a degree.
- At least one year of professional experience with Linux operating systems.
- Experience with Terraform, Ansible, or comparable infrastructure tools.
- Experience with containerization technologies such as OCI containers and Kubernetes.
- Scripting experience in Bash, Python, or similar languages.
- Development experience in Python, C++, or Go.
- Preferred experience includes Python-based development frameworks, managing Kubernetes clusters, Linux boot and system configuration, distributed databases and data modeling, large-scale server automation, TCP/IP networking, cloud virtualization, and NVIDIA GPU deployment stacks.
- Preferred knowledge includes testing, continuous integration, build and deployment technologies, continuous monitoring, Bazel, Makefiles, and performance optimization.
- An active Top Secret, Top Secret SCI, or DOE Level Q clearance is preferred; the role requires successfully obtaining and maintaining a Top Secret clearance.
- Must be willing to work extended hours and weekends and travel domestically and globally as needed.
Benefits
- Base salary is listed in two levels ranging from $125,000 to $195,000 annually, with a 10% clearance differential up to an additional $20,000 annually; compensation figures are excluded from benefit details.
- Eligible employees may receive company stock or long-term cash awards, discretionary bonuses, and access to an Employee Stock Purchase Plan.
- Benefits include medical, vision, dental, 401(k), short- and long-term disability, life insurance, paid parental leave, discounts, three weeks of paid vacation, paid holidays, and paid sick leave.
- The role requires obtaining and maintaining a Top Secret security clearance and may require extended hours, weekend work, and future domestic or global travel.
- ITAR eligibility requires U.S. citizenship or national status, lawful permanent residence, refugee or asylee status, or eligibility to obtain required U.S. Department of State authorizations.
Vacancy posted 13 hours ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer, AI Infrastructure in Palo Alto, CA vacancy
$27.12 - $51.93 per hour
...Team What the Role Entails Role Summary We are seeking a motivated Site Reliability Engineer (SRE) Intern to join our AI Compute team, supporting the daily operations of AI infrastructure. In this role, you will work closely with internal business teams and external...SuggestedHourly payFull timeInternship$165k - $265k
...goal of enabling human life on Mars.SR. SITE RELIABILITY ENGINEER (STARSHIELD) At SpaceX we’re... ...focused on Starshield's software and GPU infrastructure, you will design, operate and scale the... ...validate, and productize solutions for AI clusters (100k+ GPU scale)Develop...SuggestedPermanent employmentTemporary workImmediate startWeekend work- ...Build the Future Workforce Wand turns AI into labor. It enables humans and AI... ...the world’s first Agentic Labor Infrastructure enabling governments and global enterprises... ...to establish, lead, and scale our Site Reliability Engineering function. This role combines...SuggestedShift work
$262k - $364k
...within the AViD ecosystem have reliability and uptime appropriate to... ...and performance.Build creative engineering solutions to operations and infrastructure problems, including AI-powered automation and agentic... ...in a strategic way.Site Reliability Engineering (SRE)...Suggested$222k - $300.5k
....Job OverviewAbout the TeamIntuit's Infrastructure and Site Reliability organization owns the operational backbone... .... The Fintech Platform Systems Engineering team builds and operates the AWS-... ...defining priority for this role is AI Ops: embedding AI-driven, autonomous...SuggestedWorldwideShift work- ...Site Reliability Engineer There are NO limits to your career: come shape the future and be part... ...software engineering and applies them to infrastructure and operations problems. The main... ...Programming in Python supported by Gen AI tooling to accelerate development of...Immediate startRemote workWorldwide
- ...Site Reliability Engineer III There's nothing more exciting than being at the center of a rapidly... ...Chief Data & Analytics Office (CDAO) AI/ML & Data Platforms team, you will solve... ...solutions. Through code and cloud infrastructure, you will configure, maintain, monitor...Work at office
- ...JOB DESCRIPTION Elevate your engineering prowess to unprecedented... ...yourself among the top echelon in site reliability. As a Senior Lead Site... ...JPMorgan Chase within the Infrastructure Platforms and Foundational... ...Uses enterprise-authorized AI capabilities within the work...
- ...the Role We're looking for a Senior Site Reliability Engineer who is equally at home writing production software and running the infrastructure it lives on — and who wants to take ownership... ...policies appropriate to a healthcare AI platform Partner with engineers...Shift work
$217.57k - $260k
...explicitly states otherwise, all roles are on-site five days per week at one of our... ...me, we embrace the thoughtful use of AI tools in our daily work and there are... ...Role Overview The Staff Site Reliability Engineer, Infrastructure role is building a high-scale...Full timeTemporary workWork at officeRemote workFlexible hoursShift work- ...Lead Site Reliability Engineer Assume a critical role in defining the future of a globally recognized... ...in their applications Implements infrastructure, configuration, and network as code... ...Uses enterprise-authorized AI capabilities within the work environment...
$168k - $270.25k
NVIDIA DGX Cloud is developing and managing large-scale GPU infrastructure for AI research and production workloads. We are looking for Senior Reliability Engineers to help build the automation, tooling, and operational systems that make GPU clusters reliable, scalable...Full timeRemote work$168k - $270.25k
Site Reliability Engineering (SRE) at NVIDIA is an engineering discipline to design, build and maintain... ...that power NVIDIA’s next-generation AI-driven enterprise products and services... ...Go, with a focus on automation and infrastructure-as-code.Experience with infrastructure...Full time$272k - $431.25k
NVIDIA is looking for a Cloud Site Reliability Engineering Architect to work in IPP's (Infrastructure, Planning and Process) Cloud Infrastructure Team. IPP is a global... ...bottlenecks and optimize the speed and cost efficiency of AI development and testing systems.Leading software...Full timeWork experience placementWorldwide- Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This... ...inference services, powered by the Wafer-Scale Engine (WSE). This team will help deliver world-class, ultra-reliable inference infrastructure for leading model builders such as OpenAI and...Shift work
$168k - $270.25k
....Join our team at NVIDIA as a Senior Site reliability engineer focused on HPC storage and play a crucial... ...with engineering teams to align infrastructure with their evolving needs, document best... ...for an existing vacancy. NVIDIA uses AI tools in its recruiting processes....Full time$152k - $241.5k
...You’ll harness the power of AI to deliver groundbreaking solutions... ...and network fabrics.Use IaC(Infrastructure‑as‑Code) and config... ...lifecycle management, fleet reliability/auto-healing, E2E observability... ..., or Ruby.Mentored other engineers and influenced technical direction...Full time$148k - $235.75k
...into the unlimited potential of AI to define the next era of... ....Join our team of innovative engineers who are building an AI Data Center... ..., high-volume telemetry into reliable, job-centric insights and... ...automation.Manage deployment infrastructure and packaging (Helm + Terraform...Full time$230k - $250k
...foundation for autonomous networking, giving engineers and AI agents the ability to know the... ...been done.Forward is looking for a Site Reliability EngineerAbout the Role This is not a... ...will work closely with engineering, infrastructure, and product to ensure our platform...Night shift$145k - $175k
...care. Requirements As a Senior Site Reliability Engineer at Commence, you will own the reliability... ..., implement, and own observability infrastructure including metrics, logging, tracing,... ..., streaming systems). Exposure to AI/ML infrastructure and the reliability...Full timeRemote work$180k - $230k
...most critical constraint in AI’s growth trajectory: immediate... ...bottleneck in the AI infrastructure race. While leading tech companies... ...for a Senior SRE to own the reliability, scalability, and... ...closely with platform and data engineering to keep high-throughput, data...Work at officeLocal areaImmediate startRemote work3 days per week$276.1k - $311.4k
...Wayve is the leading developer of Embodied AI technology. Our advanced AI software and... ...its charter, hiring its founding engineers, establishing the operating model, and creating... ...the technical strategy that makes reliability a first-class property of the software running...Permanent employmentFull timeWork at officeWork from home- ...Senior SRE Team: Infra Reliability · SF Bay Area / Remote (US) You'll own the GPU infrastructure Luma's research and product run on... ...for a first-principles Linux engineer. You'll be the final escalation... ...large-scale GPU clusters for AI/ML training or inference. Familiarity...Work experience placementRemote work
$248k - $396.75k
Site Reliability Engineering (SRE) at NVIDIA is an engineering discipline focused on designing, building... ...systems, networking, cloud infrastructure, Kubernetes, databases, capacity management... ...the technical direction of NVIDIA’s AI Platform Runtime and lead reliability...Full time- ...Job Title : Senior Site Reliability Engineer Location : Santa Clara, CA Contract ENGAGEMENT SUMMARY The Candidate will provide SRE services for AI platforms and supporting infrastructure with emphasis on reliability engineering, incident response...Contract work
$132.6k - $214.5k
...Integrity, and Inclusion. We weave AI into the fabric of... ...collaborate closely with our engineering teams to develop innovative... ...particularly GCP, to optimize our infrastructure, leveraging cloud-native... ...the product and ensure the reliability and availability of our services...Full timeWork at officeVisa sponsorshipWork visa- ...Powered by the Illumio AI Security Graph, our... ...cyber resilience for the infrastructure, systems, and... ...running. Location: 5 on-site days a week in Sunnyvale... ...Team's Vision: Our Engineering team is shaping the... ...experienced Senior Site Reliability Engineer (SRE) with a...Work experience placementImmediate start
- ...Integrity, and Inclusion. We weave AI into the fabric of everything we do... ...Lead, mentor, and develop a team of Site Reliability/Production Engineers, providing technical direction,... ...health of critical Cortex services and infrastructure. Drive improvements in...Full timeWork at officeVisa sponsorshipWork visa
$155k - $260k
...news and information powered by advanced AI, recommendation systems, and adtech.... ...team to fulfill our mission: building the infrastructure layer for content intelligence.If you’re... ...RequirementsRequired5+ years of backend engineering experience building and operating distributed...Full timeLocal areaWork from home$160.36k - $240.54k
...driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets,... .... Rowe Price, and other leading investors.About the RoleThe ML Infrastructure team is responsible for building & improving the core...Immediate startFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer, AI Infrastructure. Be the first to apply!
Related searches
- site reliability engineer sre Palo Alto, CA
- site reliability engineer Palo Alto, CA
- lead infrastructure engineer Palo Alto, CA
- infrastructure engineer Palo Alto, CA
- infrastructure developer Palo Alto, CA
- remote infrastructure engineer Palo Alto, CA
- senior infrastructure engineer Palo Alto, CA
- IT site lead Palo Alto, CA
- site safety Palo Alto, CA
- website content developer Palo Alto, CA

