Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Platform Reliability Leader for AI Infrastructure

Etched.ai, Inc.

Etched.ai, Inc. in San Jose is seeking a Head of Platform Product Reliability to lead reliability engineering across server, rack and datacenter platform products. You will define reliability strategy, qualification methodologies, accelerated stress testing programs and long-term reliability standards for complex AI infrastructure systems, partnering across Platform Engineering, Mechanical, Firmware, Manufacturing and Supply Chain to ensure exceptional reliability at scale. #J-18808-Ljbffr Etched.ai, Inc.

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Platform Reliability Leader for AI Infrastructure in San Jose, CA vacancy
  • Etched is seeking a highly technical Head of Platform Product Reliability to lead reliability engineering across server, rack, and datacenter...  ...methodologies, and long-term standards for complex AI infrastructure systems. You will work with Platform Engineering, Mechanical... 
    Platform

    The Consensus

    San Jose, CA
    3 days ago
  •  ...Intuit in Mountain View seeks a Director of Product Management for Core Platform and Operational Excellence to shape the vision for our platform and its reliability at scale. You will lead a ~25-30 person team and drive multi-year roadmaps with a strong focus on developer... 
    Platform

    Jobleads-US

    Mountain View, CA
    4 days ago
  • $149.52k - $175.9k

     ...is seeking a strategic and visionary infrastructure leader with deep experience in high-availability...  ...a strong track record of driving platform strategy, operational excellence, modernization...  ..., operational efficiencies, and reliability engineering practices.Manage... 
    Platform
    Full time
    Work experience placement
    Local area
    3 days per week

    US Bank

    Cupertino, CA
    22 hours ago
  • Socket.dev in Cupertino, CA seeks a Senior Infrastructure, SRE & AI Platforms Manager to set the long-term technical strategy and roadmap for global...  ..., and AI workload orchestration, driving high reliability, performance, and cost efficiency while fostering automation... 
    Platform

    Socket.dev

    Cupertino, CA
    3 days ago
  • Etched is seeking a highly technical Head of Platform Product Reliability in San Jose. You will lead reliability engineering across server, rack and datacenter platform products from architecture to fleet deployment. The role defines reliability strategy, qualification... 
    Platform

    Delos

    San Jose, CA
    4 days ago
  • $230k - $250k

     ...giving engineers and AI agents the ability to...  ...building a groundbreaking platform that transforms how...  ...vendor environment.Global leaders like Goldman Sachs,...  ...is looking for a Site Reliability EngineerAbout the Role...  ...closely with engineering, infrastructure, and product to ensure... 
    Platform
    Night shift

    Forward Networks

    Santa Clara, CA
    2 days ago
  • Incedo Inc. seeks a Director of IT and Infrastructure to lead on‑premises, hybrid, and cloud...  ...internal systems and customer‑facing platforms. The role combines strategic leadership...  ...and engineering to accelerate delivery and reliability. #J-18808-Ljbffr Incedo Inc.
    Platform

    Incedo Inc.

    San Jose, CA
    2 days ago
  •  ...generation computing experiences—from AI and data centers, to PCs,...  ...a AI Research Scientist - Infrastructure Engineer, Reinforcement...  ...turning fragile notebooks into reliable systems that RSI, generalizing...  ...record in machine learning (ML) platforms with deep systems expertise... 
    Platform

    AMD

    Santa Clara, CA
    5 days ago
  • $136k - $218.5k

     ...of the GPU to breakthroughs in AI, high-performance computing, robotics...  ..., and gaming. Our Emulation Infrastructure team builds scalable hardware development platforms and compilation flows that...  ...vendors.Improve flow performance, reliability, scalability, usability, and... 
    Platform
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $224k - $356.5k

     ...seamlessly runs them on NVIDIA DRIVE Platforms inside the vehicle.The NvSci...  ...data and sensor processing infrastructure that forms a core element of...  ...high software quality and reliable schedulesDirectly collaborate...  ...vacancy. NVIDIA uses AI tools in its recruiting processes... 
    Platform
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • Senior Site Reliability Engineer ILocationSan Jose, Costa Rica - RemoteSummary...  ...production supportVersed in Infrastructure as Code practices using...  ...its Intelligent Operations Platform. Built to address the...  ...readiness—combining agentic AI-powered SIEM and log analytics... 
    Platform
    Flexible hours

    Sumo Logic

    San Jose, CA
    5 days ago
  • $224k - $356.5k

     ...the unlimited potential of AI to define the next era of computing...  ..., set, and construct the infrastructure and AI direction that...  ...proves it all works. When our platforms are fast and trustworthy, every...  ...partnersOwn platform reliability: define SLIs and SLOs, instrument... 
    Platform
    Full time
    Immediate start

    Nvidia

    Santa Clara, CA
    1 day ago
  • $175k - $250k

     ...Staff Infrastructure Engineer Figure is an AI robotics company developing autonomous general-purpose humanoid...  ...to deliver highly available, reliable, and automated systems. Responsibilities...  ...Extensive experience with cloud platforms (Azure, AWS, GCP) and on-prem... 
    Platform
    Full time

    Figure

    San Jose, CA
    3 days ago
  • $168k - $264.5k

     ...the unlimited potential of AI to define the next era of computing...  ...seeks a senior Site Reliability Engineer (SRE) to join our Santa...  ...Akamai CDN, WAF, and cloud infrastructure.Author, test, and activate...  ...knowledge of the Kubernetes Platform, deployments, and cloud-native... 
    Platform
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $148.32k - $203.94k

     ...including high-growth ones in AI datacenters, automated...  ...size, lower power, and better reliability. With more than 4 billion devices...  ...a hands-on Principal Infrastructure Hardware Engineer to architect...  ...design, and deliver system platforms supporting characterization,... 
    Platform

    SiTime

    Santa Clara, CA
    1 day ago
  •  ...The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers...  ...’s multi-tenant cloud networking platform and SDN infrastructureOperate and improve...  ...teams to improve service reliability and deployment workflowsDeploy and... 
    Platform
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    4 days ago
  • $146.7k - $339.3k

     ...and drive Zoom's Enterprise Infrastructure portfolio - covering data sovereignty...  ...at the intersection of AI, real-time media, and...  ...solutions that power secure, reliable, and compliant collaboration...  ...operate at the intersection of platform infrastructure and customer-facing... 
    Platform
    Work at office
    Remote work

    MAVEN

    San Jose, CA
    2 days ago
  •  ...with strong expertise in LLM infrastructure, model deployment, and high-performance...  ...scalable enterprise GenAI platforms across GPU infrastructure and...  ...cost efficiency. Develop AI platform services and APIs...  ...scalable, secure, and reliable AI solutions. Core Technologies... 
    Platform
    Temporary work

    2T Consulting

    Santa Clara, CA
    5 days ago
  • $136k - $218.5k

     ...the validation and automation infrastructure to characterize them, and...  ...using modern tooling—including AI—without losing rigor. What...  ...firmware/software, process/reliability, and operations teams to co-...  ...post-silicon phases, System/Platform level understanding, tester-... 
    Platform
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  •  ...is building the future of AI powered digital infrastructure. We are a fast-growing, well...  ...efficiency, and real-time platform management. Our technology...  ...evolved into a recognized leader in the convergence of AI,...  ...on the security and reliability of the digital world. If you... 
    Platform

    Axiado Corporation

    San Jose, CA
    22 hours ago
  •  ...The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers...  ..., upgrades, and scaling.Own the reliability, performance, and security of Kubernetes...  ..., networking, and RBAC across the platform.Lead incident response, root-cause... 
    Platform
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    5 days ago
  •  ...Job Title : Senior Site Reliability Engineer Location : Santa Clara, CA Contract ENGAGEMENT SUMMARY The Candidate will provide SRE services for AI platforms and supporting infrastructure with emphasis on reliability engineering, incident response... 
    Platform
    Contract work

    VDart Inc

    Santa Clara, CA
    2 days ago
  • $152k - $241.5k

     ...contributing member of the OEM AI Factory SA team. Our...  ...and Site Reliability Engineering. You will...  ...solutions integrated in their platforms. NVIDIA certified...  ...work with data center infrastructure experience, from hardware...  ...how a global technology leader stays at the cutting... 
    Platform
    Full time
    Work at office

    Nvidia

    Santa Clara, CA
    2 days ago
  • $256k - $414k

     ...GeForce NOW is the global leader in cloud gaming,...  ...networking for GPU-based cloud infrastructure. This role is critical...  ...cloud gaming workloads, AI/ML training, and inference platforms by delivering ultra-low-...  ...-throughput, and highly reliable interconnects across data... 
    Platform
    Full time
    Local area

    Nvidia

    Santa Clara, CA
    5 days ago
  •  ...The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers...  ...and automate the validation of platform quality.Design, build, and maintain...  ...services, workloads, and platform reliability.You6+ years of experience in a SRE,... 
    Platform
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    22 hours ago
  • $200k - $322k

     ...the unlimited potential of AI to define the next era of computing...  ...the performance of our infrastructure both on-prem and cloud. Join...  ...for performance and reliability at global scale, covering automation...  ...experience in compute platform engineering with a focus on... 
    Platform
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    22 hours ago
  •  ..., The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers...  ...track storage SLOs/SLIs, and ensure reliable deployment and maintenance of...  ...experience with enterprise or HPC storage platforms: Vast Data, Weka, NetApp, or Lustre.... 
    Platform
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    5 days ago
  •  ...The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers...  ....About the RoleLambda’s Core Cloud Platform powers compute provisioning and...  .... We are looking for a Senior Site Reliability Engineer to improve the reliability... 
    Platform
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    2 days ago
  • $152k - $241.5k

     ...the unlimited potential of AI to define the next era of computing...  ...for next-generation O-RAN infrastructure. You will work at the...  ...with hardware, architecture, platform, networking, and cellular-software...  ..., latency, throughput, and reliability.Influence future SoC,... 
    Platform
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    1 day ago
  • $207k - $300k

     ...building and developing large-scale infrastructure, distributed systems or networks,...  ...processing and Generative AI agents.Experience with server platform or pod.Preferred qualifications:Master...  ...excellence, such as system reliability, debuggability, and long-term maintainability... 
    Platform

    Google

    San Jose, CA
    5 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Platform Reliability Leader for AI Infrastructure. Be the first to apply!