Senior GPU HPC Platform Reliability Engineer
Jobleads-US
A leading AI research company in San Francisco is seeking a software engineer for its Fleet High Performance Computing team. In this role, you'll ensure the reliability and uptime of the compute fleet, working with automation systems and monitoring tools. Ideal candidates have experience managing server environments and proficiency in languages like Python or Go. Join us to innovate in AI technology while maintaining high system efficiency. #J-18808-Ljbffr Jobleads-US
$250k
...infrastructure provider building a next-generation GPU platform designed for AI training, experimentation,... ...States. The company is looking for a Senior / Staff Site Reliability Engineer to support and scale large-scale HPC and cloud environments powering GPU-...SeniorPermanent employmentRemote work- Atoms is hiring an HPC Network Engineer in San Francisco to design and manage high-performance networks connecting GPU compute. The ideal candidate will have expertise in network design... ...hands-on experience with various switch platforms. This role is critical in optimizing low...Senior
- Perplexity is building a self-serve compute platform that lets inference engineers run training jobs and inference services without worrying about GPU provisioning or cluster configuration... ...own the platform surface and drive reliability, observability, and performance...Senior
- ...is earned by shipping excellence. We seek engineers with strong intrinsic drive, a true... ...We’re looking for a systems engineer with HPC or parallel programming experience to help... ...knowledge of high-performance systems to optimize GPU performance at the bleeding edge of AI....SuggestedFull timeWork at office
- Berkeley Lab is hiring a System Infrastructure / Platform Engineer to build and manage high-performance computing systems. This role involves working with HPC systems, Linux infrastructure, and collaborating with engineers and researchers to develop scalable solutions....Senior
$215k - $275k
...raised to date.About the role:Anyscale is looking for a Senior Site Reliability Engineer to join the Infrastructure team. Anyscale aims to provide... ...the critical infrastructure that powers Anyscale’s cloud platform.You will have the opportunity to work on open-source Ray,...SeniorWork at office$232k - $319k
...are too, let's talk.The Infrastructure Platform and Shared Services TeamOkta authenticates... ...scale the service with great people and reliable, cost-effective, and efficient... ...serviceAccelerate the velocity of SRE and product engineering by developing robust platforms, powerful...SeniorPermanent employmentLocal areaWorldwideFlexible hours- Apple Service Engineering (ASE) seeks a senior SRE software engineer to own the architectural direction of Kubernetes internals powering Apple services... ...will define controllers and namespace management, raise reliability, and contribute to upstream Kubernetes. The role...Senior
$175k - $250k
...250,000.00/yr Job Title: Senior Cloud Infrastructure Engineer Location: San Francisco, CA... ...flexibility. Their platform allows users to connect modular... ...scalability, performance, and reliability across environments. What... ...Manage and automate GPU compute clusters using tools...SeniorFull timeRemote workRelocationRelocation package$300k
...building out their AI and cloud platform, powered by thousands of H1... ...inference. As a Platform Engineer/Senior Site Reliability Engineer, you’ll own the... ...performance, and automation of this GPU-powered infrastructure,... ...-performance computing (HPC) or AI/ML training...SeniorPermanent employment$220k - $290k
...Description:Calico seeks a Senior / Staff Cloud Engineer to lead the execution and operation... ...machine learning (ML) platform. As the technical authority... ...node-level GPU/TPU issues, manage scheduling... ...sciencesExperience managing Kubernetes for HPC or ML workloadsExperience...Senior- ...worldwide.We’re a team of engineers, clinicians, and innovators... ...DescriptionPrimary Function of PositionAs a Senior Systems GPU Engineer - AI & Robotics,... ...systems to ensure high reliability and performance.• Linux... ...roadmap for the robotics platform—from foundational models to...SeniorLocal areaWorldwideFlexible hours
$148.7k - $201.2k
...multi-disciplinary team of scientists, engineers, and technicians, on a mission to develop... ...computer.We are looking to hire an HPC Platform Engineer to develop, automate, and maintain... ...and simulation requirements into reliable deliverables (including cluster orchestration...Local areaFlexible hours$245k - $295k
...Crusoe.About the RoleWe are seeking a Senior Manager, Infrastructure Platform Engineering to lead a team building core... ...scale compute infrastructure into reliable, secure, and efficiently allocatable... ...with the operational challenges of GPU clusters, AI training, and...SeniorTemporary workImmediate start$139.4k - $205k
...on business impact. We are a highly senior team composed of former pioneers from... ...are seeking a highly motivated Senior Reliability & Test Engineer to join our team. This individual will... ...and validation of our unmanned platforms at the system and component levels.You...SeniorHourly payWork at officeLocal areaRemote workFlexible hours- ...building the AI Acceleration Cloud, an end-to-end platform for the full generative AI lifecycle, combining the fastest LLM inference engine with state-of-the-art AI cloud... ...critical insights into system performance and GPU utilization, and proactively resolving issues...Senior
- ...deliver top-notch technology products.As a Senior Lead Software Engineer at JPMorgan Chase within the Corporate Sector, Infrastructure Platforms team, you are an integral part of an... ...skillsFoundational understanding of NVIDIA GPU infrastructure software (e.g., DCGM, BCM,...SeniorFor contractors
- About the RoleWe’re looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You’ll partner with engineers and data scientists to build, automate, and maintain the...Senior
$152.5k - $205k
...: CRCL) is one of the world’s leading internet financial platform companies, building the foundation of a more open, global... ...everyone is a stakeholder.What you’ll be responsible for:As a Senior Site Reliability Engineer on Circle’s platform team, you’ll design, build, and...SeniorFlexible hours- DescriptionWe are looking for a Senior or Staff level Site Reliability Engineer to strengthen the reliability, scalability, and operational maturity of our platform in San Francisco, California. This role will focus on improving service health, refining observability,...Senior
$15k
...beautiful modern office, daily catered lunches, and more.As a Senior Cluster Site Reliability Engineer (SRE), you will help scale our research compute... ...and performant. Your work will provide a world-class HPC platform for researchers to focus on cutting-edge machine learning...SeniorWork at officeLocal areaRemote work$117k - $209.33k
...3Position OverviewWant to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable,... ...closely with product engineering, security, compliance, platform, and infrastructure teams to ensure services are reliable...SeniorFull timeFor contractors$152.5k - $205k
...NYSE: CRCL) is one of the world’s leading internet financial platform companies, building the foundation of a more open, global... ...everyone is a stakeholder.What you’ll be responsible forThe Site Reliability Engineer builds and maintains shared platform capabilities, common...SeniorFlexible hours$127k - $249k
The TeamPlatform Engineering is the department within SRE that is responsible for a range... ...critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager... ..., CSIExpertise in cloud infrastructure platforms, including AWS, GCP, or...SeniorWork at officeLocal areaRemote workWorldwideFlexible hours$167.7k - $245.2k
...agents behave as intended, improving reliability and reducing risks. This unified... ...enhanced observability and control.As a Senior Site Reliability Engineer (SRE), you will build, operate, and... ...Splunk Agent Observability's deployment platform and production infrastructure. You...SeniorFull timeTemporary workLocal areaFlexible hours2 days per week$165k - $225.6k
...operations across the company. From core infrastructure to enterprise platforms, we partner across functions to drive scale, reliability, and innovation through technology.The Senior Site Reliability Engineer OpportunityReporting to the Manager, Site Reliability Engineering...SeniorPermanent employmentLocal areaWorldwideFlexible hours- ...unified payments and financial platform for global businesses.... ...what’s next.About the teamThe Engineering team at Airwallex is a diverse... ...together to build scalable, reliable, and secure products that empower... ...services.What you’ll doAs a Senior Site Reliability Engineer, you...SeniorTemporary workLocal areaWorldwide
$148.5k - $223.9k
...the future of Salesforce.Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in San Francisco. Working closely with... ...that blend automation, observability, and AI-powered platforms. You will not only respond to incidents but...SeniorFull timeWorldwideWeekend work$156.86k - $191.72k
...seeking a System Infrastructure / Platform Engineer to help build and manage HPC systems and Linux-based infrastructure... ...-edge technologies such as CPU/GPU clusters, parallel storage, high-... ...Kubernetes, balancing innovation with reliability, performance, and security at...Permanent employmentFull timeRemote workFlexible hours$160k - $220k
...We're all in on this mission. If you are too, let's talk.Senior Database Reliability Engineer (DBRE) Experience Level: Mid-Senior (4+ years PostgreSQL... ...-critical systems. You will work closely with SRE, Platform, and Engineering teams to ensure performance, reliability...SeniorPermanent employmentWork at officeLocal areaWorldwideFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior GPU HPC Platform Reliability Engineer. Be the first to apply!
- platform engineer San Francisco, CA
- platform engineering manager San Francisco, CA
- client platform engineer San Francisco, CA
- platform developer San Francisco, CA
- senior platform engineer San Francisco, CA
- data platform engineer San Francisco, CA
- senior reliability engineer San Francisco, CA
- network reliability engineer San Francisco, CA
- reliability engineer San Francisco, CA
- sr reliability engineer San Francisco, CA
