Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Cloud Platform Engineer

SambaNova Systems

The era of pervasive AI has arrived. In this era, organizations will use generative AI to unlock hidden value in their data, accelerate processes, reduce costs, drive efficiency and innovation to fundamentally transform their businesses and operations at scale. SambaNova Suite™ is the first full-stack, generative AI platform, from chip to model, optimized for enterprise and government organizations. Powered by the intelligent SN40L chip, the SambaNova Suite is a fully integrated platform, delivered on-premises or in the cloud, combined with state-of-the-art open-source models that can be easily and securely fine-tuned using customer data for greater accuracy. Once adapted with customer data, customers retain model ownership in perpetuity, so they can turn generative AI into one of their most valuable assets. About SambaNova Systems Join the company that’s building the future of AI computing. At SambaNova, we are disrupting the AI and high-performance computing space with our integrated hardware and software platform. Our DataScale systems and SambaFlow software are pushing the boundaries of what’s possible with generative AI and large language models. We are a team of passionate innovators tackling some of the world’s most challenging computational problems. The Role As a Senior Cloud Site Reliability Engineer (SRE) specializing in our AI Inferencing Service, you will be the guardian of its reliability, performance, and scalability. You will bridge the gap between software development and operations, applying an engineering mindset to solve operational challenges. Your primary focus will be ensuring our inference endpoints have exceptional uptime, low-latency response times, and efficient resource utilization, directly impacting the experience of our customers and the success of our AI products. This role includes participating in a shared on‑call rotation to maintain 24/7 service reliability. What You’ll Do Service Ownership & On‑Call: Take shared ownership of the production inferencing service, including its availability, latency, performance, efficiency, change management, monitoring, emergency response, and capacity planning across multiple regions. This includes implementing and supporting AI infrastructure in new regions, such as Asia, Europe, and Latin America, to support the growth of our business. Participate in a balanced on‑call rotation to provide 24/7 support for the service. On‑Call & Work‑Life Balance We believe a sustainable on‑call schedule is critical for long‑term success and team health. Our on‑call philosophy is built on the following principles: Balanced Rotation: The on‑call rotation is shared equally across the team, typically following a primary/secondary (follow‑the‑sun) model to ensure no single person bears a disproportionate burden. Focus on Prevention: We invest heavily in automation, robust testing, and system design to prevent pages before they happen. The goal of on‑call is not to heroically fight fires, but to manage rare, complex failures and use those learnings to make the system more resilient. Actionable Alerts: We have a strict policy against alert fatigue. Alerts must be actionable and require immediate human intervention. Incident Management: Lead the response to incidents affecting the inferencing service, driving blameless post‑mortems and implementing corrective actions to prevent recurrence. Monitoring & Alerting: Develop and maintain advanced monitoring, alerting, and dashboarding (using tools like Prometheus, Grafana, Datadog) to gain deep insights into service health, model performance (e.g., latency, throughput, error rates), and accelerator utilization. A key responsibility is ensuring alerts are actionable and have a low false‑positive rate, minimizing on‑call fatigue. Performance & Scalability: Proactively identify and eliminate performance bottlenecks. Design and implement auto‑scaling policies to handle variable inference loads cost‑effectively. Use insights from on‑call incidents to drive improvements that enhance system stability and scalability. Infrastructure as Code (IaC): Manage and evolve our cloud infrastructure (on AWS, GCP, and/or Azure along with on‑prem) using tools like Terraform and Ansible, ensuring it is secure, repeatable, and scalable. CI/CD & Automation: Champion automation by building and improving CI/CD pipelines for the seamless and safe deployment of new model versions and service updates. A core goal is to automate manual toil identified during on‑call shifts, reducing future operational overhead. Capacity Planning: Forecast infrastructure needs based on product roadmaps and usage trends. Work with finance and engineering teams to manage cloud costs and optimize spending. SLOs & SLIs: Define, measure, and report on Service Level Objectives (SLOs) and Indicators (SLIs) for the inferencing platform, using data to drive prioritization and reliability investments. What We’re Looking For (Must‑Haves) Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience. 5-8+ years of experience in a Site Reliability Engineer, DevOps, or related role supporting a large‑scale, customer‑facing service in a public cloud environment (AWS, GCP, Azure). Strong programming/scripting skills in languages like Python, Go, or Java. Proven experience with containerization and orchestration technologies (Docker, Kubernetes). Deep understanding of monitoring and observability principles and tools (e.g., Prometheus, Grafana, ELK Stack, Datadog). Solid experience with Infrastructure as Code (e.g., Terraform, CloudFormation). Familiarity with CI/CD principles and tools (e.g., Jenkins, GitHub Actions, ArgoCD). Excellent problem‑solving skills and a systematic approach to troubleshooting complex distributed systems. What Will Make You Stand Out (Nice‑to‑Haves) Experience in a hybrid environment bridging cloud and on‑premise/data center infrastructure. Direct experience supporting ML/AI inferencing services in production. Familiarity with GPU‑accelerated computing and optimizing workloads for NVIDIA GPUs for purposes of mapping to RDUs. Knowledge of model serving frameworks like vLLM, SGLang or Ray. Understanding of MLOps principles and practices. Experience with managing and tuning databases (SQL or NoSQL) and caching systems (Redis, Memcached). Strong Linux/Unix system administration fundamentals. Why SambaNova? Massive Impact: You will be a key part of a critical platform with high visibility and direct impact on our product and engineers. Cutting‑Edge Technology: Work with a world‑class team on one of the most advanced AI stacks in the industry. Autonomy and Growth: We trust you to make technical decisions. This is a greenfield opportunity to build something remarkable from the ground up. Competitive Compensation: Including equity, excellent benefits, and a flexible work environment. EEO Policy SambaNova Systems is an Equal Opportunity/Affirmative Action Employer. All qualified applicants will receive consideration for employment without regard basis of age (40 and over), color, disability, gender identity, genetic information, marital status, military or veteran status, national origin/ancestry, race, religion, creed, sex (including pregnancy, childbirth, breastfeeding), sexual orientation, and any other applicable status protected by federal, state, or local laws. Benefits Summary for US-Based, Full‑Time Employment Positions SambaNova offers a competitive total rewards package, including the base salary, plus equity and benefits. We cover 95% premium coverage for employee medical insurance, and 77% premium coverage for dependents and offer a Health Savings Account (HSA) with employer contribution. We also offer Dental, Vision, Short/Long term Disability, Basic Life, Voluntary Life, and AD&D insurance plans in addition to Flexible Spending Account (FSA) options like Health Care, Limited Purpose, and Dependent Care. Our library of well‑being benefits available to you and your dependents includes a full subscription to Headspace, Gympass+ membership with access to physical gyms, One Medical membership, counseling services with an Employee Assistance Program, and much more. #J-18808-Ljbffr

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Senior Cloud Platform Engineer in San Jose, CA vacancy
  •  ...scalable AI solutions. The ideal candidate possesses a strong technical background in AI and programming with experience in various cloud platforms. You'll have the opportunity to innovate with technologies like Python and AWS and craft AI solutions that will improve banking... 
    Cloud
    Senior

    Capital One

    San Jose, CA
    5 days ago
  • Apple Inc. is seeking an experienced backend engineer to build the next generation of its advertising platforms. You will collaborate with product management to define...  ...on low-latency, high-availability systems in the cloud, and develop secure data-processing pipelines for... 
    Cloud
    Senior

    Apple

    Cupertino, CA
    17 hours ago
  • $168k - $322k

    NVIDIA in Santa Clara is hiring a Senior AI Platform Engineer to build and maintain next-gen AI-powered enterprise products. This role requires 10+ years of experience in cloud, platform, or SRE roles. Responsibilities include architecting infrastructure, mentoring, and... 
    Cloud
    Senior

    NVIDIA

    Santa Clara, CA
    1 day ago
  •  ...seeking an Associate Principal Software Engineer to build secure, scalable AI-based software...  ...solutions. You will work on developing platform features and managing APIs while collaborating...  ...in full stack development, LLM APIs, and cloud deployment. A strong understanding of... 
    Cloud
    Senior

    Saviynt

    Milpitas, CA
    1 day ago
  • $147k - $237.5k

     ...Clara, CA is seeking a skilled developer proficient in Golang and Python to enhance cloud applications and deliver security features. This role involves collaborating with senior engineers throughout the software development lifecycle and participating in operational... 
    Cloud
    Senior

    Palo Alto Networks

    Santa Clara, CA
    4 days ago
  • $224k - $356.5k

    NVIDIA is seeking an experienced Full-Stack Engineer in Santa Clara, California, to focus on the GeForce Now Platform. The role involves designing and developing real-time applications for both desktop and mobile environments. Candidates should possess extensive experience... 
    Cloud
    Senior

    NVIDIA

    Santa Clara, CA
    1 day ago
  •  ...Overview Senior Data Platform Engineer - Direct-Hire/FTE - Remote (US) This is a hands-on Senior Data Platform Engineer role that will require strong...  ...and related stacks Working experience in at least one cloud service (AWS, GCP, or Azure), preferably AWS Streaming (Kafka... 
    Cloud
    Senior
    Full time
    Work experience placement
    Local area
    Remote work
    Flexible hours

    INSPYR Solutions

    San Jose, CA
    2 days ago
  • $144k - $174k

    Omnicell seeks an Engineer III to join the Platform Infrastructure team, focusing on building and enhancing cloud-native solutions. The role involves designing robust cloud infrastructure, maintaining CI/CD pipelines, and supporting application readiness. Ideal candidates... 
    Cloud
    Senior
    Remote job

    Omnicell

    Milpitas, CA
    1 day ago
  •  ...Senior Platform Engineer Lambda, the superintelligence cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make compute as ubiquitous as... 
    Cloud
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Corporation

    San Jose, CA
    3 hours ago
  • $176k - $276k

     ...exciting time to join us! NVIDIA invites applications for a Senior DevOps Platform Engineer skilled in Platform and Release Engineering to join the...  ...engineering at scale. Background in BareMetal and hybrid cloud (AWS, GCP, Azure) environment management. Familiarity with... 
    Cloud
    Senior

    NVIDIA

    Santa Clara, CA
    6 days ago
  • $176k - $276k

    Cloud Foundations Reliability (CFR) is part of NVIDIA’s Global Network Infrastructure...  ..., and operate the Kubernetes-based platform and shared services used to provision,...  ...environments. We are looking for a hands‑on senior engineer to own the lifecycle and automation of... 
    Cloud
    Senior
    Weekend work

    NVIDIA

    Santa Clara, CA
    2 days ago
  •  ...NVIDIA is looking to hire a deeply technical, creative, and Senior AI Platform Engineer to build, support, and maintain the next generation of AI‑...  ...This role will give you the opportunity to collaborate with Cloud and AI/ML teams in a multifaceted and agile environment.... 
    Cloud
    Senior

    NVIDIA AI

    Santa Clara, CA
    1 day ago
  • $203.3k - $305.6k

    Overview Summary The Apple Services Engineering team is one of the most exciting...  ...of opportunities. The Apple Cloud Services Infrastructure team is seeking a senior engineering program manager to drive the development of our AI platform software stack. You’ll be the... 
    Cloud
    Senior
    Relocation package
    Flexible hours

    Apple

    Cupertino, CA
    4 days ago
  • A leading technology company is seeking a Senior System Software Engineer in Santa Clara, CA. The engineer will design and build a cloud-native scientific computing platform, optimizing algorithms and infrastructure for intensive workloads. This role requires over 10 years... 
    Cloud
    Senior

    Jobs for Humanity

    Santa Clara, CA
    1 day ago
  • $165k - $242k

    A cloud service provider is seeking a Senior Software Engineer II for their Inference team in Sunnyvale, California. In this role, you'll lead design reviews, implement optimizations, and improve service reliability. The ideal candidate has extensive experience with distributed... 
    Cloud
    Senior

    CoreWeave

    Sunnyvale, CA
    4 days ago
  • Coupang is seeking a Senior Staff Backend Engineer specializing in Cloud Infrastructure in Mountain View, California. This role involves enhancing system efficiencies...  ...candidate will have extensive experience in cloud platforms, DevOps tools, and a strong technical background.... 
    Cloud
    Senior

    Coupang

    Mountain View, CA
    1 day ago
  • $126k - $204.5k

    Palo Alto Networks, Inc. is seeking a Senior Software Engineer to lead technical strategy in the Cortex Vulnerability Intelligence team. This...  ...5 years of experience, proficiency in Go, and strong cloud platform knowledge (preferably GCP). The position is located in Santa... 
    Cloud
    Senior

    Palo Alto Networks, Inc.

    Santa Clara, CA
    3 days ago
  • CoreWeave is hiring a Systems Engineer for Next Gen VDI to own cloud infrastructure build-out and OS lifecycle for a new remote compute platform. You’ll design from scratch, integrate Teleport and Okta SSO, and manage image pipelines, patch cadence, and security tooling... 
    Cloud
    Senior
    Remote work

    CoreWeave

    Sunnyvale, CA
    5 days ago
  • $147k - $237.5k

     ...Alto Networks in Santa Clara is seeking a Principal Software Engineer to design, develop, and optimize robust backend systems for AI...  ...engineering, specifically with distributed systems and cloud platforms. The role requires a solution-oriented mindset in a dynamic environment... 
    Cloud
    Senior

    Palo Alto Networks

    Santa Clara, CA
    5 days ago
  • IonQ is seeking a Senior Staff DevOps Engineer for the Quantum Platform Network Security Engineering Team to build, secure, and operate scalable infrastructure for cloud‑managed SaaS products with on‑premises components. You will drive secure delivery across the CI/CD stack... 
    Cloud
    Senior

    IonQ Inc.

    Santa Clara, CA
    1 day ago
  •  ...Cupertino, California is seeking a backend software engineer to help build and scale the Ad Exchange platform. You will lead cross-functional teams to deliver...  ...experts to support diverse ad use cases while leveraging public cloud technologies and #J-18808-Ljbffr Apple Inc.
    Cloud
    Senior

    Apple Inc.

    Cupertino, CA
    3 days ago
  •  ...Madegood) is seeking a dedicated Enterprise Data & Generative AI Engineer to bolster its data capabilities as part of its...  ...background in data engineering, integrating enterprise data with cloud lakehouse platforms. You will collaborate with various teams to implement... 
    Cloud
    Senior

    Riverside Natural Foods Ltd. (Home of Madegood)

    Sunnyvale, CA
    1 day ago
  • Crusoe Energy Systems is seeking a Senior Data Engineer to join their central Data Science and Engineering team. This...  ...you to architect and build foundational data platform infrastructure that supports Crusoe's AI and cloud operations. Your role includes designing scalable... 
    Cloud
    Senior
    Full time

    Crusoe Energy Systems

    Sunnyvale, CA
    1 day ago
  •  ...in Sunnyvale, CA is seeking a seasoned individual for a leadership role in cloud application projects. You will lead the design and development of complex projects, mentoring junior engineers while ensuring the highest standards in software development. The ideal candidate... 
    Cloud
    Senior

    http:/www.ubertal.com

    Sunnyvale, CA
    1 day ago
  • Singlestore in Sunnyvale is seeking a Senior/Principal Software Engineer to lead the technical design and implementation of core capabilities for their cloud platform. The role requires over 6 years of experience in system-level software development, primarily in Golang... 
    Cloud
    Senior

    Singlestore

    Sunnyvale, CA
    6 days ago
  • Palo Alto Networks is looking for a visionary Sr. Engineering Manager, Sales Cloud to redefine their Salesforce engineering portfolio. You will lead...  ...and manage the technical architecture of the Sales Cloud platform, enhancing sales intelligence and automation. Qualified... 
    Cloud
    Senior
    Work at office

    Palo Alto Networks

    Santa Clara, CA
    3 days ago
  • $179k - $219k

     ...seeking a highly motivated full-stack software engineer based in Santa Clara, California. This...  ...designing and enhancing the FortiSASE platform, with responsibilities including integrating...  ...software development within network or cloud security, proficient in programming... 
    Cloud
    Senior

    Fortinet

    Santa Clara, CA
    1 day ago
  • $280k - $350k

    Inworld AI is seeking a Staff / Principal Platform Engineer in Mountain View, CA. This role involves owning cloud infrastructure, enhancing software deployment with AI-powered tooling, and maintaining high-performance systems. Candidates should have 8-10 years of software... 
    Cloud
    Senior
    Full time
    Work at office

    Inworld AI

    Mountain View, CA
    1 day ago
  • $109k - $160k

    CoreWeave is looking for a Senior Engineer specializing in database and stream processing in Sunnyvale, CA. You will design and implement solutions for data flow, improve data platform performance, and ensure compliance with data regulations. Ideal candidates will have... 
    Cloud
    Senior
    Remote job

    CoreWeave

    Sunnyvale, CA
    1 day ago
  • $126.7k - $150k

     ...power a foundational data and services platform that unifies enterprise systems into a canonical...  ...systems. Responsibilities Framework Engineering: Build the reusable plumbing for...  ...and safe recovery from mid‑sync failures. Cloud Native: Experience deploying and managing... 
    Cloud
    Senior
    Local area

    Archer

    San Jose, CA
    5 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Cloud Platform Engineer. Be the first to apply!