Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Cloud Platform Engineer

External SambaNova Systems

The era of pervasive AI has arrived. In this era, organizations will use generative AI to unlock hidden value in their data, accelerate processes, reduce costs, drive efficiency and innovation to fundamentally transform their businesses and operations at scale.

SambaNova Suite™ is the first full-stack, generative AI platform, from chip to model, optimized for enterprise and government organizations. Powered by the intelligent SN40L chip, the SambaNova Suite is a fully integrated platform, delivered on-premises or in the cloud, combined with state-of-the-art open-source models that can be easily and securely fine-tuned using customer data for greater accuracy. Once adapted with customer data, customers retain model ownership in perpetuity, so they can turn generative AI into one of their most valuable assets.

About SambaNova Systems

Join the company that's building the future of AI computing. At SambaNova, we are disrupting the AI and high-performance computing space with our integrated hardware and software platform. Our DataScale systems and SambaFlow software are pushing the boundaries of what's possible with generative AI and large language models. We are a team of passionate innovators tackling some of the world's most challenging computational problems.

The Role

As a Cloud Platform Engineer, you will be specializing in our AI Inferencing Service and will be the guardian of its reliability, performance, and scalability. You will bridge the gap between software development and operations, applying an engineering mindset to solve operational challenges. Your primary focus will be ensuring our inference endpoints have exceptional uptime, low-latency response times, and efficient resource utilization, directly impacting the experience of our customers and the success of our AI products. This role includes participating in a shared on-call rotation to maintain 24/7 service reliability.


What You'll Do

Service Ownership & On-Call: Take shared ownership of the production inferencing service, including its availability, latency, performance, efficiency, change management, monitoring, emergency response, and capacity planning across multiple regions. This includes implementing and supporting AI infrastructure in new regions, such as Asia, Europe, and Latin America, to support the growth of our business. Participate in a balanced on-call rotation to provide 24/7 support for the service.

On-Call & Work-Life Balance

We believe a sustainable on-call schedule is critical for long-term success and team health. Our on-call philosophy is built on the following principles:

  • Balanced Rotation: The on-call rotation is shared equally across the team, typically following a primary/secondary (follow-the-sun) model to ensure no single person bears a disproportionate burden.
  • Focus on Prevention: We invest heavily in automation, robust testing, and system design to prevent pages before they happen. The goal of on-call is not to heroically fight fires, but to manage rare, complex failures and use those learnings to make the system more resilient.
  • Actionable Alerts: We have a strict policy against alert fatigue. Alerts must be actionable and require immediate human intervention.
  • Incident Management: Lead the response to incidents affecting the inferencing service, driving blameless post-mortems and implementing corrective actions to prevent recurrence.
  • Monitoring & Alerting: Develop and maintain advanced monitoring, alerting, and dashboarding (using tools like Prometheus, Grafana, Datadog) to gain deep insights into service health, model performance (e.g., latency, throughput, error rates), and accelerator utilization. A key responsibility is ensuring alerts are actionable and have a low false-positive rate, minimizing on-call fatigue.
  • Performance & Scalability: Proactively identify and eliminate performance bottlenecks. Design and implement auto-scaling policies to handle variable inference loads cost-effectively. Use insights from on-call incidents to drive improvements that enhance system stability and scalability.
  • Infrastructure as Code (IaC): Manage and evolve our cloud infrastructure (on AWS, GCP, and/or Azure along with on-prem) using tools like Terraform and Ansible, ensuring it is secure, repeatable, and scalable.
  • CI/CD & Automation: Champion automation by building and improving CI/CD pipelines for the seamless and safe deployment of new model versions and service updates. A core goal is to automate manual toil identified during on-call shifts, reducing future operational overhead.
  • Capacity Planning: Forecast infrastructure needs based on product roadmaps and usage trends. Work with finance and engineering teams to manage cloud costs and optimize spending.
  • SLOs & SLIs: Define, measure, and report on Service Level Objectives (SLOs) and Indicators (SLIs) for the inferencing platform, using data to drive prioritization and reliability investments.
What We're Looking For (Must-Haves)
  • Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
  • 3-5+ years of experience in a Site Reliability Engineer, DevOps, or related role supporting a large-scale, customer-facing service in a public cloud environment (AWS, GCP, Azure).
  • Strong programming/scripting skills in languages like Python, Go, or Java.
  • Proven experience with containerization and orchestration technologies (Docker, Kubernetes).
  • Deep understanding of monitoring and observability principles and tools (e.g., Prometheus, Grafana, ELK Stack, Datadog).
  • Solid experience with Infrastructure as Code (e.g., Terraform, CloudFormation).
  • Familiarity with CI/CD principles and tools (e.g., Jenkins, GitHub Actions, ArgoCD).
  • Excellent problem-solving skills and a systematic approach to troubleshooting complex distributed systems.
What Will Make You Stand Out (Nice-to-Haves)
  • Experience in a hybrid environment bridging cloud and on-premise/data center infrastructure.
  • Direct experience supporting ML/AI inferencing services in production.
  • Familiarity with GPU-accelerated computing and optimizing workloads for NVIDIA GPUs for purposes of mapping to RDUs.
  • Knowledge of model serving frameworks like vLLM, SGLang or Ray.
  • Understanding of MLOps principles and practices.
  • Experience with managing and tuning databases (SQL or NoSQL) and caching systems (Redis, Memcached).
  • Strong Linux/Unix system administration fundamentals.
Why SambaNova?
  • Massive Impact: You will be a key part of a critical platform with high visibility and direct impact on our product and engineers.
  • Cutting-Edge Technology: Work with a world-class team on one of the most advanced AI stacks in the industry.
  • Autonomy and Growth: We trust you to make technical decisions. This is a greenfield opportunity to build something remarkable from the ground up.
  • Competitive Compensation: Including equity, excellent benefits, and a flexible work environment.

Submission Guidelines Please note that in order to be considered an applicant for any position at SambaNova Systems, you must submit an application form for each position for which you believe you are qualified.


EEO Policy SambaNova Systems is an Equal Opportunity/Affirmative Action Employer. All qualified applicants will receive consideration for employment without regard basis of age (40 and over), color, disability, gender identity, genetic information, marital status, military or veteran status, national origin/ancestry, race, religion, creed, sex (including pregnancy, childbirth, breastfeeding), sexual orientation, and any other applicable status protected by federal, state, or local laws.

Benefits Summary for US-Based, Full-Time Employment Positions
SambaNova offers a competitive total rewards package, including the base salary, plus equity and benefits. We cover 95% premium coverage for employee medical insurance, and 77% premium coverage for dependents and offer a Health Savings Account (HSA) with employer contribution. We also offer Dental, Vision, Short/Long term Disability, Basic Life, Voluntary Life, and AD&D insurance plans in addition to Flexible Spending Account (FSA) options like Health Care, Limited Purpose, and Dependent Care. Our library of well-being benefits available to you and your dependents includes a full subscription to Headspace, Gympass+ membership with access to physical gyms, One Medical membership, counseling services with an Employee Assistance Program, and much more.
Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Cloud Platform Engineer in San Jose, CA vacancy
  • $184k - $287.5k

     ...You will be a part of NVIDIA GeForce NOW cloud team that allows users to play high-quality...  ...even at high resolutions. As an engineer on the GeForce NOW team, you will have the...  ...definition of the next generation cloud platform. You will be at the forefront of transforming... 
    Cloud
    Full time

    Nvidia

    Santa Clara, CA
    13 hours ago
  • $176k - $276k

    Cloud Foundations Reliability (CFR) is part of NVIDIA’s Global Network Infrastructure (GNI...  ..., and operate the Kubernetes-based platform and shared services used to provision, monitor...  ....We are looking for a hands-on senior engineer to own the lifecycle and automation of the... 
    Cloud
    Full time
    Remote work
    Weekend work

    Nvidia

    Santa Clara, CA
    3 days ago
  • $70 - $75 per hour

     ...Meghana GorusuCompany: SRI Tech SolutionsData Platform EngineerLocation - San Jose, CA ( 4 Days...  ...visualization and analytics dashboards.Cloud & Data...  ...lakesData warehousesSnowflake, RedshiftData Engineering & QualityUnderstanding of: Data models,... 
    Cloud
    Hourly pay

    SRI Tech

    San Jose, CA
    13 hours ago
  • Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens...  ...work from home day is currently Tuesday.Engineering at Lambda is responsible for building...  ...management, networking, and RBAC across the platform.Lead incident response, root-cause... 
    Cloud
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    3 days ago
  • $100k

     ...innovations in software models, compilers, platforms, networking, and semiconductors. Our...  ...You AreStrong backend or infrastructure engineer with experience building and operating platforms...  ...data center infrastructure differs from cloud-native environments at scale.How... 
    Cloud
    Permanent employment

    Tenstorrent

    Santa Clara, CA
    3 days ago
  • $184k - $287.5k

     ...at the forefront of technological advancement. We’re hiring a Senior Backend/Platform Engineer to build and maintain the core infrastructure behind NVIDIA Brev. You’ll develop reliable cloud services, control planes, and execution environments that enable developers to... 
    Cloud
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $196.8k - $262.7k

     ...the way the world shops and sells. Our platform empowers millions of buyers and sellers...  ...for all.About the team and the role:The Engineering Systems Tools team powers how eBay engineers...  ...Server and/or GitHub Enterprise Cloud at scale, including repository and organization... 
    Cloud
    Immediate start
    Remote work
    Visa sponsorship

    eBay

    San Jose, CA
    2 days ago
  • $245k - $350k

     ...leveraging the world’s largest security data lake to power our cloud-native Zero Trust Exchange platform. This innovation protects our customers from...  ...Jose, reporting to the Director,Software Development Engineering ZIdentity team. You will be responsible for design and... 
    Cloud
    Full time
    Work at office
    Local area
    3 days per week

    Zscaler

    San Jose, CA
    3 days ago
  • $104.5k - $234.6k

    Leads platform projects crossing multiple teams, evolving runtimes or middleware patterns...  ...at a company leading the way in AI and cloud solutions that impact billions of lives....  ...lifecycle; provides guidance and coaching to engineers to drive improvements.Utilizes advanced... 
    Cloud
    Temporary work
    Flexible hours
    Shift work

    Oracle Corporation

    Santa Clara, CA
    2 days ago
  •  ...Overview: • Platform Architecture: • Lead the architectural design and hands-on implementation...  ...considerations • Mentor and guide the engineering team through the challenges of...  ...(Python, Ruby) • Experience building cloud native, containerized applications running... 
    Cloud

    Purple Drive

    San Jose, CA
    2 days ago
  • $255k - $340k

    Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens...  ...home day is currently Tuesday.Hardware Engineering at Lambda is responsible for building...  ..., proof-of-concept for state-of-the-art platforms, new product introduction (NPI), and fleet... 
    Cloud
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    3 days ago
  • $151.8k - $265.35k

     ...create exceptional content effortlessly. The AI for Engineering team builds a scalable, production-grade AI platform that powers creativity across design, imaging,...  ...coordination.Proven expertise in building scalable, cloud-native, microservices-based architectures with... 
    Cloud
    Full time
    Temporary work
    Local area
    Worldwide

    Adobe Systems

    San Jose, CA
    13 hours ago
  •  ...Role - Snowflake Platform Engineer Bright Vision Technologies is a forward-thinking software development company dedicated to building innovative...  ...BRIGHT VISION Job Description: Environment: Snowflake Data Cloud, SQL, Snowflake Warehousing & Performance Tuning, Data... 
    Cloud
    Full time
    H1b
    Local area
    Remote work

    GrabJobs

    San Jose, CA
    2 days ago
  • $178k - $321k

     ...working AI-native capability: a multi-agent platform (Hive Mind), agentic workflows, data...  ...agent runtime, and the harness: a resilient cloud platform, the agentic runtime and...  ...just implement it. This is a two-person engineering team: you deploy, debug, and hotfix your... 
    Cloud

    OKX

    San Jose, CA
    2 days ago
  •  ...is 100% remote anywhere in the US Overview: We are hiring a Platform Engineer to lead security tooling and enablement within Platform Engineering...  ...and supply chain security, artifact and container security, cloud security posture, and policy-driven workflows Build,... 
    Cloud
    Remote work

    GrabJobs

    San Jose, CA
    3 days ago
  •  ...operations. Bitdeer also offers advanced cloud capabilities to customers with high demand...  ...bounded contexts of the NeoCloud SRE platform — the multi-region substrate that observes...  ...monitor. Alert, Correlation & SLO: alert-engine-framework, alert-correlation, slo-... 
    Cloud
    Full time
    Contract work
    Local area

    Bitdeer Technologies Group

    San Jose, CA
    a month ago
  • $96.8k - $306.4k

    Drives cross-group platform initiatives (e.g., identity, config, API gateways) with significant...  ...at a company leading the way in AI and cloud solutions that impact billions of lives....  ...software development lifecycle; coaches engineers across teams or units to drive... 
    Cloud
    Temporary work
    Flexible hours
    Shift work

    Oracle Corporation

    Santa Clara, CA
    13 hours ago
  • $147k - $237.5k

     ...drives great outcomes.Job SummaryYour Career Our ATP Cloud team is at the forefront of integrating AI into...  ...incidents at scale. As a Principal Software Engineer, you will own the technical vision for our AI-powered platform, shaping how agentic workflows and LLM-driven automation... 
    Cloud
    Full time
    Work at office

    Palo Alto Networks

    Santa Clara, CA
    13 hours ago
  • $157.2k - $254.1k

     ...the world's most comprehensive AI security platform. Organizations are increasingly building...  ...a Principal Machine Learning Inference Engineer, you will serve as a technical authority...  ...large-scale distributed systems on a major cloud platform (GCP, AWS, Azure, or OCI).Proven... 
    Cloud
    Full time
    Work at office

    Palo Alto Networks

    Santa Clara, CA
    2 days ago
  • $110k - $187k

     ...15% of sales back into R&D. Our expert teams of physicists, engineers, data scientists and problem-solvers work together with the world...  .... Job Description/Preferred Qualifications The Cloud Platform Engineer will build, automate, secure, and operate cloud-based... 
    Cloud
    Minimum wage
    Flexible hours

    KLA

    Milpitas, CA
    5 days ago
  • $176.1k - $308.2k

     ...action on what data.Veza's Access Graph platform maps an organization's entire identity ecosystem...  ..., and agentic identities across SaaS, cloud, on-prem, and custom applications.With...  ...cloud environments, and AI agents. For engineers joining Veza today, this means the scale... 
    Cloud
    Work at office
    Remote work
    Flexible hours

    ServiceNow

    Santa Clara, CA
    2 days ago
  •  ...deterministic performance, and flexible network evolution for AI, cloud, telecom, and edge infrastructure. Designed for real world...  .... Role Overview We are seeking a Silicon Photonics PIC Platform Engineer to lead engagement with external silicon photonics foundries... 
    Cloud
    Full time
    Flexible hours

    Arycs Technologies, Inc.

    Los Gatos, CA
    a month ago
  •  ...Vice President, Endpoint Platform Engineering About the Company Pioneering cybersecurity solutions platform Industry Computer &...  ...reporting ~24x7 monitoring ~ vulnerability assessment ~ cloud security ~ managed cloud monitoring ~ and managed risk... 
    Cloud

    Confidential

    San Jose, CA
    4 days ago
  • $92.5k - $209.5k

     ...Job Description Owns moderately complex components within platform services or SDKs; leads team-level improvements to integration frameworks...  ...Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives. True innovation... 
    Cloud
    Temporary work
    Immediate start
    Flexible hours
    Shift work

    Oracle

    Santa Clara, CA
    1 day ago
  • $173.5k - $331.05k

     ...mission to transform the future of creativity by building SDKs and platform libraries that power data‑driven insights and AI‑enabled experiences across Creative Cloud. We are seeking a software engineer with strong development and computer science fundamentals and experience... 
    Cloud
    Full time
    Temporary work
    Local area
    Worldwide

    Adobe Systems

    San Jose, CA
    4 days ago
  • Overview Senior Data Platform Engineer - Direct-Hire/FTE - Remote (US) This is a hands-on Senior Data Platform Engineer role that will require...  ...Hadoop and related stacks Working experience in at least one cloud service (AWS, GCP, or Azure), preferably AWS Streaming (Kafka... 
    Cloud
    Full time
    Work experience placement
    Local area
    Remote work
    Flexible hours

    INSPYR Solutions

    San Jose, CA
    13 hours ago
  • $274k - $304k

     ...AI Platform Engineer - Training & Inference Saviynt's AI-powered identity platform manages and governs human and non-human access to all...  ...with cost-aware fallback between self-hosted SLMs and cloud LLMs • Build RL training infrastructure: define Flyte workflows... 
    Cloud

    Saviynt

    Milpitas, CA
    4 days ago
  • $122.44k - $232.19k

     ...will be joining the Intel Government Technologies Customer Engineering team as a Hardware Platform Applications Engineer (PAE). This is an exciting...  ...every person on Earth. Harnessing the capability of the cloud, the ubiquity of the Internet of Things, the latest advances... 
    Cloud
    Full time
    Internship
    Local area
    Immediate start
    Shift work

    Intel

    Santa Clara, CA
    13 hours ago
  • $240k - $260k

     ...identity security, delivering an AI-powered platform that governs and secures access to...  ...AI Platform. You define the standards ML engineers and scientists build on, and ensure every...  ...and Vector database: operate Pgvector (Cloud SQL) for POC and Qdrant on GKE for... 
    Cloud

    Saviynt

    Milpitas, CA
    a month ago
  • $216k - $345k

     ...set strategy and drive execution across a portfolio of platforms, manage a team of experienced engineers and technical program managers, oversee vendor...  ...tradeoffs. ~ Strong command of modern data architecture: cloud-native production environments, event pipelines,... 
    Cloud
    Full time

    NVIDIA

    Santa Clara, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Cloud Platform Engineer. Be the first to apply!