Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Observability Platform Engineer — Scale & Reliability

SpaceXAI

SpaceXAI is seeking a Member of Technical Staff for Observability to design and operate the core observability platform at scale. You will own metrics, logs, tracing, and alerting capabilities that empower engineers to monitor and optimize services across the fleet. You will work on high-impact telemetry pipelines and APIs, partnering with infra and product teams to embed observability into internal platforms and ensure reliability and performance under massive load. #J-18808-Ljbffr SpaceXAI

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Observability Platform Engineer — Scale & Reliability in Palo Alto, CA vacancy
  •  ...is seeking a Senior Software Engineer (Observability Solutions) to architect and lead...  ...for a scalable observability platform. You will shape system design,...  ...collaboration across teams to deliver reliable, high-volume services for enterprise-scale workloads. #J-18808-Ljbffr... 
    Suggested

    Apple

    Sunnyvale, CA
    4 days ago
  • $167.7k - $245.2k

     ...as intended, improving reliability and reducing risks....  ...applications with enhanced observability and control.As a Senior Site Reliability Engineer (SRE), you will build,...  ...'s deployment platform and production infrastructure...  ...happen on a global scale. Because our solutions... 
    Suggested
    Full time
    Temporary work
    Local area
    Flexible hours
    2 days per week

    CISCO Systems

    Palo Alto, CA
    1 day ago
  • $165k - $190k

     ...the leading SaaS security platform, trusted by global enterprises...  ...ahead, Obsidian is scaling rapidly toward long-term growth...  ...at Obsidian ensures that engineering excellence translates into...  ...around scalability, reliability, observability, and cost efficiencyCollaborate... 
    Suggested
    Work from home

    Obsidian Security

    Palo Alto, CA
    4 days ago
  • $232k - $263k

     ...leading SaaS security platform, trusted by global...  ...fundraise ahead, Obsidian is scaling rapidly toward long-...  ....Sr. Staff Site Reliability EngineerAs a Sr. Staff...  ...DevOps and Platform Engineering leadership, shaping a...  ...guide architecture for observability, detection, and... 
    Suggested
    Work from home

    Obsidian Security

    Palo Alto, CA
    1 day ago
  • $186.9k - $267.7k

     ...as intended, improving reliability and reducing risks....  ...applications with enhanced observability and control.As a Staff Site Reliability Engineer (SRE), you will...  ...Agent Observability's platform. You will define the long...  ...practices for operating large-scale cloud and on-prem... 
    Suggested
    Full time
    Temporary work
    Local area
    Flexible hours
    2 days per week

    CISCO Systems

    Palo Alto, CA
    2 days ago
  • Parallel Web Systems in the Bay Area is seeking engineers to build the data foundations powering...  ...that ingest and transform data at scale, build storage and serving layers, and establish quality, lineage, and observability to keep the data trustworthy. This fully in... 
    Work at office

    Parallel Web Systems

    Palo Alto, CA
    2 days ago
  •  ...Description Job Description Site Reliability Engineer Onsite- Bay Area, CA...  ...using Grafana and other observability tools. Ensure high...  ...reliability, and uptime across platforms. Handle infrastructure maintenance, upgrades, and scaling. Administer and improve... 

    Amiri Recruiting

    Mountain View, CA
    more than 2 months ago
  •  ...it’s already operating at scale inside some of the world’s...  ..., lead, and scale our Site Reliability Engineering function. This role combines...  ...monitoring, alerting, and observability practices. Design and implement...  ...and productionise data platforms and ML workloads. Partner closely... 
    Shift work

    Wand AI

    Palo Alto, CA
    2 days ago
  • $170k - $250k

     ...Site Reliability Engineer (SRE) Location: San Francisco, CA / Palo Alto...  ...next-generation GPU cloud platform for enterprises, startups,...  ...engineering to build the automation, observability, and platform...  ...multi-cloud GPU marketplace at scale. What You Will Do... 
    Work at office
    Visa sponsorship
    Flexible hours

    Recruiting from Scratch

    Palo Alto, CA
    2 days ago
  • $170k - $230k

     ...Site Reliability Engineer (SRE) Palo Alto / San Francisco Bay Area About...  ...is an AI infrastructure platform built to make GPU compute more...  ...will shape how Mithril scales its platform across a heterogeneous...  ...will build the automation, observability, and tooling that allows... 
    Work at office
    Local area
    1 day per week

    Mithril

    Palo Alto, CA
    3 days ago
  •  ...Staff to design, build, and operate the Data Platform infrastructure powering real-time ML pipelines and analytics at petabyte scale. You will own distributed systems handling...  ...focusing on scalability, performance, and reliability across Kafka, Spark, Flink, and Trino.... 

    SpaceXAI

    Palo Alto, CA
    2 days ago
  •  ...right place. As a Principal Engineer, Developer Platform Engineering at...  ...ways to answer that question reliably at any time. You will influence...  ...patterns under extreme loadBuild observability-driven capacity...  ...and operational outcomes at scale (e.g., AI-driven PR review... 
    Early shift

    JP Morgan Chase

    Palo Alto, CA
    17 hours ago
  • $272k - $431.25k

     ...Principal Systems Software Engineer at NVIDIA is an...  ...build and maintain large scale production systems with...  ...cloud services run maximum reliability and uptime as promised to...  ...aspects of large scale Observability & Telemetry collection platform with a focus on performance... 

    NVIDIA Gruppe

    Santa Clara, CA
    2 days ago
  • $186k - $232.5k

     ...small group of exceptional engineers as the function scales. Responsibilities Lead...  ...failure. Surface weaknesses in observability, process, testing, and...  ...in SRE, production/platform engineering, or systems engineering...  ...). Fluency in modern reliability and post-incident... 
    Full time
    Contract work

    Rivian and Volkswagen Group Technologies

    Palo Alto, CA
    2 days ago
  • $184k - $287.5k

     ...seeking a Senior System Software Engineer to lead the evolution of our next-generation Data & Observability Platform. We serve and collaborate...  ..., and ensure platform reliability.What you’ll be doing:Architect...  ...capable of handling massive scale. You will solve global latency... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $176k - $276k

     ...applications for a Senior DevOps Platform Engineer skilled in Platform and...  ...analytics workloads at scale using NVIDIA Data Center GPUs...  ...engineering rigor by setting up reliable release workflows,...  ...automation frameworks.Own observability and monitoring infrastructure... 
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  •  ...we're looking to grow as we scale Mino to millions of users....  ...future. The Role As a Senior Platform Engineer, you'll own the backend...  ...moving product team to ship reliable, performant features to end...  ...engineering team Implement observability, monitoring, and alerting... 
    Remote work
    Flexible hours

    Saga.xyz

    Los Altos, CA
    2 days ago
  • $115.5k - $189.75k

     ...software development platform for software-defined vehicles...  ...into actionable engineering insights.  The Release...  ...qualification, large-scale simulation-based...  ...and implement scalable, reliable internal services used...  ...ensuring maintainability, observability, and performance at... 
    Full time
    Temporary work
    Work at office
    Flexible hours

    Woven

    Palo Alto, CA
    17 hours ago
  •  ...delivering an AI-powered platform that governs and...  .... As a Staff Platform Engineer, you will play a critical...  ...leadership role. You will own reliability for major platform...  ...enhance centralized Observability and Monitoring...  ...Work on a large-scale, cloud-native SaaS platform... 

    Saviynt

    Milpitas, CA
    7 days ago
  •  ...looking for a Senior Site Reliability Engineer to lead the strategic evolution...  ...support our next 10x of scale. You will be the primary authority...  ...cost-efficiency.Advanced Observability: Define our monitoring and...  ...New Relic (or similar platforms) to create meaningful dashboards... 
    Full time
    Work at office
    2 days per week

    LeanData

    Santa Clara, CA
    2 days ago
  • $148k - $235.75k

     ...our team of innovative engineers who are building an AI Data Center AIOps platform that turns raw, high-volume telemetry into reliable, job-centric insights...  ...of reliability for an observability/AIOps platform: SLOs/SLIs...  ...deploying, debugging, scaling) for telemetry-heavy... 
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $128.6k - $184.9k

     ...powers our global cloud platform. As a team of six engineers distributed across the US...  ...strong focus on automation, reliability, and operational...  ...QualificationsExperience supporting large-scale infrastructure...  ...Knowledge of monitoring, observability, and reliability engineering... 
    Permanent employment
    Full time
    Temporary work
    Local area
    Worldwide
    Flexible hours

    CISCO Systems

    Milpitas, CA
    1 day ago
  • $152k - $241.5k

     ...generation of our global services platform. At NVIDIA, you’ll keep...  ...supporting large‑scale HPC clusters using Slurm,...  ...lifecycle management, fleet reliability/auto-healing, E2E observability or data-driven operations...  ..., or Ruby.Mentored other engineers and influenced technical... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $165k - $280k

     ...enabling human life on Mars.SR. SITE RELIABILITY ENGINEER (STARLINK)At SpaceX we’re leveraging our...  ...-region environment Manage petabyte scale bare metal compute clusters Closely collaborate...  ...and analytics using open source platforms such as Apache Kafka, Spark, HBase,... 
    Permanent employment
    Temporary work
    Worldwide
    Weekend work

    SpaceX

    Palo Alto, CA
    3 days ago
  • $90k - $180k

     ...About the RoleThis Senior Site Reliability Engineer position works on-site out...  ....net — a remote monitoring platform designed to help doctors,...  ...health and behavior at scale. Automate away manual operational...  ..., and cluster operations.Observability platform experience with... 
    Remote work

    Abbott

    Sunnyvale, CA
    4 days ago
  • $168k - $270.25k

     ...looking for a Senior Site Reliability Engineer (SRE) to join its GeForce Now...  ...tools to improve the SRE Observability. Be part of the Kubernetes...  ...consulting, developing software platforms and frameworks, capacity...  ...experience working on large scale distributed micro services... 
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $230k - $250k

     ...autonomous networking, giving engineers and AI agents the...  ...a groundbreaking platform that transforms how...  ...is looking for a Site Reliability EngineerAbout the Role...  ...about availability, observability, incident response, and...  ...team as the company scales — this is a... 
    Night shift

    Forward Networks

    Santa Clara, CA
    2 days ago
  • $210.6k - $305.1k

     ...and operates our US GovCloud platform. This team is responsible...  ...helping customers deploy at scale while also delivering AI-powered...  ..., Collaboration, and Observability portfolios Your Impact  As part...  ...led a distributed team of 5+ engineers, can demonstrate strong technical... 
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    Los Altos, CA
    4 days ago
  • Elevate your engineering prowess to unprecedented levels by joining...  ...the top echelon in site reliability. As a Senior Lead Site Reliability...  ...within the Infrastructure Platforms and Foundational Services (...  ...to create and implement observability and reliability designs for... 

    JP Morgan Chase

    Palo Alto, CA
    2 days ago
  • $184k - $287.5k

    At NVIDIA, Site Reliability Engineering provides a rare chance to define, develop, and support large-scale production systems with high efficiency and availability. This...  ..., or Ruby.Hands-on experience with observability platforms (e.g., Prometheus, Grafana).Strong communication... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Observability Platform Engineer — Scale & Reliability. Be the first to apply!