Observability Platform Engineer — Scale & Reliability
SpaceXAI
SpaceXAI is seeking a Member of Technical Staff for Observability to design and operate the core observability platform at scale. You will own metrics, logs, tracing, and alerting capabilities that empower engineers to monitor and optimize services across the fleet. You will work on high-impact telemetry pipelines and APIs, partnering with infra and product teams to embed observability into internal platforms and ensure reliability and performance under massive load. #J-18808-Ljbffr SpaceXAI
- ...is seeking a Senior Software Engineer (Observability Solutions) to architect and lead... ...for a scalable observability platform. You will shape system design,... ...collaboration across teams to deliver reliable, high-volume services for enterprise-scale workloads. #J-18808-Ljbffr...Suggested
$167.7k - $245.2k
...as intended, improving reliability and reducing risks.... ...applications with enhanced observability and control.As a Senior Site Reliability Engineer (SRE), you will build,... ...'s deployment platform and production infrastructure... ...happen on a global scale. Because our solutions...SuggestedFull timeTemporary workLocal areaFlexible hours2 days per week$165k - $190k
...the leading SaaS security platform, trusted by global enterprises... ...ahead, Obsidian is scaling rapidly toward long-term growth... ...at Obsidian ensures that engineering excellence translates into... ...around scalability, reliability, observability, and cost efficiencyCollaborate...SuggestedWork from home$232k - $263k
...leading SaaS security platform, trusted by global... ...fundraise ahead, Obsidian is scaling rapidly toward long-... ....Sr. Staff Site Reliability EngineerAs a Sr. Staff... ...DevOps and Platform Engineering leadership, shaping a... ...guide architecture for observability, detection, and...SuggestedWork from home$186.9k - $267.7k
...as intended, improving reliability and reducing risks.... ...applications with enhanced observability and control.As a Staff Site Reliability Engineer (SRE), you will... ...Agent Observability's platform. You will define the long... ...practices for operating large-scale cloud and on-prem...SuggestedFull timeTemporary workLocal areaFlexible hours2 days per week- Parallel Web Systems in the Bay Area is seeking engineers to build the data foundations powering... ...that ingest and transform data at scale, build storage and serving layers, and establish quality, lineage, and observability to keep the data trustworthy. This fully in...Work at office
- ...Description Job Description Site Reliability Engineer Onsite- Bay Area, CA... ...using Grafana and other observability tools. Ensure high... ...reliability, and uptime across platforms. Handle infrastructure maintenance, upgrades, and scaling. Administer and improve...
- ...it’s already operating at scale inside some of the world’s... ..., lead, and scale our Site Reliability Engineering function. This role combines... ...monitoring, alerting, and observability practices. Design and implement... ...and productionise data platforms and ML workloads. Partner closely...Shift work
$170k - $250k
...Site Reliability Engineer (SRE) Location: San Francisco, CA / Palo Alto... ...next-generation GPU cloud platform for enterprises, startups,... ...engineering to build the automation, observability, and platform... ...multi-cloud GPU marketplace at scale. What You Will Do...Work at officeVisa sponsorshipFlexible hours$170k - $230k
...Site Reliability Engineer (SRE) Palo Alto / San Francisco Bay Area About... ...is an AI infrastructure platform built to make GPU compute more... ...will shape how Mithril scales its platform across a heterogeneous... ...will build the automation, observability, and tooling that allows...Work at officeLocal area1 day per week- ...Staff to design, build, and operate the Data Platform infrastructure powering real-time ML pipelines and analytics at petabyte scale. You will own distributed systems handling... ...focusing on scalability, performance, and reliability across Kafka, Spark, Flink, and Trino....
- ...right place. As a Principal Engineer, Developer Platform Engineering at... ...ways to answer that question reliably at any time. You will influence... ...patterns under extreme loadBuild observability-driven capacity... ...and operational outcomes at scale (e.g., AI-driven PR review...Early shift
$272k - $431.25k
...Principal Systems Software Engineer at NVIDIA is an... ...build and maintain large scale production systems with... ...cloud services run maximum reliability and uptime as promised to... ...aspects of large scale Observability & Telemetry collection platform with a focus on performance...$186k - $232.5k
...small group of exceptional engineers as the function scales. Responsibilities Lead... ...failure. Surface weaknesses in observability, process, testing, and... ...in SRE, production/platform engineering, or systems engineering... ...). Fluency in modern reliability and post-incident...Full timeContract work$184k - $287.5k
...seeking a Senior System Software Engineer to lead the evolution of our next-generation Data & Observability Platform. We serve and collaborate... ..., and ensure platform reliability.What you’ll be doing:Architect... ...capable of handling massive scale. You will solve global latency...Full time$176k - $276k
...applications for a Senior DevOps Platform Engineer skilled in Platform and... ...analytics workloads at scale using NVIDIA Data Center GPUs... ...engineering rigor by setting up reliable release workflows,... ...automation frameworks.Own observability and monitoring infrastructure...Full time- ...we're looking to grow as we scale Mino to millions of users.... ...future. The Role As a Senior Platform Engineer, you'll own the backend... ...moving product team to ship reliable, performant features to end... ...engineering team Implement observability, monitoring, and alerting...Remote workFlexible hours
$115.5k - $189.75k
...software development platform for software-defined vehicles... ...into actionable engineering insights. The Release... ...qualification, large-scale simulation-based... ...and implement scalable, reliable internal services used... ...ensuring maintainability, observability, and performance at...Full timeTemporary workWork at officeFlexible hours- ...delivering an AI-powered platform that governs and... .... As a Staff Platform Engineer, you will play a critical... ...leadership role. You will own reliability for major platform... ...enhance centralized Observability and Monitoring... ...Work on a large-scale, cloud-native SaaS platform...
- ...looking for a Senior Site Reliability Engineer to lead the strategic evolution... ...support our next 10x of scale. You will be the primary authority... ...cost-efficiency.Advanced Observability: Define our monitoring and... ...New Relic (or similar platforms) to create meaningful dashboards...Full timeWork at office2 days per week
$148k - $235.75k
...our team of innovative engineers who are building an AI Data Center AIOps platform that turns raw, high-volume telemetry into reliable, job-centric insights... ...of reliability for an observability/AIOps platform: SLOs/SLIs... ...deploying, debugging, scaling) for telemetry-heavy...Full time$128.6k - $184.9k
...powers our global cloud platform. As a team of six engineers distributed across the US... ...strong focus on automation, reliability, and operational... ...QualificationsExperience supporting large-scale infrastructure... ...Knowledge of monitoring, observability, and reliability engineering...Permanent employmentFull timeTemporary workLocal areaWorldwideFlexible hours$152k - $241.5k
...generation of our global services platform. At NVIDIA, you’ll keep... ...supporting large‑scale HPC clusters using Slurm,... ...lifecycle management, fleet reliability/auto-healing, E2E observability or data-driven operations... ..., or Ruby.Mentored other engineers and influenced technical...Full time$165k - $280k
...enabling human life on Mars.SR. SITE RELIABILITY ENGINEER (STARLINK)At SpaceX we’re leveraging our... ...-region environment Manage petabyte scale bare metal compute clusters Closely collaborate... ...and analytics using open source platforms such as Apache Kafka, Spark, HBase,...Permanent employmentTemporary workWorldwideWeekend work$90k - $180k
...About the RoleThis Senior Site Reliability Engineer position works on-site out... ....net — a remote monitoring platform designed to help doctors,... ...health and behavior at scale. Automate away manual operational... ..., and cluster operations.Observability platform experience with...Remote work$168k - $270.25k
...looking for a Senior Site Reliability Engineer (SRE) to join its GeForce Now... ...tools to improve the SRE Observability. Be part of the Kubernetes... ...consulting, developing software platforms and frameworks, capacity... ...experience working on large scale distributed micro services...Full time$230k - $250k
...autonomous networking, giving engineers and AI agents the... ...a groundbreaking platform that transforms how... ...is looking for a Site Reliability EngineerAbout the Role... ...about availability, observability, incident response, and... ...team as the company scales — this is a...Night shift$210.6k - $305.1k
...and operates our US GovCloud platform. This team is responsible... ...helping customers deploy at scale while also delivering AI-powered... ..., Collaboration, and Observability portfolios Your Impact As part... ...led a distributed team of 5+ engineers, can demonstrate strong technical...Full timeTemporary workLocal areaFlexible hours- Elevate your engineering prowess to unprecedented levels by joining... ...the top echelon in site reliability. As a Senior Lead Site Reliability... ...within the Infrastructure Platforms and Foundational Services (... ...to create and implement observability and reliability designs for...
$184k - $287.5k
At NVIDIA, Site Reliability Engineering provides a rare chance to define, develop, and support large-scale production systems with high efficiency and availability. This... ..., or Ruby.Hands-on experience with observability platforms (e.g., Prometheus, Grafana).Strong communication...Full time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Observability Platform Engineer — Scale & Reliability. Be the first to apply!
- data platform engineer Palo Alto, CA
- client platform engineer Palo Alto, CA
- platform engineering manager Palo Alto, CA
- platform developer Palo Alto, CA
- senior platform engineer Palo Alto, CA
- platform engineer Palo Alto, CA
- platform product manager Palo Alto, CA
- platform manager Palo Alto, CA
- director of digital platform Palo Alto, CA
- power platform Palo Alto, CA




