Senior AI and HPC Observability Engineer
$152k - $241.5kNVIDIA
NVIDIA is a pioneer in accelerated computing, known for inventing the GPU and driving breakthroughs in gaming, computer graphics, high-performance computing, and artificial intelligence. Our technology powers everything from generative AI to autonomous systems, and we continue to shape the future of computing through innovation and collaboration. Within this mission, our team, Managed AI Superclusters (MARS) builds and scales the infrastructure, platforms, and tools that enable researchers and engineers to develop the next generation of AI/ML systems. By joining us, you’ll help design solutions that power some of the world’s most advanced computing workloads.Observability is at the heart of this transformation. We are looking for a strong AI & HPC Observability Engineer to build and scale next-generation Observability and Telemetry platforms. You will design and develop high-throughput, reliable telemetry pipelines and modern data infrastructure. This role requires solid distributed systems fundamentals, production-grade coding, and a passion for operational excellence.What You Will Be Doing:Design and scale observability platforms handling high-volume metrics, logs, and traces across distributed environmentsBuild high-performance backend services for telemetry ingestion, processing, and routingDevelop and extend OpenTelemetry collectors, processors, exporters, and instrumentation librariesBuild and optimize metrics pipelines using large-scale time-series storage systemsDesign and operate real-time and batch telemetry pipelines using streaming and distributed data technologiesImprove platform reliability, performance, and cost efficiency through tuning, capacity planning, and system optimizationDevelop monitoring, alerting, and service reliability frameworks to ensure platform health and performanceCollaborate with platform engineering, infrastructure, and site reliability teams to deliver production-grade observability solutionsWhat We Need to see:Bachelor’s degree in Computer Science, Computer Engineering, or related field or equivalent experience5+ years of experience building backend or distributed systems in production environmentsStrong programming skills in Python, Go, or Java, with experience developing production-quality softwareHands-on experience with modern observability architectures, including metrics, logs, and tracesSolid experience with PromQL and time-series data systemsExperience building or operating distributed data pipelines using technologies such as Kafka, Spark, or FlinkExperience working with Kubernetes and cloud-native infrastructureStrong understanding of distributed systems, concurrency, and fault-tolerant system design. Strong debugging, performance tuning, and production operations skillsWays To Stand Out from The Crowd:Proven experience designing and scaling observability platforms for AI, GPU, or HPC environmentsHands-on expertise with OpenTelemetry, Prometheus, Kafka, and high-volume distributed telemetry pipelinesStrong background in data engineering, time-series data modeling, and real-time performance tuningExperience integrating observability with AI/ML pipelines, GPU workload monitoring, or intelligent alertingDemonstrated use of statistical or machine learning techniques for anomaly detection, correlation, or predictive insightsYour base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 152,000 USD - 241,500 USD for Level 3, and 184,000 USD - 287,500 USD for Level 4.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until March 6, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa Clara; US, WA, SeattleType: Full time
- Atlassian is seeking a Systems Engineer to build observability for AI-assisted software development. The role includes designing pipelines and working directly with Git to collect AI-generated code data. Ideal candidates will have systems engineering experience, particularly...Senior
$147k - $202k
Secure Every Identity, from AI to HumanIdentity is the key to... ...let's talk.The Auth0 Platform Observability team owns the observability... ...looking for an Observability Engineer to help ensure that our Product... ...in these areas.As a Senior Engineer on this team, you will...SeniorLocal areaWorldwideFlexible hours$200k - $270k
.... To usher in this new era, we seek AI-native thinkers across every function... ...unlock the power of AI. As a Senior Forward Deployed Engineer, Applied AI on our Cortex AI team, you... ...environments. Build the safety guardrails, observability, and human-review workflows that...Senior$142k - $220.5k
Job DescriptionA Senior Engineer on the AI Enablement team designs and delivers AI platform capabilities that let Nordstrom Technology use Generative... ...with developer tools, cloud services, LLM providers, and observability systems.Own systems and designs that span multiple...SeniorFull timeTemporary work$170k - $220k
...a hands-on, high-agency Site Reliability Engineer to help shape and scale the reliability layer... .... Build safe, repeatable, and observable workflows.GitHub Operations: Manage GitHub... ...and improve deployment confidence. Using AI tools to generate code is encouraged and...Senior- Invoca is seeking experienced engineers to own critical stages of its AI-powered engineering lifecycle. You will define success metrics, drive improvements, and establish safe, observable ways to integrate AI into workflows. The role emphasizes business impact, reliability...SeniorRemote job
$200k - $275k
Secure Every Identity, from AI to HumanIdentity is the key to... ...identity, platform, and security engineering teams, write production code... ...on the patterns.Engage senior leadership. Brief the CISO, CIO... ...runtime decisions.Build evals and observability. Authorization decision...SeniorLocal areaWorldwideFlexible hoursShift work$126.2k - $264.1k
What You’ll DoAs a Senior Principal Network Reliability Engineer, you will:Provide technical leadership for the deployment... ...initiatives supporting AI infrastructure, new OCI Regions, backbone... ..., operational readiness, observability, capacity planning, and lifecycle...SeniorTemporary workFlexible hours- The Trade Desk in Bellevue is looking for a talented engineer to join our Service Excellence team. You will build and maintain automation... ...should have strong debugging instincts and familiarity with observability concepts. The position offers a competitive benefits package...Senior
- Zoox is seeking a Senior Software Engineer focused on High Performance Computing in Seattle. This role involves enhancing Zoox's HPC infrastructure for machine learning workflows. Candidates should have experience in designing distributed storage systems and proficiency...Senior
- A technology startup in Seattle is seeking a Senior Rust Engineer to own the observability of their SDK. This role involves designing and building a telemetry pipeline for their Rust core across various platforms. Ideal candidates have strong experience in Rust, telemetry...SeniorFlexible hours
$127k - $249k
The TeamPlatform Engineering is the department within SRE that is responsible for a range of... ...facing edge and internal service mesh), and observability and alerting systems.The Fleet... ...have redefined the data platform for the AI era, enabling builders to create, transform...SeniorWork at officeLocal areaRemote workWorldwideFlexible hours$134.25k - $214.8k
...company where you matter.Your ImpactAre you an engineer who gets excited about the challenge of making complex distributed systems observable — not just instrumenting them, but... ...infrastructure engineeringExperience with agentic AI tooling or building LLM-powered developer...SeniorWork experience placementWork at officeRemote work- Crusoe Cloud is revolutionizing HPC by offering sustainable, low-cost GPU compute power. As a Senior Cloud Support Engineer, you will be the primary technical support contact, helping... ...goals and accelerate development in AI, physics simulations, and computational biology...Senior
$165k - $225.6k
Secure Every Identity, from AI to HumanIdentity is the key to unlocking the... ...and innovation through technology.The Senior Site Reliability Engineer OpportunityReporting to the Manager,... ...GitHub Actions).Containerization & Observability: Strong hands-on experience managing...SeniorPermanent employmentLocal areaWorldwideFlexible hours$145k - $193.75k
...are uniting human teams with AI agents. By orchestrating the... ...every day.Corporate Systems Engineering builds and operates the software... ...auditable automation.As a Senior Software Engineer I (Automation... ...human reviewersExpertise in observability platforms (e.g., Datadog, CloudWatch...SeniorFull timeTemporary workWork at officeLocal areaRemote work$127k - $249k
We are looking for an experienced Senior or Staff Engineer for our SRE, InfraSec team, to guide the security... ..., such as runtime scanning, security observability, CSPM, and moreCloud Expertise: Strong... ...redefined the data platform for the AI era, enabling builders to create,...SeniorLocal areaRemote workWorldwideFlexible hours- Otter in Seattle is seeking a software engineer to own Revenue Recapture, an automated tool that recovers funds... ...and Support to turn customer feedback into improvements, while focusing on reliability, testing, and observability across the stack. #J-18808-Ljbffr Otter.aiSenior
$105.4k - $207.8k
...power the next generation of AI-enabled cyber defense?If yes,... ...looking for a hands-on Data Engineer to build and operate the governed... ...2/31/2026.Work you'll doAs a Senior Consultant, Strategy, Growth... ..., and connectors, with observability and runbooks for production support...SeniorLocal areaVisa sponsorship$210.6k - $305.1k
...even the ones they don’t own. Powered by AI and an unmatched set of cloud, internet... ...Networking, Security, Collaboration, and Observability portfolios Your Impact As part of this... ...You have led a distributed team of 5+ engineers, can demonstrate strong technical vision...SeniorFull timeTemporary workLocal areaFlexible hours- ...usher in this new era, we seek AI-native thinkers across... ...work gets done.Staff Software Engineer - External Observability PlatformLocation: Bellevue,... ..., and skilled at mentoring senior engineers while holding... ...high-performance computing (HPC) or handling global installations...Work at office3 days per week
- NICE is seeking a Senior AI Software Engineer to shape infrastructure automation and AI-driven operations across a global cloud platform. You... ...ll work with Terraform/CDK/Pulumi, Kubernetes, and modern observability tools while collaborating with cross-functional teams on...Senior
- ...hire people in any country where we have a legal entity. Overview We're building the observability layer for AI‑assisted software development, measuring how much of an engineering org's code is actually written by AI coding agents, and attributing every line back to...SeniorWork at officeLocal area
$157.9k - $213.6k
Amazon Web Services (AWS) is seeking an experienced Senior Compliance Engineer — Environmental to drive global product compliance for hazardous substances... ...across all AWS markets, including hardware supporting AI/ML infrastructure. This role sits within the Global Trade...SeniorLocal areaFlexible hours- ...infrastructure that powers high-volume, real-time business operations across multiple systems and platforms. We are seeking a Principal AI Engineer to define architecture and technical strategy for highly scalable AI systems. The role covers leading model serving, optimizing...SeniorRemote job
$138k - $241.4k
AI agents are opening up a whole new way to build on the cloud. Developers can describe what they want and watch an agent... ...and their agents love to build.We are looking for a Senior Developer Experience Engineer to define and drive what great agent-driven development looks...SeniorFlexible hours$153k - $204k
...CoreWeave is The Essential Cloud for AI™. Built for pioneers by... ...What You'll Do: The Systems Engineering team owns the host software... ...extending that framework into HPC verification, Slurm-on-Kubernetes... .... About the role: As a Senior Software Engineer on the Systems...SeniorPermanent employmentFull timeTemporary workCasual workLive inWork at officeFlexible hours$134.5k - $265.1k
....Work you will do:As a Cyber Forward Deployed Engineer (FDE) Sr Consultant, you will drive delivery of cybersecurity and AI-enabled solutions at the intersection of client... ...goals.5+ Years experience working with senior client stakeholders to translate requirements...SeniorLocal areaVisa sponsorship- ...of e-commerce customer service is intelligent, efficient, and AI-driven. Our team is dedicated to replacing traditional human-agent... ...-Bachelor and above with majors in computer science, computer engineering, statistics, applied mathematics, data science or other related...SeniorWork experience placement
- We Are:The Global AI Infrastructure team is at the center of enabling... ...workflow orchestration with observability, identity, and policy controls... ...CUDA along with LLM inference engines (TensorRT-LLM), production... ...of 1,000+ GPU clusters for AI, HPC, and agentic AI workloads with...Full timeWork experience placementLive inWork at officeLocal area
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior AI and HPC Observability Engineer. Be the first to apply!
- ai research engineer Seattle, WA
- machine learning ai engineer Seattle, WA
- ai developer Seattle, WA
- senior ai engineer Seattle, WA
- ai engineer Seattle, WA
- ai ml engineer Seattle, WA
- ai prompt engineer Seattle, WA
- ai engineer remote Seattle, WA
- srs distribution Seattle, WA
- senior operations coordinator Seattle, WA

