Senior AI and HPC Observability Engineer
$152k - $241.5kNVIDIA
NVIDIA is a pioneer in accelerated computing, known for inventing the GPU and driving breakthroughs in gaming, computer graphics, high-performance computing, and artificial intelligence. Our technology powers everything from generative AI to autonomous systems, and we continue to shape the future of computing through innovation and collaboration. Within this mission, our team, Managed AI Superclusters (MARS) builds and scales the infrastructure, platforms, and tools that enable researchers and engineers to develop the next generation of AI/ML systems. By joining us, you’ll help design solutions that power some of the world’s most advanced computing workloads.Observability is at the heart of this transformation. We are looking for a strong AI & HPC Observability Engineer to build and scale next-generation Observability and Telemetry platforms. You will design and develop high-throughput, reliable telemetry pipelines and modern data infrastructure. This role requires solid distributed systems fundamentals, production-grade coding, and a passion for operational excellence.What You Will Be Doing:Design and scale observability platforms handling high-volume metrics, logs, and traces across distributed environmentsBuild high-performance backend services for telemetry ingestion, processing, and routingDevelop and extend OpenTelemetry collectors, processors, exporters, and instrumentation librariesBuild and optimize metrics pipelines using large-scale time-series storage systemsDesign and operate real-time and batch telemetry pipelines using streaming and distributed data technologiesImprove platform reliability, performance, and cost efficiency through tuning, capacity planning, and system optimizationDevelop monitoring, alerting, and service reliability frameworks to ensure platform health and performanceCollaborate with platform engineering, infrastructure, and site reliability teams to deliver production-grade observability solutionsWhat We Need to see:Bachelor’s degree in Computer Science, Computer Engineering, or related field or equivalent experience5+ years of experience building backend or distributed systems in production environmentsStrong programming skills in Python, Go, or Java, with experience developing production-quality softwareHands-on experience with modern observability architectures, including metrics, logs, and tracesSolid experience with PromQL and time-series data systemsExperience building or operating distributed data pipelines using technologies such as Kafka, Spark, or FlinkExperience working with Kubernetes and cloud-native infrastructureStrong understanding of distributed systems, concurrency, and fault-tolerant system design. Strong debugging, performance tuning, and production operations skillsWays To Stand Out from The Crowd:Proven experience designing and scaling observability platforms for AI, GPU, or HPC environmentsHands-on expertise with OpenTelemetry, Prometheus, Kafka, and high-volume distributed telemetry pipelinesStrong background in data engineering, time-series data modeling, and real-time performance tuningExperience integrating observability with AI/ML pipelines, GPU workload monitoring, or intelligent alertingDemonstrated use of statistical or machine learning techniques for anomaly detection, correlation, or predictive insightsYour base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 152,000 USD - 241,500 USD for Level 3, and 184,000 USD - 287,500 USD for Level 4.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until March 6, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa Clara; US, WA, SeattleType: Full time
$146k - $194k
...powered by Lattice OS, an AI-powered operating... ...About the TeamThe Thermal Engineering team's work is essential... ...are looking for a Senior Thermal Engineer in Costa... ...management of expeditionary AI/HPC data centers—air-... ...our candidates. We've observed a rise in sophisticated...SeniorFull timeWork experience placementImmediate start$100k - $150k
...is the vertically integrated AI cloud engineered for AI. We own and operate... ...the Role (Job Purpose) Senior Infrastructure Support Engineers... ...3+ years hands-on with GPU, HPC, or large-scale data centre... ...from north-south. ~ Observability and incident response....SeniorFull timeRemote workFlexible hours$147k - $202k
...Secure Every Identity, from AI to Human Identity is the key... ...talk. The Auth0 Platform Observability team owns the observability tooling... ...looking for an Observability Engineer to help ensure that our... ...development in these areas. As a Senior Engineer on this team, you...SeniorLocal areaFlexible hours$185k - $210k
...CoreWeave is the AI Hyperscaler™, delivering a cloud platform of cutting edge... ...innovation. About the role The Observability department plays a pivotal role in... ...Intelligence. We are seeking senior observability engineers with specializations in logging and tracing...SeniorTemporary workCasual workWork at officeRemote workFlexible hours$142k - $220.5k
Job DescriptionA Senior Engineer on the AI Enablement team designs and delivers AI platform capabilities that let Nordstrom Technology use Generative... ...with developer tools, cloud services, LLM providers, and observability systems.Own systems and designs that span multiple...SeniorFull timeTemporary work$170k - $220k
...a hands-on, high-agency Site Reliability Engineer to help shape and scale the reliability layer... .... Build safe, repeatable, and observable workflows.GitHub Operations: Manage GitHub... ...and improve deployment confidence. Using AI tools to generate code is encouraged and...Senior$200k - $275k
Secure Every Identity, from AI to HumanIdentity is the key to... ...identity, platform, and security engineering teams, write production code... ...on the patterns.Engage senior leadership. Brief the CISO, CIO... ...runtime decisions.Build evals and observability. Authorization decision...SeniorLocal areaWorldwideFlexible hoursShift work$126.2k - $264.1k
What You’ll DoAs a Senior Principal Network Reliability Engineer, you will:Provide technical leadership for the deployment... ...initiatives supporting AI infrastructure, new OCI Regions, backbone... ..., operational readiness, observability, capacity planning, and lifecycle...SeniorTemporary workFlexible hours$144.8k - $261.45k
...security problems that don't scale and engineer them away.As a Senior Security Automation Engineer, you will build automation and AI-assisted frameworks to scale Adobe's web... ...bar through architecture, testing, CI/CD, observability, documentation, and mentoring....SeniorFull timeTemporary workLocal areaWorldwide$127k - $249k
The TeamPlatform Engineering is the department within SRE that is responsible for a range of... ...facing edge and internal service mesh), and observability and alerting systems.The Fleet... ...have redefined the data platform for the AI era, enabling builders to create, transform...SeniorWork at officeLocal areaRemote workWorldwideFlexible hours$191k - $253k
...is powered by Lattice OS, an AI-powered operating system that... ...TEAM: The Corporate Technology Engineering team is responsible for... ...re seeking a highly motivated Senior Software Engineer to join a fast... ...security of our candidates. We've observed a rise in sophisticated...SeniorFull timeWork experience placementImmediate startRotating shift$145k - $193.75k
...are uniting human teams with AI agents. By orchestrating the... ...every day.Corporate Systems Engineering builds and operates the software... ...auditable automation.As a Senior Software Engineer I (Automation... ...human reviewersExpertise in observability platforms (e.g., Datadog, CloudWatch...SeniorFull timeTemporary workWork at officeLocal areaRemote work$134.25k - $214.8k
...company where you matter.Your ImpactAre you an engineer who gets excited about the challenge of making complex distributed systems observable — not just instrumenting them, but... ...infrastructure engineeringExperience with agentic AI tooling or building LLM-powered developer...SeniorWork experience placementWork at officeRemote work$159.2k - $301.6k
The Opportunity We are looking for a Senior AI Systems Engineer with deep C++ expertise to help build the next generation of AI-enabled product... ...and how to design systems that make AI useful, reliable, observable, and secure. What You’ll Do AI-Native Systems Architecture...SeniorFull timeTemporary workLocal areaRemote workWorldwide$232k - $319k
Secure Every Identity, from AI to HumanIdentity is the key to unlocking the potential... ...on Edge networking, K8s platform, Observability, automation platform & tooling. What you... ...serviceAccelerate the velocity of SRE and product engineering by developing robust platforms, powerful...SeniorPermanent employmentLocal areaWorldwideFlexible hours$182k - $242k
...Description CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers,... ...closely with the internal and customer engineering teams, offering valuable insights from the... ...focusing on building distributed systems or HPC/cloud services, with an expertise focused...SeniorPermanent employmentFull timeTemporary workCasual workWork at officeFlexible hours$191k - $253k
...of systems is powered by Lattice OS, an AI-powered operating system that turns thousands... ...the board. ABOUT THE JOB As an engineer on the Robotics Data Foundation team, you... ...the security of our candidates. We've observed a rise in sophisticated phishing and fraudulent...SeniorFull timeWork experience placementImmediate start$188k - $275k
...Description CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers,... ...at What You'll Do: The Field Engineering organization at CoreWeave is dedicated to... ...driving InfiniBand/RoCE fabric validation and HPC performance benchmarking, defining how we...SeniorPermanent employmentFull timeContract workTemporary workCasual workWork at officeFlexible hours- ...most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating... ...That is magic at work, and it’s what we show up for every day. Observability Engineering owns the breadth and depth of observability and telemetry at...Full timeWork at officeRemote work
- ...AZX in Seattle is seeking a Software Engineer to build models and implement an agentic, formally-grounded solution that addresses client issues with strict compliance and tariffs. You'll bridge research and production, packaging, testing, CI, and documentation, delivering...Senior
$157.9k - $213.6k
Amazon Web Services (AWS) is seeking an experienced Senior Compliance Engineer — Environmental to drive global product compliance for hazardous substances... ...across all AWS markets, including hardware supporting AI/ML infrastructure. This role sits within the Global Trade...SeniorLocal areaFlexible hours$113.4k - $221.75k
The OpportunityAdobe is reimagining the future of creativity, where generative AI and human imagination work together. As a Quality Engineer on the Agentic Harness team, you’ll help the harness behave the way creators expect. You’ll verify that the agent’s behavior, tool...SeniorFull timeTemporary workLocal areaWorldwide- ...of e-commerce customer service is intelligent, efficient, and AI-driven. Our team is dedicated to replacing traditional human-agent... ...-Bachelor and above with majors in computer science, computer engineering, statistics, applied mathematics, data science or other related...SeniorWork experience placement
- We Are:The Global AI Infrastructure team is at the center of enabling... ...workflow orchestration with observability, identity, and policy controls... ...CUDA along with LLM inference engines (TensorRT-LLM), production... ...of 1,000+ GPU clusters for AI, HPC, and agentic AI workloads with...Full timeWork experience placementLive inWork at officeLocal area
- ...A leading data and AI infrastructure provider is seeking a Senior Staff Technical Program Manager for Reliability to enhance the reliability and performance... ...leading programs in partnership with senior engineering leaders, requiring over 10 years of experience in cloud...Senior
$134.5k - $265.1k
....Work you will do:As a Cyber Forward Deployed Engineer (FDE) Sr Consultant, you will drive delivery of cybersecurity and AI-enabled solutions at the intersection of client... ...goals.5+ Years experience working with senior client stakeholders to translate requirements...SeniorLocal areaVisa sponsorship- ...resisted innovation. Join us!About the RoleWe’re looking for a Senior Design Systems Engineer II to shape and scale the visual and interaction... ...JavaScript, and modern front-end tooling.Expert use of modern AI-assisted tools (e.g., Cursor, GitHub Copilot, ChatGPT) Proficiency...Senior
$133.5k - $194k
...enabling rapid campaign and landing page deploymentMaintain strong engineering standards while supporting marketing speedEstablish performance... ...SEO best practicesAI-Enabled Marketing InnovationIntegrate AI agents and LLM APIs to power:Dynamic content generationPersonalized...Senior- RDQ126R35At Databricks, observability and governance are what turn a massive, multi-tenant data and AI platform into one customers can... ...scale. We are looking for a senior technical leader to drive strategy... ...these surfaces, raising the engineering bar of the combined team,...SeniorWorldwide
$192k - $240k
...gain real-time visibility, and control spend effortlessly. Brex’s AI-native automation and world-class service eliminate manual... ...the tools, resources, and support you need to grow your career.Engineering at BrexEngineering at Brex is about building systems that scale...SeniorWork at officeRemote workWork from home
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior AI and HPC Observability Engineer. Be the first to apply!
- ai engineer Seattle, WA
- ai engineer remote Seattle, WA
- ai prompt engineer Seattle, WA
- ai developer Seattle, WA
- senior ai engineer Seattle, WA
- machine learning ai engineer Seattle, WA
- ai ml engineer Seattle, WA
- senior operations technician Seattle, WA
- senior operations associate Seattle, WA
- senior cloud service delivery manager Seattle, WA



