Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Infrastructure Engineer - AI, Automation, Observability and Monitoring

$200k - $322k

NVIDIA

NVIDIA is looking for a skilled and motivated Senior Infrastructure Engineer to join our dynamic team. You will contribute to innovative solutions and optimize our operations using brand new technology to agentize workflows to building agents for different domains of infrastructure. This role offers an outstanding opportunity to work with groundbreaking technologies and contribute to NVIDIA's innovation. The ideal candidate will have extensive experience understanding complex engineering challenges, crafting flawless solutions, and successfully implementing end-to-end projects.What you’ll be doing:Lead the design, development, and deployment of world-class software solutions, ensuring flawless execution and adherence to industry standards.Architect and implement software systems that meet ambitious performance, scalability, and reliability requirements.Collaborate with multi-functional teams to determine project requirements, providing mentorship and encouraging a culture of inclusion and collaboration.Advocate for the use of standard methodologies in software development, such as code reviews, testing, and continuous integration, to ensure high-quality deliverables.Drive innovation by exploring and integrating new technologies and methodologies that improve our engineering capabilities.Mentor and guide junior engineers, fostering an encouraging and collaborative environment, passionate about continuous learning and growth.What we need to see:12+ years of demonstrable experience in software engineering roles, focusing on team leadership and project management.Proven track record of crafting and implementing software solutions that solve complex business problems.Expertise in programming languages such as C++, Python, and Java, with a deep understanding of software development principles.Claude and Codex - being able to write agents, skills and harnessesExperience coding integrations to tools like BigPanda, ITMP, Prometheus etc.Proven track record with automation and software development and deliveryExperience working with Fortune 500 companies across diverse industries.Master’s or Ph.D. in a relevant field (e.g., Computer Science, Software Engineering) or equivalent experience.Ways to stand out from the crowd:Proficiency in brand new technologies like AI, machine learning, and data-driven solutions.Proven ability to develop innovative and effective solutions using advanced software engineering techniques.Outstanding team leadership and mentoring skills, with a passion for encouraging a collaborative and innovative environment.NVIDIA provides competitive salaries and comprehensive benefits. Our experienced and dedicated team is expanding rapidly due to outstanding growth. If you are an engineering leader looking for ambitious challenges, apply now to be a key player in driving our organization's success.Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 200,000 USD - 322,000 USD.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until September 13, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa ClaraType: Full time

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Senior Infrastructure Engineer - AI, Automation, Observability and Monitoring in Santa Clara, CA vacancy
  • $153k - $204k

     ...Essential Cloud for AI™. Built for...  ...combines superior infrastructure performance with deep...  ...talented and experienced Senior Software Engineer to join our Network...  ...services. As a Datapath Monitoring/Observability Engineer, you will...  ...or Python for automation and scripting.... 
    Senior
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    Sunnyvale, CA
    10 days ago
  • $152k - $241.5k

    Senior Systems Software Engineer, Observability and Telemetry Platform at NVIDIA is...  ...manual work through automation, performance...  ...scale, real time monitoring, logging and alertingEngage...  ...experience with Infrastructure automation,...  ...vacancy. NVIDIA uses AI tools in its recruiting... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    5 hours ago
  • $184k - $287.5k

     ...every GeForce NOW engineer to deliver...  ...StackStorm event-driven automation, HashiCorp Vault...  ...heals, you build the infrastructure behind it.What...  ...CI/CD, observability, and automation systems...  ...and VaultImplement monitoring solutions including...  ...use of AI-assisted tools (Claude... 
    Senior
    Full time
    Local area

    Nvidia

    Santa Clara, CA
    3 days ago
  •  ...Senior Platform Engineer, Observability and AIOps Synopsys is the leader in...  ...to rapidly innovate AI-powered products....  ...improve how complex infrastructure is observed, understood...  ...for observability, automation, and operational...  ...that enhance monitoring, alerting, incident... 
    Senior

    Synopsys

    Sunnyvale, CA
    11 hours ago
  • $182k - $242k

     ...Essential Cloud for AI™. Built for pioneers...  ...CoreWeave combines superior infrastructure performance with deep...  ...We're looking for a Senior Engineer to be a driving force...  ...insights and observability systems that make our...  ...OpenTelemetry) used to monitor and diagnose infrastructure... 
    Senior
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    Sunnyvale, CA
    11 days ago
  •  ...scale high-performance monitoring platforms for RDMA,...  ..., and TCP/IP infrastructure. Build backend telemetry services, observability dashboards, alerts,...  ...functionally, improve engineering practices, automate operational...  ...experience includes AI/ML, HPC, storage, GPU... 
    Senior
    Full time

    Clockwork.io

    Palo Alto, CA
    3 days ago
  • $141.3k - $360.7k

     ...technology and engineering, keeping the customer...  ...-oriented Senior Software Engineer...  ...pipelines — the automated test frameworks,...  ...workflows, and infrastructure as code that let...  ...work with agentic AI as your default...  ...Establish evaluation, observability, and monitoring for the signals... 
    Senior
    Work at office
    Local area
    Remote work
    Monday to Thursday
    Flexible hours

    Roku

    San Jose, CA
    1 day ago
  •  ...We are seeking a Senior Network Engineer to support,...  ...enterprise network infrastructure. This role owns...  ...capture and network monitoring tools to...  ...experience with network automation frameworks and...  ...Experience with AIOps or AI-assisted network monitoring, observability, and analytics... 
    Senior
    Remote work

    Advantest

    San Jose, CA
    4 days ago
  •  ...discovery to powering AI and the...  ...We are seeking a Senior Network Engineer to join the AMD...  ..., optimization, automation, and production...  ...backend network infrastructure for GPU clusters...  ...platforms, and observability systems.You will...  ...production deployment, monitoring, incident... 
    Senior

    AMD

    San Jose, CA
    2 days ago
  • $184k - $287.5k

     ...are looking for a Senior System Software Engineer, Software Defined Networking...  ...for NVIDIA's AI Clouds hosting GPU-...  ..., CI/CD, observability, and incident response...  ...OVN, OpenFlow)Build Infrastructure-as-a-Service virtual...  ...network observability — monitoring, telemetry,... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $224k - $356.5k

     ...are looking for a Senior System Software Engineer who sees the big picture...  ...methodology.Drive automation, monitoring, and performance...  ...of cloud infrastructure and distributed system...  ...tolerance, scalability, observability).Demonstrated...  ...initiatives.Exposure to AI-assisted... 
    Senior
    Full time
    Local area

    Nvidia

    Santa Clara, CA
    2 days ago
  • $155.42k - $205.9k

     ...DescriptionAbout the Team: The AI Validation Platform team...  ...proud to serve as the infrastructure platform for teams...  ...Role: We are seeking a Senior ML Infrastructure engineer to help build and scale...  ...Drive the development of monitoring, observability, and metrics to ensure reliability... 
    Senior
    Full time
    Local area
    Remote work
    Work from home
    Relocation
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    6 hours ago
  • $176k - $276k

    Production engineering is a field that involves...  ...for HPC and AI/ML workloads.Storage...  ...mindset focused on automating storage...  ...storage performance monitoring, automated fault...  ...production storage infrastructure by supervising availability...  ...Experience with observability and tracing tools... 
    Senior
    Full time
    Flexible hours

    Nvidia

    Santa Clara, CA
    1 day ago
  • $141.3k - $360.7k

     ...Design and develop automated test frameworks for...  ...services. Use agentic AI workflows to author...  ...CI/CD workflows and infrastructure as code using...  ...Establish evaluation, observability, and monitoring for test reliability...  ...with development, data engineering, and product teams.... 
    Senior
    Full time
    Work at office
    Local area
    Remote work
    Monday to Friday
    Monday to Thursday
    Flexible hours

    Roku

    San Jose, CA
    16 days ago
  •  ...world's largest AI chip, 56 times larger...  ...the Wafer-Scale Engine (WSE). This team...  ...inference infrastructure for leading model...  ...pipelines, shared observability common tooling. This...  ...their pain points, automate their toil, and mentor...  ...SLOs, or drift monitoring.Prior work on... 
    Shift work

    Cerebras Systems

    Sunnyvale, CA
    1 day ago
  • $146.7k - $339.3k

     ...leads global network infrastructure and data center...  ...0,000 employees and AI/ML training infrastructure...  ...distributed network engineers and data center...  ...global team we design, automate, and optimize IT infrastructure...  ...network performance monitoring and observability platforms. Designing... 
    Senior
    Full time
    Work at office
    Remote work

    Zoom

    San Jose, CA
    6 hours ago
  • $140k - $224.25k

     ...Cloud provides the infrastructure and software...  ...train, and deploy AI at scale. As...  ...are looking for a Senior Software Engineer to design and build...  ...software and automated decision-making....  ...data.Establish monitoring, data-quality controls...  ...integration, observability, and production... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    5 hours ago
  •  ...foundation for physical AI — a unified...  ...looking for a Senior AI Engineer to design, build...  ...the agentic infrastructure that powers our...  ...and production monitoring. You understand...  ...feedback loops, observability, and lifecycle control...  ...delivery practices: automated testing (unit,... 
    Senior
    Full time

    Dexmate

    Santa Clara, CA
    1 day ago
  • $184k - $287.5k

    We are seeking a Senior DevOps / Cloud Simulation Infrastructure Engineer to own the complete...  ...to high-scale, automated cloud validation...  ...structural validation, AI-driven runtime...  ...to quality gates.Observability & CI/CD:...  ...verification, and monitoring pipelines. Implement... 
    Senior
    Full time
    Local area

    Nvidia

    Santa Clara, CA
    3 days ago
  •  ...world's largest AI chip, 56 times larger...  ...the TeamThe Core Infrastructure team builds the...  ...that power engineering workflows across...  ...beyond scripting or automation. Our frameworks act...  ...dependencies, and observability across large and...  ...cost efficiency, monitoring, and operational... 
    Senior

    Cerebras Systems

    Sunnyvale, CA
    1 day ago
  • $147k - $237.5k

     ...Python, etc) Solid understanding of infrastructure and cloud environments (...  ...willingness and ability to leverage AI tools and agents to improve productivity...  ...activities Experience with observability frameworks and operational monitoring for cloud-native services (... 
    Senior

    Jobleads-US

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

     ...NVIDIA's DGX Cloud AI Efficiency Team means...  ...to the infrastructure that powers our innovative...  ...infrastructure software engineer to join our team....  ...of AI systems.As a senior DGX Cloud AI Infrastructure...  ....Experience with observability platforms for monitoring and logging (e.g.,... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    1 day ago
  • $262k - $364k

     ...reproducible evaluation infrastructure and automated test harnesses...  ...patterns.Build the metrics engines and dashboards to...  ..., ML platforms, or observability infrastructure,...  ...distributed tracing, semantic monitoring, and telemetry...  ...most comprehensive AI platform in the industry... 
    Senior

    Google

    Sunnyvale, CA
    1 day ago
  • $184k - $287.5k

     ...compute farm, and we need an automation engineer to own it end to end. You...  ...equivalent experience.6+ years in infrastructure engineering with strong...  ...resemble production.Observability instincts — you instrument...  ...existing vacancy. NVIDIA uses AI tools in its recruiting processes... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $180k - $240k

     ...role We are seeking a Senior AI Infrastructure Engineer to design, build, and...  ...Agentic Infrastructure & Automation Self-Healing AI Infrastructure...  ..., CrewAI, or AutoGen) to monitor GPU cluster health,...  ...Lake Monitoring & Observability System Metrics:... 
    Senior
    Odd job
    Work at office

    Gatik AI

    Santa Clara, CA
    3 days ago
  • $153.2k - $234.1k

     ...software stack through intelligent automation, AI-enabled engineering workflows, and data-driven validation...  ...such as Jira, GitHub, dashboards, observability platforms, and cloud services into...  ...establishing governance, evaluation, monitoring, and security practices for AI-... 
    Senior
    Full time
    Local area
    Work from home
    Flexible hours

    General Motors

    Sunnyvale, CA
    2 days ago
  • $178k - $321k

     ...early but working AI-native capability...  ...data and AI infrastructure everything else depends...  ...is a two-person engineering team: you deploy,...  ..., Drive-synced automation, locally scheduled...  ...code, CI/CD, and observability throughout. Treat...  ...) and continuous-monitoring data foundation:... 
    Senior

    OKX

    San Jose, CA
    3 days ago
  •  ...Cloud, is a leader in AI cloud infrastructure serving tens of thousands...  ...We are looking for a Senior Site Reliability Engineer to improve the...  ...Kubernetes, infrastructure automation, observability, deployment systems, and...  ...systems.Build monitoring, alerting, and tracing... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    1 day ago
  •  ...Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of...  ...of our network through monitoring, failover, and redundancyContribute to automation of network configuration...  ...-call rotation for Network Engineering teamYouHave 10+ years of... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    3 days ago
  • $150k - $217k

     ...in partnership with other engineering teams.Lead and improve the...  ...debug and optimize code and automate routine tasks.Systematic problem...  ...are safe and efficient by monitoring network performance,...  ...products and services.The AI and Infrastructure team is redefining what’s possible... 
    Senior
    Worldwide

    Google

    Sunnyvale, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Infrastructure Engineer - AI, Automation, Observability and Monitoring. Be the first to apply!