Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Software Engineer - AI Infrastructure Performance Insights & Observability

$182k - $242k

CoreWeave

Job Description

Job Description

CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability. Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025. Learn more at

About this role

We're looking for a Senior Engineer to be a driving force on CoreWeave's Benchmarking & Performance team, with a focus on building the performance insights and observability systems that make our AI infrastructure legible at every layer, from individual GPUs and NVLink/InfiniBand fabrics up through distributed training and inference workloads. You will own how we detect, diagnose, and surface performance signals across every data center in our global infrastructure, turning billions of raw telemetry events into real-time insight that engineers, product teams, and executives can act on with confidence.

This is not a straightforward data engineering or BI role. The role will focus on building the observability and insight tooling that lets us answer, in near real time, whether a GPU fleet, a fabric, or a training run is performing the way it should, and why it isn't when it's not. If you're energized by building the systems that turn raw infrastructure telemetry into trusted, actionable performance intelligence, and you want that work to sit closer to the hardware and the workload than to a dashboard, this role was built for you.

What you'll do

  • Performance Insights & Observability - Design and build the systems that continuously assess AI infrastructure health and performance: GPU utilization and efficiency, interconnect (NVLink, InfiniBand, RoCE) fabric behavior, distributed training and inference throughput, and hardware degradation signals. Build the detection and diagnosis logic that surfaces anomalies and regressions before they become incidents, not just dashboards that report on them after the fact.
  • Time-Series & Metrics Infrastructure - Own and extend our time-series database (TSDB) layer as the backbone of real-time observability. Write and optimize PromQL/MetricsQL queries that power alerting, anomaly detection, and trend analysis across thousands of GPUs and hundreds of benchmark runs. Bridge streaming metrics and batch-analytical workloads so engineers get sub-second answers during live incidents and analysts get complete historical context for root cause work.
  • Fabric & GPU Telemetry - Build and validate the pipelines and metrics that make network fabric and GPU-level behavior observable and comparable across racks, clusters, and hardware generations, including gray failure detection, congestion and error-rate signals, and health scoring that holds up under audit.
  • Data Lake Architecture (in support of insight work) - Design and build the performance data lake that underpins the above: table formats (Apache Iceberg, Parquet, Avro), hot/cold tiering, and schema evolution for latency distributions, throughput metrics, GPU utilization, cost-per-token, and hardware health signals. This is foundational infrastructure, not the end product.
  • Query Optimization & Performance - Profile and tune query engines against columnar and time-series stores so that the observability layer meets its own strict P99 latency and freshness SLAs. Benchmark the benchmarking infrastructure itself.
  • BI & Reporting (secondary) - Where needed, build self-service views (Grafana, Looker, or similar) for engineers, product managers, and executives, but as a downstream output of the insight and observability work above, not the primary deliverable.

Who you are

  • 5+ years of experience building distributed systems, observability platforms, or performance engineering tooling, ideally for infrastructure or ML systems rather than general-purpose BI.
  • Strong coding in Python or Go (C++ a plus) and deep familiarity with networked systems, GPU infrastructure, and performance analysis.
  • Hands-on experience with Kubernetes at production scale, CI/CD, and observability stacks (Prometheus, Grafana, OpenTelemetry) used to monitor and diagnose infrastructure, not just report on it.
  • Working knowledge of time-series databases and fluency in PromQL or MetricsQL for building real-time alerting and anomaly detection, not only historical dashboards.
  • Familiarity with data lake architectures and modern table formats (Iceberg, Parquet, Avro) sufficient to support an insights platform, though this is not the primary skill this role is hiring for.
  • Comfortable working close to hardware and workload behavior: GPU utilization patterns, interconnect fabric health, distributed training/inference performance characteristics.
  • Strong communicator comfortable collaborating with cross-functional teams and external partners.

Nice to have

  • Experience with time-series databases, LSM-based storage engines, or custom telemetry pipelines.
  • Experience running MLPerf submissions or similar large-scale audited benchmarks.
  • Contributions to OSS projects such as Apache Iceberg, Apache Spark, Trino, llm-d, vLLM, or PyTorch.
  • Direct experience benchmarking or monitoring large GPU fleets or multi-region clusters.
  • Experience with CUDA kernels, NCCL/SHARP, RDMA/NUMA, or GPU interconnect topologies.
  • Familiarity with data cataloging, lineage tools, or data governance frameworks.

Why CoreWeave?

Own the data that tells the truth about performance. You'll build the analytical foundations the industry relies on to measure—and achieve—state-of-the-art results, working hand-in-hand with world-class partners and communities. If you love turning vast streams of telemetry into clear, trusted insights that move markets and drive engineering excellence, we'd love to talk.

At CoreWeave, we work hard, have fun, and move fast! We're in an exciting stage of hyper-growth that you will not want to miss out on. We're not afraid of a little chaos, and we're constantly learning. Our team cares deeply about how we build our product and how we work together, which is represented through our core values:

  • Be Curious at Your Core
  • Act Like an Owner
  • Empower Employees
  • Deliver Best-in-Class Client Experiences
  • Achieve More Together

We support and encourage an entrepreneurial outlook and independent thinking. We foster an environment that encourages collaboration and provides the opportunity to develop innovative solutions to complex problems. As we get set for take off, the growth opportunities within the organization are constantly expanding. You will be surrounded by some of the best talent in the industry, who will want to learn from you, too. Come join us!

The base salary range for this role is $182,000 to $242,000. The starting salary will be determined based on job-related knowledge, skills, experience, and market location. We strive for both market alignment and internal equity when determining compensation. In addition to base salary, our total rewards package includes a discretionary bonus, equity awards, and a comprehensive benefits program (all based on eligibility). 

What We Offer

The range we've posted represents the typical compensation range for this role. To determine actual compensation, we review the market rate for each candidate which can include a variety of factors. These include qualifications, experience, interview performance, and location.

In addition to a competitive salary, we offer a variety of benefits to support your needs. The benefits below reflect our US-based offerings for full-time employees; for roles in other locations, benefits vary and are shared during the hiring process. These include:

  • Medical, dental, and vision insurance - 100% paid for by CoreWeave
  • Company-paid Life Insurance 
  • Voluntary supplemental life insurance 
  • Short and long-term disability insurance 
  • Flexible Spending Account
  • Health Savings Account
  • Tuition Reimbursement 
  • Ability to Participate in Employee Stock Purchase Program (ESPP)
  • Mental Wellness Benefits through Spring Health 
  • Family-Forming support provided by Carrot
  • Paid Parental Leave 
  • Flexible, full-service childcare support with Kinside
  • 401(k) with a generous employer match
  • Flexible PTO
  • Catered lunch each day in our office and data center locations
  • A casual work environment
  • A work culture focused on innovative disruption

California Applicants

California Consumer Privacy Act 

Equal Opportunity & Accommodations

CoreWeave is an equal opportunity employer, committed to fostering an inclusive and supportive workplace. All qualified applicants and candidates will receive consideration for employment without regard to race, color, religion, sex, disability, age, sexual orientation, gender identity, national origin, veteran status, or genetic information.

As part of this commitment and consistent with the Americans with Disabilities Act (ADA) , CoreWeave will ensure that qualified applicants and candidates with disabilities are provided reasonable accommodations for the hiring process, unless such accommodation would cause an undue hardship. If reasonable accommodation is needed, please contact: View email address on us.fitly.work.

Export Control Compliance

This position requires access to export controlled information. To conform to U.S. Government export regulations applicable to that information, applicant must either be (A) a U.S. person, defined as a (i) U.S. citizen or national, (ii) U.S. lawful permanent resident (green card holder), (iii) refugee under 8 U.S.C. § 1157, or (iv) asylee under 8 U.S.C. § 1158, (B) eligible to access the export controlled information without a required export authorization, or (C) eligible and reasonably likely to obtain the required export authorization from the applicable U.S. government agency. CoreWeave may, for legitimate business reasons, decline to pursue any export licensing process.

Vacancy posted a month ago
Similar jobs that could be interesting for youBased on the Senior Software Engineer - AI Infrastructure Performance Insights & Observability in Sunnyvale, CA vacancy
  • $182k - $242k

     ...The Essential Cloud for AI™. Built for pioneers by...  ...CoreWeave combines superior infrastructure performance with deep technical...  ...We're looking for a Senior Engineer to be a driving force on...  ...building the performance insights and observability systems that make our AI... 
    Senior
    Performance
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    Sunnyvale, CA
    17 days ago
  • $180k - $215k

     ...intelligence (AI) powered technology...  ...self-driving software to develop,...  ...looking for a Senior Software Engineer to build infrastructure, tools, and...  ...quality, system performance, and test coverage...  ...actionable insights from logs,...  ...and Planning observability. Optimize performance... 
    Senior
    Performance
    Visa sponsorship

    Kodiak

    Mountain View, CA
    a month ago
  • $153k - $204k

     ...The Essential Cloud for AI™. Built for pioneers by pioneers...  ...combines superior infrastructure performance with deep technical expertise...  ...talented and experienced Senior Software Engineer to join our Network...  ...As a Datapath Monitoring/Observability Engineer, you will focus... 
    Senior
    Performance
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    Sunnyvale, CA
    a month ago
  • $183k - $275k

     ...opportunity for AI to drive...  ...of the software and hardware...  ...layer. This performance simulation...  ...will own the infrastructure that makes...  ...much more. Engineers across the...  ...& Observability: Build monitoring...  ...actionable insights through dashboards...  ...ve briefed senior engineering... 
    Senior
    Performance
    Temporary work
    Immediate start
    Flexible hours

    Nuro

    Mountain View, CA
    15 days ago
  • $153k - $204k

     ...Essential Cloud for AI™. Built for...  ...combines superior infrastructure performance with deep technical...  ...to deliver reliable insights and data that security...  ...About the Role As a Senior Software Engineer on the Security Infrastructure...  ...shape security observability and detection... 
    Senior
    Performance
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    Sunnyvale, CA
    9 days ago
  • $174k - $252k

     ...maintaining, or launching software products, and 1...  ...experience with performance, large scale...  ...'s software engineers develop the next-...  ...by the Technical Infrastructure team to keep it running...  ...possible.The AI and Infrastructure...  ...capabilities and insights by delivering AI... 
    Senior
    Performance
    Worldwide

    Google

    Sunnyvale, CA
    3 hours ago
  • $140k - $210k

     ...Clockwork.io – Software Driven Fabrics...  ...systems engineers who share a vision...  ...computing. As AI workloads grow...  ..., traditional infrastructure struggles to meet...  ...demands of performance, reliability,...  ...cross-stack observability to catch and...  ...experienced Senior Software Engineer... 
    Senior
    Performance

    Clockwork.io

    Palo Alto, CA
    18 days ago
  • $175.6k - $263.4k

     ...Job Area: Engineering Group, Engineering...  ...'s CAD Infrastructure team develops...  ...across CPU, DSP, AI, SoC, Physical...  ...Verification, and Software organizations....  ...are seeking a Senior Full Stack...  ...Troubleshoot performance, scalability,...  ...Kibana, Splunk, or observability platforms.... 
    Senior
    Performance
    Work experience placement
    Work from home

    Qualcomm

    Santa Clara, CA
    10 days ago
  •  ...Synopsys is the leader in engineering solutions from...  ...to rapidly innovate AI-powered products. We...  ...improve how complex infrastructure is observed, understood, and operated...  ...working across software, systems, and operations...  ..., reliable, high-performance systems and services... 
    Senior
    Performance

    Synopsys Inc

    Sunnyvale, CA
    23 days ago
  • $193.93k - $291.15k

     ...profound opportunity for AI to drive positive...  ...driving, and the ML Infrastructure team builds and operates...  ..., orchestration, observability, and cost management...  ...Science, Electrical Engineering, or a closely related...  ...ability to reason about performance, failure modes, and reliability... 
    Senior
    Performance
    Work experience placement
    Immediate start
    Flexible hours

    Nuro

    Mountain View, CA
    24 days ago
  • $164.2k - $205.2k

     ...the company. Building unified infrastructure on top of heterogeneous...  ...companies experience. As an engineer on this team, you won’t just...  ...instance qualification for AI training and stateful execution...  ...include eligibility for annual performance bonus, equity, and the... 
    Senior
    Performance
    Local area
    Worldwide

    Databricks

    Mountain View, CA
    20 days ago
  • $198k - $326k

     ...gain valuable insights every day....  ...it will be performed both from...  ...LinkedIn's AI model training...  ..., feature engineering and serving...  ..., compute software, and...  ...Model Training Infrastructure: As an engineer...  ...and observability, and updating...  ...level: Mid-Senior LevelIndustry... 
    Senior
    Performance
    For contractors
    Work at office
    Flexible hours

    Linkedin

    Sunnyvale, CA
    3 hours ago
  • $184k - $287.5k

     ...forefront of the generative AI revolution, building the software and systems that power...  ...We are looking for a Senior Software Engineer to lead the bring-up,...  .... You will lead deep performance and reliability...  ...large-scale AI clusters, infrastructure, and end-to-end workloads... 
    Senior
    Performance
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    1 day ago
  • $140k - $224.25k

     ...Cloud provides the infrastructure and software platform that...  ...train, and deploy AI at scale. As demand...  ...are looking for a Senior Software Engineer to design and build...  ...the reliability, performance, and scalability of...  ...system integration, observability, and production operations... 
    Senior
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

     ...unlimited potential of AI to define the next...  ...is looking for a Senior GPU Platforms Engineer to be part of our innovative...  ...and platform software, driving the next era...  ...and deliver results.Observability stack expert: Design and develop high-performance, distributed observability... 
    Senior
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    21 hours ago
  • $184k - $287.5k

     ...NVIDIA's DGX Cloud AI Efficiency Team...  ...to the infrastructure that powers our...  ...infrastructure software engineer to join our team...  ...AI systems.As a senior DGX Cloud AI Infrastructure...  ...Experience with observability platforms for...  ...Intelligence, High-Performance Computing, and... 
    Senior
    Performance
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $262k - $364k

     ...distributed team of engineers.Facilitate...  ...enhance large-scale software solutions.Build next...  ...generative AI tools or LLM interfaces...  ....Experience with Performance optimization of...  ...engineering is a critical infrastructure service for...  ...capabilities and insights by delivering AI... 
    Senior
    Performance
    Worldwide

    Google

    Sunnyvale, CA
    2 days ago
  •  ...is the leader in engineering solutions from silicon...  ...rapidly innovate AI-powered products....  ...found them—more observable, more resilient,...  ...implement hybrid infrastructure solutions...  ...clear operational insight.  The Impact...  ...without sacrificing performance.  Increase engineering... 
    Senior
    Performance
    Work at office
    Relocation

    Synopsys Inc

    Sunnyvale, CA
    a month ago
  • $193.93k - $291.15k

     ...profound opportunity for AI to drive positive change...  ...generalists where ML and systems engineering converge to push autonomy performance forward. As a Senior Perception ML Data Infrastructure Engineer, you will own...  ...in Computer Science , Software Engineering, Robotics, or... 
    Senior
    Performance
    Immediate start
    Flexible hours

    Nuro

    Mountain View, CA
    a month ago
  •  ...we're building the AI-native social...  ...-Meta product and engineering leaders, we've raised...  ...'re looking for a Senior Software Engineer, Infrastructure to own the reliability...  ..., and build the observability, alerting, and on-...  ...Improve performance, cost efficiency,... 
    Senior
    Performance
    Work at office
    Remote work
    Flexible hours
    Shift work

    Nectar Social

    Palo Alto, CA
    2 days ago
  • $119.8k - $234.7k

     ...than 25%Profession: Software...  ...in a cloud-first, AI-powered world. This...  ...integrating AI-driven insights and automation to...  ...Microsoft Entra. Develop performant, reliable, and...  ...scenarios. Drive engineering excellence through...  ...tests, and ensure observability and reliability for... 
    Senior
    Performance
    Ongoing contract
    Local area
    Remote work
    Flexible hours
    3 days per week

    Microsoft

    Mountain View, CA
    21 hours ago
  •  ...Description We’re seeking a Senior Software Engineer to join our Infrastructure team within Platform....  ...availability and performance. • Technical...  ...applications. • Experience with observability platforms such as...  ...API partners. We use AI-assisted development tools... 
    Senior
    Performance

    James Consultancy Services , LLC

    Mountain View, CA
    3 days ago
  • $152k - $241.5k

     ...and operates large-scale GPU infrastructure for AI workloads. We are looking for Software Engineers with SRE or Production...  ...duties, incident response, observability, and durable solutions.Ability...  ...Spectrum-X, or GPU cluster performance validation.Experience building... 
    Senior
    Performance
    Permanent employment
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $300 per month

     ...vertically integrated AI infrastructure company built from...  ...and be part of a high-performing team that believes...  ...Center Infrastructure Engineering (DCIE) team is...  ...deployment maintenance, observability, critical environment...  ...and motivated Software Engineer to join Crusoe... 
    Senior
    Performance
    Temporary work

    Crusoe

    Sunnyvale, CA
    23 days ago
  • $139k - $242k

     ...Essential Cloud for AI™. Built for pioneers...  ...CoreWeave combines superior infrastructure performance with deep technical...  ..., high-reliability engine of fleet management....  ...engineering, building the software that keeps our global...  ...the Role:  As a Senior Software Engineer on... 
    Senior
    Performance
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    Sunnyvale, CA
    a month ago
  • $152k - $287.5k

     ...into the unlimited potential of AI to define the next era of...  ...the world. NVIDIA seeks a driven Software Engineer to advance our Kubernetes and AI Observability. What you will be doing:...  ...development. ~ Understanding of performance, security, and reliability in complex... 
    Senior
    Performance
    Full time
    Work experience placement

    NVIDIA

    Santa Clara, CA
    3 days ago
  • $92.5k - $209.5k

     ...Oracle Cloud Infrastructure (OCI) delivers...  ...edge, data, and AI capabilities....  ...are seeking a Software Developer 3 to...  ...with experienced engineers, product...  ...participate in performance and resilience...  ...improve service observability. Help investigate...  ...managers, and senior engineers. ~... 
    Senior
    Performance
    Temporary work
    Flexible hours

    Oracle

    Santa Clara, CA
    2 days ago
  • $193.93k - $352.29k

     ...opportunity for AI to drive positive...  ...fungible is the infrastructure that decides whether...  ...Nuro's own engineering organization, under...  ...the gateway and observability layer underneath....  ...~5+ years of software engineering experience...  ...eligible for an annual performance bonus, equity,... 
    Senior
    Performance
    Immediate start
    Flexible hours

    Nuro

    Mountain View, CA
    5 days ago
  • $114.4k - $124.8k

     ...experience in Site Reliability Engineering, Systems Engineering, or Infrastructure Operations in large-...  ..., or workflows with AI agent tooling, LLM...  ...availability, optimizing performance, and managing incidents...  ...remediation Implement observability platforms using Datadog... 
    Senior
    Performance
    Hourly pay
    Full time
    Temporary work

    CYNET SYSTEMS

    Santa Clara, CA
    4 days ago
  • $125.7k - $194.8k

     ...all started when engineer Fred Luddy wrote...  ...ServiceNow is the AI control tower for...  ...member of the core infrastructure team, you will be...  ...etc. Improve observability and reliability...  ...metrics to measure performance of microservices...  ...analyzing AI-driven insights, or exploring AI’... 
    Performance
    Full time
    Work at office
    Immediate start
    Remote work
    Flexible hours

    ServiceNow

    Mountain View, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Software Engineer - AI Infrastructure Performance Insights & Observability. Be the first to apply!