Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Infrastructure Software Engineer

$2,500 per month

Etched

Job Description

Job Description

About Etched

Etched is building hardware for frontier intelligence. We co-design chips, racks, software, and manufacturing to deliver best-in-class throughput and latency across both prefill and decode workloads. Our first products are heavily focused on inference . Backed by hundreds of millions from top-tier investors and staffed by leading engineers, Etched is redefining the infrastructure layer for the fastest growing industry in history.

Job Summary

Building cutting-edge model-specific ASICs requires crafting custom infrastructure and toolchains to support ultra-fast, reliable, and scalable development across the stack - from simulation to silicon. We build this infrastructure as software - and we engineer it with the same best practices we apply to our products. We use the same rigor, design discipline, and quality standards and testing as we do to our ASIC, software, and platform.

You will lead the development and adoption of next-generation infrastructure tooling, enabling Etched ASIC, Software, and Platform engineers to iterate faster, build more reliably, and push the boundaries of AI performance. This includes building and scaling our hybrid high-performance compute (HPC) cluster, optimized for massively parallel CI, EDA workflows, Emulation, and hardware-aware job execution.

You’ll also architect and implement a state-of-the-art observability stack with LLM integration and a strong emphasis on streaming health and performance telemetry, log aggregation, distributed tracing, insight generation, synthetic testing, and smart alerting - across CI pipelines, simulation clusters, and service endpoints.

This role demands a strong software engineering mindset, quality instincts, and deep understanding of systems. It’s not just about writing scripts - it’s about writing code that builds and manages infrastructure with precision, repeatability, and intent.

Key responsibilities

  • Design and build the orchestration layers that drive our hybrid high-performance clusters—enabling simulation, synthesis, and continuous integration of AI ASICs at unprecedented scale.

  • Develop and maintain a fully programmable infrastructure control plane to ensure reproducibility, auditability, and rapid iteration across the entire stack.

  • Create tools and abstractions that empower engineers to harness massive parallelism without worrying about the underlying complexity..

  • Prototype and execute workload orchestration and migration strategies between on-premise and cloud environments, balancing performance, storage availability and replication, uptime, and cost across heterogeneous hardware and compute backends.

  • Implement real-time telemetry, tracing systems that surface insights from millions of metrics, enabling proactive debugging and system optimization.

  • Build a full observability stack that includes dashboards, alerting, automated responses, and a synthetic testing framework to proactively test infrastructure performance and reliability for various application and data flows, ensuring we remain proactive against issues impacting development and productivity workflows.

Representative projects

  • Design and deploy a fully automated, scalable hybrid HPC cluster, combining bare-metal servers and switches with cloud instances, provisioned through MaaS and orchestrated via SLURM and Kubernetes, optimized for mixed EDA workloads and parallel CI pipelines.

  • Develop a real-time observability system for ASIC toolchain jobs and distributed builds, integrating Prometheus, Grafana, and VictoriaMetrics with streaming telemetry, tracing, and alerting to detect performance regressions before they hit silicon.

  • Architect and implement a programmable infrastructure-as-code control plane, using Terraform, Ansible, and Puppet, to version, audit, and redeploy every layer of Etched's development stack with deterministic reproducibility.

  • Create a zero-downtime interactive development environment that provisions and connects Jupyter and VS Code sessions to GPUs and high-memory nodes via a secure zero-trust network, abstracting away cluster state and machine failures.

  • Prototype and evaluate dynamic workload migration strategies between on-premise and cloud environments to optimize for latency, reliability, and cost across simulation and synthesis pipelines.

  • Design a synthetic testing and fault injection framework to validate the behavior of infrastructure under high-load, degraded hardware, and intermittent network partitions - before they happen in production.

You may be a good fit if you

  • Are a systems-minded software engineer who loves building foundational platforms, working close to the metal and cloud, solving high-leverage problems at scale.

  • Are a deeply technical engineer who treats infrastructure as a software problem - prioritizing clean abstractions, version control,small change lists, easy roll backs, testing, and long-term maintainability over ad hoc configuration.

  • Have strong programming skills in languages such as Python, Go, Rust, and C++, and are comfortable building production-grade tooling.

  • Possess expert-level knowledge of Linux, virtualization, containerization, and CI/CD pipelines, with a deep understanding of how to debug, optimize, and scale complex systems.

  • Are familiar with Infrastructure as Code tools like OpenTofu, Ansible, or Puppet, and enjoy designing declarative, reproducible infrastructure systems.

  • Understand and use PromQL and other telemetry/query languages and have used LLM to extract insight from real-time metrics, and know how to architect and tune observability stacks.

  • Have a track record of debugging and resolving difficult hardware-software integration problems across bare-metal systems, networks, and distributed workloads.

  • Can lead and mentor technical teams, guiding design decisions and helping others develop sound engineering instincts.

  • Have 8+ years of experience in infrastructure engineering, systems programming, or backend software development - ideally in environments where performance, scale, or hardware interaction mattered.

  • Are driven by curiosity, take initiative, and have an innate sense of ownership — you thrive in uncharted territory, design for edge cases, and love making systems more powerful, reliable, and elegant.

Strong candidates may also have experience with

  • Familiarity with Bazel build system

  • Deep understanding of ASIC development flows, especially those involving Synopsys, Cadence, and Verilator, including how EDA tools interact with infrastructure for simulation, synthesis, and verification.

  • Hands-on experience architecting systems with AWS, GCP, or Azure, including hybrid on-prem/cloud deployments, workload migration strategies, and cloud-native orchestration tooling.

  • Experience monitoring, provisioning, and debugging bare-metal servers, network hardware, and high-performance storage systems in rack-scale environments.

  • Comfortable in profiling and optimizing compute environments for single-threaded latency, memory-bound workloads, or I/O throughput, especially in the context of simulation or CI performance.

  • Proficiency building or operating telemetry systems at scale using Prometheus, Grafana, Loki, VictoriaMetrics, and tools for distributed tracing, log aggregation, and real-time alerting across heterogeneous mediums (SMS, email, push alerts, etc.)

Benefits

  • Medical, dental, and vision packages with generous premium coverage

    • $500 per month credit for waiving medical benefits

  • Housing subsidy of $2,500 per month for those living within walking distance of the office

  • Relocation support for those moving to San Jose (Santana Row)

  • Various wellness benefits covering fitness, mental health, and more

  • Daily lunch + dinner in our office

  • Unlimited compute budget subject to ROI justification

How we’re different

Etched believes in the Bitter Lesson. We are the first inference-focused frontier AI system. Our addressable market is the entirety of inference, unlike many of our competitors.

We are a fully in-person team in San Jose (Santana Row), and greatly value engineering skills. We do not have boundaries between engineering and research, and we expect all of our technical staff to contribute to both and work across disciplines as needed.

Compensation Range: $150K - $250K

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Infrastructure Software Engineer in San Jose, CA vacancy
  • $165.2k - $223.6k

    AWS Infrastructure Services owns the design, planning, delivery, and operation of all AWS global infrastructure. In other...  ...people who want to help. You’ll join a diverse team of software, hardware, and network engineers, supply chain specialists, security experts,... 
    Suggested
    Full time
    Internship
    Local area
    Worldwide
    Flexible hours

    Amazon.com Services LLC

    Santa Clara, CA
    10 hours ago
  • $116.5k - $173.25k

     ...Job Summary: This role contributes software solutions under guidance across phases...  ...Software Development Lifecycle (SDLC). The engineer works collaboratively with senior...  ...designing. We spend our days scaling our infrastructure and building new features to meet and exceed... 
    Suggested
    Full time
    Work at office
    Local area
    Immediate start
    Flexible hours

    Paypal

    San Jose, CA
    1 day ago
  • $388k

     ...created, and delivered to a global audience. Engineering teams within Netflix work hard every...  ...experiences in an ever-growing complex software landscape. Our Content & Business...  ...problems rather than managing infrastructure. Currently, this platform is foundational... 
    Suggested
    Hourly pay
    Full time
    Work experience placement
    Immediate start
    Flexible hours

    Netflix

    Los Gatos, CA
    1 day ago
  •  ...robots as easy to build and deploy as software. Today, robotics is fragmented, slow, and...  ...Role We are looking for a Senior AI Engineer to design, build, and ship AI-powered...  ...across the full stack — from the agentic infrastructure that powers our robot operations, to... 
    Suggested
    Full time

    Dexmate

    Santa Clara, CA
    1 day ago
  • $388k

     ...Our Content & Business Products (CBP) Engineering teams build the products and services to...  ...teams build on, across the full software stack — from GraphQL/RESTful APIs and embeddable...  ...orchestration and other related infrastructure platform domains. Deep understanding... 
    Suggested
    Hourly pay
    Full time
    Immediate start
    Remote work
    Flexible hours

    Netflix

    Los Gatos, CA
    1 day ago
  •  ...and continuous learning are core to how we work. The Role We're looking for a Senior Software Engineer to join our Platform team and build the foundational infrastructure that powers Gridmatic. Our platform challenges are shaped by the nature of energy markets: forecasts... 
    Full time

    Gridmatic

    Cupertino, CA
    1 day ago
  • $150.4k - $277.6k

     ...Platform And Frameworks Software Engineer, Sear The SPEAR team in Apple's Security Engineering & Architecture organization is hiring...  ...team will deliver well-designed, robust, and maintainable infrastructure and mitigations that eliminate whole classes of software vulnerabilities... 
    Relocation

    Apple

    Cupertino, CA
    1 day ago
  • $139k - $257.55k

     ...League content as trusted, structured, and AI-ready knowledge for internal and external clients. We are looking for a Senior Software Engineer (P40) who will help design and build these knowledge management and GenAI services , enabling product teams to create high-... 
    Temporary work
    Local area
    Worldwide

    Adobe

    San Jose, CA
    1 day ago
  •  ...class leadership team: Our Heads of AI, Engineering, and Product bring extensive experience...  .... Now we need an experienced software engineer to make these systems scale. You...  ...alongside applied AI scientists and an AI infrastructure engineer to transform code into reliable... 

    Kai Cyber, Inc.

    San Jose, CA
    2 days ago
  • $79.1k - $166.1k

     ...candidate with a masters degree in Computer Engineering, Electrical Engineering, Computer...  ...exposure in embedded systems, systems software, firmware, or hardware-adjacent software...  ...companion software that protect Oracle Cloud Infrastructure servers at scale. We work across... 
    Full time
    Internship
    Flexible hours

    Oracle

    Santa Clara, CA
    4 days ago
  • $165.2k - $223.6k

    AWS Infrastructure Services owns the design, planning, delivery, and operation of all AWS global infrastructure. In other...  ...people who want to help. You’ll join a diverse team of software, hardware, and network engineers, supply chain specialists, security experts,... 
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    a month ago
  • $120.75k - $161k

     ...Platform team to design, develop, and maintain the large-scale software platforms that serve millions of users globally. In this...  ...tolerance, horizontal scalability, and load balancing. Performance Engineering: Drive the scaling, optimization, and innovation of the Data Platform... 
    Permanent employment
    Full time
    Work at office
    3 days per week

    Eightfold

    Santa Clara, CA
    1 day ago
  • $224k - $356.5k

     ...outstanding team at NVIDIA to craft the future of computing! As a Senior Software Engineer - Traffic & Networking, you'll have the chance to define and deliver innovative solutions for our cloud infrastructure. Collaborate with top engineers, architects, and managers to build... 
    Full time

    Nvidia

    Santa Clara, CA
    11 days ago
  • $184k - $287.5k

     ...At NVIDIA, we are redefining the future of technology, and our Senior Software Engineer, Networking role offers a uniquely ambitious opportunity to contribute to world-class innovations. If you thrive in a collaborative and inclusive environment, are driven to succeed,... 
    Full time

    Nvidia

    Santa Clara, CA
    a month ago
  • $127.1k - $185k

     ...re looking for a talented early-career engineer to join our team that owns the network...  ...distributed AI/ML systems. You'll work on software that enables the world's largest AI...  ...computing, networking, and machine learning infrastructure - building the systems that power the... 
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    22 days ago
  • $175k - $263k

     ...leave your mark, come join us. THE ROLE Join the Systems Software team to architect and deliver the core software powering the...  ...interact directly with ASIC capabilities. Availability-Focused Engineering: Understanding of non-disruptive upgrade (NDU) technologies... 
    Work at office
    Flexible hours

    Everpure, Inc.

    Santa Clara, CA
    3 days ago
  • $184k - $287.5k

     ...generative AI revolution, building the software and systems that power the world’s most...  .... We are looking for a Senior Software Engineer to lead the bring-up, triage, benchmarking...  ...debugging of large-scale AI clusters, infrastructure, and end-to-end workloads, setting the... 
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    a month ago
  • $152k - $241.5k

     ...than raw traffic scale. We want hands-on engineers with strong systems fundamentals,...  ...Attestation Cloud platform, verification flows, software development kits and command-line tools...  ...systems, systems software, SDKs, infrastructure platforms, or customer-facing APIs.Demonstrated... 
    Full time
    Local area
    Remote work

    Nvidia

    Santa Clara, CA
    4 days ago
  • $143k - $197.44k

     ...and services ranging from Level 2 to Level 4, addressing mobility, logistics, and urban services. WeRide.ai is looking for a Software Engineer to build powerful and efficient world-leading cloud infra platforms for autonomous driving, inlcuding PaaS platforms, such as... 

    WeRide.ai

    San Jose, CA
    24 days ago
  • $193.3k - $261.5k

    AWS Infrastructure Services owns the design, planning, delivery, and operation of all AWS global infrastructure. In other...  ...people who want to help. You’ll join a diverse team of software, hardware, and network engineers, supply chain specialists, security experts,... 
    Internship
    Local area
    Flexible hours

    Amazon

    Santa Clara, CA
    a month ago
  • $140k - $224.25k

    NVIDIA DGX Cloud provides the infrastructure and software platform that enables enterprises to build, train, and deploy AI at scale. As demand for...  ...GPU infrastructure.We are looking for a Senior Software Engineer to design and build the systems that connect customer demand... 
    Full time

    Nvidia

    Santa Clara, CA
    7 days ago
  • $224k - $356.5k

     ...Autonomous Vehicles Platform team is looking for a Senior System Software Engineer. As part of our team, you will work on our Autonomous...  ...teams across the stack: hardware, embedded software, and cloud infrastructure.What you'll be doing:Drive AI-based automotive systems from... 
    Full time

    Nvidia

    Santa Clara, CA
    7 days ago
  • $193.3k - $261.5k

     ...and SystemC models of these custom SoCs — that let software teams start development months before silicon arrives...  ...of first silicon. We're looking for a software engineer to build and own the models and infrastructure that make this possible.What you'll do:- Build and... 
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    a month ago
  • $280k - $380k

     ...We focus on reliability and automation, engineering systems that perform under stress and...  ...everything you can, and turning complex infrastructure into reliable, well‑documented systems,...  ...(Site Reliability Engineering) Senior Software Engineer to join our dynamic team. The... 
    Work at office
    Local area
    Remote work
    Monday to Thursday
    Flexible hours

    Roku

    San Jose, CA
    a month ago
  • $388k

     ...merge creativity, intuition and cutting-edge technology. Come be a part of what’s next. Netflix's Ads Engineering team is building the technology and infrastructure behind our advertising business and in-house ad tech ecosystem. We are creating highly performant... 
    Hourly pay
    Full time
    Local area
    Immediate start
    Flexible hours

    Netflix

    Los Gatos, CA
    1 day ago
  • $132k - $198.45k

     ...investors. About the Team We're looking for a Fullstack Software Engineer to join the Fleet Platform & Operations Tooling team. You'...  ...work at the intersection of fleet operations and software infrastructure. You'll own features end-to-end — from building responsive... 
    Immediate start
    Flexible hours

    Nuro

    San Jose, CA
    18 days ago
  • $136.3k - $231.7k

     ...into R&D. Our expert teams of physicists, engineers, data scientists and problem-solvers...  ...and the brightest research scientist, software engineers, application development engineers...  ..., platform integration, and infrastructure support in enterprise or production environments... 
    Minimum wage
    Full time
    Flexible hours

    KLA-Tencor

    Milpitas, CA
    12 days ago
  • $168k - $270.25k

     ...The NVIDIA Operations organization is seeking an experienced software engineering professional for the position of System Data Architect. As a member of our team you will be an integral part of building data platforms and tools. You will support initiatives for the Data... 

    NVIDIA

    Santa Clara, CA
    5 days ago
  • $184k - $287.5k

     ...We are looking for a Senior Software Engineer to become part of our storage management plane team. The management plane is a web-based...  ...capabilities to handle and supervise our distributed storage infrastructure. Our team is continually dedicated to acquiring and implementing... 
    Full time

    Nvidia

    Santa Clara, CA
    a month ago
  • $45 - $60 per hour

     ...such as Java, Python, Go, etc., with a certain programming foundation and good learning ability. # Have a certain understanding of software R&D processes and DevOps concepts, and those who understand the working modes and needs of the QA and RD teams are preferred. #... 
    Hourly pay
    Internship
    Local area

    TikTok

    San Jose, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Infrastructure Software Engineer. Be the first to apply!