Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Inference Software Engineer

$2,000 per month
Full-time

Etched

About Etched

Etched is building AI chips that are hard-coded for individual model architectures. Our first product (Sohu) only supports transformers, but has an order of magnitude more throughput and lower latency than a B200. With Etched ASICs, you can build products that would be impossible with GPUs, like real-time video generation models and extremely deep & parallel chain-of-thought reasoning agents.

Job Summary

Etched’s Inference SW team enables optimal mapping of models to Sohu’s dataflow architecture and serving requests across multiple chips, hosts and racks. We are seeking a highly skilled and motivated engineer to join our team as we work towards enabling Mixture-of-Experts (MoE) architectures on Sohu systems. You’ll build SW enabling frontier inference performance to satisfy exponentially growing serving demand.

This role is for a general contributor and will be expected to contribute to all parts of our stack. We also have more specialized needs for this team posted on the site.

Key responsibilities

  • Support porting state-of-the-art models to our architecture. Help build programming abstractions and testing capabilities to rapidly iterate on model porting

  • Scale and enhance Sohu’s runtime, including multi-node inference, intra-node execution, state management, and robust error handling

  • Optimize routing and communication layers using Sohu’s collectives

  • Develop tools for performance profiling and debugging, identifying bottlenecks and correctness issues

You may be a good fit if you have

  • Proficiency in Rust and/or C++

  • Good familiarity with PyTorch and/or JAX.

  • Good familiarity with transformers architectures

  • Ported applications to non-standard or accelerator hardware platforms.

  • Solid systems knowledge, including Linux internals, accelerator architectures (e.g., GPUs, TPUs), and high-speed interconnects (e.g., NVLink, InfiniBand)

Strong candidates may also have experience with

  • Developed low-latency, high-performance applications using both kernel-level and user-space networking stacks.

  • Deep understanding of distributed systems concepts, algorithms, and challenges, including consensus protocols, consistency models, and communication patterns.

  • Solid grasp of large language model architectures, particularly Mixture-of-Experts (MoE).

  • Experience analyzing performance traces and logs from distributed systems and ML workloads.

  • Built applications with extensive SIMD (Single Instruction, Multiple Data) optimizations for performance-critical paths.

  • Familiar with cluster orchestration tools (e.g., Kubernetes, Slurm) and ML platforms (e.g., Ray, Kubeflow)

  • Experience designing and implementing CI/CD pipelines for MLOps workflows.

Benefits

  • Full medical, dental, and vision packages, with generous premium coverage

  • Housing subsidy of $2,000/month for those living within walking distance of the office

  • Daily lunch and dinner in our office

  • Relocation support for those moving to West San Jose

Compensation Range

  • $175,000 - $275,000

How we’re different

Etched believes in the Bitter Lesson . We think most of the progress in the AI field has come from using more FLOPs to train and run models, and the best way to get more FLOPs is to build model-specific hardware. Larger and larger training runs encourage companies to consolidate around fewer model architectures, which creates a market for single-model ASICs.

We are a fully in-person team in West San Jose, and greatly value engineering skills. We do not have boundaries between engineering and research, and we expect all of our technical staff to contribute to both as needed.

Vacancy posted 9 hours ago
Similar jobs that could be interesting for youBased on the Inference Software Engineer in San Jose, CA vacancy
  • $2,000 per month

     ...architecture and design of the Sohu host software stack Implement high-performance,...  ...handling continuous batching and real time inference Implement inference-time...  ...person team in Cupertino, and greatly value engineering skills. We do not have boundaries between... 
    Suggested
    Full time
    Work at office
    Relocation package

    Etched

    Cupertino, CA
    3 hours ago
  • $152k - $241.5k

    We are now looking for a Senior Software Engineer for Quantized Inference! NVIDIA is seeking software engineers to accelerate the discovery and deployment of efficient inference recipes for LLMs. A recipe defines which operators are transformed into low‑precision or sparsified... 
    Suggested

    NVIDIA

    Santa Clara, CA
    4 days ago
  • $136.8k - $259.2k

     ...Software Engineer Graduate (Inference Infrastructure) - 2026 Start (PHD) Location: San Jose Team: Technology Employment Type: Regular The Inference Infrastructure team is the creator and open-source maintainer of AIBrix, a Kubernetes-native control plane for... 
    Suggested
    Full time
    Temporary work

    Pangleglobal

    San Jose, CA
    3 hours ago
  • $136.8k - $259.2k

     ...A leading technology company is looking for a Software Engineer Graduate to join the Inference Infrastructure team in San Jose. This role involves designing and building large-scale cluster management systems and collaborating across teams for LLM inference solutions.... 
    Suggested
    Full time

    Pangleglobal

    San Jose, CA
    3 hours ago
  • $245k - $325k

     ...SambaNova Systems is seeking a Director of Software Engineering to lead the SambaStack platform engineering team. This role involves ensuring the delivery of reliable AI inference services while managing a high-performing group of engineers. Candidates should have over... 
    Suggested
    Full time

    jobs.frontdoordefense.com - Jobboard

    San Jose, CA
    3 hours ago
  • $229.9k - $262.4k

     ...Sr. Lead AI Engineer (Inference Optimization, FM hosting, AI Platform) Overview: At Capital One, we are creating responsible and reliable...  ...One. Design, develop, test, deploy, and support AI software components including foundation model training, large language... 
    Full time
    Part time
    Local area

    Capital One

    San Jose, CA
    2 days ago
  • $197.3k - $225.1k

    Lead AI Engineer (FM Hosting, LLM Inference) Overview At Capital One, we are creating responsible and reliable AI systems, changing banking for good...  ...One. Design, develop, test, deploy, and support AI software components including foundation model training, large language... 
    Full time
    Part time
    Local area

    Capital One

    San Jose, CA
    2 days ago
  •  ...d-Matrix in Santa Clara is seeking a Principal System Software Engineer - AI Inference Execution to help productize the SW stack for our AI compute engine. You will develop and maintain AI deployment software, collaborating across ML, compiler, and hardware teams to optimize... 
    Full time
    3 days per week

    d-Matrix

    Santa Clara, CA
    3 hours ago
  • $224k - $356.5k

     ...application is built. We are seeking a deeply technical software manager to lead production AI inference for NVIDIA Inference Microservices (NIM), the...  ...-ready software stack, combining optimized inference engines, model profiles/recipes, validated runtime configurations... 
    Full time

    NVIDIA

    Santa Clara, CA
    3 hours ago
  • $165k - $242k

     ...A cloud service provider is seeking a Senior Software Engineer II for their Inference team in Sunnyvale, California. In this role, you'll lead design reviews, implement optimizations, and improve service reliability. The ideal candidate has extensive experience with distributed... 
    Full time

    CoreWeave

    Sunnyvale, CA
    3 hours ago
  •  ...technology. We are at the forefront of software and hardware innovation, pushing the boundaries...  ...the US/Canada. The role: Software Engineer, Developer and Qualification Tools...  ...diagnostic tools for d-Matrix' cutting edge AI inference accelerators. You will be responsible... 
    Full time

    D-matrix

    Santa Clara, CA
    9 hours ago
  • $177.69k - $341.73k

    Responsibilities The Machine Learning (ML) System sub-team combines system engineering and the art of machine learning to develop and maintain massively distributed ML training and inference system/services around the world, providing high-performance, highly reliable,... 
    Full time

    ByteDance

    San Jose, CA
    3 hours ago
  •  ...one of their most valuable assets. About the Role SambaNova is hiring Software Engineers for SambaNova’s SambaStack platform. We are helping enterprises and service providers host their own AI inference platforms for end users, powered by our state-of-the-art RDU (... 
    Full time
    Temporary work
    Local area
    Flexible hours

    SambaNova Systems

    San Jose, CA
    1 day ago
  •  ...products that bring generative AI into the physical world. As a Software Engineer, you will play a central role in developing the platforms,...  ...span the software-hardware boundary, combining low‑latency inference pipelines, robust cloud infrastructure, and tightly... 
    Immediate start

    Siemens

    Santa Clara, CA
    1 day ago
  • $241.8k - $409.2k

     ...GPGPU Software Architect/ Principal Engineer XPENG is a leading smart technology company at the forefront of innovation, integrating advanced AI and...  ...the single source of truth. Benchmark includes MLPerf Inference, Stable Diffusion XL, and 70B LLM Cross-functional Collaboration... 
    Full time

    XPENG

    Santa Clara, CA
    3 hours ago
  • $165.2k - $223.6k

     ...Amazon Web Services (AWS) is building a central pipeline of Software Development Engineer (SDE) talent for anticipated roles in 2026. This...  ...fundamentals, including transformer architecture, training/inference lifecycles, and optimization techniques Knowledge of Python... 
    Internship
    Local area
    Flexible hours
    Day shift

    Amazon

    Cupertino, CA
    2 days ago
  •  ...backend services and APIs that support model inference, orchestration, and tool execution...  ...into working features Collaborate with ML engineers to integrate, evaluate, and...  ...scalability 1-3+ years of professional software engineering experience Strong experience... 
    Full time

    Eightfold

    Santa Clara, CA
    9 hours ago
  •  ...week in-office collaboration. We are looking for a Senior Software Engineer to build scalable AI infrastructure. As a Senior Software...  ...optimizations. Integrate training artifacts into on-robot inference stacks. Requirements Bachelor’s or Master’s degree... 
    Full time
    Work at office
    Visa sponsorship

    RoboForce

    Milpitas, CA
    9 hours ago
  • $180k - $220k

    Senior Software Engineer - ML/LLM Serving This range is provided by Alldus. Your actual pay will be based on your skills and experience —...  ...and optimize infrastructure that powers the deployment and inference of machine learning models across varied customer environments... 
    Full time
    Flexible hours

    Alldus

    San Jose, CA
    3 hours ago
  • $56.25 - $173 per hour

     ...opportunity for an accomplished, creative, Senior (possibly Staff) Software Engineer to support their product development needs. This is largely...  ...interviewing at HealthCare Recruiters International by 2x Inferred from the description for this job Medical insurance Vision... 
    Full time
    Summer work
    Internship
    Work at office
    Immediate start
    1 day per week

    HealthCare Recruiters International

    San Jose, CA
    3 hours ago
  • $152k - $204k

     ...CRWV) in March 2025. Learn more at What You'll Do: Senior engineers are area owners who lead designs, raise engineering standards,...  ...orchestration, and hardware teams to evolve our Kubernetes-native inference platform and meet strict P99 SLAs at scale. About the role:... 
    Permanent employment
    Temporary work
    Casual work
    Work at office
    Flexible hours
    Shift work

    CoreWeave

    Sunnyvale, CA
    3 days ago
  • $240k - $260k

     ...Job Description Job Description AI Platform Engineer – Training & Inference Saviynt's AI-powered identity platform manages and governs human and non-human access to all of an organization's applications, data, and business processes. Customers trust Saviynt to... 

    Saviynt

    Milpitas, CA
    7 days ago
  •  ...NVIDIA Corporation is seeking a Manager, Software Engineering to lead production AI inference for NVIDIA Inference Microservices. This role involves managing a team responsible for the deployment of optimized AI inference solutions and ensuring high-quality software releases... 
    Full time

    NVIDIA Corporation

    Santa Clara, CA
    3 hours ago
  • $181.1k - $318.4k

     ...Building and scaling these features requires not just world-class engineering, but a deep understanding of how institutions around the world...  ...experience. ~12+ years of industry experience as a software engineer, including 3+ years as a tech lead/architect. ~ Demonstrated... 
    Full time
    Relocation

    Apple

    Cupertino, CA
    3 hours ago
  •  ...Job Title: Fullstack Software Engineer As a Fullstack Software Engineer, you will design, develop, and maintain automation software solutions for semiconductor manufacturing at TSMC Arizona. You will work closely with multidisciplinary teams to create high-performance... 
    Work experience placement
    Monday to Friday

    TSMC

    San Jose, CA
    3 days ago
  • $388k

     ...apart:\n\n * Prior experience in AndroidTV and building and operating end-to-end systems\n * Prior experience in QoE metrics-driven software development and deployment\n * Prior experience in embedded development including identity and security \n\n\n\n\nGenerally, our... 
    Hourly pay
    Full time
    Immediate start
    Worldwide
    Flexible hours

    Netflix

    Los Gatos, CA
    9 hours ago
  • $150k - $275k

     ...in San Jose is seeking a highly skilled Supercomputing Engineer specialized in networking. This role involves developing high-performance networking solutions and optimizing software communication across inference nodes. Candidates should have strong C/C++ skills and experience... 
    Full time
    Relocation package

    Etched

    San Jose, CA
    3 hours ago
  • $187.74k - $190k

     ...A technology solutions firm in San Jose is seeking a Software Engineer to design software systems tailored to user needs. The role involves developing applications, maintaining databases, and collaborating with stakeholders. Candidates should have a Master's degree in... 
    Permanent employment
    Full time

    Nextgentechinc

    San Jose, CA
    3 hours ago
  • $120.75k - $161k

     ...Platform team to design, develop, and maintain the large-scale software platforms that serve millions of users globally. In this...  ...tolerance, horizontal scalability, and load balancing. Performance Engineering: Drive the scaling, optimization, and innovation of the Data Platform... 
    Permanent employment
    Full time
    Work at office
    3 days per week

    Eightfold

    Santa Clara, CA
    9 hours ago
  • $2,000 per month

     ...intelligence. We co-design chips, racks, software, and manufacturing to deliver best-in-...  ...first products are heavily focused on inference . Backed by hundreds of millions from top...  ...top-tier investors and staffed by leading engineers, Etched is redefining the infrastructure... 
    Work at office
    Relocation package

    ETCHED LLC

    San Jose, CA
    5 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Inference Software Engineer. Be the first to apply!