Inference Software Engineer
$2,000 per monthEtched
About Etched
Etched is building AI chips that are hard-coded for individual model architectures. Our first product (Sohu) only supports transformers, but has an order of magnitude more throughput and lower latency than a B200. With Etched ASICs, you can build products that would be impossible with GPUs, like real-time video generation models and extremely deep & parallel chain-of-thought reasoning agents.
Job Summary
Etched’s Inference SW team enables optimal mapping of models to Sohu’s dataflow architecture and serving requests across multiple chips, hosts and racks. We are seeking a highly skilled and motivated engineer to join our team as we work towards enabling Mixture-of-Experts (MoE) architectures on Sohu systems. You’ll build SW enabling frontier inference performance to satisfy exponentially growing serving demand.
This role is for a general contributor and will be expected to contribute to all parts of our stack. We also have more specialized needs for this team posted on the site.
Key responsibilities
Support porting state-of-the-art models to our architecture. Help build programming abstractions and testing capabilities to rapidly iterate on model porting
Scale and enhance Sohu’s runtime, including multi-node inference, intra-node execution, state management, and robust error handling
Optimize routing and communication layers using Sohu’s collectives
Develop tools for performance profiling and debugging, identifying bottlenecks and correctness issues
You may be a good fit if you have
Proficiency in Rust and/or C++
Good familiarity with PyTorch and/or JAX.
Good familiarity with transformers architectures
Ported applications to non-standard or accelerator hardware platforms.
Solid systems knowledge, including Linux internals, accelerator architectures (e.g., GPUs, TPUs), and high-speed interconnects (e.g., NVLink, InfiniBand)
Strong candidates may also have experience with
Developed low-latency, high-performance applications using both kernel-level and user-space networking stacks.
Deep understanding of distributed systems concepts, algorithms, and challenges, including consensus protocols, consistency models, and communication patterns.
Solid grasp of large language model architectures, particularly Mixture-of-Experts (MoE).
Experience analyzing performance traces and logs from distributed systems and ML workloads.
Built applications with extensive SIMD (Single Instruction, Multiple Data) optimizations for performance-critical paths.
Familiar with cluster orchestration tools (e.g., Kubernetes, Slurm) and ML platforms (e.g., Ray, Kubeflow)
Experience designing and implementing CI/CD pipelines for MLOps workflows.
Benefits
Full medical, dental, and vision packages, with generous premium coverage
Housing subsidy of $2,000/month for those living within walking distance of the office
Daily lunch and dinner in our office
Relocation support for those moving to West San Jose
Compensation Range
$175,000 - $275,000
How we’re different
Etched believes in the Bitter Lesson . We think most of the progress in the AI field has come from using more FLOPs to train and run models, and the best way to get more FLOPs is to build model-specific hardware. Larger and larger training runs encourage companies to consolidate around fewer model architectures, which creates a market for single-model ASICs.
We are a fully in-person team in West San Jose, and greatly value engineering skills. We do not have boundaries between engineering and research, and we expect all of our technical staff to contribute to both as needed.
$2,000 per month
...architecture and design of the Sohu host software stack Implement high-performance,... ...handling continuous batching and real time inference Implement inference-time... ...person team in Cupertino, and greatly value engineering skills. We do not have boundaries between...SuggestedFull timeWork at officeRelocation package$152k - $241.5k
We are now looking for a Senior Software Engineer for Quantized Inference! NVIDIA is seeking software engineers to accelerate the discovery and deployment of efficient inference recipes for LLMs. A recipe defines which operators are transformed into low‑precision or sparsified...Suggested$136.8k - $259.2k
...Software Engineer Graduate (Inference Infrastructure) - 2026 Start (PHD) Location: San Jose Team: Technology Employment Type: Regular The Inference Infrastructure team is the creator and open-source maintainer of AIBrix, a Kubernetes-native control plane for...SuggestedFull timeTemporary work$136.8k - $259.2k
...A leading technology company is looking for a Software Engineer Graduate to join the Inference Infrastructure team in San Jose. This role involves designing and building large-scale cluster management systems and collaborating across teams for LLM inference solutions....SuggestedFull time$245k - $325k
...SambaNova Systems is seeking a Director of Software Engineering to lead the SambaStack platform engineering team. This role involves ensuring the delivery of reliable AI inference services while managing a high-performing group of engineers. Candidates should have over...SuggestedFull time$229.9k - $262.4k
...Sr. Lead AI Engineer (Inference Optimization, FM hosting, AI Platform) Overview: At Capital One, we are creating responsible and reliable... ...One. Design, develop, test, deploy, and support AI software components including foundation model training, large language...Full timePart timeLocal area$197.3k - $225.1k
Lead AI Engineer (FM Hosting, LLM Inference) Overview At Capital One, we are creating responsible and reliable AI systems, changing banking for good... ...One. Design, develop, test, deploy, and support AI software components including foundation model training, large language...Full timePart timeLocal area- ...d-Matrix in Santa Clara is seeking a Principal System Software Engineer - AI Inference Execution to help productize the SW stack for our AI compute engine. You will develop and maintain AI deployment software, collaborating across ML, compiler, and hardware teams to optimize...Full time3 days per week
$224k - $356.5k
...application is built. We are seeking a deeply technical software manager to lead production AI inference for NVIDIA Inference Microservices (NIM), the... ...-ready software stack, combining optimized inference engines, model profiles/recipes, validated runtime configurations...Full time$165k - $242k
...A cloud service provider is seeking a Senior Software Engineer II for their Inference team in Sunnyvale, California. In this role, you'll lead design reviews, implement optimizations, and improve service reliability. The ideal candidate has extensive experience with distributed...Full time- ...technology. We are at the forefront of software and hardware innovation, pushing the boundaries... ...the US/Canada. The role: Software Engineer, Developer and Qualification Tools... ...diagnostic tools for d-Matrix' cutting edge AI inference accelerators. You will be responsible...Full time
$177.69k - $341.73k
Responsibilities The Machine Learning (ML) System sub-team combines system engineering and the art of machine learning to develop and maintain massively distributed ML training and inference system/services around the world, providing high-performance, highly reliable,...Full time- ...one of their most valuable assets. About the Role SambaNova is hiring Software Engineers for SambaNova’s SambaStack platform. We are helping enterprises and service providers host their own AI inference platforms for end users, powered by our state-of-the-art RDU (...Full timeTemporary workLocal areaFlexible hours
- ...products that bring generative AI into the physical world. As a Software Engineer, you will play a central role in developing the platforms,... ...span the software-hardware boundary, combining low‑latency inference pipelines, robust cloud infrastructure, and tightly...Immediate start
$241.8k - $409.2k
...GPGPU Software Architect/ Principal Engineer XPENG is a leading smart technology company at the forefront of innovation, integrating advanced AI and... ...the single source of truth. Benchmark includes MLPerf Inference, Stable Diffusion XL, and 70B LLM Cross-functional Collaboration...Full time$165.2k - $223.6k
...Amazon Web Services (AWS) is building a central pipeline of Software Development Engineer (SDE) talent for anticipated roles in 2026. This... ...fundamentals, including transformer architecture, training/inference lifecycles, and optimization techniques Knowledge of Python...InternshipLocal areaFlexible hoursDay shift- ...backend services and APIs that support model inference, orchestration, and tool execution... ...into working features Collaborate with ML engineers to integrate, evaluate, and... ...scalability 1-3+ years of professional software engineering experience Strong experience...Full time
- ...week in-office collaboration. We are looking for a Senior Software Engineer to build scalable AI infrastructure. As a Senior Software... ...optimizations. Integrate training artifacts into on-robot inference stacks. Requirements Bachelor’s or Master’s degree...Full timeWork at officeVisa sponsorship
$180k - $220k
Senior Software Engineer - ML/LLM Serving This range is provided by Alldus. Your actual pay will be based on your skills and experience —... ...and optimize infrastructure that powers the deployment and inference of machine learning models across varied customer environments...Full timeFlexible hours$56.25 - $173 per hour
...opportunity for an accomplished, creative, Senior (possibly Staff) Software Engineer to support their product development needs. This is largely... ...interviewing at HealthCare Recruiters International by 2x Inferred from the description for this job Medical insurance Vision...Full timeSummer workInternshipWork at officeImmediate start1 day per week$152k - $204k
...CRWV) in March 2025. Learn more at What You'll Do: Senior engineers are area owners who lead designs, raise engineering standards,... ...orchestration, and hardware teams to evolve our Kubernetes-native inference platform and meet strict P99 SLAs at scale. About the role:...Permanent employmentTemporary workCasual workWork at officeFlexible hoursShift work$240k - $260k
...Job Description Job Description AI Platform Engineer – Training & Inference Saviynt's AI-powered identity platform manages and governs human and non-human access to all of an organization's applications, data, and business processes. Customers trust Saviynt to...- ...NVIDIA Corporation is seeking a Manager, Software Engineering to lead production AI inference for NVIDIA Inference Microservices. This role involves managing a team responsible for the deployment of optimized AI inference solutions and ensuring high-quality software releases...Full time
$181.1k - $318.4k
...Building and scaling these features requires not just world-class engineering, but a deep understanding of how institutions around the world... ...experience. ~12+ years of industry experience as a software engineer, including 3+ years as a tech lead/architect. ~ Demonstrated...Full timeRelocation- ...Job Title: Fullstack Software Engineer As a Fullstack Software Engineer, you will design, develop, and maintain automation software solutions for semiconductor manufacturing at TSMC Arizona. You will work closely with multidisciplinary teams to create high-performance...Work experience placementMonday to Friday
$388k
...apart:\n\n * Prior experience in AndroidTV and building and operating end-to-end systems\n * Prior experience in QoE metrics-driven software development and deployment\n * Prior experience in embedded development including identity and security \n\n\n\n\nGenerally, our...Hourly payFull timeImmediate startWorldwideFlexible hours$150k - $275k
...in San Jose is seeking a highly skilled Supercomputing Engineer specialized in networking. This role involves developing high-performance networking solutions and optimizing software communication across inference nodes. Candidates should have strong C/C++ skills and experience...Full timeRelocation package$187.74k - $190k
...A technology solutions firm in San Jose is seeking a Software Engineer to design software systems tailored to user needs. The role involves developing applications, maintaining databases, and collaborating with stakeholders. Candidates should have a Master's degree in...Permanent employmentFull time$120.75k - $161k
...Platform team to design, develop, and maintain the large-scale software platforms that serve millions of users globally. In this... ...tolerance, horizontal scalability, and load balancing. Performance Engineering: Drive the scaling, optimization, and innovation of the Data Platform...Permanent employmentFull timeWork at office3 days per week$2,000 per month
...intelligence. We co-design chips, racks, software, and manufacturing to deliver best-in-... ...first products are heavily focused on inference . Backed by hundreds of millions from top... ...top-tier investors and staffed by leading engineers, Etched is redefining the infrastructure...Work at officeRelocation package
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Inference Software Engineer. Be the first to apply!
- software engineer full time San Jose, CA
- software system engineer San Jose, CA
- consulting software engineer San Jose, CA
- software engineer travel San Jose, CA
- real time software engineer San Jose, CA
- network software engineer San Jose, CA
- senior software engineer remote San Jose, CA
- entry level software engineer remote San Jose, CA
- software engineer intern San Jose, CA
- software developer fintech San Jose, CA















