Inference Software Engineer
$2,000 per monthEtched
Overview
Etched is building AI chips that are hard-coded for individual model architectures. Our first product (Sohu) supports transformers, delivering an order of magnitude more throughput and lower latency than GPUs. With Etched ASICs, you can build products that would be impossible with GPUs, like real-time video generation models and extremely deep and parallel chain-of-thought reasoning agents.
Key responsibilities
- Contribute to the architecture and design of the Sohu host software stack
- Implement high-performance, modular code across the complete Etched software stack, consisting of a mix of Rust, C++ and Python.
- Interface with firmware and drivers teams delivering highest-performance HW/SW stack.
- Work with AI model researchers and product-facing teams building out the Etched serving front-end.
Representative projects
- Build scheduling logic for handling continuous batching and real time inference
- Implement inference-time acceleration techniques such as speculative decoding, tree search, KV cache sharing, etc.
- Implement distributed networking primitives for efficient multi-server inference
You may be a good fit if you have
- Experience with C++ and Python
- Familiarity with transformer model architectures and inference serving stacks (vLLM, SGLang, etc.) or experience working in distributed inference/training environments
- Experience working cross-functionally in large software and hardware organizations
Strong candidates may also have
- Experience with Rust
- Familiarity with GPU kernels, the CUDA compilation stack and related tools, or other hardware accelerators
- Understanding of distributed systems, networking, and parallel programming
- Full medical, dental, and vision packages, with 100% of premium covered
- Housing subsidy of $2,000/month for those living within walking distance of the office
- Daily lunch and dinner in our office
- Relocation support for those moving to Cupertino
How we’re different
Etched believes in the Bitter Lesson. We think most of the progress in the AI field has come from using more FLOPs to train and run models, and the best way to get more FLOPs is to build model-specific hardware. Larger and larger training runs encourage companies to consolidate around fewer model architectures, which creates a market for single-model ASICs.
We are a fully in-person team in Cupertino, and greatly value engineering skills. We do not have boundaries between engineering and research, and we expect all of our technical staff to contribute to both as needed.
Equal employment opportunity
Etched is an Equal Employment Opportunity employer; we do not discriminate on the basis of any protected group status under any applicable law.
#J-18808-Ljbffr- ...architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud... ...high-speed inference. About the Role We're hiring a Software Engineer to help contribute to projects on our Inference Platform...SuggestedFull time
$184k - $287.5k
We are seeking highly skilled and motivated software engineers to join us and build AI inference systems that serve large-scale models with extreme efficiency. You’ll architect and implement high-performance inference stacks, optimize GPU kernels and compilers, drive industry...SuggestedFull time$152k - $241.5k
We are now looking for a Senior Software Engineer for Deep Learning Inference! Would you like to make a big impact in Deep Learning by helping build a state-of-the-art inference framework for accelerating Deep Learning models, especially Large Language Models, on NVIDIA...SuggestedFull time$152k - $241.5k
...some of the world’s most challenging problems. We're seeking talented and motivated engineers to join our TensorRT team in developing the industry-leading deep learning inference software for NVIDIA AI accelerators. As a Senior Software Engineer in the TensorRT team, you...SuggestedFull time$193.3k - $261.5k
We develop AWS Neuron, the complete software stack for Trainium, Amazon's custom cloud-scale machine learning accelerators. Join... ...on the Trainium hardware.As a Sr. Software Development Engineer on the Inference Model Enablement team, you will onboard and optimize state...SuggestedInternshipLocal areaFlexible hours$165.2k - $223.6k
...Web Services (AWS) builds AWS Neuron, the software development kit used to accelerate deep... ...and JAX enabling unparalleled ML inference and training performance.The Inference Enablement... ...the hardware-software boundary, our engineers build systematic infrastructure, innovate...Work experience placementInternshipLocal areaFlexible hours$152k - $241.5k
...technology for safety-critical applications? Join NVIDIA's TensorRT team as a Senior Software Engineer, and be at the forefront of technology, enabling high-performance AI inference solutions for automotive safety and other specialized platforms. Your expertise will help...Full time- ...the transformation of technology. We are at the forefront of software and hardware innovation, pushing the boundaries of what is... ...3 days per week.The role: Principal System Software Engineer, AI Inference ExecutionWhat you will do:The role requires you to be part...3 days per week
$139k - $204k
...What You’ll Do Senior engineers are area owners who lead designs, raise engineering standards, and deliver measurable improvements... ...orchestration, and hardware teams to evolve our Kubernetes‑native inference platform and meet strict P99 SLAs at scale. About The Role...Permanent employmentTemporary workCasual workWork at officeRemote workFlexible hoursShift work- ...NVIDIA is seeking a Senior Software Engineer for Deep Learning Inference to help build a state-of-the-art inference framework on NVIDIA GPUs, accelerating large language models. You will join the TensorRT Workflows team and tackle scalable, real-time inferencing challenges...
$165k - $242k
...Apply for the Senior Software Engineer II, Inference role at CoreWeave. CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence...Permanent employmentTemporary workCasual workWork at officeRemote workFlexible hoursShift work$139k - $204k
What You’ll Do Senior engineers are area owners who lead designs, raise engineering standards, and deliver measurable improvements to... ...orchestration, and hardware teams to evolve our Kubernetes‑native inference platform and meet strict P99 SLAs at scale. About The Role...Permanent employmentTemporary workCasual workWork at officeRemote workFlexible hoursShift work$2,000 per month
...chain-of-thought reasoning agents. Job Summary Etched’s Inference SW team enables optimal mapping of models to Sohu’s dataflow... ...hosts and racks. We are seeking a highly skilled and motivated engineer to join our team as we work towards enabling Mixture-of-...Full timeWork at officeRelocation package$92k - $135k
...CRWV) in March 2025. Learn more at What You'll Do: Join the Inference team to ship production features that improve latency,... ...practices, and grow quickly with mentorship from experienced engineers. About the role: Implement well-scoped features and fixes...Permanent employmentFull timeTemporary workCasual workInternshipWork at officeFlexible hours$160.36k - $240.54k
...components. Develop observability to track ML model lifecycles from data generation to on-road validation.Maintain an in-house ML inference platform to serve large language models efficiently.Maintain an in-house ML compiler platform to compile, deploy, and validate Nuro...Immediate startFlexible hours$152k - $241.5k
...company”.We are looking for an AI & Deep Learning Compiler Engineer. NVIDIA is hiring software engineers for its Deep Learning & AI Compiler (DLC) team... ...of AI, our DLC has been the backbone of NVIDIA’s inference engine, spanning across data centers, personal devices,...Full timeRemote work- ...to deliver industry-leading training and inference speeds; over 10 times faster than GPU-... ....About The RoleWe're hiring a Principal Engineer for our Inference Cloud Platform. This team... ...10+ years of experience in software engineering, with substantial individual...
$170k - $216k
...solutions to speed up developer velocity. We’re looking for a software engineer to join the team to build and maintain the critical data and... ...Software Engineer. You will: Develop Waymo's inference platform to make it scalable, high throughput, and low...Full timeRemote work- ...to deliver industry-leading training and inference speeds; over 10 times faster than GPU-... ...inference.About the RoleWe're hiring a Staff Engineer to own major areas of the architecture... ...Qualifications8+ years of experience in software engineering, with substantial individual...
$135k - $210k
...is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars. SOFTWARE ENGINEER, INFERENCE (AI DATA ENGINEERING) The application software team is the central nervous system of SpaceX – we create mission critical...Permanent employmentFull timeTemporary workRemote workWorldwideWeekend work$224k - $356.5k
We are now looking for a Senior System Software Engineer to work on Dynamo. NVIDIA is hiring software engineers for its GPU-accelerated deep... ...processing. We are a fast-paced team building Generative AI inference platform to make design and deployment of new AI models...Full time- NVIDIA is seeking a highly capable Engineering Manager to lead the next generation of LLM/VLM inference software. You will architect and guide a team of engineers, interfacing with researchers and GPU architects to deliver production-grade software that sets the standard...
$165k - $242k
...A cloud service provider is seeking a Senior Software Engineer II for their Inference team in Sunnyvale, California. In this role, you'll lead design reviews, implement optimizations, and improve service reliability. The ideal candidate has extensive experience with distributed...$139k - $204k
...CoreWeave is seeking a Senior Engineer to lead designs and enhance engineering standards within their Kubernetes-native inference platform. Responsibilities include driving architecture, defining SLIs/SLOs, and mentoring engineers, with 3-8 years of experience preferred...Remote workFlexible hours- NVIDIA seeks a Senior Product Manager for AI Platform Inference (Finance) in Santa Clara to lead tooling, SDKs, and libraries enabling... ...years in technical product management, knowledge of inference software and GenAI concepts, and strong communication. Equity and benefits...
- Cerebras Systems is seeking a Software Engineer to build and maintain high-performance, low-latency inference infrastructure. You will deploy scalable inference services, optimize auto-scaling, and integrate with Docker/Kubernetes in production environments. You will collaborate...
- ...NVIDIA is seeking a Senior Agentic AI Software Engineer to advance agentic AI systems and workloads from scalable research to production-grade solutions. You will build agentic components, analyze inference dynamics, and collaborate with teams owning evaluation pipelines...
- ...Staff Data Engineer, Full StackPrimary Skills: Data Engineering (Expert), Python (Expert... ...conversational AI systems.Knowledge of causal inference, time-series forecasting, and advanced... ...in Agile environments with modern software engineering and DevOps practices.Strong...Contract workRemote work
$216k - $414k
We are looking for Software Engineering Manager to lead the development efforts for the Triton Inference Server team! Academic and commercial groups around the world are using GPUs to power a revolution in deep learning, enabling breakthroughs in problems from image classification...$230k - $250k
Cerebras Systems is seeking a Sr. Member of Technical Staff in Sunnyvale, CA. This role involves designing resilient software features for cloud-based AI inference, leveraging AWS tools and services. Candidates should have a Master’s degree in Computer Science and experience...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Inference Software Engineer. Be the first to apply!
- senior robotics software engineer Cupertino, CA
- software system engineer Cupertino, CA
- intel software engineer Cupertino, CA
- software developer fintech Cupertino, CA
- software development engineer aws Cupertino, CA
- information technology software engineer Cupertino, CA
- software developer internship no experience Cupertino, CA
- ngo software engineer Cupertino, CA
- software engineer travel Cupertino, CA
- software engineer internship Cupertino, CA




