Inference Intern
Etched
Job Description
Job Description
About Etched
Etched is building hardware for frontier intelligence. We co-design chips, racks, software, and manufacturing to deliver best-in-class throughput and latency across both prefill and decode workloads. Our first products are heavily focused on inference . Backed by hundreds of millions from top-tier investors and staffed by leading engineers, Etched is redefining the infrastructure layer for the fastest growing industry in history.
Job Summary
We are seeking talented Fall '26, Spring '27, and Summer '27 Inference Architecture interns to join our team and contribute to the design of next-generation AI accelerators. This role focuses on developing and optimizing compute architectures that deliver exceptional performance and efficiency for inference workloads. You will work on cutting-edge architectural problems and performance modeling over the course of your internship.
Key responsibilities
Support porting state-of-the-art models to our architecture. Help build programming abstractions and testing capabilities to rapidly iterate on model porting.
Assist in building, enhancing, and scaling our runtime, including multi-node inference, intra-node execution, state management, and robust error handling.
Contribute to optimizing routing and communication layers using our collectives.
Utilize performance profiling and debugging tools to identify bottlenecks and correctness issues.
Develop and leverage a deep understanding of our architecture to co-design both HW instructions and model architecture operations to maximize model performance
Implement high-performance software components for the Model Toolkit
You may be a good fit if you have
Progress towards a Bachelor’s, Master’s, or PhD degree in computer science, computer engineering, applied mathematics, or a related field
Proficiency in Python, C++
Understanding of performance-sensitive or complex distributed software systems, e.g. Linux internals, accelerator architectures (e.g. GPUs, TPUs), Compilers, or high-speed interconnects (e.g. NVLink, InfiniBand).
Ported applications to non-standard accelerator hardware or hardware platforms.
Deep knowledge of transformer model architectures and/or inference serving stacks (vLLM, SGLang, etc.)
Strong candidates may have some experience with
Proficiency in Rust
Low-latency, high-performance applications using both kernel-level and user-space networking stacks.
Deep understanding of distributed systems concepts, algorithms, and challenges, including consensus protocols, consistency models, and communication patterns.
Solid grasp of Transformer architectures, particularly Mixture-of-Experts (MoE).
Built applications with extensive SIMD (Single Instruction, Multiple Data) optimizations for performance-critical paths.
Familiarity with PyTorch or JAX.
Math competitions (AIME, AMC, etc)
We encourage you to apply even if you do not believe you meet every qualification.
Program details
12-week paid internship
Generous housing support for those relocating
Daily lunch and dinner in our office
Based at our office in San Jose, CA
Direct mentorship from industry leaders and world-class engineers
Opportunity to work on one of the most important problems of our time
For any questions, contact View email address on ziprecruiter.com.
How we’re different
Etched believes in the Bitter Lesson. We are the first inference-focused frontier AI system. Our addressable market is the entirety of inference, unlike many of our competitors.
We are a fully in-person team in San Jose (Santana Row), and greatly value engineering skills. We do not have boundaries between engineering and research, and we expect all of our technical staff to contribute to both and work across disciplines as needed.
$195.2k - $361.2k
...models fast on the hardware people actually own. You optimize inference engines (llama.cpp, vLLM) for constrained local and edge environments... ...learn / grow intoCuriosity is required. You will develop:The internals of modern inference engines and where the milliseconds...InternshipFull timeLocal areaImmediate startShift work- ...and decode workloads. Our first products are heavily focused on inference . Backed by hundreds of millions from top-tier investors and... ...industry in history. Job Summary Our electrical engineering interns will work on both hands-on product design and the development...InternshipSummer internshipWork at officeRelocation
$184k - $287.5k
We are seeking highly skilled and motivated software engineers to join us and build AI inference systems that serve large-scale models with extreme efficiency. You’ll architect and implement high-performance inference stacks, optimize GPU kernels and compilers, drive industry...SuggestedFull time- ...Frontier Group sits at the leading edge of what’s possible with LLM inference on heterogeneous hardware. Our charter spans the full stack:... ...out of them at inference time.• Are familiar with the internals of open-source inference frameworks (vLLM, SGLang, TensorRT-LLM...Suggested
$193.3k - $261.5k
...accelerators. Join us to optimize the latest models to run really fast on the Trainium hardware.As a Software Development Engineer on the Inference Model Enablement team, you will onboard and optimize state-of-the-art open-source and customer LLMs, both dense and MoE, for...InternshipLocal areaFlexible hours- ...and decode workloads. Our first products are heavily focused on inference . Backed by hundreds of millions from top-tier investors and... ...looking for Summer '26, Fall '26, Spring '27, and Summer '27 interns. You may be a good fit if you have Progress towards...InternshipSummer workSummer internshipWork at officeRelocation
- ...and decode workloads. Our first products are heavily focused on inference . Backed by hundreds of millions from top-tier investors and... ...are looking for Summer '26, Fall '26, Spring '27, and Summer '27 interns. Key responsibilities: Interview Scheduling & Logistics: Schedule...InternshipSummer workSummer internshipWork at officeRelocation
$165.2k - $223.6k
...popular ML frameworks like PyTorch and JAX enabling unparalleled ML inference and training performance.The Inference Enablement and... ...participate in design discussions, code review, and communicate with internal and external stakeholders. You will work cross-functionally to...InternshipWork experience placementLocal areaFlexible hours- ...and decode workloads. Our first products are heavily focused on inference . Backed by hundreds of millions from top-tier investors and... ...strategies, and supply chain management, we are looking for a Finance intern to tackle strategic financial challenges and execute on our day...InternshipSummer internshipWork at officeRelocation
$182.5k - $260.5k
...roleAs a Senior Staff Machine Learning Scientist, you own the inference and optimization layer that makes AI in agentic workflows fast... ...C++ for low-level interop is a plus.Solid grasp of transformer internals and the levers that move real inference performance and cost:...$21.69 per hour
...of Santa Clara and Mission College are partnering for a Student Internship Program for Summer 2026. The City is looking for Student Intern II's to perform work in various departments Citywide. As part of this program, Student Intern II's are required to be currently...InternshipHourly paySummer workSummer internship$20 - $24 per hour
...successful InsurTech companies, certified as a Great Place to Work®! We’re looking for an eager and motivated Office Administrator Intern to join our team and gain hands-on experience in a fast-paced, professional environment. This role provides exposure to day-to-day...InternshipHourly payFull timePart timeWork experience placementSummer internshipWork at officeLocal areaVisa sponsorshipMonday to ThursdayFlexible hours$25 - $35 per hour
...opportunity. Program dates are: ~ May 26, 2026 - July 31, 2026 ~ June 15, 2026 - August 21, 2026 ~*We will not be able to accommodate interns outside of these two program dates. This role will be based at Archer HQ in San Jose, CA Relocation and housing will not...InternshipHourly payFull timeSummer internshipLocal areaRelocation$152k - $241.5k
...applications? Join NVIDIA's TensorRT team as a Senior Software Engineer, and be at the forefront of technology, enabling high-performance AI inference solutions for automotive safety and other specialized platforms. Your expertise will help shape the performance, functionality,...Full time$60 per hour
...Ellis Technologies, Inc. is seeking a Researcher Intern for Summer 2026 in San Jose, California. This role involves building impactful features and conducting cutting-edge research in vision and graphics. PhD candidates in Computer Science or related fields with a strong...InternshipHourly paySummer work$124k - $195.5k
NVIDIA is recruiting a Senior Inference Performance Engineer to push NVIDIA's performance limits on large-scale AI inference benchmarks... ....Profile workloads using Nsight Systems, kernel traces, and internal analysis tools. Use roofline and speed-of-light analysis to find...Full time$152k - $241.5k
...optimizations for next-generation NVIDIA GPUs.Advance the state-of-the-art: Solve complex compilation problems for AI workloads (both inference and training) and successfully transition these breakthroughs into enterprise and consumer products.Collaborate on hardware/...Full time- ...perspectives. Join us as we shape the future of AI and beyond. Together, we advance your career. THE ROLE:We are looking for a Senior GPU Inference Performance Engineer to own end-to-end performance analysis of GPU-accelerated AI inference workloads. You will profile, diagnose,...
$184k - $287.5k
...in a fast-growing technology company that leads the AI revolution.What you will be doing:Implement language and multimodal model inference as part of NVIDIA Inference Microservices (NIMs).Contribute new features, fix bugs and deliver production code to TRT-LLM, NVIDIA’...Full time$184k - $287.5k
We're now looking for a Sr. Inference Engineer, for GPU Kernel Optimization! What does it take to push every LLM inference operation to its performance ceiling? Our LLM Inference Performance Analysis and Optimization team builds the answer from the ground up. We develop...Full time- ...and decode workloads. Our first products are heavily focused on inference . Backed by hundreds of millions from top-tier investors and... ...history. Job Summary Our mechanical/thermal engineering interns will work within the mechanical and thermal engineering...InternshipSummer workSummer internshipWork at officeRelocation
- ...and decode workloads. Our first products are heavily focused on inference . Backed by hundreds of millions from top-tier investors and... ...team. We are looking for Fall '26, Spring '27, and Summer '27 interns. This role requires business acumen, analytical skills, strong...InternshipSummer internshipWork at officeRelocationShift work
$152k - $241.5k
...are now looking for a Senior Software Engineer for Quantized Inference! NVIDIA is seeking software engineers to accelerate the discovery... ...ML layers affect execution timeFamiliarity with PyTorch internals (custom ops, autograd, export) or equivalent frameworkExperience...Full time$152k - $241.5k
...problems. We're seeking talented and motivated engineers to join our TensorRT team in developing the industry-leading deep learning inference software for NVIDIA AI accelerators. As a Senior Software Engineer in the TensorRT team, you will be responsible for designing and...Full time$152k - $241.5k
We are now looking for a Senior Software Engineer for Deep Learning Inference! Would you like to make a big impact in Deep Learning by helping build a state-of-the-art inference framework for accelerating Deep Learning models, especially Large Language Models, on NVIDIA...Full time$152k - $241.5k
NVIDIA is the platform upon which every new AI‑powered application is built. We are seeking a Senior Software Engineer - AI Inference to advance open‑source LLM serving by contributing directly to upstream inference engines like vLLM and SGLang-ensuring they run best‑in...Full timeRemote work- ...accelerating deep learning models, and enabling RL training and SOTA LLM and Multimodal inference at scale across multi-GPU and multi-node systems. You will collaborate across internal GPU software teams and engage with open-source communities to integrate and optimize...
$163k - $253k
...technical presentations and documentation. • Occasional domestic and international travel (<10%).What You BringPh.D. in Computer Science,... ..., including LLMs, DLRMs, and large-scale training and inference systems. • Demonstrated ability to translate workload characteristics...Work at officeFlexible hours$148.7k - $201.2k
...organization, the team is tasked with developing the next generation of EC2 Supercomputers, optimized for high-performance training and inference workloads. We are looking for an experienced System development engineer to drive development for new EC2 machine learning...InternshipLocal areaWork from homeWorldwideFlexible hours- ...and decode workloads. Our first products are heavily focused on inference . Backed by hundreds of millions from top-tier investors and... ...industry in history. Job Summary We’re hiring for a GTM intern - someone who will help us build the operational backbone of...InternshipSummer internshipWork at officeRelocation
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Inference Intern. Be the first to apply!




