Member of Technical Staff — Inference-Multi-HardwarePalo Alto, CA
RadixArk
About The Role RadixArk is seeking a Member of Technical Staff: Accelerator Systems to push the limits of performance for frontier AI systems. Most performance engineering assumes a single vendor's stack. This role assumes none. You\'ll bring up, optimize, and maintain SGLang, Miles, and the RadixArk infrastructure stack across NVIDIA and AMD GPUs, Google TPUs, modern server CPUs, and a growing set of emerging AI accelerators. That means porting kernels and runtimes onto unfamiliar hardware, designing the abstractions that keep one codebase fast on all of it. You will be working directly with silicon and our partners, often on pre-release platforms with immature tooling. This is one of the broadest technical roles at RadixArk. The problem changes shape with every new platform: a memory hierarchy that punishes your last set of assumptions, a compiler that fuses differently, a collective library that doesn\'t exist yet. We\'re looking for engineers who find that appealing rather than exhausting, and who can go deep on a new architecture fast without losing the performance instincts they built on the last one. Requirements 4+ years of experience in systems, performance, or ML infrastructure engineering Deep expertise in at least one accelerator programming model (CUDA, ROCm/HIP, Pallas/XLA, Triton, or a vendor SDK), with demonstrated ability to pick up new ones quickly. Strong understanding of accelerator architecture: memory hierarchy, bandwidth limits, occupancy, and the tradeoffs between them Experience writing or optimizing high-performance kernels for ML workloads Experience with distributed execution and communication libraries (NCCL, RCCL, MPI, or equivalents) Proficiency in C++ and Python Strong debugging and profiling skills at the system level, including on platforms where the tooling is incomplete or unreliable Track record of performance work that shipped into production Strong Plus Experience bringing up ML workloads on new silicon. Hands-on depth in more than one vendor ecosystem Experience with compiler stacks (XLA, MLIR, TVM, Triton) or building compiler passes and IR transformations Experience designing hardware abstraction layers or portable kernel interfaces Quantization and mixed-precision work across differing numeric formats and hardware support levels Experience with distributed inference systems (SGLang, vLLM) or training/RL frameworks (Miles, Megatron, veRL, TorchTitan) CPU inference optimization (AVX-512/AMX, oneDNN, NUMA-aware execution) Experience optimizing collective communication at scale, or scaling workloads to 1000+ accelerators Contributions to kernel, compiler, or ML systems open source Direct collaboration with silicon vendors or cloud partners on technical evaluations Background in HPC or other performance-critical systems Responsibilities Bring up RadixArk\'s inference and training systems on new accelerator platforms and drive them to competitive performance Design hardware abstractions that let a single codebase stay fast across vendors without forking Port and optimize kernels across programming models and memory architectures Build cross-platform benchmarking, profiling, and regression detection so performance claims hold up on every target Debug numerical divergence and correctness gaps between platforms Work with vendor engineering teams on pre-release hardware, compiler and driver issues, and roadmap feedback Partner with kernel, runtime, distributed systems, and product engineers to land performance wins end to end Serve as the internal source of truth on what each platform is actually good at Contribute hardware-specific optimizations, benchmarks, and portability work back to open-source SGLang and Miles About RadixArk RadixArk is an infrastructure-first company built by engineers who\'ve shipped production AI systems, created SGLang (30K+ GitHub stars, the fastest open LLM serving engine), and developed Miles (our large-scale RL framework). We\'re on a mission to democratize frontier-level AI infrastructure by building world-class open systems for inference and training. Our team has optimized kernels serving billions of tokens daily, designed distributed training systems coordinating 10,000+ GPUs, and contributed to infrastructure that powers leading AI companies and research labs. We\'re backed by well-known infrastructure investors and partner with Nvidia, Google, and frontier AI labs. Join us in building infrastructure that gives real leverage back to the AI community. Compensation We offer competitive compensation with meaningful equity, comprehensive benefits, and flexible work arrangements. Compensation depends on location, experience, and level. Equal Opportunity RadixArk is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more. #J-18808-Ljbffr RadixArk
- About The Role RadixArk is seeking a Member of Technical Staff — Diffusion Model to advance the frontier of generative modeling. You will work... ...training Collaborate with systems teams to scale training and inference Translate research ideas into practical production...SuggestedFlexible hours
- About The Role RadixArk is seeking a Member of Technical Staff, Developer Technology (DevTech) to make LLM inference and training dramatically faster, cheaper, and more accessible... ...bottlenecks from kernels to distributed multi-node systems. Go deep in one or two focus...SuggestedFlexible hours
- About The Role RadixArk is hiring a Member of Technical Staff — CI Engineer to own the infrastructure that keeps SGLang moving. Our CI system... ...every commit to one of the fastest‑growing open‑source LLM inference engines. When CI is green and fast, 100+ contributors ship...SuggestedFlexible hoursNight shift
- ...RadixArk is seeking a Developer Advocate to build and engage our technical community around SGLang, Miles, and our open source... ...level AI infrastructure by building world-class open systems for inference and training. Our team has optimized kernels serving billions...SuggestedFlexible hours
- RadixArk is seeking a Member of Technical Staff — Inference to push the limits of large-scale AI inference. You will work on the core systems that serve frontier models at scale, optimizing performance, latency, throughput, and cost across thousands of GPUs. This role...SuggestedWorldwideFlexible hours
$180k - $250k
Member of Technical Staff -- TPU Systems (JAX / XLA / PALLAS) About the Role RadixArk is looking for a TPU Systems Engineer to build high-performance inference and training systems using JAX, XLA, and Pallas. You'll push large-model workloads to their limits on TPU hardware...Full timeFlexible hours- Member of Technical Staff, ML Inference Engineering Sanas is pioneering the future of human communication. Founded by a team of Stanford researchers and... ...GPU performance for high-throughput AI workloads across multi-node training and inference Analyze and improve latency...
- Member of Technical Staff — Kernel / Compiler / Communication About the Role RadixArk is seeking a Member of Technical Staff — Kernel / Compiler... .... This role is critical to scaling training and inference across thousands of GPUs, where microseconds and memory bandwidth...Flexible hours
- ...abstraction. The Role As a Member of Technical Staff, AI Training Platform, you... ...infrastructure. Design and scale multi-node distributed training... ...performance tuning, LLM inference, CUDA, Triton, NCCL, vLLM,... ...when working from our Palo Alto office. #J-18808-Ljbffr Unconventional...Work at officeFlexible hours
$150k
...Member of Technical Staff Location: Palo Alto, CA Company Stage of Funding: Seed Stage AI Startup ($27M Raised) Office Type: Onsite (5 Days Per Week) Salary: $150,000 Base + Significant Equity ($250K-$750K+ Total Annual Compensation) Company Description...Work at officeVisa sponsorshipRelocation package- ...Job Description MosaixSoft, Inc. is recruiting for our Los Altos, CA office: Member of the Technical Staff (job code #37659). Design, architect, and implement software projects by studying and predicting information needs, system flow, resource usage, and the scalability...Work at office
$180k
Member of Technical Staff - Multimodal Understanding SpaceXAI’s mission is to create AI systems that can... ...pre‑training, post‑training, inference, data processing, and tokenization at... ...inference optimization, GPU utilization, multi‑GPU/TPU setups, hardware co‑design). Deep...Temporary work- We are looking for a Member of Technical Staff with strong Python skills and a passion for building scalable platforms for AI and ML workloads. As MTS, you'll influence strategic decisions, partner closely with the founding team, and play a critical role in shaping Activeloop...
- ...their teammates. ABOUT THE ROLE: The RL infrastructure team is looking for an engineer to help with low precision RL training and inference. RESPONSIBILITIES: Design and optimize our inference stack for all shapes of RL workloads at xAI, from small scale ablations to...
$180k
...you will develop and manage training and inference clusters, as well as highly reliable... ...This is an in‑person role based in Palo Alto, CA or Washington, DC, with up to 50% travel... ...curiosity, and enthusiasm for tackling complex technical challenges in secure environments....Temporary work- Member of Technical Staff, LLM Post-Training, Applied Sanas is pioneering the future of human communication... ...— through fine‑tuning, alignment, and inference optimization — into models that... ...LoRA adapters within one model instance (Multi‑LoRA), or models distributed across...
- RadixArk is hiring a Performance Engineer in Palo Alto, CA — someone who can push LLM inference and training systems to the limit across real production... ...efficiency Partner with customers and cloud partners on deep technical evaluations Contribute performance insights back to...Flexible hours
- Member of Technical Staff Physical AI (Robotics / World Models) Palo Alto, CA About Orbifold AI Orbifold AI advances the frontier of physical AI and world model companies through rigorous evaluation and curated, real-world data. We work directly with leading robotics and...Shift work
$200k - $400k
Member of Technical Staff — Cluster Infrastructure & Supercomputing About the Role RadixArk is looking... ...powers frontier-level AI training and inference. You will design and operate highly reliable... ...systems Experience debugging complex multi-layer issues across hardware, OS,...$29.15 - $43.73 per hour
...in Dearborn, Mich., and Palo Alto, Calif. Meet the team: The Mission... ...Analyst is a senior-level technical role responsible for executing... ...coaching and feedback to team members to ensure efficiency and a culture... ...exceptional ability to follow multi‑step instructions and maintain...Hourly payPermanent employmentContract workLong distanceShift workRotating shift$120k - $155k
...Location: San Francisco Bay Area (offices in San Francisco and Los Altos, CA). Responsibilities Network and take part in industry and VC... ...ambiguity and risk. Ability to work collaboratively with team members. Track record of successfully completing projects with minimal...Work experience placementLocal area- ...services/APIs) Own projects full lifecycle Design discussions & technical scoping Implementation & testing Post-launch iteration... ...Location: Hybrid - we’re in the office in Palo Alto, CA near the Caltrain station from Tuesday to Thursday , with Mondays...Work at officeRemote workVisa sponsorshipMonday to Friday
$180k
...and scale robust, high-performance systems that power immersive, multi-modal media interactions—leveraging cutting-edge AI to enable... ...PREFERRED SKILLS AND EXPERIENCE: Experience with real-time systems, inference serving, or multi-modal data processing at scale....Temporary workWorldwide$300k - $350k
Member of Technical Staff Level 1 - Engineering Bellevue | Hybrid NTT DATA AIVista, Inc., a wholly owned subsidiary... ...Models (LLMs) AI infrastructure and inference systems Data processing and knowledge pipelines Agent protocols and multi‑agent architectures Demonstrated ability...Work experience placementLocal areaFlexible hours- AI Startup in Stealth | Mountain View, CA AI Startup in Stealth is a full-stack semiconductor company founded by pioneers... ...and strategic partners. We are looking for an exceptional Member of Technical Staff to help design, build, and scale core components of our next...
$140k - $200k
...for building high-performance AI systems. About the Role As a Member of Technical Staff, Infra, you'll own the scalability and reliability of the... ...dental, vision, PTO, flexible work schedule). This position is based onsite in Mountain View, CA. #J-18808-Ljbffr Abaka AIFlexible hours$140k - $220k
...deeply curious Wants to own features from design to development to deployment to maintenance Is willing to put the work in to solve the hardest of problems Location: Palo Alto, CA • Base Salary Range: $140,000/yr to $220,000/yr + Equity + Benefits #J-18808-Ljbffr Pylon$180k
...‑scale orchestration and virtualization — to make training and inference at xAI as fast, reliable, and scalable as possible. This is a... ...across GPU memory hierarchy, networking fabric, filesystems, and multi‑GPU operations Create and maintain infrastructure‑as‑code,...Temporary work- RadixArk is hiring a Member of Technical Staff — CI Engineer to own the infrastructure that keeps SGLang moving. Our CI system runs 300+ GPU tests... ...commit to one of the fastest-growing open-source LLM inference engines. When CI is green and fast, 100+ contributors ship...Flexible hoursNight shift
- ...About the Role As a Member of Technical Staff [Research] at NeoCognition , you’ll be part of the core team advancing the frontier of LLM agents... ...opportunities. Stay abreast of emerging work in reasoning, multi-agent systems, RLHF, tool use, and LLM fine-tuning — and contribute...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Member of Technical Staff — Inference-Multi-HardwarePalo Alto, CA. Be the first to apply!
- systems support technician Palo Alto, CA
- help desk technical support Palo Alto, CA
- support analyst Palo Alto, CA
- technical analyst Palo Alto, CA
- technical support assistant Palo Alto, CA
- help desk assistant Palo Alto, CA
- IT assistant Palo Alto, CA
- tech aide Palo Alto, CA
- IT help desk technician Palo Alto, CA
- mri tech aide Palo Alto, CA


