Member of Technical Staff — Inference-Multi-HardwarePalo Alto, CA
RadixArk
About The Role RadixArk is seeking a Member of Technical Staff: Accelerator Systems to push the limits of performance for frontier AI systems. Most performance engineering assumes a single vendor's stack. This role assumes none. You\'ll bring up, optimize, and maintain SGLang, Miles, and the RadixArk infrastructure stack across NVIDIA and AMD GPUs, Google TPUs, modern server CPUs, and a growing set of emerging AI accelerators. That means porting kernels and runtimes onto unfamiliar hardware, designing the abstractions that keep one codebase fast on all of it. You will be working directly with silicon and our partners, often on pre-release platforms with immature tooling. This is one of the broadest technical roles at RadixArk. The problem changes shape with every new platform: a memory hierarchy that punishes your last set of assumptions, a compiler that fuses differently, a collective library that doesn\'t exist yet. We\'re looking for engineers who find that appealing rather than exhausting, and who can go deep on a new architecture fast without losing the performance instincts they built on the last one. Requirements 4+ years of experience in systems, performance, or ML infrastructure engineering Deep expertise in at least one accelerator programming model (CUDA, ROCm/HIP, Pallas/XLA, Triton, or a vendor SDK), with demonstrated ability to pick up new ones quickly. Strong understanding of accelerator architecture: memory hierarchy, bandwidth limits, occupancy, and the tradeoffs between them Experience writing or optimizing high-performance kernels for ML workloads Experience with distributed execution and communication libraries (NCCL, RCCL, MPI, or equivalents) Proficiency in C++ and Python Strong debugging and profiling skills at the system level, including on platforms where the tooling is incomplete or unreliable Track record of performance work that shipped into production Strong Plus Experience bringing up ML workloads on new silicon. Hands-on depth in more than one vendor ecosystem Experience with compiler stacks (XLA, MLIR, TVM, Triton) or building compiler passes and IR transformations Experience designing hardware abstraction layers or portable kernel interfaces Quantization and mixed-precision work across differing numeric formats and hardware support levels Experience with distributed inference systems (SGLang, vLLM) or training/RL frameworks (Miles, Megatron, veRL, TorchTitan) CPU inference optimization (AVX-512/AMX, oneDNN, NUMA-aware execution) Experience optimizing collective communication at scale, or scaling workloads to 1000+ accelerators Contributions to kernel, compiler, or ML systems open source Direct collaboration with silicon vendors or cloud partners on technical evaluations Background in HPC or other performance-critical systems Responsibilities Bring up RadixArk\'s inference and training systems on new accelerator platforms and drive them to competitive performance Design hardware abstractions that let a single codebase stay fast across vendors without forking Port and optimize kernels across programming models and memory architectures Build cross-platform benchmarking, profiling, and regression detection so performance claims hold up on every target Debug numerical divergence and correctness gaps between platforms Work with vendor engineering teams on pre-release hardware, compiler and driver issues, and roadmap feedback Partner with kernel, runtime, distributed systems, and product engineers to land performance wins end to end Serve as the internal source of truth on what each platform is actually good at Contribute hardware-specific optimizations, benchmarks, and portability work back to open-source SGLang and Miles About RadixArk RadixArk is an infrastructure-first company built by engineers who\'ve shipped production AI systems, created SGLang (30K+ GitHub stars, the fastest open LLM serving engine), and developed Miles (our large-scale RL framework). We\'re on a mission to democratize frontier-level AI infrastructure by building world-class open systems for inference and training. Our team has optimized kernels serving billions of tokens daily, designed distributed training systems coordinating 10,000+ GPUs, and contributed to infrastructure that powers leading AI companies and research labs. We\'re backed by well-known infrastructure investors and partner with Nvidia, Google, and frontier AI labs. Join us in building infrastructure that gives real leverage back to the AI community. Compensation We offer competitive compensation with meaningful equity, comprehensive benefits, and flexible work arrangements. Compensation depends on location, experience, and level. Equal Opportunity RadixArk is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more. #J-18808-Ljbffr RadixArk
- Member of Technical Staff — Kernel / Compiler / Communication RadixArk is seeking a deeply technical engineer who pushes the limits of performance... .... This role is critical to scaling training and inference across thousands of GPUs, where microseconds and memory bandwidth...SuggestedFlexible hours
- About The Role RadixArk is seeking a Member of Technical Staff — Diffusion Model to advance the frontier of generative modeling. You will work... ...training Collaborate with systems teams to scale training and inference Translate research ideas into practical production...SuggestedFlexible hours
$324k - $396k
...communication skills. They should be able to concisely and accurately share knowledge with their teammates. Member of Technical Staff (X.AI LLC; Palo Alto, CA): Introduce innovative techniques and analyses to theAI field to facilitate breakthroughs in quantitative reasoning...SuggestedRemote work- About The Role RadixArk is seeking a Member of Technical Staff, Developer Technology (DevTech) to make LLM inference and training dramatically faster, cheaper, and more accessible... ...bottlenecks from kernels to distributed multi-node systems. Go deep in one or two focus...SuggestedFlexible hours
- About The Role RadixArk is seeking a Member of Technical Staff — Training to build and scale the systems that train frontier AI models. You will... ...post-training systems. Experience working on training / inference correctness or other precision-related problem Experience...SuggestedFlexible hours
- About The Role RadixArk is hiring a Member of Technical Staff — CI Engineer to own the infrastructure that keeps SGLang moving. Our CI system... ...every commit to one of the fastest‑growing open‑source LLM inference engines. When CI is green and fast, 100+ contributors ship...Flexible hoursNight shift
- RadixArk is seeking a Member of Technical Staff — Inference to push the limits of large-scale AI inference. You will work on the core systems that serve frontier models at scale, optimizing performance, latency, throughput, and cost across thousands of GPUs. This role...WorldwideFlexible hours
$180k
SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging...Temporary work$148.5k - $223.9k
Senior Member of Technical Staff - AI ResearchSkip to main content#Senior Member... ...Flexiblelocations: California - Palo Alto: California - San... ...tool use, planning, memory, multi-step reasoning)** *... ...training, evaluation, and inference pipelines** *Infrastructure...Work at office$150k - $230k
...early-stage AI startup is looking for a Member of Technical Staff to help build and scale real AI-... ...and communication platforms). Optimize inference costs and performance across large language... ...This role is fully on-site in Palo Alto, California , Monday through Friday. Remote...Full timeInternshipImmediate startRemote workVisa sponsorshipMonday to Friday$13 per hour
...deeply curious Wants to own features from design to development to deployment to maintenance Is willing to put the work in to solve the hardest of problems Location: Palo Alto , CA Base Salary Range: $140,000/yr to $220,000/yr + Equity + Benefits #J-18808-Ljbffr Pylon$13 per hour
...deeply curious Wants to own features from design to development to deployment to maintenance Is willing to put the work in to solve the hardest of problems Location: Palo Alto , CA Base Salary Range: $140,000/yr to $220,000/yr + Equity + Benefits #J-18808-Ljbffr Pylon- Job Description MosaixSoft, Inc. is recruiting for our Los Altos, CA office: Member of the Technical Staff (job code #37659). Design, architect, and implement software projects by studying and predicting information needs, system flow, resource usage, and the scalability...Work at office
$324k - $396k
About the Role Member of Technical Staff (X.AI LLC; Palo Alto, CA): Introduce innovative techniques and analyses to the AI field to facilitate breakthroughs in quantitative reasoning and language understanding. Stabilize large language model training, pipeline parallelism...Remote work- ...their teammates. ABOUT THE ROLE: The RL infrastructure team is looking for an engineer to help with low precision RL training and inference. RESPONSIBILITIES: Design and optimize our inference stack for all shapes of RL workloads at xAI, from small scale ablations to...
$180k
Member of Technical Staff - Multimodal Understanding About xAI xAI’s mission is to create AI systems that... ...pre‑training, post‑training, inference, data processing, and tokenization at... ...inference optimisation, GPU utilisation, multi‑GPU/TPU setups, hardware co‑design)....Temporary work- We are looking for a Member of Technical Staff with strong Python skills and a passion for building scalable platforms for AI and ML workloads. As MTS, you'll influence strategic decisions, partner closely with the founding team, and play a critical role in shaping Activeloop...
- Member of Technical Staff, ProductsProducts · Palo Alto · full-time Build the product surface through which enterprises actually use AI agents. About Sycamore... ..., real impact: Trust architectures, memory systems, multi-agent coordination. The foundational layer that makes...Full timeContract work
- RadixArk is hiring a Performance Engineer in Palo Alto, CA — someone who can push LLM inference and training systems to the limit across real production... ...efficiency Partner with customers and cloud partners on deep technical evaluations Contribute performance insights back to...Flexible hours
$180k
...you will develop and manage training and inference clusters, as well as highly reliable... ...This is an in‑person role based in Palo Alto, CA or Washington, DC, with up to 50% travel... ...curiosity, and enthusiasm for tackling complex technical challenges in secure environments....Temporary work$230k
...industry-leading training and inference speeds and empowers machine... ...multiple openings for Sr. Member of Technical Staff. Title: Sr. Member of Technical... ...techniques (e.g., multi-threading, asynchronous processing... ...E Arques Avenue, Sunnyvale, CA 94085 Telecommuting permitted...Remote work$120k - $155k
...Location: San Francisco Bay Area (offices in San Francisco and Los Altos, CA). Responsibilities Network and take part in industry and VC... ...ambiguity and risk. Ability to work collaboratively with team members. Track record of successfully completing projects with minimal...Work experience placementLocal area$29.15 - $43.73 per hour
...in Dearborn, Mich., and Palo Alto, Calif. Meet the team: The Mission... ...Analyst is a senior-level technical role responsible for executing... ...coaching and feedback to team members to ensure efficiency and a culture... ...exceptional ability to follow multi‑step instructions and maintain...Hourly payPermanent employmentContract workLong distanceShift workRotating shift- ...services/APIs) Own projects full lifecycle Design discussions & technical scoping Implementation & testing Post-launch iteration... ...Authorization Location: Hybrid - we’re in the office in Palo Alto, CA near the Caltrain station from Tuesday to Thursday , with Mondays...Work at officeRemote workVisa sponsorshipMonday to Friday
- ...AI Startup in Stealth | Mountain View, CA AI Startup in Stealth is a full-stack semiconductor company founded by pioneers... ...investors and strategic partners. We are looking for an exceptional Member of Technical Staff to help design, build, and scale core components of our next...
$180k
...and scale robust, high-performance systems that power immersive, multi-modal media interactions—leveraging cutting-edge AI to enable... ...PREFERRED SKILLS AND EXPERIENCE: Experience with real-time systems, inference serving, or multi-modal data processing at scale....Temporary workWorldwide- Member Of Technical Staff - Extreme-Scale Sparse Linear Algebra, Domain Decomposition & GPU Solver Architecture Vinci | Full-Time | Remote / Hybrid... ...GPU-accelerated sparse linear algebra (CUDA + HIP) Multi-GPU and distributed execution paradigms You think about:...Full timeRemote work
$180k
...‑scale orchestration and virtualization — to make training and inference at xAI as fast, reliable, and scalable as possible. This is a... ...across GPU memory hierarchy, networking fabric, filesystems, and multi‑GPU operations Create and maintain infrastructure‑as‑code,...Temporary work- Member of Technical Staff, Infrastructure Build and operate the secure execution substrate for enterprise agents, customer applications, and Sycamore... ...that you have personally operated and recovered important multi-tenant systems. What we are looking for 5-12 years of...
$180k - $250k
Member of Technical Staff -- TPU Systems (JAX / XLA / PALLAS) About the Role RadixArk is looking for a TPU Systems Engineer to build high-performance inference and training systems using JAX, XLA, and Pallas. You'll push large-model workloads to their limits on TPU hardware...Full timeFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Member of Technical Staff — Inference-Multi-HardwarePalo Alto, CA. Be the first to apply!
- life support technician Palo Alto, CA
- personal computer support technician Palo Alto, CA
- systems support technician Palo Alto, CA
- technical support analyst Palo Alto, CA
- user support analyst Palo Alto, CA
- help desk technical support Palo Alto, CA
- technical support specialist Palo Alto, CA
- IT assistant Palo Alto, CA
- help desk assistant Palo Alto, CA
- work from home technical support specialist Palo Alto, CA


