Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Member of Technical Staff — Inference-Multi-HardwarePalo Alto, CA

RadixArk

About The Role RadixArk is seeking a Member of Technical Staff: Accelerator Systems to push the limits of performance for frontier AI systems. Most performance engineering assumes a single vendor's stack. This role assumes none. You\'ll bring up, optimize, and maintain SGLang, Miles, and the RadixArk infrastructure stack across NVIDIA and AMD GPUs, Google TPUs, modern server CPUs, and a growing set of emerging AI accelerators. That means porting kernels and runtimes onto unfamiliar hardware, designing the abstractions that keep one codebase fast on all of it. You will be working directly with silicon and our partners, often on pre-release platforms with immature tooling. This is one of the broadest technical roles at RadixArk. The problem changes shape with every new platform: a memory hierarchy that punishes your last set of assumptions, a compiler that fuses differently, a collective library that doesn\'t exist yet. We\'re looking for engineers who find that appealing rather than exhausting, and who can go deep on a new architecture fast without losing the performance instincts they built on the last one. Requirements 4+ years of experience in systems, performance, or ML infrastructure engineering Deep expertise in at least one accelerator programming model (CUDA, ROCm/HIP, Pallas/XLA, Triton, or a vendor SDK), with demonstrated ability to pick up new ones quickly. Strong understanding of accelerator architecture: memory hierarchy, bandwidth limits, occupancy, and the tradeoffs between them Experience writing or optimizing high-performance kernels for ML workloads Experience with distributed execution and communication libraries (NCCL, RCCL, MPI, or equivalents) Proficiency in C++ and Python Strong debugging and profiling skills at the system level, including on platforms where the tooling is incomplete or unreliable Track record of performance work that shipped into production Strong Plus Experience bringing up ML workloads on new silicon. Hands-on depth in more than one vendor ecosystem Experience with compiler stacks (XLA, MLIR, TVM, Triton) or building compiler passes and IR transformations Experience designing hardware abstraction layers or portable kernel interfaces Quantization and mixed-precision work across differing numeric formats and hardware support levels Experience with distributed inference systems (SGLang, vLLM) or training/RL frameworks (Miles, Megatron, veRL, TorchTitan) CPU inference optimization (AVX-512/AMX, oneDNN, NUMA-aware execution) Experience optimizing collective communication at scale, or scaling workloads to 1000+ accelerators Contributions to kernel, compiler, or ML systems open source Direct collaboration with silicon vendors or cloud partners on technical evaluations Background in HPC or other performance-critical systems Responsibilities Bring up RadixArk\'s inference and training systems on new accelerator platforms and drive them to competitive performance Design hardware abstractions that let a single codebase stay fast across vendors without forking Port and optimize kernels across programming models and memory architectures Build cross-platform benchmarking, profiling, and regression detection so performance claims hold up on every target Debug numerical divergence and correctness gaps between platforms Work with vendor engineering teams on pre-release hardware, compiler and driver issues, and roadmap feedback Partner with kernel, runtime, distributed systems, and product engineers to land performance wins end to end Serve as the internal source of truth on what each platform is actually good at Contribute hardware-specific optimizations, benchmarks, and portability work back to open-source SGLang and Miles About RadixArk RadixArk is an infrastructure-first company built by engineers who\'ve shipped production AI systems, created SGLang (30K+ GitHub stars, the fastest open LLM serving engine), and developed Miles (our large-scale RL framework). We\'re on a mission to democratize frontier-level AI infrastructure by building world-class open systems for inference and training. Our team has optimized kernels serving billions of tokens daily, designed distributed training systems coordinating 10,000+ GPUs, and contributed to infrastructure that powers leading AI companies and research labs. We\'re backed by well-known infrastructure investors and partner with Nvidia, Google, and frontier AI labs. Join us in building infrastructure that gives real leverage back to the AI community. Compensation We offer competitive compensation with meaningful equity, comprehensive benefits, and flexible work arrangements. Compensation depends on location, experience, and level. Equal Opportunity RadixArk is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more. #J-18808-Ljbffr RadixArk

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Member of Technical Staff — Inference-Multi-HardwarePalo Alto, CA in Palo Alto, CA vacancy
  • Member of Technical Staff — Kernel / Compiler / Communication RadixArk is seeking a deeply technical engineer who pushes the limits of performance...  .... This role is critical to scaling training and inference across thousands of GPUs, where microseconds and memory bandwidth... 
    Suggested
    Flexible hours

    RadixArk

    Palo Alto, CA
    1 day ago
  • About The Role RadixArk is seeking a Member of Technical Staff — Diffusion Model to advance the frontier of generative modeling. You will work...  ...training Collaborate with systems teams to scale training and inference Translate research ideas into practical production... 
    Suggested
    Flexible hours

    RadixArk

    Palo Alto, CA
    1 day ago
  • $324k - $396k

     ...communication skills. They should be able to concisely and accurately share knowledge with their teammates. Member of Technical Staff (X.AI LLC; Palo Alto, CA): Introduce innovative techniques and analyses to theAI field to facilitate breakthroughs in quantitative reasoning... 
    Suggested
    Remote work

    Neura Market

    Palo Alto, CA
    1 day ago
  • About The Role RadixArk is seeking a Member of Technical Staff, Developer Technology (DevTech) to make LLM inference and training dramatically faster, cheaper, and more accessible...  ...bottlenecks from kernels to distributed multi-node systems. Go deep in one or two focus... 
    Suggested
    Flexible hours

    RadixArk

    Palo Alto, CA
    5 days ago
  • About The Role RadixArk is seeking a Member of Technical Staff — Training to build and scale the systems that train frontier AI models. You will...  ...post-training systems. Experience working on training / inference correctness or other precision-related problem Experience... 
    Suggested
    Flexible hours

    RadixArk

    Palo Alto, CA
    3 days ago
  • About The Role RadixArk is hiring a Member of Technical Staff — CI Engineer to own the infrastructure that keeps SGLang moving. Our CI system...  ...every commit to one of the fastest‑growing open‑source LLM inference engines. When CI is green and fast, 100+ contributors ship... 
    Flexible hours
    Night shift

    RadixArk

    Palo Alto, CA
    1 day ago
  • RadixArk is seeking a Member of Technical Staff — Inference to push the limits of large-scale AI inference. You will work on the core systems that serve frontier models at scale, optimizing performance, latency, throughput, and cost across thousands of GPUs. This role... 
    Worldwide
    Flexible hours

    Dormont Manufacturing Co

    Palo Alto, CA
    4 days ago
  • $180k

    SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging...
    Temporary work

    Neura Market

    Palo Alto, CA
    3 days ago
  • $148.5k - $223.9k

    Senior Member of Technical Staff - AI ResearchSkip to main content#Senior Member...  ...Flexiblelocations: California - Palo Alto: California - San...  ...tool use, planning, memory, multi-step reasoning)** *...  ...training, evaluation, and inference pipelines** *Infrastructure... 
    Work at office

    Salesforce, Inc.

    Palo Alto, CA
    5 days ago
  • $150k - $230k

     ...early-stage AI startup is looking for a Member of Technical Staff to help build and scale real AI-...  ...and communication platforms). Optimize inference costs and performance across large language...  ...This role is fully on-site in Palo Alto, California , Monday through Friday. Remote... 
    Full time
    Internship
    Immediate start
    Remote work
    Visa sponsorship
    Monday to Friday

    Jobr

    Palo Alto, CA
    19 hours ago
  • $13 per hour

     ...deeply curious Wants to own features from design to development to deployment to maintenance Is willing to put the work in to solve the hardest of problems Location: Palo Alto , CA Base Salary Range: $140,000/yr to $220,000/yr + Equity + Benefits #J-18808-Ljbffr Pylon

    Pylon

    Palo Alto, CA
    2 days ago
  • $13 per hour

     ...deeply curious Wants to own features from design to development to deployment to maintenance Is willing to put the work in to solve the hardest of problems Location: Palo Alto , CA Base Salary Range: $140,000/yr to $220,000/yr + Equity + Benefits #J-18808-Ljbffr Pylon

    Pylon

    Palo Alto, CA
    2 days ago
  • Job Description MosaixSoft, Inc. is recruiting for our Los Altos, CA office: Member of the Technical Staff (job code #37659). Design, architect, and implement software projects by studying and predicting information needs, system flow, resource usage, and the scalability... 
    Work at office

    MosaixSoft, Inc.

    Los Altos, CA
    4 days ago
  • $324k - $396k

    About the Role Member of Technical Staff (X.AI LLC; Palo Alto, CA): Introduce innovative techniques and analyses to the AI field to facilitate breakthroughs in quantitative reasoning and language understanding. Stabilize large language model training, pipeline parallelism... 
    Remote work

    Xai

    Palo Alto, CA
    4 days ago
  •  ...their teammates. ABOUT THE ROLE: The RL infrastructure team is looking for an engineer to help with low precision RL training and inference. RESPONSIBILITIES: Design and optimize our inference stack for all shapes of RL workloads at xAI, from small scale ablations to... 

    Pantera Capital

    Palo Alto, CA
    1 day ago
  • $180k

    Member of Technical Staff - Multimodal Understanding About xAI xAI’s mission is to create AI systems that...  ...pre‑training, post‑training, inference, data processing, and tokenization at...  ...inference optimisation, GPU utilisation, multi‑GPU/TPU setups, hardware co‑design).... 
    Temporary work

    xAI

    Palo Alto, CA
    5 days ago
  • We are looking for a Member of Technical Staff with strong Python skills and a passion for building scalable platforms for AI and ML workloads. As MTS, you'll influence strategic decisions, partner closely with the founding team, and play a critical role in shaping Activeloop... 

    S27a

    Mountain View, CA
    2 days ago
  • Member of Technical Staff, ProductsProducts · Palo Alto · full-time Build the product surface through which enterprises actually use AI agents. About Sycamore...  ..., real impact: Trust architectures, memory systems, multi-agent coordination. The foundational layer that makes... 
    Full time
    Contract work

    Sycamore

    Palo Alto, CA
    4 days ago
  • RadixArk is hiring a Performance Engineer in Palo Alto, CA — someone who can push LLM inference and training systems to the limit across real production...  ...efficiency Partner with customers and cloud partners on deep technical evaluations Contribute performance insights back to... 
    Flexible hours

    RadixArk

    Palo Alto, CA
    3 days ago
  • $180k

     ...you will develop and manage training and inference clusters, as well as highly reliable...  ...This is an in‑person role based in Palo Alto, CA or Washington, DC, with up to 50% travel...  ...curiosity, and enthusiasm for tackling complex technical challenges in secure environments.... 
    Temporary work

    Pantera Capital

    Palo Alto, CA
    3 days ago
  • $230k

     ...industry-leading training and inference speeds and empowers machine...  ...multiple openings for Sr. Member of Technical Staff. Title: Sr. Member of Technical...  ...techniques (e.g., multi-threading, asynchronous processing...  ...E Arques Avenue, Sunnyvale, CA 94085 Telecommuting permitted... 
    Remote work

    Cerebras Systems

    Sunnyvale, CA
    5 days ago
  • $120k - $155k

     ...Location: San Francisco Bay Area (offices in San Francisco and Los Altos, CA). Responsibilities Network and take part in industry and VC...  ...ambiguity and risk. Ability to work collaboratively with team members. Track record of successfully completing projects with minimal... 
    Work experience placement
    Local area

    VC Stack

    Los Altos, CA
    3 days ago
  • $29.15 - $43.73 per hour

     ...in Dearborn, Mich., and Palo Alto, Calif. Meet the team: The Mission...  ...Analyst is a senior-level technical role responsible for executing...  ...coaching and feedback to team members to ensure efficiency and a culture...  ...exceptional ability to follow multi‑step instructions and maintain... 
    Hourly pay
    Permanent employment
    Contract work
    Long distance
    Shift work
    Rotating shift

    Doist

    Palo Alto, CA
    1 day ago
  •  ...services/APIs) Own projects full lifecycle Design discussions & technical scoping Implementation & testing Post-launch iteration...  ...Authorization Location: Hybrid - we’re in the office in Palo Alto, CA near the Caltrain station from Tuesday to Thursday , with Mondays... 
    Work at office
    Remote work
    Visa sponsorship
    Monday to Friday

    Sierra Ventures

    Palo Alto, CA
    2 days ago
  •  ...AI Startup in Stealth | Mountain View, CA AI Startup in Stealth is a full-stack semiconductor company founded by pioneers...  ...investors and strategic partners. We are looking for an exceptional Member of Technical Staff to help design, build, and scale core components of our next... 

    DensityAI

    Mountain View, CA
    2 days ago
  • $180k

     ...and scale robust, high-performance systems that power immersive, multi-modal media interactions—leveraging cutting-edge AI to enable...  ...PREFERRED SKILLS AND EXPERIENCE: Experience with real-time systems, inference serving, or multi-modal data processing at scale.... 
    Temporary work
    Worldwide

    SpaceXAI

    Palo Alto, CA
    a month ago
  • Member Of Technical Staff - Extreme-Scale Sparse Linear Algebra, Domain Decomposition & GPU Solver Architecture Vinci | Full-Time | Remote / Hybrid...  ...GPU-accelerated sparse linear algebra (CUDA + HIP) Multi-GPU and distributed execution paradigms You think about:... 
    Full time
    Remote work

    Vinci4d

    Palo Alto, CA
    3 days ago
  • $180k

     ...‑scale orchestration and virtualization — to make training and inference at xAI as fast, reliable, and scalable as possible. This is a...  ...across GPU memory hierarchy, networking fabric, filesystems, and multi‑GPU operations Create and maintain infrastructure‑as‑code,... 
    Temporary work

    xAI

    Palo Alto, CA
    1 day ago
  • Member of Technical Staff, Infrastructure Build and operate the secure execution substrate for enterprise agents, customer applications, and Sycamore...  ...that you have personally operated and recovered important multi-tenant systems. What we are looking for 5-12 years of... 

    Sycamore

    Palo Alto, CA
    4 days ago
  • $180k - $250k

    Member of Technical Staff -- TPU Systems (JAX / XLA / PALLAS) About the Role RadixArk is looking for a TPU Systems Engineer to build high-performance inference and training systems using JAX, XLA, and Pallas. You'll push large-model workloads to their limits on TPU hardware... 
    Full time
    Flexible hours

    RadixArk

    Palo Alto, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Member of Technical Staff — Inference-Multi-HardwarePalo Alto, CA. Be the first to apply!