Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Member of Technical Staff — Performance

RadixArk

RadixArk is hiring a Performance Engineer in Palo Alto, CA — someone who can push LLM inference and training systems to the limit across real production workloads. You’ll work on the performance-critical path of SGLang, Miles, and the RadixArk infrastructure stack: latency, throughput, GPU utilization, memory efficiency, scheduling, batching, kernel behavior, distributed execution, and cost-per-token. This is not a generic benchmarking role. You’ll be working on the systems that determine whether frontier-scale AI workloads are actually usable, affordable, and reliable in production. Our customers care about real numbers: P99 latency, TTFT, tokens/sec/GPU, throughput under long-context workloads, cost-per-million tokens, RL rollout efficiency, and training-inference consistency. You’ll help us measure, debug, and improve these systems across NVIDIA, AMD, Google TPU, and cloud partner environments. This role is for someone who loves performance debugging, understands that small systems details can create massive product impact, and wants to work at the frontier of AI infrastructure. What You'll Do Analyze and improve performance across SGLang, Miles, and RadixArk production deployments Benchmark LLM inference and training workloads across GPUs, TPUs, and cloud environments Optimize latency, throughput, memory usage, batching, scheduling, routing, and GPU utilization Investigate performance regressions in real customer environments Work closely with kernel, runtime, distributed systems, and product engineers Build internal tooling for profiling, tracing, benchmarking, and regression detection Translate customer workload characteristics into concrete performance tuning strategies Help define performance metrics that matter commercially, including cost-per-token and serving efficiency Partner with customers and cloud partners on deep technical evaluations Contribute performance insights back to open-source SGLang and Miles What We're Looking For Strong systems engineering background, especially in performance-critical software Experience with GPU systems, distributed systems, inference serving, ML runtimes, or high-performance computing Familiarity with profiling tools, performance debugging, tracing, and benchmark methodology Comfort working with Python and C++ Experience with CUDA, Triton, Pallas, ROCm, XLA, or kernel-level optimization is a strong plus Understanding of LLM inference concepts such as batching, KV cache, prefill/decode, speculative decoding, MoE, long context, and P99 latency Ability to debug messy real-world performance issues across software, hardware, and infrastructure layers Strong communication skills — you should be able to explain performance tradeoffs to both engineers and customers Prior experience with production AI infrastructure, cloud GPU environments, or open-source ML systems is a plus About RadixArk RadixArk is an infrastructure-first company built by engineers who've shipped production AI systems, created SGLang, and developed Miles, our large-scale RL framework. We're on a mission to democratize frontier-level AI infrastructure by building world-class open systems for inference and training. Our team has optimized kernels serving billions of tokens daily, designed distributed training systems coordinating 10,000+ GPUs, and contributed to infrastructure that powers leading AI companies and research labs. We're backed by well-known infrastructure investors and partner with Nvidia, Google, AWS, and frontier AI labs. Join us in building infrastructure that gives real leverage back to the AI community. Compensation We offer competitive compensation with meaningful equity, comprehensive health benefits, and flexible work arrangements. Compensation is determined by location, level, and experience. RadixArk is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more. #J-18808-Ljbffr RadixArk

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Member of Technical Staff — Performance in Palo Alto, CA vacancy
  •  ...About the Role As a Member of Technical Staff [Research] at NeoCognition , you’ll be part of the core team advancing the frontier of LLM agents...  ...experiences. Design and execute experiments, benchmark performance, and analyze model behaviors to identify failure modes and... 
    Performance

    NeoCognition Inc.

    Palo Alto, CA
    13 hours ago
  •  ...About the Role As a Member of Technical Staff [Platform] at NeoCognition , you’ll design and build the internal systems that power everything...  ...prototype to production. Implement and monitor observability and performance metrics across services and pipelines. Collaborate closely... 
    Performance

    NeoCognition

    Palo Alto, CA
    5 days ago
  •  ...MosaixSoft, Inc. is recruiting for our Los Altos, CA office: Member of the Technical Staff (job code #37659). Design, architect, and implement...  ...test coverage. Analyze and debug issues of various nature (performance issues, software bugs, problems related to high resource... 
    Performance
    Work at office

    MosaixSoft, Inc.

    Los Altos, CA
    5 days ago
  •  ...SuperIntelligence, xAI, Apple and Intel. What You’ll Do As Member of the Technical Staff - Software at Architect, you’ll build the core platform...  ...design. You’ll own the full stack, developing intuitive, high‑performance systems using TypeScript, Python, and modern frameworks,... 
    Performance

    Architect Labs

    Palo Alto, CA
    5 days ago
  • RadixArk is seeking a Member of Technical Staff — Inference to push the limits of large-scale AI inference. You will work on the core systems that serve frontier models at scale, optimizing performance, latency, throughput, and cost across thousands of GPUs. This role... 
    Performance
    Worldwide
    Flexible hours

    Dormont Manufacturing Co

    Palo Alto, CA
    3 days ago
  • Member of Technical Staff — Kernel / Compiler / Communication About the Role RadixArk is seeking a Member of Technical Staff — Kernel / Compiler / Communication to push the limits of performance for frontier AI systems. You will work at the lowest layers of the stack... 
    Performance
    Flexible hours

    RadixArk

    Palo Alto, CA
    3 days ago
  •  ...SuperIntelligence, xAI, Apple and Intel. What You'll Do As a Founding Member of the Technical Staff on the RTL Design team at Architect, you'll own the AI-...  ...on research and development on energy-efficient, high-performance HW accelerators on your block-of-expertise. What We... 
    Performance

    Architect

    Palo Alto, CA
    5 days ago
  • RadixArk is seeking a Member of Technical Staff — Training to build and scale the systems that train frontier AI models. You will work on large...  .... This role sits at the intersection of ML, systems, and performance engineering. Your work will directly impact how next-... 
    Performance
    Flexible hours

    RadixArk

    Palo Alto, CA
    5 days ago
  • We are looking for a Member of Technical Staff with strong Python skills and a passion for building scalable platforms for AI and ML workloads...  ...build scalable software infrastructure optimized for high-performance AI and ML workloads, ensuring robustness and maintainability... 
    Performance

    S27a

    Mountain View, CA
    1 day ago
  • About the Role RadixArk is looking for a Member of Technical Staff (Cluster / Platform) to architect and scale the core compute platform that...  ...inference. You will design and operate highly reliable, high-performance GPU/TPU clusters, build next-generation scheduling and... 
    Performance
    Flexible hours

    RadixArk

    Palo Alto, CA
    1 day ago
  •  ...and strategic partners. We are looking for an exceptional Member of Technical Staff to help design, build, and scale core components of our...  ...silicon, systems, runtime, and infrastructure. Emphasis on performance, scalability, and reliability. AI Compute Systems Development... 
    Performance

    DensityAI

    Mountain View, CA
    3 days ago
  •  ...SuperIntelligence, xAI, Apple and Intel. What You’ll Do As a Founding Member of the Technical Staff at Architect, you'll be at the forefront of training AI...  ...fine-tuning and evaluation, ensuring that theoretical performance translates into production-ready implementations. This is... 
    Performance

    Architect Labs

    Palo Alto, CA
    5 days ago
  • Member of Technical Staff, LLM Post-Training, Applied Sanas is pioneering the future of human communication. Founded by a team of Stanford researchers...  ..., alignment, and inference optimization — into models that perform reliably in high‑stakes, real‑world, on‑premise... 
    Performance

    Sanas

    Palo Alto, CA
    3 days ago
  • Member of Technical Staff Physical AI (Robotics / World Models) Palo Alto, CA About Orbifold AI Orbifold AI advances the frontier of physical...  ...This is a highly applied role focused on real-world system performance , not purely theoretical research. Key Responsibilities... 
    Performance
    Shift work

    Bonfirevc

    Palo Alto, CA
    5 days ago
  • Member of Technical Staff — Accelerator Systems About the Role RadixArk is seeking a Member of Technical Staff: Accelerator Systems to push the limits of performance for frontier AI systems. Most performance engineering assumes a single vendor's stack. This role assumes... 
    Performance
    Flexible hours

    RadixArk

    Palo Alto, CA
    21 hours ago
  •  ...think everyone should be able to make them. As a Frontend Member of Technical Staff (MTS), you’ll help make that a reality. You’ll own the creation...  ...— we care about taste, not just correctness Optimize for performance at real consumer scale (tens of millions of users, high-... 
    Performance

    Astrocade

    Palo Alto, CA
    2 days ago
  •  ...SuperIntelligence, xAI, Apple and Intel. What You’ll Do As a Founding Member of the Technical Staff on the RTL Design team at Architect, you’ll own the AI-...  .... Python : Strong skills for design automation, performance modeling, regression infrastructure, and tooling. PPA... 
    Performance

    Kindredventures

    Palo Alto, CA
    3 days ago
  • $180k

    Member of Technical Staff, Pre-training Data Infrastructure xAI’s mission is to create AI systems that can accurately understand the universe...  ...troubleshooting complex distributed data processing systems for maximum performance. Building bespoke data processing systems from scratch.... 
    Performance
    Temporary work
    Relocation

    xAI

    Palo Alto, CA
    5 days ago
  • $180k

     ...in reasoning and utility. Architect high-performance systems for personalized, reliable...  ...interview (“phone interview”) during which a member of our team will ask some basic questions...  ...the main process, which consists of 2 technical interviews and 1 project deep-dive interview... 
    Performance
    Temporary work

    Pantera Capital

    Palo Alto, CA
    3 days ago
  • $180k - $250k

    Member of Technical Staff -- TPU Systems (JAX / XLA / PALLAS) About the Role RadixArk is looking for a TPU Systems Engineer to build high-performance inference and training systems using JAX, XLA, and Pallas. You'll push large-model workloads to their limits on TPU hardware... 
    Performance
    Full time
    Flexible hours

    RadixArk

    Palo Alto, CA
    3 days ago
  • $120k - $200k

     ...curated datasets, or full-cycle data engineering, Abaka AI provides the foundation for building high-performance AI systems. About the Role As a Member of Technical Staff, Platform, you'll build full-stack product features for our Expert Talent platform end-to-end—from a... 
    Performance
    Flexible hours

    Embedding VC

    Mountain View, CA
    5 days ago
  • $500 per month

     ...think everyone should be able to make them. As a Frontend Member of Technical Staff (MTS), you’ll help make that a reality. You’ll own the creation...  ...- we care about taste, not just correctness Optimize for performance at real consumer scale (tens of millions of users, high-... 
    Performance
    Work at office
    Flexible hours

    Astrocade

    Palo Alto, CA
    1 day ago
  • Member of Technical Staff — Developer Technology About the Role RadixArk is seeking a Member of Technical Staff, Developer Technology (DevTech...  ...of how modern AI is served and trained: SGLang is a high-performance inference engine that serves trillions of tokens daily across... 
    Performance
    Flexible hours

    RadixArk

    Palo Alto, CA
    3 days ago
  • $300k - $350k

    Member of Technical Staff Level 1 - Engineering Bellevue | Hybrid NTT DATA AIVista, Inc., a wholly owned subsidiary of NTT DATA, is based in Silicon...  ...architectural and implementation decisions that balance performance, scalability, reliability, cost, and customer requirements... 
    Performance
    Work experience placement
    Local area
    Flexible hours

    NTT DATA AIVista

    Palo Alto, CA
    2 days ago
  •  ...Automate infrastructure provisioning, configuration, monitoring, and alerting using Infrastructure as Code (IaC) principles.Drive performance tuning, cost optimization, and reliability improvements across the entire stack.Collaborate closely with researchers and product... 
    Performance

    Odyssey

    Palo Alto, CA
    12 hours ago
  •  ...Native Take features from initial concept and technical design through production launch Solve complex frontend performance, scalability, and reliability challenges Build...  ...years of frontend/software engineering experience Staff/L6-level engineering capability Deep... 
    Performance
    H1b

    Jobtailor

    Palo Alto, CA
    2 days ago
  • $180k

     ...and reward models tailored to image/video/audio quality and coherence. Implement efficient algorithms for state‑of‑the‑art model performance, including real‑time inference, distillation, and scalable serving for visual content. Develop scalable data collection and processing... 
    Performance
    Temporary work

    Xai

    Palo Alto, CA
    5 days ago
  •  ...infrastructure. Responsibilities As a senior member of the AI Infrastructure engineering...  ...degree in Computer Science or relevant technical field involving coding or equivalent...  ...systems, computer networks, and high-performance applications ~6+ years' experience delivering... 
    Performance
    Flexible hours

    Oracle

    Redwood City, CA
    2 days ago
  • Member of Technical Staff, ML Inference Engineering Sanas is pioneering the future of human communication. Founded by a team of Stanford researchers...  ...the level of the engineers working alongside them. Performance Optimization Optimize system and GPU performance for high... 
    Performance

    Sanas

    Palo Alto, CA
    3 days ago
  •  ...harness powering Perplexity’s flagship answer experience. Improve single- and multi-agent orchestration, context management, performance, and reliability. Translate user needs and quantitative insights into agent improvements. Build excellent observability and... 
    Performance
    Full time

    Perplexity

    Palo Alto, CA
    12 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Member of Technical Staff — Performance. Be the first to apply!