Member of Technical Staff — Performance
RadixArk
RadixArk is hiring a Performance Engineer in Palo Alto, CA — someone who can push LLM inference and training systems to the limit across real production workloads. You’ll work on the performance-critical path of SGLang, Miles, and the RadixArk infrastructure stack: latency, throughput, GPU utilization, memory efficiency, scheduling, batching, kernel behavior, distributed execution, and cost-per-token. This is not a generic benchmarking role. You’ll be working on the systems that determine whether frontier-scale AI workloads are actually usable, affordable, and reliable in production. Our customers care about real numbers: P99 latency, TTFT, tokens/sec/GPU, throughput under long-context workloads, cost-per-million tokens, RL rollout efficiency, and training-inference consistency. You’ll help us measure, debug, and improve these systems across NVIDIA, AMD, Google TPU, and cloud partner environments. This role is for someone who loves performance debugging, understands that small systems details can create massive product impact, and wants to work at the frontier of AI infrastructure. What You'll Do Analyze and improve performance across SGLang, Miles, and RadixArk production deployments Benchmark LLM inference and training workloads across GPUs, TPUs, and cloud environments Optimize latency, throughput, memory usage, batching, scheduling, routing, and GPU utilization Investigate performance regressions in real customer environments Work closely with kernel, runtime, distributed systems, and product engineers Build internal tooling for profiling, tracing, benchmarking, and regression detection Translate customer workload characteristics into concrete performance tuning strategies Help define performance metrics that matter commercially, including cost-per-token and serving efficiency Partner with customers and cloud partners on deep technical evaluations Contribute performance insights back to open-source SGLang and Miles What We're Looking For Strong systems engineering background, especially in performance-critical software Experience with GPU systems, distributed systems, inference serving, ML runtimes, or high-performance computing Familiarity with profiling tools, performance debugging, tracing, and benchmark methodology Comfort working with Python and C++ Experience with CUDA, Triton, Pallas, ROCm, XLA, or kernel-level optimization is a strong plus Understanding of LLM inference concepts such as batching, KV cache, prefill/decode, speculative decoding, MoE, long context, and P99 latency Ability to debug messy real-world performance issues across software, hardware, and infrastructure layers Strong communication skills — you should be able to explain performance tradeoffs to both engineers and customers Prior experience with production AI infrastructure, cloud GPU environments, or open-source ML systems is a plus About RadixArk RadixArk is an infrastructure-first company built by engineers who've shipped production AI systems, created SGLang, and developed Miles, our large-scale RL framework. We're on a mission to democratize frontier-level AI infrastructure by building world-class open systems for inference and training. Our team has optimized kernels serving billions of tokens daily, designed distributed training systems coordinating 10,000+ GPUs, and contributed to infrastructure that powers leading AI companies and research labs. We're backed by well-known infrastructure investors and partner with Nvidia, Google, AWS, and frontier AI labs. Join us in building infrastructure that gives real leverage back to the AI community. Compensation We offer competitive compensation with meaningful equity, comprehensive health benefits, and flexible work arrangements. Compensation is determined by location, level, and experience. RadixArk is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more. #J-18808-Ljbffr RadixArk
- ...About the Role As a Member of Technical Staff [Research] at NeoCognition , you’ll be part of the core team advancing the frontier of LLM agents... ...experiences. Design and execute experiments, benchmark performance, and analyze model behaviors to identify failure modes and...Performance
- ...About the Role As a Member of Technical Staff [Platform] at NeoCognition , you’ll design and build the internal systems that power everything... ...prototype to production. Implement and monitor observability and performance metrics across services and pipelines. Collaborate closely...Performance
- ...MosaixSoft, Inc. is recruiting for our Los Altos, CA office: Member of the Technical Staff (job code #37659). Design, architect, and implement... ...test coverage. Analyze and debug issues of various nature (performance issues, software bugs, problems related to high resource...PerformanceWork at office
- ...SuperIntelligence, xAI, Apple and Intel. What You’ll Do As Member of the Technical Staff - Software at Architect, you’ll build the core platform... ...design. You’ll own the full stack, developing intuitive, high‑performance systems using TypeScript, Python, and modern frameworks,...Performance
- RadixArk is seeking a Member of Technical Staff — Inference to push the limits of large-scale AI inference. You will work on the core systems that serve frontier models at scale, optimizing performance, latency, throughput, and cost across thousands of GPUs. This role...PerformanceWorldwideFlexible hours
- Member of Technical Staff — Kernel / Compiler / Communication About the Role RadixArk is seeking a Member of Technical Staff — Kernel / Compiler / Communication to push the limits of performance for frontier AI systems. You will work at the lowest layers of the stack...PerformanceFlexible hours
- ...SuperIntelligence, xAI, Apple and Intel. What You'll Do As a Founding Member of the Technical Staff on the RTL Design team at Architect, you'll own the AI-... ...on research and development on energy-efficient, high-performance HW accelerators on your block-of-expertise. What We...Performance
- RadixArk is seeking a Member of Technical Staff — Training to build and scale the systems that train frontier AI models. You will work on large... .... This role sits at the intersection of ML, systems, and performance engineering. Your work will directly impact how next-...PerformanceFlexible hours
- We are looking for a Member of Technical Staff with strong Python skills and a passion for building scalable platforms for AI and ML workloads... ...build scalable software infrastructure optimized for high-performance AI and ML workloads, ensuring robustness and maintainability...Performance
- About the Role RadixArk is looking for a Member of Technical Staff (Cluster / Platform) to architect and scale the core compute platform that... ...inference. You will design and operate highly reliable, high-performance GPU/TPU clusters, build next-generation scheduling and...PerformanceFlexible hours
- ...and strategic partners. We are looking for an exceptional Member of Technical Staff to help design, build, and scale core components of our... ...silicon, systems, runtime, and infrastructure. Emphasis on performance, scalability, and reliability. AI Compute Systems Development...Performance
- ...SuperIntelligence, xAI, Apple and Intel. What You’ll Do As a Founding Member of the Technical Staff at Architect, you'll be at the forefront of training AI... ...fine-tuning and evaluation, ensuring that theoretical performance translates into production-ready implementations. This is...Performance
- Member of Technical Staff, LLM Post-Training, Applied Sanas is pioneering the future of human communication. Founded by a team of Stanford researchers... ..., alignment, and inference optimization — into models that perform reliably in high‑stakes, real‑world, on‑premise...Performance
- Member of Technical Staff Physical AI (Robotics / World Models) Palo Alto, CA About Orbifold AI Orbifold AI advances the frontier of physical... ...This is a highly applied role focused on real-world system performance , not purely theoretical research. Key Responsibilities...PerformanceShift work
- Member of Technical Staff — Accelerator Systems About the Role RadixArk is seeking a Member of Technical Staff: Accelerator Systems to push the limits of performance for frontier AI systems. Most performance engineering assumes a single vendor's stack. This role assumes...PerformanceFlexible hours
- ...think everyone should be able to make them. As a Frontend Member of Technical Staff (MTS), you’ll help make that a reality. You’ll own the creation... ...— we care about taste, not just correctness Optimize for performance at real consumer scale (tens of millions of users, high-...Performance
- ...SuperIntelligence, xAI, Apple and Intel. What You’ll Do As a Founding Member of the Technical Staff on the RTL Design team at Architect, you’ll own the AI-... .... Python : Strong skills for design automation, performance modeling, regression infrastructure, and tooling. PPA...Performance
$180k
Member of Technical Staff, Pre-training Data Infrastructure xAI’s mission is to create AI systems that can accurately understand the universe... ...troubleshooting complex distributed data processing systems for maximum performance. Building bespoke data processing systems from scratch....PerformanceTemporary workRelocation$180k
...in reasoning and utility. Architect high-performance systems for personalized, reliable... ...interview (“phone interview”) during which a member of our team will ask some basic questions... ...the main process, which consists of 2 technical interviews and 1 project deep-dive interview...PerformanceTemporary work$180k - $250k
Member of Technical Staff -- TPU Systems (JAX / XLA / PALLAS) About the Role RadixArk is looking for a TPU Systems Engineer to build high-performance inference and training systems using JAX, XLA, and Pallas. You'll push large-model workloads to their limits on TPU hardware...PerformanceFull timeFlexible hours$120k - $200k
...curated datasets, or full-cycle data engineering, Abaka AI provides the foundation for building high-performance AI systems. About the Role As a Member of Technical Staff, Platform, you'll build full-stack product features for our Expert Talent platform end-to-end—from a...PerformanceFlexible hours$500 per month
...think everyone should be able to make them. As a Frontend Member of Technical Staff (MTS), you’ll help make that a reality. You’ll own the creation... ...- we care about taste, not just correctness Optimize for performance at real consumer scale (tens of millions of users, high-...PerformanceWork at officeFlexible hours- Member of Technical Staff — Developer Technology About the Role RadixArk is seeking a Member of Technical Staff, Developer Technology (DevTech... ...of how modern AI is served and trained: SGLang is a high-performance inference engine that serves trillions of tokens daily across...PerformanceFlexible hours
$300k - $350k
Member of Technical Staff Level 1 - Engineering Bellevue | Hybrid NTT DATA AIVista, Inc., a wholly owned subsidiary of NTT DATA, is based in Silicon... ...architectural and implementation decisions that balance performance, scalability, reliability, cost, and customer requirements...PerformanceWork experience placementLocal areaFlexible hours- ...Automate infrastructure provisioning, configuration, monitoring, and alerting using Infrastructure as Code (IaC) principles.Drive performance tuning, cost optimization, and reliability improvements across the entire stack.Collaborate closely with researchers and product...Performance
- ...Native Take features from initial concept and technical design through production launch Solve complex frontend performance, scalability, and reliability challenges Build... ...years of frontend/software engineering experience Staff/L6-level engineering capability Deep...PerformanceH1b
$180k
...and reward models tailored to image/video/audio quality and coherence. Implement efficient algorithms for state‑of‑the‑art model performance, including real‑time inference, distillation, and scalable serving for visual content. Develop scalable data collection and processing...PerformanceTemporary work- ...infrastructure. Responsibilities As a senior member of the AI Infrastructure engineering... ...degree in Computer Science or relevant technical field involving coding or equivalent... ...systems, computer networks, and high-performance applications ~6+ years' experience delivering...PerformanceFlexible hours
- Member of Technical Staff, ML Inference Engineering Sanas is pioneering the future of human communication. Founded by a team of Stanford researchers... ...the level of the engineers working alongside them. Performance Optimization Optimize system and GPU performance for high...Performance
- ...harness powering Perplexity’s flagship answer experience. Improve single- and multi-agent orchestration, context management, performance, and reliability. Translate user needs and quantitative insights into agent improvements. Build excellent observability and...PerformanceFull time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Member of Technical Staff — Performance. Be the first to apply!
- mri tech aide Palo Alto, CA
- service desk assistant Palo Alto, CA
- end user support technician Palo Alto, CA
- operations support technician Palo Alto, CA
- help desk technical support Palo Alto, CA
- technical assistant Palo Alto, CA
- support analyst Palo Alto, CA
- technical associate Palo Alto, CA
- life support technician Palo Alto, CA
- tech assistant Palo Alto, CA


