Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Member of Technical Staff — Reliability-CI Infrastructure

RadixArk

RadixArk is hiring a Member of Technical Staff — CI Engineer to own the infrastructure that keeps SGLang moving. Our CI system runs 300+ GPU tests across NVIDIA, AMD, Intel, and Ascend hardware pools, gating every commit to one of the fastest-growing open-source LLM inference engines. When CI is green and fast, 100+ contributors ship with confidence. When it isn’t, the entire project stalls. That bottleneck is your problem to solve. You won’t just maintain pipelines — you’ll architect them. You’ll replace brittle static thresholds with regression-based detection, harden runners against supply-chain attacks from fork PRs, and cut cycle times so contributors get feedback in minutes, not hours. You’ll work directly with core maintainers, hardware partners, and the open-source community to keep the system that gates every merge request trustworthy, fast, and secure. This is not a role for someone who wants to write CI YAML and walk away. It’s for an engineer who treats CI infrastructure the way we treat serving infrastructure — as a system worth designing well. What You’ll Do Own CI reliability end-to-end — triage failures, distinguish real regressions from flaky tests and infra issues, keep main green Build regression-based CI — replace hardcoded static thresholds with automated baseline comparison (metrics pipeline, durable storage, detection logic) Cut CI time — right-size eval suites, deduplicate server startups, separate PR smoke tests from nightly full runs Improve developer experience — faster feedback, clearer failure messages, workflow orchestration Requirements 3+ years operating CI/CD at scale (GitHub Actions, Buildkite, Jenkins, GitLab CI, or similar) Deep Linux, Docker, GPU computing knowledge Strong Bash and Python Security mindset — CI supply chain risks, fork PR attack vectors, runner hardening NVIDIA GPU drivers, CUDA, NCCL, InfiniBand/RDMA experience in CI contexts Familiarity with ML inference workloads (model loading, KV cache, quantization) Nice to Have Large open-source project CI experience (100+ contributors) AMD ROCm or Intel XPU CI pipelines What Success Looks Like Day 20 — Full CI landscape understood, daily triage taken over, top recurring flaky tests fixed, PR CI time reduced 30%+ Day 40 — Regression-based checks live on nightly CI, ephemer al runner prototype deployed, runner isolation in place Day 60 — Zero flaky tests. Main CI 100% green when no real regression exists About RadixArk RadixArk is an infrastructure-first company built by engineers who’ve shipped production AI systems, created SGLang (20K+ GitHub stars, the fastest open LLM serving engine), and developed Miles (our large-scale RL framework). We’re on a mission to democratize frontier-level AI infrastructure by building world-class open systems for inference and training. Our team has optimized kernels serving billions of tokens daily, designed distributed training systems coordinating 10,000+ GPUs, and contributed to infrastructure that powers leading AI companies and research labs. We’re backed by well-known infrastructure investors and partner with Nvidia, Google, AWS, and frontier AI labs. Join us in building infrastructure that gives real leverage back to the AI community. Compensation We offer competitive base with meaningful equity, comprehensive health benefits, and flexible work arrangements. Compensation is determined by location, level, and experience. RadixArk is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more. #J-18808-Ljbffr RadixArk

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Member of Technical Staff — Reliability-CI Infrastructure in Palo Alto, CA vacancy
  •  ...superintelligence. We are seeking a Software Engineer - Infrastructure (5+ years experience) to lead the productization and reliability of our core systems. You will architect the...  ...systems and integration and deployment (e.g., CI/CD). You will own the full development... 
    Suggested

    Ricursive Intelligence

    Palo Alto, CA
    5 days ago
  • Member of Technical Staff, Infrastructure Build and operate the secure execution substrate for enterprise agents, customer applications, and Sycamore’...  ...per-tenant failure isolation and no downtime is not. Reliability. How the platform behaves over time: service level objectives... 
    Suggested

    Sycamore

    Palo Alto, CA
    1 day ago
  • $180k

    Member Of Technical Staff - Cloud Infrastructure SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in...  ...manage training and inference clusters, as well as highly reliable applications, across bare metal, classified cloud,... 
    Suggested
    Temporary work

    SpaceXAI

    Palo Alto, CA
    5 days ago
  •  ...hardware and robot systems to the infrastructure and state-of-the-art...  ...We're looking for a senior or staff-level Research Engineer or ML...  ...our robot- learning pipeline reliable, reproducible, and measurable...  ...system design, automated testing, CI, code review, and production-... 
    Suggested

    Socket.dev

    Mountain View, CA
    5 days ago
  • $180k

     ...AI supercomputers from the ground up. As part of the Compute Infrastructure team, you will own both the raw GPU supercomputer and the platform...  ...— to make training and inference at xAI as fast, reliable, and scalable as possible. This is a broad, high‑impact role... 
    Suggested
    Temporary work

    xAI

    Palo Alto, CA
    3 days ago
  •  ...Research), and Fei‑Fei Li (Godmother of AI). We’re building the infrastructure for a new era of interactive entertainment. About the Role...  ...MTS, you’ll own the systems that make all of that work: reliable, cost‑efficient, and fast at real scale. We benchmark everything... 

    Astrocade

    Palo Alto, CA
    3 days ago
  • We are looking for a Member of Technical Staff with strong Python skills and a...  ...in shaping Activeloop's AI infrastructure. What You Will Be Doing...  ...performance demands. Implement CI/CD Pipelines: Develop...  ...business objectives. Enhance Reliability: Ensure platform... 

    S27a

    Mountain View, CA
    4 days ago
  • $180k

    Member of Technical Staff, Pre-training Data Infrastructure xAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. All employees... 
    Temporary work
    Relocation

    xAI

    Palo Alto, CA
    3 days ago
  • About the Role As a Member of Technical Staff [Platform] at NeoCognition , you’ll...  ...You’ll create the tooling, infrastructure, and developer experience...  ...environments (cloud compute, storage, CI/CD pipelines, observability...  ...across teams. Build reliable systems for data management... 

    NeoCognition

    Palo Alto, CA
    1 day ago
  • About the Role As a Member of Technical Staff [Agent] at NeoCognition , you’ll be...  ...from backend APIs and data infrastructure to front-end interfaces and...  ...applications scalable, reliable, and delightful to use. You...  ...engineering excellence — testing, CI/CD, documentation, and... 

    NeoCognition

    Palo Alto, CA
    5 days ago
  •  ...AI silicon, systems, software, and infrastructure to build the next generation of full...  .... We are looking for an exceptional Member of Technical Staff to help design, build, and scale core...  ...Emphasis on performance, scalability, and reliability. AI Compute Systems Development:... 

    DensityAI

    Mountain View, CA
    3 days ago
  •  ...and scale a high-throughput data infrastructure that processes and manages...  ...with strong guarantees around reliability, latency, and cost efficiency...  ...end, from design to production Staff-level candidates are expected to define technical direction and own architectural... 
    Immediate start

    Rhoda AI

    Palo Alto, CA
    3 days ago
  • Member Of Technical Staff - Extreme-Scale Sparse Linear Algebra, Domain Decomposition & GPU Solver Architecture...  ..., we are building the AI-enabled infrastructure that modern hardware programs use to...  ...Multi-GPU scaling experience Strong CI, regression, and correctness... 
    Full time
    Remote work

    Vinci4d

    Palo Alto, CA
    5 days ago
  • About The Role RadixArk is seeking a Member of Technical Staff — Training to build and scale the...  ...on large-scale distributed training infrastructure for LLMs and generative models, pushing...  ...of scale, efficiency, accuracy and reliability across 10k, or 100k+ of GPUs. This role... 
    Flexible hours

    RadixArk

    Palo Alto, CA
    5 days ago
  • RadixArk is seeking a Member of Technical Staff — Inference to push the limits of large-scale AI inference...  ...of systems engineering, ML infrastructure, and performance optimization. Your...  ...bottlenecks across the stack Drive reliability and scalability of inference infrastructure... 
    Worldwide
    Flexible hours

    Dormont Manufacturing Co

    Palo Alto, CA
    1 day ago
  • $120k - $200k

     ...support global partners with fast, reliable, and scalable data solutions. Our...  .... About the Role As a Member of Technical Staff, Platform, you'll build full-stack product...  ..., and observability of assessment infrastructure. Qualifications 1+... 
    Full time
    Flexible hours

    Abaka AI

    Mountain View, CA
    7 days ago
  • $175k - $350k

     ...collaborate across research, product, and infrastructure teams to enable rapid iteration, high reliability, and secure delivery of novel AI...  ...engineering productivity—CI/CD pipelines, service templates,...  ...to assess fit and alignment. Technical Interview - A deep dive with an... 
    Full time

    Inflection AI

    Palo Alto, CA
    3 days ago
  •  ...They’re developing advanced infrastructure that enables organisations to...  ...services, infrastructure, and reliability, working closely with a strong engineering team on technically challenging problems. Design,...  ...(Terraform) Familiarity with CI/CD, DevOps practices, and monitoring... 
    Flexible hours
    3 days per week

    DeepRec.ai

    Palo Alto, CA
    4 days ago
  • RadixArk is seeking a Member of Technical Staff — Training to build and scale the systems that train...  ...on large-scale distributed training infrastructure for LLMs and generative models,...  ...the limits of scale, efficiency, and reliability across thousands of GPUs. This role... 
    Flexible hours

    RadixArk

    Palo Alto, CA
    3 days ago
  • $148.5k - $223.9k

    Senior Member of Technical Staff - AI ResearchSkip to main content#Senior Member of Technical Staff...  ...AI to operate accurately and reliably.** *Critically evaluate code (Human...  ...evaluation, and inference pipelines** *Infrastructure & Deployment** *Experience deploying... 
    Work at office

    Salesforce, Inc.

    Palo Alto, CA
    2 days ago
  •  ...Role We are seeking a Sr. Member of Technical Staff to design and develop software...  ...workflows, improve system reliability through automation, and...  ...documentation for infrastructure configurations, inference...  ...PostgreSQL, Redis, and NFS. CI/CD and version control: Jenkins... 

    Cerebras Systems, Inc.

    Sunnyvale, CA
    1 day ago
  • About the Role As a Member of Technical Staff [Research] at NeoCognition , you’ll be part of the core...  ...that can reason, plan, and act reliably in the real world. We are an AI research...  ..., Mistral, or similar) and training infrastructure. Publications in top-tier AI venues... 

    NeoCognition Inc.

    Palo Alto, CA
    1 day ago
  • $300k - $350k

     ...Member of Technical Staff Level 1 – Engineering Bellevue | Hybrid NTT DATA AIVista, Inc., a wholly...  ...AI capabilities, including agentic infrastructure, model customization, and agentic...  ...that balance performance, scalability, reliability, cost, and customer requirements.... 
    Full time
    Work experience placement
    Local area
    Flexible hours

    NTT DATA AIVista

    Palo Alto, CA
    3 days ago
  •  ...projects full lifecycle Design discussions & technical scoping Implementation & testing Post-...  ...to system architecture - ensure reliable, maintainable, scalable stack Balance velocity...  ...Auth0 Backend: NestJS Prisma Postgres Infrastructure: AWS CDK Vercel Dev Tools: Cursor... 
    Work at office
    Remote work
    Visa sponsorship
    Monday to Friday

    ProductNow

    Palo Alto, CA
    3 days ago
  • $180k

     ...Own backend engineering for scalable, low-latency voice infrastructure and model integrations. Collaborate directly with Grok Voice...  ...teams to deliver end-to-end experiences. Drive performance, reliability, and quality of voice interactions at global scale. Move... 
    Temporary work

    SpaceXAI

    Palo Alto, CA
    a month ago
  • $180k

     ...implement scalable systems to support Grok's AI-driven media experiences, ensuring high performance, reliability, and low-latency at global scale. Architect robust infrastructure for real-time multi-modal interactions, including handling generation requests, media... 
    Temporary work
    Worldwide

    SpaceXAI

    Palo Alto, CA
    a month ago
  • About The Role RadixArk is seeking a Member of Technical Staff: Accelerator Systems to push the limits of performance for frontier AI systems...  ..., optimize, and maintain SGLang, Miles, and the RadixArk infrastructure stack across NVIDIA and AMD GPUs, Google TPUs, modern... 
    Flexible hours

    RadixArk

    Palo Alto, CA
    3 days ago
  • $180k

    Member of Technical Staff - Multimodal Understanding About xAI xAI’s mission is to create AI systems that can accurately understand the universe...  ..., large‑scale pre‑training, post‑training/alignment, infrastructure/scaling, evaluation, tooling/demos, and end‑to‑end product... 
    Temporary work

    xAI

    Palo Alto, CA
    2 days ago
  • $180k - $250k

    Member of Technical Staff -- TPU Systems (JAX / XLA / PALLAS) About the Role RadixArk is looking for a TPU Systems Engineer to build high-performance...  ...TPU hardware, working on SGLang-JAX and other critical infrastructure that enables efficient deployment of frontier models on... 
    Full time
    Flexible hours

    RadixArk

    Palo Alto, CA
    1 day ago
  • $180k

     ...high-performance systems for personalized, reliable interactions at global scale. Build and...  ...(“phone interview”) during which a member of our team will ask some basic questions...  ...enter the main process, which consists of 2 technical interviews and 1 project deep-dive... 
    Temporary work

    Pantera Capital

    Palo Alto, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Member of Technical Staff — Reliability-CI Infrastructure. Be the first to apply!