Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Member of Technical Staff - Infrastructure

Gimlet Labs

About Us Gimlet is building the next generation of AI infrastructure: large-scale AI datacenters and the orchestration platform that coordinates them. The future of AI will require vastly more compute than exists today. But as AI workloads become more complex and new hardware architectures emerge, simply deploying more GPUs isn't enough. The challenge is making increasingly diverse compute work together. Gimlet's platform intelligently partitions and routes workloads across heterogeneous hardware, enabling step-function improvements in performance and efficiency. Customers deploy through production-grade APIs without needing to think about hardware selection, placement, or optimization. We work with foundation labs, hyperscalers, and AI-native companies to power production workloads at massive scale and help define the infrastructure layer for the future of AI. About this Role We are looking for an Infrastructure platform Engineer to design, build, and operate the cluster infrastructure behind Gimlet’s heterogeneous inference cloud. Unlike traditional cloud platforms built around a single hardware ecosystem, Gimlet's infrastructure spans multiple accelerator vendors and architectures. Infrastructure engineers play a key role in bringing new hardware platforms online, building the operational abstractions that make heterogeneous infrastructure manageable at scale, and ensuring new silicon can serve production workloads reliably from day one. This role is highly hands‑on. You will work across bare metal, Linux, Kubernetes or cluster schedulers, high‑speed networking, observability, provisioning, and incident response. You will partner closely with distributed systems, runtime, compiler, and hardware teams to ensure Gimlet’s infrastructure can support demanding AI workloads at production scale. What you will work on Design, deploy, and operate large‑scale CPU, GPU, and accelerator clusters powering production AI inference. Build automation for provisioning, configuration, upgrades, validation, and lifecycle management. Design and scale provisioning systems for heterogeneous bare‑metal infrastructure across multiple datacenters and hardware vendors. Operate cluster scheduling, resource allocation, isolation, quotas, and utilization systems. Debug complex production issues across Linux, networking, storage, drivers, firmware, and orchestration layers. Build and operate high‑performance networking infrastructure, including RDMA‑enabled environments and accelerator interconnects. Build observability for cluster health, capacity, performance, failures, and workload behavior. Improve reliability, availability, and recovery across multi‑node production systems. Work with distributed systems and runtime teams to support low‑latency, high‑throughput inference workloads. Evaluate and integrate new hardware platforms, accelerators, networking technologies, and datacenter designs. Create runbooks, operational standards, and incident response practices as the fleet scales. You may be a good fit if Experience in infrastructure, cluster engineering, platform engineering, SRE, HPC, or distributed systems. Deep Linux systems experience, including debugging performance, networking, storage, processes, and kernel‑level issues. Experience operating Kubernetes, Slurm, Nomad, or similar orchestration and scheduling systems. Strong automation skills using tools such as Terraform, Ansible, Helm, Python, Go, or equivalent. Experience with GPU or accelerator infrastructure, including drivers, firmware, CUDA/ROCm stacks, or hardware validation. Familiarity with high‑performance networking such as InfiniBand, RoCE, high‑speed Ethernet, or datacenter fabrics. Strong operational judgment: you know how to build systems that are observable, recoverable, and boring in production. Comfort working in a fast‑moving startup environment with high ownership and ambiguity. Strong candidates may also have Experience building or operating AI inference, training, HPC, or neocloud infrastructure. Experience with bare‑metal provisioning, PXE/iPXE, image pipelines, BIOS/firmware management, or rack bring‑up. Experience with multi‑tenant cluster isolation, quota systems, fair scheduling, or usage accounting. Experience debugging distributed workload performance across compute, memory, network, and storage bottlenecks. Experience building observability platforms using technologies such as Prometheus, OpenTelemetry, Grafana, or similar tooling. Familiarity with heterogeneous hardware environments across NVIDIA, AMD, Intel, ARM, or emerging accelerators. #J-18808-Ljbffr

Vacancy posted 22 hours ago
Similar jobs that could be interesting for youBased on the Member of Technical Staff - Infrastructure in San Francisco, CA vacancy
  •  ...observe their code. We are responsible for designing, building, and scaling core infrastructure that powers a high-volume data platform for AI applications. We are looking for team members who love building enabling systems that empower our engineers and power our rapidly... 
    Suggested
    Work at office

    LlamaIndex

    San Francisco, CA
    22 hours ago
  • $275k - $315k

     ...amount of untrusted, freshly generated kernels on real silicon, quickly and safely, which is why we're hiring a Member of Technical Staff for Sandbox Infrastructure to build the layer that makes it possible: a serverless GPU container service across NVIDIA, AMD, TPU and... 
    Suggested
    Full time
    Work at office
    Relocation
    Relocation package

    San Francisco Tensor Company

    San Francisco, CA
    2 days ago
  •  ...enterprises that integrate LLMs into their products. The team is 5 people with a research and product focus. As a Member of Technical Staff on our infrastructure team, you'll own the cloud systems that serve our compression API end-to-end. You'd get to build global low-... 
    Suggested
    Visa sponsorship

    The Token Company

    San Francisco, CA
    1 day ago
  • $250k

     ...compute platform building the next generation of agentic infrastructure for GPU-intensive workloads. Operating across the full technology...  .... This opportunity offers the chance to join as a Member of Technical Staff at a pivotal stage in the company's growth. You'll help... 
    Suggested
    Full time
    San Francisco, CA
    a month ago
  • Member of Technical Staff - Infrastructure Security We're partnering with a frontier AI research company that is building next-generation open-weight foundation models with the mission of making advanced AI broadly accessible. Their team includes researchers, engineers... 
    Suggested

    Xcede

    San Francisco, CA
    1 day ago
  •  ...of AI. Our mission: make intelligence open and accessible to all. Role Overview Reflection.AI is looking for a Member of Technical Staff - Infrastructure Security to secure our geographically diverse multi‑cloud Kubernetes and cloud environments. In this role, you’ll... 
    Work at office
    Visa sponsorship

    Visa Hunt

    San Francisco, CA
    3 days ago
  • $180k - $250k

    Job Title Member of Technical Staff, Backend Salary $180k-$250k + Equity Company Description Well-funded AI infrastructure startup Job Description Join a high-growth team building the core infrastructure powering how modern AI applications process and understand documents... 

    Jack & Jill

    San Francisco, CA
    4 days ago
  • # Founding Member of Technical Staff, AI Infrastructure**Location:** San Francisco / Bay Area preferred. Remote exceptional for the right person.We look for fast learners with high agency, AI-native workflows, clear technical communication, and evidence-backed judgment... 
    Full time
    Remote work

    Touchdown Labs, Inc.

    San Francisco, CA
    1 day ago
  • $200k - $400k

     ...chance or handed off to an algorithm. We're building the infrastructure to understand human behavior at scale and to represent humans...  ...research is even possible to run. About the Role As a Member of Technical Staff in Research Infrastructure, you will build the platform... 
    Live in
    Flexible hours

    Simile

    San Francisco, CA
    2 days ago
  • $200k

    Member of Technical Staff, Supercomputing Platform & Infrastructure Magic’s mission is to build safe AGI that accelerates humanity’s progress on the world’s most important problems. We believe the most promising path to safe AGI lies in automating research and code generation... 
    Relocation
    Visa sponsorship

    Magic AI, Inc

    San Francisco, CA
    16 hours ago
  •  ...people take ownership, grow together, and share both the challenges and the wins. What You'll Do Build the supercomputing infrastructure that runs our agents. Our agents tackle long-horizon, high-performance workloads, and you'll design the cloud compute,... 
    Work at office
    Remote work
    Flexible hours

    Asari AI

    San Francisco, CA
    4 days ago
  • $150k - $300k

    Building Open Superintelligence Infrastructure Prime Intellect is building the open superintelligence stack - from frontier agentic models...  ...Solutions Architect for GPU Infrastructure, you'll be the technical expert who transforms customer requirements into production‑ready... 

    Prime Intellect

    San Francisco, CA
    1 day ago
  • $175k - $260k

     ...we are moving fast enough that the people joining now are building the playbook everyone who comes after them will run. As an Infrastructure Engineer, you will build and scale the systems that power our rapid growth. You will ensure our infrastructure is highly available... 
    Work at office
    Visa sponsorship

    Slope

    San Francisco, CA
    3 days ago
  •  ...curating the world's highest-quality training datasets — spanning video, audio, images, text, and 3D. We combine exabyte-scale data infrastructure and novel multimodal understanding techniques that push the frontier of foundation models. Video alone makes up 80% of internet... 

    Sieve

    San Francisco, CA
    2 days ago
  • $150k - $265k

     ...human again. Mission We're building the platform for the future of voice technology. Our market edge is extensible, reliable infrastructure designed for the full complexity of voice interactions. 18 months, 150k developers, adding 1000 every day. Give it a try here... 
    Full time
    Shift work

    Vapi

    San Francisco, CA
    16 hours ago
  • $300k

     ...from human expertise. About the role This role is for strong infrastructure engineers who can build the systems layer for RL at scale: distributed...  ...building tools, platforms, or services used by other technical users. Strong judgment around technical tradeoffs: when to... 
    Work at office
    Local area

    Vmax AI Corp

    San Francisco, CA
    2 days ago
  •  ...JAX) and their underlying system architectures Strong engineering skills: performant, maintainable code and the ability to debug complex codebases Bonus: contributions to open-source inference or systems infrastructure (e.g. vLLM, SGLang, Triton) #J-18808-Ljbffr Uncover

    Uncover

    San Francisco, CA
    2 days ago
  • About Mandolin Nearly every disease will become treatable in our lifetimes. Mandolin is laying the clinical and financial infrastructure to get groundbreaking treatments to patients faster, powered by AI agents. Mandolin partners closely with the largest healthcare institutions... 
    Local area

    Mandolin

    San Francisco, CA
    16 hours ago
  •  ...Horowitz, GIC, Goldman Sachs, KKR, Visa, and others. Technical Skills Develop and maintain infrastructure that powers digital asset custody, trading, staking,...  ...to solve problems, and assist or teach other team members when possible. You may be a fit for this role if you... 
    Worldwide

    Crypto Pro Network

    San Francisco, CA
    1 day ago
  • $250k - $300k

     ...next step in your career? Join one of the most interesting infrastructure companies in the AI space right now. Founded by two of the most...  ...Skills / Must Have: A track record of impressive technical work you can speak to in depth; the years matter less than the... 
    Full time
    Remote work
    San Francisco, CA
    a month ago
  • $250k - $300k

     ...champions growth and development? Join one of the most exciting AI infrastructure companies in the market, building a platform that deploys and...  ...Skills / Must Have: ~ A track record of impressive technical work you can speak to in depth, the years matter less than the... 
    Full time
    Remote work
    San Francisco, CA
    a month ago
  •  ...users create characters, worlds, stories, and relationships with AI, and making that feel fast, reliable, and alive takes serious infrastructure. We are looking for an engineer who wants to help own that whole stack. We run more of our own than most companies our size.... 

    janitorAI

    San Francisco, CA
    1 day ago
  •  ...pipelines. Familiarity with Python and with cloud or compute infrastructure. Walden Robotics offers a competitive total compensation...  ...a team of exceptional professionals who combine world-class technical skills with creative vision, grounded in humility and collaboration... 
    Work from home
    Flexible hours

    Walden Robotics

    San Francisco, CA
    4 days ago
  •  ...Strong background with production systems at real scale. Data Infrastructure Development: Extensive experience developing cloud-based data...  ...and translating their needs into durable infrastructure. Technical Judgment: The ability to lead a work stream and make high-leverage... 
    Work from home
    Flexible hours

    Walden Robotics

    San Francisco, CA
    16 hours ago
  •  ..., drug discovery, and particle physics at institutions like DeepMind, Waymo, Cruise, Insitro, Nabla Bio, and CERN. We look for infrastructure engineers who are excited to tackle unsolved problems. Training an LPM means scaling novel architectures over multimodal physical... 

    Uncover

    San Francisco, CA
    2 days ago
  •  ...lifecycle Implement the platform-level quality and monitoring tooling that data and research teams build their checks on Scale infrastructure to improve engineering velocity and ensure reliability, with monitoring and alerting to match Work across the full data... 
    Immediate start

    Causal Labs

    San Francisco, CA
    4 days ago
  • $150k - $350k

     ...Member of Technical Staff | Distributed Systems San Francisco - Onsite $150k-$350k base + equity I'm working with a small, deeply technical $80 million Series A AI infrastructure company , building an inference cloud for agentic workloads that can partition and orchestrate... 

    Acceler8 Talent

    San Francisco, CA
    16 hours ago
  • $150k - $350k

     ...Member of Technical Staff - Distributed Systems San Francisco, CA- 5 days per week onsite $150,000–$350,000 + equity The Opportunity Join a rapidly growing AI infrastructure company building a multi-silicon cloud platform for fast, efficient inference. The future of AI... 

    Acceler8 Talent

    San Francisco, CA
    16 hours ago
  •  ...requirements, and very few precedents to copy from. About the Role Members of Technical Staff (MTS) are the senior engineers who build the platform...  ...is systems engineering at its core. Multi‑tenant data infrastructure across very different portcos. Event‑driven pipelines... 

    BEACON SOFTWARE COMPANY

    San Francisco, CA
    22 hours ago
  • $227.5k - $401k

     ...motivated individuals who tackle unique technical challenges at scale and solve them...  ...financial technology sector. As a Member of Technical Staff , you will operate with a high degree...  ...Experience working with on-premise infrastructure . Why Adyen? This is an exceptional... 
    Work at office
    Immediate start
    Relocation
    Flexible hours

    Adyen

    San Francisco, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Member of Technical Staff - Infrastructure. Be the first to apply!