Member of Technical Staff, Sandbox Infrastructure
$275k - $315kSan Francisco Tensor Company
At SF Tensor, we're building the future of high-performance compute We firmly believe that the future of AI depends on the unglamorous: rethinking and rebuilding the stack, all the way down. From the hardware underneath it to the compiler targeting it and the cloud running it. Right now those three things fight each other and that friction shows up as a tax on every researcher trying to build something ambitious. We're here to axe that tax and make compute faster, cheaper and more available. When we succeed, compute will be portable enough that "which cloud, which chip" stop being something you worry about. To achieve this, we are building our Kernel Optimizer, which takes code and finds its fastest possible form for whatever vendor and cluster topology you point it at, automatically, as well as the Model Foundry which manages the runs, makes research easier and moves workloads across clouds and chips as prices and availability ship, instead of leaving you locked into whatever vendor you signed with first. We're backed by Susa Ventures, Y Combinator, along with some great funds and angels including Max Mullen and Paul Graham, as well as founders and executives at Neuralink, Notion and AMD. We're looking for researchers, engineers and organizations who agree with the basic premise: you don't get the next leap in AI without a leap in compute first. About the Role We build the fastest GPU compiler in the world. Most compilers have to preserve correctness at every transform, constraining how far they can search, while we prove correctness at the end instead, allowing us to search a far wider space, with agents, with RL, with anything that works and still guarantee the result. It's why we hold #1 on NVIDIA's own kernel benchmark across hundreds of production kernels. For our search to work, we need to run an enormous amount of untrusted, freshly generated kernels on real silicon, quickly and safely, which is why we're hiring a Member of Technical Staff for Sandbox Infrastructure to build the layer that makes it possible: a serverless GPU container service across NVIDIA, AMD, TPU and Trainium, at a scale and fidelity nobody sells off the shelf. It needs to run three things: (1) our compiler measures every candidate program inside it, so a noisy or unfair sandbox directly corrupts the reward signal the search learns from, (2) our own post-training runs inside it (e.g., when we RL a model on AMD kernel engineering, every rollout is in a sandbox) and (3) customer workloads that require isolation, such as their RL rollouts, run on it too. We already do this on NVIDIA and AMD, including having rolled AMD GPU support in gVisor from scratch . The work now is depth and breadth, adding support for more vendors, higher fidelity instrumentation, faster cold starts and more features. A lot of our sandboxes need to run on spot pools without losing works, they can span multiple GPUs and sometimes multiple nodes, they need to survive failure and preemption as well as live-migration (which our stack does while keeping sockets intact so that instances can relocate mid-flight while still streaming data in or out without a hiccup). What You'll Do You'll extend our sandboxing stack to new vendors and accelerators You'll build and maintain GPU virtualization below the runtime, including gVisor work at the driver and ioctl level You'll make sandboxes first‑class citizens on spot capacity, which means preemption‑aware scheduling, checkpointing and rescheduling You'll support multi‑GPU and multi‑node sandboxes, including the interconnect (NVLink, NVSwitch) and RDMA paths (InfiniBand, RoCE) paths those require You'll own live migration end‑to‑end, including our socket‑preserving migration You'll guarantee measurement and profiling fidelity as well as their isolation You'll work directly with the compiler, post‑training and kernel teams to ensure their throughput is not capped by sandboxes What We're Looking For Someone with strong low-level systems engineering background: Linux kernel internals, containers, namespaces, cgroups, syscall interception or hypervisors Someone with experience in GPU systems engineering: drivers, runtimes or scheduling on accelerator fleets Someone comfortable with distributed systems failure modes: preemption, partial failure, checkpoint/restore and dealing with states you can't afford to loose Someone proficient in Go, C/C++ or Rust Someone with a strong bias toward building the thing yourself when no vendor supports what you need Nice to Have Someone who's worked directly with gVisor, Firecracker, Kata, QEMU/KVM or similar Someone who's worked directly on CRIU, -live migration or connection‑preserving failover work Someone familiar with NCCL/RCCL, RDMA, InfiniBand or vendor interconnects Someone who's run large fleets on spot or other preemptible capacity Someone familiar with bare‑metal provisioning, hypervisors or fleet management at scale Someone with a security background in isolation boundaries and untrusted code execution Why Join Us Most "serverless GPU" products stop at running a container on a GPU. Our's runs arbitrary code across four vendors, on spot capacity, across multiple nodes with migration that keeps live sockets open and timings clean enough to use as an RL reward. It's a hard systems problem and it's directly load-bearing, because the faster and more reliable the sandboxes, the more programs our compiler can search and the faster our models train the same week. We're a small team operating at frontier scale. We pre-trained foundation models on 4,000 AMD GPUs as a team of three, designed and brought up GB300 NVL72 clusters and designed a TOP500 supercomputer. We believe that hard problems get solved in person and most of our work happens at our office in San Francisco. We offer relocation assistance and, where possible, we'd like you here as often as possible. The base salary range for this full-time position is $275,000-$315,000, plus meaningful equity and benefits. #J-18808-Ljbffr San Francisco Tensor Company
$150k - $300k
...Building Open Superintelligence Infrastructure Prime Intellect is building... ...stack: environments, secure sandboxes, verifiable evals, and our... ...reliable at scale. Core Technical Responsibilities Infrastructure... ...and encourage team members to contribute to the broader...SuggestedFull timeWork at officeRemote workVisa sponsorshipRelocation packageFlexible hours- ...and the wins. What You'll Do Build the supercomputing infrastructure that runs our agents. Our agents tackle long-horizon, high-... ...and you'll design the cloud compute, distributed systems, and sandboxed tooling that keeps them reliable, efficient, and ready to scale...SuggestedWork at officeRemote workFlexible hours
- ...enterprises that integrate LLMs into their products. The team is 5 people with a research and product focus. As a Member of Technical Staff on our infrastructure team, you'll own the cloud systems that serve our compression API end-to-end. You'd get to build global low-...SuggestedVisa sponsorship
$150k - $300k
Building Open Superintelligence Infrastructure Prime Intellect is building the open superintelligence... ...-training stack: environments, secure sandboxes, verifiable evals, and our async RL... ...for GPU Infrastructure, you'll be the technical expert who transforms customer...Suggested- ...observe their code. We are responsible for designing, building, and scaling core infrastructure that powers a high-volume data platform for AI applications. We are looking for team members who love building enabling systems that empower our engineers and power our rapidly...SuggestedWork at office
$250k
...compute platform building the next generation of agentic infrastructure for GPU-intensive workloads. Operating across the full technology... .... This opportunity offers the chance to join as a Member of Technical Staff at a pivotal stage in the company's growth. You'll help...Full time- ...DeepMind, OpenAI, Google Brain, Meta, Character.AI, Anthropic and beyond. Role Overview Reflection.AI is looking for a Member of Technical Staff - Infrastructure Security to secure our geographically diverse multi-cloud Kubernetes and cloud environments. In this role, you’ll...Full timeRelocation package
- Member of Technical Staff - Infrastructure Security We're partnering with a frontier AI research company that is building next-generation open-weight foundation models with the mission of making advanced AI broadly accessible. Their team includes researchers, engineers...
- # Founding Member of Technical Staff, AI Infrastructure**Location:** San Francisco / Bay Area preferred. Remote exceptional for the right person.We look for fast learners with high agency, AI-native workflows, clear technical communication, and evidence-backed judgment...Full timeRemote work
$200k - $400k
...chance or handed off to an algorithm. We're building the infrastructure to understand human behavior at scale and to represent humans... ...research is even possible to run. About the Role As a Member of Technical Staff in Research Infrastructure, you will build the platform...Live inFlexible hours$200k
Member of Technical Staff, Supercomputing Platform & Infrastructure Magic’s mission is to build safe AGI that accelerates humanity’s progress on the world’s most important problems. We believe the most promising path to safe AGI lies in automating research and code generation...RelocationVisa sponsorship$150k - $300k
...systems for the life sciences. About The Role We’re looking for an Infrastructure Engineer to build, operate, and scale the deployment... ...environments, own Kubernetes‑based infrastructure for Phylo services, sandboxed agent execution, compute workloads, storage mounts,...Work at office$150k - $300k
...stack: environments, secure sandboxes, verifiable evals, and our async... ...areas are: Building the infrastructure to serve LLMs efficiently at... ...RL training stack. Core Technical Responsibilities LLM Serving... ...development and encourage team members to contribute to the broader...Work at officeRemote workVisa sponsorshipRelocation packageFlexible hoursShift work- ...curating the world's highest-quality training datasets — spanning video, audio, images, text, and 3D. We combine exabyte-scale data infrastructure and novel multimodal understanding techniques that push the frontier of foundation models. Video alone makes up 80% of internet...
- About Us Gimlet is building the next generation of AI infrastructure: large-scale AI datacenters and the orchestration platform that coordinates them. The future of AI will require vastly more compute than exists today. But as AI workloads become more complex and new...
- About Mandolin Nearly every disease will become treatable in our lifetimes. Mandolin is laying the clinical and financial infrastructure to get groundbreaking treatments to patients faster, powered by AI agents. Mandolin partners closely with the largest healthcare institutions...Local area
$150k - $265k
...human again. Mission We're building the platform for the future of voice technology. Our market edge is extensible, reliable infrastructure designed for the full complexity of voice interactions. 18 months, 150k developers, adding 1000 every day. Give it a try here...Full timeShift work- ...drug discovery, and particle physics at institutions like DeepMind, Waymo, Cruise, Insitro, Nabla Bio, and CERN. We look for infrastructure engineers who are excited to tackle unsolved problems. Progress on an LPM is gated by how fast we can evaluate it: large-scale...
- Overview Anchorage Digital is building infrastructure that enables the world’s largest financial... .... Responsibilities Contribute to the technical direction of our infrastructure... ...problems, and assist or teach other team members when possible. Qualifications 2-5 years...
$250k - $300k
...next step in your career? Join one of the most interesting infrastructure companies in the AI space right now. Founded by two of the most... ...Skills / Must Have: A track record of impressive technical work you can speak to in depth; the years matter less than the...Full timeRemote work$250k - $300k
...champions growth and development? Join one of the most exciting AI infrastructure companies in the market, building a platform that deploys and... ...Skills / Must Have: ~ A track record of impressive technical work you can speak to in depth, the years matter less than the...Full timeRemote work$200k - $350k
...nation-state-level threats. We're a small technical team, we move fast, and we're growing... ...input from CISOs and senior security staff at the frontier labs, and US security and... ...the role An SL5 datacenter requires infrastructure engineering at the frontier of security...- ...scale for AI workloads. We take consumer hardware and build the infrastructure platform around it to enable consumption as elastic compute... ...Operations. In this role you will Own the architecture and technical roadmap for Mount Thor’s production network. Design and...
$150k - $300k
Building Open Superintelligence Infrastructure Prime Intellect is building... ...stack: environments, secure sandboxes, verifiable evals, and our... ...that runs the jobs. Core Technical Responsibilities Hosted Training... ...and encourage team members to contribute to the broader...Work at officeLocal areaRemote workVisa sponsorshipRelocation packageFlexible hours- ...users create characters, worlds, stories, and relationships with AI, and making that feel fast, reliable, and alive takes serious infrastructure. We are looking for an engineer who wants to help own that whole stack. We run more of our own than most companies our size....
- ...building the foundational software and infrastructure that everything else depends on. This is... ...growth. What We Look For Senior to staff-level experience in software engineering... ...source projects or other publicly visible technical work. Comfort owning ambiguous, high-...
- ...pipelines. Familiarity with Python and with cloud or compute infrastructure. Walden Robotics offers a competitive total compensation... ...a team of exceptional professionals who combine world-class technical skills with creative vision, grounded in humility and collaboration...Work from homeFlexible hours
- Role We seek experienced engineers to architect and scale the core infrastructure behind distributed training pipelines and petabyte-scale data catalogs. You\'ll work directly with researchers to accelerate experiments, develop new datasets, improve infrastructure efficiency...
- ...lifecycle Implement the platform-level quality and monitoring tooling that data and research teams build their checks on Scale infrastructure to improve engineering velocity and ensure reliability, with monitoring and alerting to match Work across the full data...Immediate start
- ...Strong background with production systems at real scale. Data Infrastructure Development: Extensive experience developing cloud-based data... ...and translating their needs into durable infrastructure. Technical Judgment: The ability to lead a work stream and make high-leverage...Work from homeFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Member of Technical Staff, Sandbox Infrastructure. Be the first to apply!
- mri tech aide San Francisco, CA
- salesforce technical analyst San Francisco, CA
- service desk assistant San Francisco, CA
- end user support technician San Francisco, CA
- operations support technician San Francisco, CA
- help desk technical support San Francisco, CA
- technical assistant San Francisco, CA
- support analyst San Francisco, CA
- technical associate San Francisco, CA
- life support technician San Francisco, CA


