Software Engineer, Infrastructure
$180k - $250kThe Consensus
fal is the generative media ecosystem powering the next generation of AI products. We build the infrastructure, tools, and model access that teams need to move from idea to production, and do it at scale without compromise. For developers and enterprises, fal is the foundation that makes generative media not just possible, but practical: a unified platform where high-performance inference, orchestration, and observability come together to unlock new categories of AI-native products.
As generative media reshapes industries across a market projected to grow by hundreds of billions over the next decade, fal is becoming the ecosystem that ambitious teams build on.
You are a hands-on engineer who builds the software and processes that keep a large fleet of GPU servers healthy and productive. You write systems and tooling for managing 1000s of servers including provisioning, health monitoring, error detection, and recovery — and when something breaks that automation can’t fix, you drive resolution with partners.
Key responsibilities
-
Build and maintain Python fleet tracking system that manages the full lifecycle of servers including contracting and procurement, target use, pricing, availability, health, RMAs, etc
-
Build server management tooling that automates provisioning, health checks, GPU diagnostics, recovery and alerting
-
Create and maintain metrics, dashboards, and alerting for hardware health across the fleet (GPU errors, disk failures, network issues, thermals)
-
Leverage AI to an extreme level to build tools and automate alerting and recovery
-
Implement and enforce OS-level security: hardening baselines, SELinux/AppArmor policies, SSH key management, vulnerability scanning, and compliance automation
-
Manage and optimize distributed and local storage systems supporting model weights, checkpoints, and temporary scratch: NVMe arrays, NFS, parallel file systems, and object storage
-
Tune Linux systems for AI workloads: kernel parameters, NUMA topology, CPU pinning, hugepages, I/O schedulers, and GPU driver stack optimization (NVIDIA drivers, CUDA, container runtimes)
-
Develop a suite of automated error detection and recovery processes
-
Work with partners to solve technical issues
Requirements
-
3+ years experience managing bare-metal and cloud based server fleets at scale (100+ nodes)
-
Strong software engineering skills in Python; you write production tooling, not scripts
-
Deep Linux systems knowledge: boot process, kernel tuning, networking, storage, systemd, cgroups, namespaces, performance profiling
-
Strong experience with configuration management and infrastructure-as-code: Ansible, Terraform, cloud-init
-
Solid understanding of storage technologies: LVM, RAID, NVMe, NFS, Lustre or GPFS, and Linux I/O stack tuning
-
Familiarity with hardware diagnostics and failure modes (GPUs, NVMe, NICs, memory)
-
Experience building internal tools or dashboards for infrastructure visibility
-
Excellent communication and ability to drive technical decisions across teams
-
Self-starter who executes quickly, takes ownership, and constantly seeks improvement
Nice to have
-
Familiarity with network configuration and diagnostics (VLAN, VXLAN, ECMP, BGP, tcpdump)
-
Experience with NVIDIA GPU infrastructure: driver management, health monitoring, DCGM, NVLink/NVSwitch diagnostics, RDMA, InfiniBand/RoCEv2
-
Experience with AMD GPUs
-
Experience with bare metal and VM provisioning (PXE/iPXE, Kickstart, libvirt, Qemu/KVM)
-
Experience with compliance frameworks relevant to cloud providers (SOC 2, ISO 27001)
Compensation
-
$180,000-250,000 plus equity + benefits
Location
-
San Francisco, CA (we are open to remote in the US for Senior and Staff levels)
What we offer at fal
-
Interesting and challenging work
-
A lot of learning and growth opportunities
-
We are offering relocation assistance to San Francisco.
-
We offer relocation assistance to San Francisco.
-
Health, dental, and vision insurance (US)
-
Regular team events and offsites
$190k - $280k
...About Sentry Bad software is everywhere, and we’re tired of it. Sentry is on a mission to help developers write better software... ...and an informative deployment pipeline. As an engineer on the Infrastructure Engineering team, you’ll help deliver on our mission: We...SuggestedHourly payFull timeWork at office$170k - $216k
...across 15+ U.S. states. The Simulation Infrastructure team creates reliable, scalable, and... ...products that evaluate the Waymo Driver's software stack at a massive scale. We solve... ...for a broad range of customers Software Engineers, Product, Data Science, System Engineering...SuggestedFull timeRemote work- ...the data collection systems, operational capability, exabyte-scale data warehouse, and software toolchain – to help our partners drive the field forward. Infrastructure engineers build the platform that support our ever-increasing scale of data collection. Sample...SuggestedFull time
- ...Exa is building a search engine from scratch to serve every AI agent. We build massive-scale infrastructure to crawl the web, train state-of-the-art embedding models to process... ...of machines Design GPU scheduling software so we max out our cluster utilization Build...SuggestedFull timeH1b
$215k - $265k
...with them on scheming mitigations. We're looking for a Software Engineer to build the platform that the rest of Apollo runs on . This... ..., and in what order. Build and maintain Apollo's cloud infrastructure . This means IaC, networking, environment management,...SuggestedFull timeWork experience placementWork at officeImmediate startVisa sponsorshipFlexible hours- ...economy has grown 6x over the last decade, software businesses have gone from not worrying... .... We're building the agentic infrastructure layer at Anrok—the systems that let AI... ...product builds on. We're looking for an engineer who pairs a strong distributed-systems...Full timeWork at officeRemote workHome officeFlexible hours3 days per week
$204k - $259k
...15+ U.S. states. The Simulation ML Infrastructure team builds scalable AI/ML infrastructure... ...systems. This role reports to an Engineering Manager. You will: Be part of... ...prefer: ~5+ years of professional software engineering experience, with at least 3...Full timeRemote work- ...coding agents that replace traditional software development by generating, testing, and... ...You'll Be Responsible For Platform & Infrastructure Maintain stability of our platform... ...~4+ years of software/platform engineering experience with production systems ~...Full timeFlexible hours
- ...reliability, performance, and security for multi‑tenant compute. What You’ll Do Design and operate secure, multi‑tenant container infrastructure with fast startup and smart autoscaling. Ship cloud deployments (Helm/Terraform) with SSO, network controls, and audit...Full timeRemote work
- ...About Flow Flow Engineering is an AI-native requirements platform for modern engineering organizations, enabling hardware... ...and speed. About the role Flow is hiring a Software Engineer with an infrastructure focus to build and scale the core platform behind Flow...Full timeFlexible hours
- ...teams for Google Workspace. What you'll do The Payments Infrastructure team builds the trust boundary between a live conversation... ...plaintext cardholder data. Make payments something other engineers can use without becoming compliance experts: drive the platform...Full timeFlexible hours
- ...Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable... ...Conviction. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As a Software Engineer at on the Training Infrastructure...Full timeFlexible hours
- ...About the Role We are seeking a Cloud Infrastructure Engineer to help design and evolve the platforms that power OpenAI’s products. In this role, you will be a hands-on technical leader, driving the architecture, scalability, reliability, and security of critical...Full time
- ...into real, working systems and build any software needed for running large-scale frontier... ...About the Role We are looking for engineers to operate the next generation of compute... ...systems engineering with hands-on infrastructure work on our largest datacenters. You will...Full time
- ...government agencies address real-world challenges. The Infrastructure Engineering team is crucial to the overall success of Hayden products:... ...Uphold Engineering Standards: Establish and enforce elite software engineering and DevOps standards, including rigorous...Full timeShift work
- ...About the Role This role broadly owns infrastructure across the stack. If it’s running in the cloud, you probably care about it. The life... ...ChatGPT and GPT-4 to hundreds of millions of users, engineered the foundations of autonomous driving, built next-generation...Full time
$250k
...product-market fit with a substantial customer pipeline already in place. Role Overview We’re looking for an ML infrastructure engineer to design and build the core systems that enable scalable, efficient training of large models for deployment and research....Full time$209k - $240k
...office workdays. About the Product Infrastructure Team: The Product Infrastructure... ...classes of problems up-front for product engineers. Solve hard technical challenges such... ...values, and enthusiastic about making software toolmaking ubiquitous, we want to hear...Full timeWork at officeLocal area- ...Background Specter is creating a software-defined "control plane" for the physical... ...will ultimately become the perception engine for a company's physical footprint, enabling... ...Specter is hiring an ML infrastructure engineer to build and scale the machine...Full time
- ...business with billions in revenue The Role Handshake is building the infrastructure layer that powers the next generation of AI agents across our platform. As a Senior Software Engineer on our Agentic Infrastructure team, you'll be at forefront of AI at...Full timeFreelanceInternshipWork at officeRemote workFlexible hours
$175k - $215k
...redefines the future of global mobility. The Waymo Logs Infrastructure powers critical decision making and model development... ...and clients You have: ~4+ years of professional software engineering experience ~ Experience working on large-scale distributed...Full timeRemote work- ...of the great threats AI presents: mass-manufactured social engineering. Countless scams, deepfakes, and other social engineering attacks... ...for an experienced backend engineer to build out the infrastructure required to rapidly scale up our engineering organization. A...Full timeWork at officeFlexible hours
$160k - $180k
...out before, during, and after playing games. The Database Infrastructure team develops and operates all of Discord’s databases and data... ...matters most to our users. Work with a talented team of engineers who have built one of the largest communication platforms in...Full time- ..., and Google Lens. Before that, Clay led the product and design teams for Google Workspace. What you’ll do As a Software Engineer, Infrastructure at Sierra, you will be responsible for designing, building, and maintaining the core systems that make our AI platform...Full timeFlexible hours
$190k - $260k
...Notion. Centralize was founded by Rachit Kataria, a founding engineer on Facebook Shops who helped scale it to 250M MAUs, and Will... ...what Centralize becomes. The Role We are hiring an infrastructure engineer to own scalability across Centralize. Our system processes...Remote jobFull timeRelocationVisa sponsorship- ...an experienced, creative, and versatile engineer who is eager to tackle the challenge of... ..., building, and scaling the core infrastructure and systems powering Lightfield's AI-driven... ...Who you are ~3+ years of experience in software development, with a strong background...Full timeWork from home
- ...Runloop.ai is building the foundational infrastructure for the next generation of AI development. We provide AI engineers and data scientists with lightning-fast, secure... ...innovation in the age of AI. The Role As a Software Engineer on our Infrastructure team, you'll...Full timeWork at officeWork from home1 day per week
- ...including stablecoins. You’ll help design, deploy and operate the infrastructure that makes this possible. This is a hands-on devops role... ...role sits at the intersection of infrastructure, platform engineering and security. You’ll define how Modern Treasury scales and...Full timeLocal areaImmediate startRemote work
- ...About the Team We’re hiring software engineers to join our broader Infrastructure organization, which supports multiple high-impact teams. Depending on your interests and experience, you could work on one of several focus areas—including Core Distributed Systems, Databases...Full time
$100k - $300k
...Abnormal AI, Zscaler Preeminent research labs like Deepmind and SAIL About the Role We're hiring a Senior Storage Infrastructure Engineer on our Core Platform team to own how we store, protect, and operate data at scale. You'll build the backup, observability,...Full time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Software Engineer, Infrastructure. Be the first to apply!
- software developer positions San Francisco, CA
- senior software engineer remote San Francisco, CA
- software engineer contract San Francisco, CA
- IT software developer San Francisco, CA
- cybersecurity software engineer San Francisco, CA
- part time software developer remote San Francisco, CA
- junior software developer internship San Francisco, CA
- junior software engineer San Francisco, CA
- software system engineer San Francisco, CA
- software engineer remote San Francisco, CA

