Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Software Engineer, Infrastructure

$180k - $250k

The Consensus

fal is the generative media ecosystem powering the next generation of AI products. We build the infrastructure, tools, and model access that teams need to move from idea to production, and do it at scale without compromise. For developers and enterprises, fal is the foundation that makes generative media not just possible, but practical: a unified platform where high-performance inference, orchestration, and observability come together to unlock new categories of AI-native products.

As generative media reshapes industries across a market projected to grow by hundreds of billions over the next decade, fal is becoming the ecosystem that ambitious teams build on.

You are a hands-on engineer who builds the software and processes that keep a large fleet of GPU servers healthy and productive. You write systems and tooling for managing 1000s of servers including provisioning, health monitoring, error detection, and recovery — and when something breaks that automation can’t fix, you drive resolution with partners.

Key responsibilities
  • Build and maintain Python fleet tracking system that manages the full lifecycle of servers including contracting and procurement, target use, pricing, availability, health, RMAs, etc

  • Build server management tooling that automates provisioning, health checks, GPU diagnostics, recovery and alerting

  • Create and maintain metrics, dashboards, and alerting for hardware health across the fleet (GPU errors, disk failures, network issues, thermals)

  • Leverage AI to an extreme level to build tools and automate alerting and recovery

  • Implement and enforce OS-level security: hardening baselines, SELinux/AppArmor policies, SSH key management, vulnerability scanning, and compliance automation

  • Manage and optimize distributed and local storage systems supporting model weights, checkpoints, and temporary scratch: NVMe arrays, NFS, parallel file systems, and object storage

  • Tune Linux systems for AI workloads: kernel parameters, NUMA topology, CPU pinning, hugepages, I/O schedulers, and GPU driver stack optimization (NVIDIA drivers, CUDA, container runtimes)

  • Develop a suite of automated error detection and recovery processes

  • Work with partners to solve technical issues

Requirements
  • 3+ years experience managing bare-metal and cloud based server fleets at scale (100+ nodes)

  • Strong software engineering skills in Python; you write production tooling, not scripts

  • Deep Linux systems knowledge: boot process, kernel tuning, networking, storage, systemd, cgroups, namespaces, performance profiling

  • Strong experience with configuration management and infrastructure-as-code: Ansible, Terraform, cloud-init

  • Solid understanding of storage technologies: LVM, RAID, NVMe, NFS, Lustre or GPFS, and Linux I/O stack tuning

  • Familiarity with hardware diagnostics and failure modes (GPUs, NVMe, NICs, memory)

  • Experience building internal tools or dashboards for infrastructure visibility

  • Excellent communication and ability to drive technical decisions across teams

  • Self-starter who executes quickly, takes ownership, and constantly seeks improvement

Nice to have
  • Familiarity with network configuration and diagnostics (VLAN, VXLAN, ECMP, BGP, tcpdump)

  • Experience with NVIDIA GPU infrastructure: driver management, health monitoring, DCGM, NVLink/NVSwitch diagnostics, RDMA, InfiniBand/RoCEv2

  • Experience with AMD GPUs

  • Experience with bare metal and VM provisioning (PXE/iPXE, Kickstart, libvirt, Qemu/KVM)

  • Experience with compliance frameworks relevant to cloud providers (SOC 2, ISO 27001)

Compensation
  • $180,000-250,000 plus equity + benefits

Location
  • San Francisco, CA (we are open to remote in the US for Senior and Staff levels)

What we offer at fal
  • Interesting and challenging work

  • A lot of learning and growth opportunities

  • We are offering relocation assistance to San Francisco.

  • We offer relocation assistance to San Francisco.

  • Health, dental, and vision insurance (US)

  • Regular team events and offsites

#J-18808-Ljbffr
Vacancy posted 12 hours ago
Similar jobs that could be interesting for youBased on the Software Engineer, Infrastructure in San Francisco, CA vacancy
  • $190k - $280k

     ...About Sentry Bad software is everywhere, and we’re tired of it. Sentry is on a mission to help developers write better software...  ...and an informative deployment pipeline. As an engineer on the Infrastructure Engineering team, you’ll help deliver on our mission: We... 
    Suggested
    Hourly pay
    Full time
    Work at office

    Sentry

    San Francisco, CA
    1 day ago
  • $170k - $216k

     ...across 15+ U.S. states. The Simulation Infrastructure team creates reliable, scalable, and...  ...products that evaluate the Waymo Driver's software stack at a massive scale. We solve...  ...for a broad range of customers Software Engineers, Product, Data Science, System Engineering... 
    Suggested
    Full time
    Remote work

    Waymo

    San Francisco, CA
    1 day ago
  •  ...the data collection systems, operational capability, exabyte-scale data warehouse, and software toolchain – to help our partners drive the field forward. Infrastructure engineers build the platform that support our ever-increasing scale of data collection. Sample... 
    Suggested
    Full time

    Xdof

    San Francisco, CA
    1 day ago
  •  ...Exa is building a search engine from scratch to serve every AI agent. We build massive-scale infrastructure to crawl the web, train state-of-the-art embedding models to process...  ...of machines Design GPU scheduling software so we max out our cluster utilization Build... 
    Suggested
    Full time
    H1b

    Exa

    San Francisco, CA
    1 day ago
  • $215k - $265k

     ...with them on scheming mitigations. We're looking for a Software Engineer to build the platform that the rest of Apollo runs on . This...  ..., and in what order. Build and maintain Apollo's cloud infrastructure . This means IaC, networking, environment management,... 
    Suggested
    Full time
    Work experience placement
    Work at office
    Immediate start
    Visa sponsorship
    Flexible hours

    Apollo Research

    San Francisco, CA
    1 day ago
  •  ...economy has grown 6x over the last decade, software businesses have gone from not worrying...  .... We're building the agentic infrastructure layer at Anrok—the systems that let AI...  ...product builds on. We're looking for an engineer who pairs a strong distributed-systems... 
    Full time
    Work at office
    Remote work
    Home office
    Flexible hours
    3 days per week

    Anrok, Inc.

    San Francisco, CA
    1 day ago
  • $204k - $259k

     ...15+ U.S. states. The Simulation ML Infrastructure team builds scalable AI/ML infrastructure...  ...systems. This role reports to an Engineering Manager.   You will:  Be part of...  ...prefer: ~5+ years of professional software engineering experience, with at least 3... 
    Full time
    Remote work

    Waymo

    San Francisco, CA
    13 hours ago
  •  ...coding agents that replace traditional software development by generating, testing, and...  ...You'll Be Responsible For Platform & Infrastructure Maintain stability of our platform...  ...~4+ years of software/platform engineering experience with production systems ~... 
    Full time
    Flexible hours

    Emergent Labs

    San Francisco, CA
    1 day ago
  •  ...reliability, performance, and security for multi‑tenant compute. What You’ll Do Design and operate secure, multi‑tenant container infrastructure with fast startup and smart autoscaling. Ship cloud deployments (Helm/Terraform) with SSO, network controls, and audit... 
    Full time
    Remote work

    Julius Ai

    San Francisco, CA
    1 day ago
  •  ...About Flow Flow Engineering is an AI-native requirements platform for modern engineering organizations, enabling hardware...  ...and speed.​ About the role Flow is hiring a Software Engineer with an infrastructure focus to build and scale the core platform behind Flow... 
    Full time
    Flexible hours

    Flow Engineering

    San Francisco, CA
    1 day ago
  •  ...teams for Google Workspace.  What you'll do The Payments Infrastructure team builds the trust boundary between a live conversation...  ...plaintext cardholder data. Make payments something other engineers can use without becoming compliance experts: drive the platform... 
    Full time
    Flexible hours

    Sierra Limited

    San Francisco, CA
    1 day ago
  •  ...Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable...  ...Conviction. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As a Software Engineer at on the Training Infrastructure... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    1 day ago
  •  ...About the Role We are seeking a Cloud Infrastructure Engineer to help design and evolve the platforms that power OpenAI’s products. In this role, you will be a hands-on technical leader, driving the architecture, scalability, reliability, and security of critical... 
    Full time

    OpenAI

    San Francisco, CA
    1 day ago
  •  ...into real, working systems and build any software needed for running large-scale frontier...  ...About the Role We are looking for engineers to operate the next generation of compute...  ...systems engineering with hands-on infrastructure work on our largest datacenters. You will... 
    Full time

    OpenAI

    San Francisco, CA
    1 day ago
  •  ...government agencies address real-world challenges. The Infrastructure Engineering team is crucial to the overall success of Hayden products:...  ...Uphold Engineering Standards: Establish and enforce elite software engineering and DevOps standards, including rigorous... 
    Full time
    Shift work

    Hayden Ai

    San Francisco, CA
    1 day ago
  •  ...About the Role This role broadly owns infrastructure across the stack. If it’s running in the cloud, you probably care about it. The life...  ...ChatGPT and GPT-4 to hundreds of millions of users, engineered the foundations of autonomous driving, built next-generation... 
    Full time

    The Generalist

    San Francisco, CA
    1 day ago
  • $250k

     ...product-market fit with a substantial customer pipeline already in place.   Role Overview We’re looking for an ML infrastructure engineer to design and build the core systems that enable scalable, efficient training of large models for deployment and research.... 
    Full time

    Epsilon Labs, Inc.

    San Francisco, CA
    1 day ago
  • $209k - $240k

     ...office workdays. About the Product Infrastructure Team: The Product Infrastructure...  ...classes of problems up-front for product engineers. Solve hard technical challenges such...  ...values, and enthusiastic about making software toolmaking ubiquitous, we want to hear... 
    Full time
    Work at office
    Local area

    Notion

    San Francisco, CA
    1 day ago
  •  ...Background Specter is creating a software-defined "control plane" for the physical...  ...will ultimately become the perception engine for a company's physical footprint, enabling...  ...Specter is hiring an ML infrastructure engineer to build and scale the machine... 
    Full time

    S.e. Specter

    San Francisco, CA
    1 day ago
  •  ...business with billions in revenue The Role Handshake is building the infrastructure layer that powers the next generation of AI agents across our platform. As a Senior Software Engineer on our Agentic Infrastructure team, you'll be at forefront of AI at... 
    Full time
    Freelance
    Internship
    Work at office
    Remote work
    Flexible hours

    Handshake

    San Francisco, CA
    13 hours ago
  • $175k - $215k

     ...redefines the future of global mobility. The Waymo Logs Infrastructure powers critical decision making and model development...  ...and clients   You have: ~4+ years of professional software engineering experience ~ Experience working on large-scale distributed... 
    Full time
    Remote work

    Waymo

    San Francisco, CA
    1 day ago
  •  ...of the great threats AI presents: mass-manufactured social engineering. Countless scams, deepfakes, and other social engineering attacks...  ...for an experienced backend engineer to build out the infrastructure required to rapidly scale up our engineering organization. A... 
    Full time
    Work at office
    Flexible hours

    Doppel

    San Francisco, CA
    1 day ago
  • $160k - $180k

     ...out before, during, and after playing games. The Database Infrastructure team develops and operates all of Discord’s databases and data...  ...matters most to our users. Work with a talented team of engineers who have built one of the largest communication platforms in... 
    Full time

    Discord

    San Francisco, CA
    1 day ago
  •  ..., and Google Lens. Before that, Clay led the product and design teams for Google Workspace.  What you’ll do As a Software Engineer, Infrastructure at Sierra, you will be responsible for designing, building, and maintaining the core systems that make our AI platform... 
    Full time
    Flexible hours

    Sierra Limited

    San Francisco, CA
    1 day ago
  • $190k - $260k

     ...Notion. Centralize was founded by Rachit Kataria, a founding engineer on Facebook Shops who helped scale it to 250M MAUs, and Will...  ...what Centralize becomes. The Role We are hiring an infrastructure engineer to own scalability across Centralize. Our system processes... 
    Remote job
    Full time
    Relocation
    Visa sponsorship

    Centralize

    San Francisco, CA
    1 day ago
  •  ...an experienced, creative, and versatile engineer who is eager to tackle the challenge of...  ..., building, and scaling the core infrastructure and systems powering Lightfield's AI-driven...  ...Who you are ~3+ years of experience in software development, with a strong background... 
    Full time
    Work from home

    Lightfield

    San Francisco, CA
    1 day ago
  •  ...Runloop.ai is building the foundational infrastructure for the next generation of AI development. We provide AI engineers and data scientists with lightning-fast, secure...  ...innovation in the age of AI. The Role As a Software Engineer on our Infrastructure team, you'll... 
    Full time
    Work at office
    Work from home
    1 day per week

    Runloop

    San Francisco, CA
    1 day ago
  •  ...including stablecoins. You’ll help design, deploy and operate the infrastructure that makes this possible. This is a hands-on devops role...  ...role sits at the intersection of infrastructure, platform engineering and security. You’ll define how Modern Treasury scales and... 
    Full time
    Local area
    Immediate start
    Remote work

    Modern Treasury

    San Francisco, CA
    1 day ago
  •  ...About the Team We’re hiring software engineers to join our broader Infrastructure organization, which supports multiple high-impact teams. Depending on your interests and experience, you could work on one of several focus areas—including Core Distributed Systems, Databases... 
    Full time

    OpenAI

    San Francisco, CA
    1 day ago
  • $100k - $300k

     ...Abnormal AI, Zscaler Preeminent research labs like Deepmind and SAIL About the Role We're hiring a Senior Storage Infrastructure Engineer on our Core Platform team to own how we store, protect, and operate data at scale. You'll build the backup, observability,... 
    Full time

    Cogent Security

    San Francisco, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Software Engineer, Infrastructure. Be the first to apply!