Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Software Engineer, Infrastructure

$180k - $250k

fal

Fal Generative Media Engineer

Fal is the generative media ecosystem powering the next generation of AI products. We build the infrastructure, tools, and model access that teams need to move from idea to production, and do it at scale without compromise. For developers and enterprises, fal is the foundation that makes generative media not just possible, but practical: a unified platform where high-performance inference, orchestration, and observability come together to unlock new categories of AI-native products.

As generative media reshapes industries across a market projected to grow by hundreds of billions over the next decade, fal is becoming the ecosystem that ambitious teams build on.

About This Role

You are a hands-on engineer who builds the software and processes that keep a large fleet of GPU servers healthy and productive. You write systems and tooling for managing 1000s of servers including provisioning, health monitoring, error detection, and recovery — and when something breaks that automation can't fix, you drive resolution with partners.

What You'll Do
  • Build and maintain Python fleet tracking system that manages the full lifecycle of servers including contracting and procurement, target use, pricing, availability, health, RMAs, etc

  • Build server management tooling that automates provisioning, health checks, GPU diagnostics, recovery and alerting

  • Create and maintain metrics, dashboards, and alerting for hardware health across the fleet (GPU errors, disk failures, network issues, thermals)

  • Leverage AI to an extreme level to build tools and automate alerting and recovery

  • Implement and enforce OS-level security: hardening baselines, SELinux/AppArmor policies, SSH key management, vulnerability scanning, and compliance automation

  • Manage and optimize distributed and local storage systems supporting model weights, checkpoints, and ephemeral scratch: NVMe arrays, NFS, parallel file systems, and object storage

  • Tune Linux systems for AI workloads: kernel parameters, NUMA topology, CPU pinning, hugepages, I/O schedulers, and GPU driver stack optimization (NVIDIA drivers, CUDA, container runtimes)

  • Develop a suite of automated error detection and recovery processes

  • Work with partners to solve technical issues

Qualifications/Nice to Have
  • 3+ years experience managing bare-metal and cloud based server fleets at scale (100+ nodes)

  • Strong software engineering skills in Python; you write production tooling, not scripts

  • Deep Linux systems knowledge: boot process, kernel tuning, networking, storage, systemd, cgroups, namespaces, performance profiling

  • Strong experience with configuration management and infrastructure-as-code: Ansible, Terraform, cloud-init

  • Solid understanding of storage technologies: LVM, RAID, NVMe, NFS, Lustre or GPFS, and Linux I/O stack tuning

  • Familiarity with hardware diagnostics and failure modes (GPUs, NVMe, NICs, memory)

  • Experience building internal tools or dashboards for infrastructure visibility

  • Excellent communication and ability to drive technical decisions across teams

  • Self-starter who executes quickly, takes ownership, and constantly seeks improvement

  • Familiarity with network configuration and diagnostics (VLAN, VXLAN, ECMP, BGP, tcpdump)

  • Experience with NVIDIA GPU infrastructure: driver management, health monitoring, DCGM, NVLink/NVSwitch diagnostics, RDMA, InfiniBand/RoCEv2

  • Experience with AMD GPUs

  • Experience with bare metal and VM provisioning (PXE/iPXE, Kickstart, libvirt, Qemu/KVM)

  • Experience with compliance frameworks relevant to cloud providers (SOC 2, ISO 27001)

What We Offer At Fal
  • Interesting and challenging work

  • A lot of learning and growth opportunities

  • We are offering relocation assistance to San Francisco.

  • Health, dental, and vision insurance (US)

  • Regular team events and offsites

Compensation
  • $180,000-250,000 plus equity + comprehensive benefits package

U.S. EQUAL EMPLOYMENT OPPORTUNITY INFORMATION:

Fal provides equal employment opportunities to applicants and employees without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, protected veteran status, disability, or any other classification protected by applicable law.

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Software Engineer, Infrastructure in San Francisco, CA vacancy
  •  ...Software Engineer At Anyscale, we're on a mission to democratize distributed computing and make it accessible to software developers...  ...is looking for a Software Engineer to join the Platform and Infrastructure team. Anyscale aims to provide the next generation of tools... 
    Suggested
    Flexible hours

    Anyscale

    San Francisco, CA
    3 days ago
  • $180k - $280k

     ...Do Design and maintain the core infrastructure across our cloud environments - own it...  ...the limits before our growth does and engineer past them. Own developer velocity. Build...  .... ~ You highly leverage AI for software development and can juggle multiple workstreams... 
    Suggested
    Work at office
    Relocation

    Pylon

    San Francisco, CA
    4 days ago
  • $168k - $213k

     ...team comes from world-class organizations like Google, Netflix, Stripe, Plaid, Brex, and more. The Role As a Senior Infrastructure Engineer , you will own the foundation that enables us to deploy secure, compliant AI systems at major financial institutions... 
    Suggested
    Flexible hours

    Bretton AI

    San Francisco, CA
    4 days ago
  •  ...Job Title: Software Engineer - Infrastructure Employment type: Full-time Location: San Francisco, CA - hybrid, 3 4 days/week on-site, with flexibility Seniority: 4+ years of software engineering building distributed systems About this role... 
    Suggested
    Full time
    3 days per week

    Apptad Inc

    San Francisco, CA
    1 day ago
  • $260k - $300k

     ...We are an applied AI lab building end-to-end software agents. We're the makers of Devin, the first AI software engineer. Our team is extremely talent-dense. Among...  ...Google DeepMind, and others. Role Mission Infrastructure Engineers at Cognition build the systems that... 
    Suggested

    Cognition AI

    San Francisco, CA
    1 day ago
  •  ...Software Engineer, Infrastructure Our mission is to automate coding. The first step in our journey is to build the best tool for professional programmers, using a combination of inventive research, design, and engineering. Our organization is very flat, and our team... 
    Work at office

    Anysphere

    San Francisco, CA
    2 days ago
  • $250k - $300k

     ...Job Description Job Description Software Engineer - Marketplace (Growth Infrastructure) Company: AfterQuery Location: San Francisco, CA (FiDi office, Monday-Friday; half day or remote on Sunday) Compensation: $250,000 - $300,000 base ($300,000 - $400,000... 
    Full time
    H1b
    Work at office
    Remote work
    Visa sponsorship
    Monday to Friday

    Transparent Search Group

    San Francisco, CA
    3 days ago
  • $160k - $300k

     ...Infrastructure Engineer We are looking for an infrastructure engineer who treats cloud infrastructure as a software problem. You will own Hebbia's AWS footprint end-to-end — accounts, networking, container orchestration, and everything defined in code — and the developer... 

    Hebbia

    San Francisco, CA
    2 days ago
  • $208k - $260k

     ...educational institutions Work together with engineers, scientists, operators, and more from...  ...Handshake AI Human data is the core infrastructure to AI advancement. Frontier AI labs...  ...The Role Handshake is hiring a Senior Software Engineer to join Backend Platform, with... 
    Full time
    Work at office
    Remote work
    Flexible hours

    Handshake

    San Francisco, CA
    8 days ago
  • $184k - $259.44k

     ...Scale AI is seeking a highly skilled and motivated Software Engineer, Frontier AI Infrastructure to join our dynamic Public Sector Engineering team. As a part of this team, you will own the model inference layer - enabling state of the art models, debugging the latest... 
    Full time
    Work at office
    3 days per week
    Early shift

    Scale AI

    San Francisco, CA
    1 day ago
  •  ...About the Team We’re hiring Software Engineers to join our broader Infrastructure organization, which supports multiple high-impact teams. Depending on your interests and experience, you could work on one of several focus areas—including Core Distributed Systems,... 

    OpenAI

    San Francisco, CA
    3 days ago
  • $300k

     ...exciting new opportunity?  Join a stealth-mode hyperscale infrastructure startup building a 300MW+ AI compute platform designed to...  ...infrastructure projects currently under development. The Principal Software Engineer will take ownership of the software architecture that... 
    Full time
    Remote work
    Flexible hours
    San Francisco, CA
    more than 2 months ago
  •  ...legacy systems with adaptive, learning software. Founded in early 2024, Serval is already...  ...IT, HR, Finance, Security, Legal, and Engineering. Our mission is to eliminate...  ...modern enterprises. As a Software Engineer, Infrastructure, you’ll build and scale the foundational... 

    Serval

    San Francisco, CA
    22 hours ago
  • $183k - $220k

     ...out of every 10 workers, but lacks good software tools to get the job done. Payments in...  ...senior member of our small but mighty engineering team, you'll work closely with our cross...  ...evolution of our backend tech stack and infrastructure. Whether iterating on core workflows,... 
    Full time
    For contractors
    Work at office
    Remote work

    Siteline

    San Francisco, CA
    10 days ago
  • $165k - $185k

     ...ship, and believe the best outcomes come from building together. About The Role We’re looking for a Senior Software Engineer, Infrastructure to own the platform that VSCO product and data teams ship on. You’ll join a small infra team that treats AWS, EKS, and... 
    Full time
    Temporary work
    Work at office
    Local area
    Worldwide
    Flexible hours

    VSCO

    San Francisco, CA
    5 days ago
  • $209k - $235k

     ...The Role We're looking for a Senior Software Engineer to join our ML Infrastructure team and support the foundational infrastructure that powers Gridmatic. Our platform challenges are shaped by the nature of energy markets: forecasts and trading decisions run on tight... 
    Full time
    Home office
    Flexible hours

    Gridmatic Inc

    San Francisco, CA
    2 days ago
  • $266k

     ...About the Team The Search Product Infrastructure team builds the systems that power search experiences across ChatGPT. We partner with...  ...efficiency at ChatGPT scale. About the Role As a senior engineer on the Search Product Infrastructure team, you will design, build... 
    Full time
    Work at office
    Relocation package

    OpenAI

    San Francisco, CA
    2 days ago
  • $150k - $200k

     ...more time for their life's work. About the Role: The Product Infrastructure team works on creating abstractions and data models that...  ...by handling entire classes of problems up-front for product engineers. Solve hard technical challenges such as designing abstractions... 
    Local area

    Notion Labs, Inc

    San Francisco, CA
    12 hours ago
  •  ...reliability, performance, and security for multi-tenant compute. What You'll Do Design and operate secure, multi-tenant container infrastructure with fast startup and smart autoscaling. Ship cloud deployments (Helm/Terraform) with SSO, network controls, and audit logging.... 
    Remote work

    Julius

    San Francisco, CA
    12 hours ago
  • What to expect This is a full-time on-site role located in San Francisco for a Senior Software Engineer, Infrastructure. You will be responsible for scaling our infrastructure, owning major architecture decisions, and building the platform on which product features are... 
    Full time

    Orion Sleep

    San Francisco, CA
    22 hours ago
  • $190k - $221k

    Senior Software Engineer, Infrastructure TRM Labs is a blockchain intelligence company committed to fighting crime and creating a safer world. By leveraging blockchain data, threat intelligence, and advanced analytics, our products empower governments, financial institutions... 
    Remote work

    TRM Labs

    San Francisco, CA
    2 days ago
  • Platform Engineer HUD is building infrastructure to create RL training data and evals for frontier AI agents, as well as a marketplace to sell these to...  ...Can write clean, maintainable code and apply strong software engineering judgment across product architecture, infrastructure... 
    Full time
    Work at office
    Remote work
    Relocation
    Visa sponsorship

    Hud (yc W25)

    San Francisco, CA
    22 hours ago
  • Software Engineer Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence...  ...Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies... 
    Flexible hours

    Baseten

    San Francisco, CA
    12 hours ago
  •  ...a proven product-market fit with a substantial customer pipeline already in place. Role Overview We're looking for an ML infrastructure engineer to design and build the core systems that enable scalable, efficient training of large models for deployment and research.... 

    Epsilon Health

    San Francisco, CA
    12 hours ago
  •  ...Impact: You will own and architect core infrastructure systems that power our platform from...  ...you'll have the opportunity to shape the engineering organization and lead major technical initiatives...  ...researchers. Overview As a Senior Software Engineer - Infrastructure at AfterQuery... 
    Local area

    AfterQuery

    San Francisco, CA
    22 hours ago
  •  ...About the Team The Applied Engineering team works across research, engineering, product, and design to bring OpenAI's technology...  .... You'll join the team responsible for running the core infrastructure that supports products like ChatGPT and the API. The systems... 
    Relocation package

    OpenAI

    San Francisco, CA
    2 days ago
  • Senior Software Engineer At Commure, we're building the AI Operating System for healthcare, the foundation that defines how care is delivered...  ...the Role We're hiring a Senior Software Engineer on the Infrastructure team to own the foundational infrastructure and internal... 
    Local area
    Immediate start

    Commure

    San Francisco, CA
    22 hours ago
  • Software Engineer Voxel's perception system is the technical core of everything we ship. Our models detect human activity, equipment interactions...  .... We're hiring a strong software engineer to own the ML Infrastructure that powers how Voxel trains and ships vision models. You'... 
    Work at office
    Flexible hours

    Voxel

    San Francisco, CA
    22 hours ago
  • $100k - $300k

     ...from Wiz, Abnormal AI, Zscaler Preeminent research labs like Deepmind and SAIL About The Role We’re hiring a Senior Storage Infrastructure Engineer on our Core Platform team to own how we store, protect, and operate data at scale. You’ll build the backup, observability,... 

    Cogent

    San Francisco, CA
    4 days ago
  • Software Engineer, Client Infrastructure Engineering · Full-time · San Francisco; New York Our mission is to automate coding. The first step in our journey is to build the best tool for professional programmers, using a combination of inventive research, design, and engineering... 
    Full time
    Work at office

    Anysphere

    San Francisco, CA
    12 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Software Engineer, Infrastructure. Be the first to apply!