Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Infrastructure Engineer (Storage)

$180k - $200k

Lightning AI

New York, New York, United States; Remote; San Francisco, California, United States; Seattle, Washington, United States Who We Are Lightning AI is the company behind PyTorch Lightning. Founded in 2019, we build an end-to-end platform for developing, training, and deploying AI systems—designed to take ideas from research to production with less friction. Through our merger with Voltage Park, a neocloud and AI Factory, Lightning AI combines developer-first software with cost-efficient, large-scale compute. Teams get the tools they need for experimentation, training, and production inference, with security, observability, and control built in. We serve solo researchers, startups, and large enterprises. Lightning AI operates globally with offices in New York City, San Francisco, Seattle, and London, and is backed by Coatue, Index Ventures, Bain Capital Ventures, and Firstminute. Our Values Move Fast : We act with speed and precision, breaking down big challenges into achievable steps. Focus : We complete one goal at a time with care, collaborating as a team to deliver features with precision. Balance : Sustained performance comes from rest and recovery. We ensure a healthy work-life balance to keep you at your best. Craftsmanship : Innovation through excellence. Every detail matters, and we take pride in mastering our craft. Minimal : Simplicity drives our innovation. We eliminate complexity through discipline and focus on what truly matters. What We're Looking For Lightning AI is seeking a Storage Infrastructure Engineer to join our Infrastructure Engineering team. In this role, you will focus on building and operating the storage systems that power large-scale AI/ML training, inference, and HPC workloads. You will work at the intersection of software, hardware, and operations—developing automation, improving reliability, and scaling distributed storage systems across our bare-metal infrastructure. You will help own the data plane of our storage infrastructure, supporting high-throughput, low-latency data access for some of the most demanding AI workloads. You’ll play a key role in managing and evolving our storage stack (including VAST and S3-compatible systems like Ceph), ensuring performance, reliability, and efficiency at scale. We’re flexible on location for this team. This role can work hybrid out of one of our US-based hubs (Seattle, NYC, or SF) or fully remote within the U.S., with occasional company and team offsites. We are not able to provide visa sponsorship for this position at this time. What You'll Do Storage Systems & Infrastructure Operate and scale distributed storage systems, including VAST and S3-compatible object storage (e.g., Ceph) Improve performance, reliability, and efficiency of storage systems supporting large-scale AI/ML workloads Troubleshoot complex storage and data path issues across hardware and software layers Optimize storage performance to support high-throughput, low-latency AI training and inference workloads Automation & Tooling Build and maintain automation for provisioning, managing, and monitoring storage infrastructure Develop Python-based tools and workflows to reduce manual operational overhead Improve lifecycle management of storage clusters, from deployment through maintenance and scaling Systems & Operations Manage and operate Linux-based systems in production, including bare-metal environments Partner with infrastructure and data center teams on hardware bring-up, upgrades, and issue resolution Support capacity planning, utilization tracking, and forecasting for storage systems Leverage monitoring and telemetry to diagnose issues and improve system performance and reliability Cross-Functional Collaboration Work closely with Infrastructure Engineering, Network Engineering, and Platform teams to integrate storage into the broader platform Contribute to design discussions around new infrastructure deployments and scaling strategies Help define best practices for operating storage systems in high-performance computing environments What You'll Need Required Qualifications 5+ years of experience in infrastructure engineering, systems engineering, or related roles Hands-on experience operating distributed storage systems (e.g., VAST, Ceph, or similar) Proficiency in Python or similar scripting/programming languages for automation Experience working with bare-metal infrastructure and hardware-oriented systems Ability to debug complex issues across system boundaries (storage, OS, hardware, networking) Experience with storage networking protocols (e.g., NFS or similar) Experience with capacity planning, monitoring, and performance tuning Ideal Experience Experience with VAST storage systems in production environments Experience operating S3-compatible object storage at scale Data center operations experience, including working with physical hardware Familiarity with AI/ML or HPC workloads and their storage requirements Background in high-performance or low-latency distributed systems Familiarity with high-performance data transfer technologies (e.g., RDMA, GPU Direct Storage) Experience supporting GPU-based workloads or large-scale compute clusters Compensation We are committed to offering competitive compensation that reflects the value each team member brings to our mission. Final offers are based on factors such as experience, skills, geographic location, and role expectations. In addition to base salary, our total rewards package for eligible roles includes a discretionary bonus, a meaningful equity component, and comprehensive benefits. The anticipated annual base salary range for this role is: $180,000 - $200,000 USD Benefits and Perks We offer a comprehensive and competitive benefits package designed to support our employees’ health, well-being, and long-term success. Benefits may vary by location, team, and role. Comprehensive medical, dental and vision coverage (U.S.); Private medical and dental insurance (U.K.) Retirement and financial wellness support (U.S.); Pension contribution (U.K.) Generous paid time off, plus holidays Paid parental leave Wellness and work-from-home stipends Flexible work environment At Lightning AI, we are committed to fostering an inclusive and diverse workplace. We believe that diverse teams drive innovation and create better products. We provide equal employment opportunities to all employees and applicants without regard to race, color, religion, gender, sexual orientation, gender identity, national origin, age, disability, veteran status, or any other protected characteristic. We are dedicated to building a culture where everyone can thrive and contribute to their fullest potential. #J-18808-Ljbffr

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Infrastructure Engineer (Storage) in San Francisco, CA vacancy
  •  ...Description You will work closely with our engineering team across the autonomy stack to...  ...systems and customer deployments. Being an infrastructure specialist, you will build out...  ...initiatives, optimize our cloud compute and storage systems, and own our administration and... 
    Suggested
    Local area

    Aerovect

    San Francisco, CA
    1 day ago
  • $170k - $230k

     ...innovative virtual payment products. Our infrastructure will be across multiple continents over...  ...services, including all of compute, storage, networking, container orchestration frameworks...  ...years of experience in infrastructure engineering or software engineering ~ Strong... 
    Suggested
    Work at office
    Local area
    Home office
    Flexible hours

    Highnote

    San Francisco, CA
    2 days ago
  •  ...Lens. Before that, Clay led the product and design teams for Google Workspace.  What you’ll do As a Forward Deployed Infrastructure Engineer at Sierra, and a founding member of this function, you will own the end-to-end lifecycle of customer deployments, while helping... 
    Suggested
    Full time
    Flexible hours

    Sierra

    San Francisco, CA
    9 days ago
  •  ...With your expertise in delivering infrastructure solutions, you are a top-performer in your...  ...winning team. As an Infrastructure Engineer II at JPMorgan Chase within the Infrastructure...  ..., networking terminology, databases, storage engineering, deployment/integration... 
    Suggested

    J.P. Morgan

    San Francisco, CA
    7 days ago
  •  ...goals. Whoever deploys frontier compute infrastructure fastest will decide whether AI expands...  ...the Low Voltage Contractor and internal engineering teams, and act as remote hands during...  ...stacked, and cabled servers, network, and storage hardware at scale. You read and... 
    Suggested
    For contractors
    Local area
    Remote work

    FluidStack

    San Francisco, CA
    7 hours ago
  • $150k - $250k

     ...Fluidstack At Fluidstack, we’re building the infrastructure for abundant intelligence. We partner...  ...The Role Infrastructure Deployment Engineer is a Subject Matter Expert (SME) in fiber...  ...racks, servers, network devices, and storage units, are installed and grounded according... 
    Full time
    For contractors
    Local area
    Remote work

    FluidStack

    San Francisco, CA
    2 days ago
  • $109k - $186k

     ...Deployment Engineer To achieve our mission, we must ensure our network deployments are designed for performance and scale, as well...  ...to scale and meet the growing demand for reliable internet infrastructure. You Will Have An Impact By Designing performant and... 

    Meter Service

    San Francisco, CA
    5 days ago
  • $226.5k - $339.7k

     ...Stripe is a financial infrastructure platform for businesses that powers millions of companies—from large enterprises to ambitious startups...  ...network issues, or collaborate with Security and Engineering teams to implement solutions that support our rapid growth.... 
    Work at office
    Local area

    Stripe

    San Francisco, CA
    2 days ago
  • $125k - $195k

     ...Infrastructure & Site Reliability Engineer The fab2 team can build anything. We design all the hardware and software needed to make chips and we make...  ...discovery, DNS, reverse proxies, TLS, S3 compatible storage, VPNs Scale our observability platform: Build systems... 
    Work at office
    Visa sponsorship
    Night shift

    Fab2

    San Francisco, CA
    3 days ago
  •  ...Join a scaling startup building systems that need to stay fast, reliable, and cost-efficient as usage ramps. We’re hiring an Infrastructure Engineer to own a Kubernetes-based platform-improving developer velocity, production reliability, and the cloud foundations we scale... 

    Harrison Clarke

    San Francisco, CA
    2 days ago
  • $210k

     ...Senior Infrastructure Engineer | San Francisco (Hybrid) | Up to $210,000 Join a growing Communications business in San Francisco, who builds device that enable resilient communications in the world’s harshest environments. As a Senior Infrastructure Engineer, you’ll join... 
    2 days per week
    3 days per week

    Edison Smart®

    San Francisco, CA
    2 days ago
  • $120k - $200k

     ...Capital, Scale Venture Partners, YC, the founders of Twilio, Affirm, ElevenLabs, and many more. About the Role As a Senior Infrastructure Engineer at Bland, you’ll help us build the backbone that enables millions of AI‑powered phone conversations. You’re not just keeping... 
    Work at office
    Night shift

    Bland Company

    San Francisco, CA
    4 days ago
  •  ...Chalk Infrastructure Engineer Chalk is building the data platform that powers the future of machine learning applications. We tear down complexity, latency, and scale barriers that have traditionally constrained ML capabilities. Our platform combines Rust-speed performance... 
    Work at office
    Flexible hours

    CHALK INC

    San Francisco, CA
    1 day ago
  •  ...Judgment Labs builds infrastructure for Agent Behavior Monitoring (ABM) . While traditional observability focuses on logging exceptions...  ...others. The Role We are looking for a Senior Infrastructure Engineer to architect and scale the deployment infrastructure that powers... 

    Judgment Labs

    San Francisco, CA
    2 days ago
  •  ...Navi Infrastructure Engineer Navi captures everything a pilot sees and hears and turns it into automated debrief intelligence. The platform is live at flight schools, business jet operators, and the U.S. Air Force — and demand is accelerating. You're the person who... 

    Navi Ai

    San Francisco, CA
    3 days ago
  •  ...List. We’re a quickly‑growing team headquartered in Union Square, San Francisco. Role Summary Weekend is seeking a Senior Infrastructure Engineer to ensure the stability, security, and scalability of our infrastructure as our AI gaming platform grows. You'll own our Kubernetes... 
    Work at office
    Remote work
    Work from home
    Relocation
    Visa sponsorship
    Flexible hours

    Weekend

    San Francisco, CA
    1 day ago
  • $7.5k

     ...working together in the office. The Founders: The founders are engineers who have run their own alternative school together. One...  ...ESA department and leading school choice advocates The Role: Infrastructure Engineer As Infrastructure Engineer, you will own every dollar... 
    Work at office
    Relocation
    Visa sponsorship
    Relocation package
    Day shift

    EQL Tech (sales & engineering talent)

    San Francisco, CA
    2 days ago
  • $130k - $240k

     ...Maxana is seeking an experienced Infrastructure Engineer for a confidential client — a fast-growing AI company. In this role you will build and maintain the platform layer supporting large-scale ML training, inference, and deployment. This is a high-impact role at the... 
    Flexible hours

    Maxana

    San Francisco, CA
    7 hours ago
  • $130k

     ...collectively earn over $3 million a day. About The Role As an Infrastructure Engineer at Mercor, you’ll build and scale the systems that power...  ...Identifying and fixing performance bottlenecks in compute, storage, and networking. What We’re Looking For Strong experience with... 
    Work at office
    Relocation package

    Mercor Inc

    San Francisco, CA
    1 day ago
  • $225k - $250k

     ...open‑source AI coding agent. Trusted by the world's largest engineering organizations and individual developers alike, Cline is model...  ...engineers everywhere they work while keeping code on your infrastructure. Teams choose Cline because it is the only coding agent that... 

    Céline

    San Francisco, CA
    7 hours ago
  •  ...Infrastructure Engineer III Need local or someone willing to relocate: Job Posting ID: ADSKJP00003698 Job Posting Title: Infrastructure Engineer III Max Pay Rate: $75/hr on w2 all inclusive Location: San Francisco, CA Client: Autodesk Job Title: Network Cloud Engineer... 
    Work experience placement
    Local area
    Relocation

    Omega Solutions Inc

    San Francisco, CA
    2 days ago
  • $150k - $250k

     ...Our client is building next-generation financial infrastructure that allows businesses to move money as easily as they move data. Their...  ...distributed workloads Collaborate closely with a small, experienced engineering team Improve internal tooling, CI/CD pipelines, and... 
    Work at office
    2 days per week
    3 days per week

    Eden Prescott

    San Francisco, CA
    2 days ago
  •  ...Job Description Senior Infrastructure Engineer serves as a key contributor in the planning, design, construction, renovation, and delivery of infrastructure across UCSF facilities. This role oversees the build‑out, expansion, repair, and modernization of infrastructure... 
    Work experience placement
    Local area
    Worldwide

    UCSF Health

    San Francisco, CA
    58 minutes ago
  •  ...paced, high‑intensity environment. If this resonates, you belong at HappyRobot. About The Role We're looking for a Site Reliability Engineer to take the lead on scaling our operational resilience as we grow. You’ll own the stability, observability, and debugging... 
    Full time
    Shift work

    Happy Robot

    San Francisco, CA
    2 days ago
  •  ...Job Overview Senior Infrastructure Engineer serves as a key contributor in the planning, design, construction, renovation, and delivery of infrastructure across UCSF facilities, overseeing the build‑out, expansion, repair, and modernization of infrastructure services to... 
    Local area

    University of California , San Francisco

    San Francisco, CA
    2 days ago
  • $180k - $250k

     ...Senior Infrastructure Engineer Title of Role: Senior Infrastructure Engineer Location: San Francisco, hybrid Company Stage of Funding: Seed - AI, Devtools, Enterprise, Data Office Type: Hybrid Salary: $180K-$250K Company Description We're representing... 
    Work at office

    Recruiting from Scratch

    San Francisco, CA
    1 day ago
  •  ...We're looking for Infrastructure Engineers who will be instrumental in building and securing the backbone of our enterprise-grade AI data platform. You’ll design systems that handle large volumes of sensitive financial data under strict security and compliance requirements... 
    Work at office

    Rowspace

    San Francisco, CA
    2 days ago
  •  ...technology and a strong can‑do attitude. About the role Most infrastructure roles ask you to maintain what exists. This one asks you to rethink...  ...problem. It is a living system, and we are looking for an engineer who sees that as an opportunity rather than a burden. As an... 
    Casual work
    Flexible hours

    Careers

    San Francisco, CA
    1 hour ago
  •  ...Title: Infrastructure Engineer Duration: Fulltime Location: Bay Area, CA Client : LS / META Required: * Bachelor's degree in computer science or related technical discipline * 4+ years of technical engineering experience... 
    Full time

    Concord IT Systems

    San Francisco, CA
    7 hours ago
  • $170k - $230k

    We need an infrastructure engineer to build the platform that makes our agents fast, reliable, and scalable. You'll own the systems that deploy, monitor, and orchestrate hundreds of concurrent agent workflows processing sensitive enterprise data. Compensation range for... 

    Optimized, Inc.

    San Francisco, CA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Infrastructure Engineer (Storage). Be the first to apply!