Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Sr. Site Reliability Engineer (SRE)

$165k - $225k

Moonlite AI

Sr. Site Reliability Engineer (SRE)

Chicago, IL or Remote

Moonlite delivers high-performance AI infrastructure for organizations running intensive computational research, large-scale model training, and demanding data processing workloads. We provide infrastructure deployed in our facilities or co-located in yours, delivering flexible on-demand or reserved compute that feels like an extension of your existing data center. Our team of AI infrastructure specialists combines bare-metal performance with cloud-native operational simplicity, enabling research teams and enterprises to deploy demanding AI workloads with enterprise-grade reliability and compliance.

Your Role:

You will be instrumental in building and operating production-grade AI infrastructure with deep Kubernetes expertise at its core. Working closely with our systems engineers, network engineers, and platform engineering team, you'll architect and operate the Kubernetes infrastructure that powers our control plane and orchestrates compute, storage, and networking at scale. This role requires deep understanding of Kubernetes internals, custom resource definitions (CRDs), storage and network integrations, and building production-grade clusters from the ground up (not just deploying in managed environments). You'll ensure enterprise-grade reliability while establishing the automation, observability, and operational practices.

Job Responsibilities
  • Kubernetes Infrastructure Engineering: Design, build, and operate production Kubernetes clusters on bare-metal infrastructure – including cluster bootstrapping, control plane architecture, etcd management, and scaling strategies for high-performance compute workloads.
  • Kubernetes Networking & CNIs: Implement and operate custom Kubernetes networking solutions with SR-IOV for high-performance GPU interconnects, multi-tenancy isolation and advanced networking policies. Configure CNI plugins and network segmentation for research workloads.
  • Custom Operators & Controllers: Develop and maintain custom Kubernetes operators and controllers for bare-metal provisioning, infrastructure lifecycle management, and resource orchestration across compute, storage, and networking domains.
  • GPU Infrastructure Integration: Deploy and optimize NVIDIA GPU operators, device plugins, and other custom scheduling logic for GPU workload placement and utilization optimization.
  • Platform Integration & Storage: Build deep integrations between Kubernetes and underlying infrastructure including CSI drivers for storage, custom admission controllers for policy enforcement, and scheduling extensions for specialized hardware placement.
  • Infrastructure Automation: Design and implement automation using Terraform, Ansible, Helm, and custom operators to orchestrate infrastructure workflows and enable deployments across multiple regions.
  • Production Operations & Reliability: Manage production bare-metal infrastructure across multiple regions. Build systems ensuring high availability, fault tolerance, and graceful degradation – establishing SLIs, SLOs, and monitoring to meet enterprise reliability commitments.
  • Observability & Incident Response: Build comprehensive monitoring, logging, and alerting using Prometheus, Grafana, and ELK stack. Lead incident response, conduct postmortems, and implement preventative measures to improve reliability and reduce MTTR.
  • Performance & Capacity Planning: Identify and resolve performance bottlenecks across infrastructure domains. Monitor utilization trends, forecast capacity needs, and optimize resource allocation for various workloads.
Requirements
  • Experience: 5+ years in SRE, DevOps, or infrastructure engineering roles with proven experience operating production infrastructure at scale.
  • Kubernetes Infrastructure Expertise: Deep hands-on experience building and operating production Kubernetes clusters on bare-metal infrastructure – not just deploying workloads in managed clusters. Must understand cluster bootstrapping, control plane architecture, etcd operations, and scaling strategies.
  • Kubernetes Internals & Integration: Strong understanding of Kubernetes internals including custom resource definitions (CRDs), operators, controllers, admission webhooks, and scheduling. Experience integrating storage (CSI drivers), networking (CNI, SR-IOV), and specialized hardware (GPU device plugins) with Kubernetes.
  • Linux Systems Experience: Strong fundamentals in Linux systems administration, performance tuning, troubleshooting, and automation in production environments.
  • Infrastructure Automation: Proficiency with infrastructure-as-code tools (Terraform, Ansible, Helm) and building automation to reduce operational overhead.
  • Networking Fundamentals: Solid understanding of networking concepts including IPAM, DNS, DHCP, VLAN/VXLAN, routing, load balancing, and experience troubleshooting network issues in production.
  • Observability & Monitoring: Experience building and maintaining comprehensive monitoring solutions using tools like Prometheus, Grafana, and centralized logging systems.
  • Reliability Practices: Understanding of SRE principles including SLIs/SLOs/SLAs, error budgets, incident management, and blameless postmortems.
  • Scripting & Automation: Strong scripting skills in Go, Python, or Bash for automation, tooling development, and operational efficiency.
  • Problem-Solving Under Pressure: Demonstrated ability to troubleshoot complex issues under pressure, manage incidents effectively, and communicate clearly during outages.
  • Collaboration & Communication: Excellent communication skills and ability to work across teams including systems engineers, network engineers, and software developers.
Preferred Qualifications
  • Experience building custom Kubernetes operators or controllers for infrastructure orchestration
  • Deep familiarity with Kubernetes networking (Calico, Cilium, Multus), service mesh technologies, and network policy management
  • Experience with GPU workload orchestration including NVIDIA GPU Operator, MIG, time-slicing, and device plugins
  • Background with advanced Kubernetes features including custom schedulers, admission controllers, and API server extensions
  • Experience with Kubernetes cluster federation or multi-cluster management
  • Knowledge of high-performance networking technologies (InfiniBand, RDMA, RoCE) and their integration with Kubernetes
  • Experience with enterprise storage systems (VAST, Lightbits, Ceph, or similar)
  • Familiarity with configuration management at scale and GitOps practices
  • Understanding of security best practices for Kubernetes and bare-metal infrastructure
  • Experience operating infrastructure in regulated industries or co-located data center environments
  • Background supporting research institutions, technical computing environments, or enterprise AI infrastructure
Key Technologies
  • Kubernetes, Linux, Terraform, Ansible, Prometheus, Grafana, ELK Stack, Go, Python, Bash, NVIDIA GPU Technologies, High-Performance Networking, Enterprise Storage Systems
Why Moonlite
  • Build Critical Research Infrastructure: Your work will directly enable quantitative research teams and AI practitioners to push the boundaries of what's possible in financial modeling and AI research.
  • Enterprise Impact: Build and operate infrastructure that supports mission-critical research and AI workloads for leading financial institutions and research organizations.
  • Technical Excellence: Join an infrastructure team focused on delivering enterprise-grade reliability while pushing the boundaries of high-performance computing capabilities.
  • Hands-On Ownership: As part of our growing infrastructure team, you'll have significant ownership over critical systems and the autonomy to influence our operational practices and technology choices.
  • Industry Leadership: Work alongside experienced infrastructure professionals who have built and operated systems for the most demanding computing environments.

We offer a competitive total compensation package combining a competitive base salary, startup equity, and industry-leading benefits. The total compensation range for this role is $165,000 – $225,000, which includes both base salary and equity. Actual compensation will be determined based on experience, skills, and market alignment. We provide generous benefits, including a 6% 401(k) match, fully covered health insurance premiums, and other comprehensive offerings to support your well-being and success as we grow together.

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Sr. Site Reliability Engineer (SRE) in United States vacancy
  •  ...candidate for this role to work on site in the specified location(s)....  ...by delivering innovative and reliable technology solutions that...  ...Bank Platform Operations and Engineering organization, you will help ensure...  ...Site Reliability Engineering (SRE) practices that enhance client... 
    Senior
    Full time
    Work at office

    The Charles Schwab Corporation

    Austin, TX
    3 days ago
  • $160k - $185k

     ...their fitness journey and revolutionized the industry along the way. And we’re just getting started!OverviewThe Sr. Manager, Site Reliability Engineering (SRE) leads the strategy, execution, and continuous improvement of reliability, availability, and performance across... 
    Senior
    Work at office
    Local area
    Remote work
    Work from home

    Planet Fitness

    Hampton, NH
    5 days ago
  •  ...Sr. Site Reliability Engineer (SRE) New York City, NY - LOCALS ONLY Hybrid, 3 days 6-Month Contract 10-15 years Our client is seeking a Senior Site Reliability Engineer (SRE) with 10–15 years of experience to support front-office trading systems in a production... 
    Senior
    Contract work
    Local area

    RIT Solutions

    New York, NY
    1 day ago
  •  ...security to responsibly propel the global lottery industry ever forward. Position Summary We are looking for a skilled Site Reliability Engineer (SRE) to enhance the stability, performance, and reliability of our production systems. The SRE will work closely with... 
    Senior
    Permanent employment
    Work experience placement
    Local area

    SCIENTIFIC GAMES

    Alpharetta, GA
    more than 2 months ago
  • $101k - $161k

     ...several prestigious awards, such as Best Engineering Team, Best Company for Diversity,...  ...DescriptionWho You'll Work WithWe’re looking for Site Reliability Engineers to join our growing Arista’s...  ...-as-a-Service (CVaaS) global SRE team. SREs at Arista combine strong software... 
    Senior

    Arista Networks

    Santa Clara, CA
    1 day ago
  •  ...Senior Site Reliability Engineer (SRE) Our client is a global technology consulting and digital solutions company that enables enterprises across industries to reimagine business models, accelerate innovation, and maximize growth by harnessing digital technologies.... 
    Senior
    Local area

    E-Solutions

    Los Angeles, CA
    17 hours ago
  •  ...capabilities successfully. Provide expert-level guidance to engineering and product teams, contributing to high-level architecture...  ...and recommending improvements for scalability, performance, reliability, and operational readiness. Partner with application teams... 
    Senior
    Remote work
    Flexible hours

    OrangePeople

    United States
    2 days ago
  • $120k - $175k

     ...Senior Site Reliability Engineer (SRE) Atlanta, GA preferred, Remote At PrizePicks, we are the fastest-growing sports company in North America, as recognized by Inc. 5000. As the leading platform for Daily Fantasy Sports, we cover a diverse range of sports leagues... 
    Senior
    Remote work
    Work visa
    Flexible hours

    PrizePicks

    Atlanta, GA
    4 days ago
  • $175k - $229k

     ...Senior Site Reliability Engineer (SRE) Instrumental builds the manufacturing acceleration platform behind the world's most complex electronics. We capture digital exhaust and engineering context from assembly lines—images, test logs, BOM data, performance, repair cycles... 
    Senior

    Instrumental Inc

    Palo Alto, CA
    3 days ago
  •  ...Description The Senior Site Reliability Engineer (SRE) will implement, secure, and operate the cloud infrastructure that supports CenCore Group's proprietary enterprise SaaS platform. This role is responsible for maintaining a scalable, highly available, secure... 
    Senior
    Work at office
    Remote work

    CenCore

    United States
    4 days ago
  •  ...engage in remote collaboration for a worldwide presence. About the Role We are looking for an experienced Senior Site Reliability Engineer (SRE) to own the reliability, availability, and operational excellence of business-critical production systems. This is a... 
    Senior
    Remote work
    Worldwide
    Home office

    Oowlish Technology

    United States
    3 days ago
  •  ...Senior Site Reliability Engineer At Swile, we believe that good products can help reduce friction in daily professional life and boost employee...  ...Brazil. Your role as a Senior Site Reliability Engineer (SRE) centers around creatively solving problems, ensuring a balance... 
    Senior
    Remote work

    Swile

    United States
    5 days ago
  • $175k - $185k

     ...Senior Site Reliability Engineer (SRE) Remote, US Branch is on a mission to empower workers with financial freedom. We do this by helping companies accelerate payments and providing working Americans with accessible, free financial services. We're committed to building... 
    Senior
    Daily paid
    Remote work
    Home office
    Flexible hours

    Branch

    United States
    1 day ago
  •  ...Senior Site Reliability Engineer (SRE) We are looking for a highly experienced and driven Senior Site Reliability Engineer to join our forward-thinking cloud development and operations team. In this role, you will contribute to the design, development, and operation... 
    Senior
    Remote work

    Mirantis

    United States
    4 days ago
  • $92.7k - $203.94k

     ...one community at a time. Position Summary The Senior Site Reliability Engineer is pivotal in ensuring the reliability, scalability, and performance...  ...engineering, infrastructure, and operations teams to embed SRE best practices, improve application resiliency, and optimize... 
    Senior
    Hourly pay
    Full time
    Temporary work
    Local area

    CVS Health

    Woonsocket, RI
    16 hours ago
  • $60 - $65 per hour

     ...Innova Solutions has a client that is immediately hiring for a Senior Site Reliability Engineer (SRE) Position Type: Full-time (Contract ) Duration: 19 Months Location: Des Moines, IA Minneapolis, MN Irving, TX As a Senior Site Reliability Engineer... 
    Senior
    Hourly pay
    Full time
    Contract work
    Temporary work
    Work experience placement
    Immediate start
    Worldwide
    Flexible hours

    Innova Solutions

    Des Moines, IA
    5 days ago
  •  ...Information Technology group delivers secure, reliable technology solutions that enable...  ...RoleAs a Senior Application Support Engineer, you will help power DTCC's global...  ...and settlement.Leveraging Site Reliability Engineering (SRE) principles, you will support a portfolio... 
    Senior
    Remote work
    Flexible hours

    DTCC- The Depository Trust & Clearing Corporation

    Boston, MA
    5 days ago
  •  ...We are seeking an experienced Site Reliability Engineer (SRE) – Microsoft Hyper-V & Private Cloud to operate highly available private cloud and Virtual Desktop Infrastructure (VDI) platforms based on Microsoft Hyper-V. This role combines deep Hyper-V expertise with modern... 
    Senior
    Temporary work
    Local area

    2T Consulting

    Jersey City, NJ
    15 days ago
  •  ...startups across the US. We’re building a pool of world-class Site Reliability Engineers for current roles and for upcoming opportunities. You will...  ...into one of our partner startups or added to our vetted SRE network for future projects. This role is ideal for engineers... 
    Senior
    Local area

    Breakout Tools

    San Francisco, CA
    4 days ago
  •  ...risk—the leading cause of cybersecurity breaches—and build safer, more resilient organizations. The Role: As a Senior Site Reliability Engineer (SRE) at Dune Security, you will play a critical role in ensuring our platform's stability, scalability, and security. You will... 
    Senior
    Full time
    Work at office

    Dune Security

    New York, NY
    5 days ago
  •  ...Lovelace is the only provider of enterprise-scale context engines capable of analyzing trillions of real-time data points...  ...~ Lovelace AI is seeking a highly skilled and motivated Site Reliability Engineer (SRE) to join our growing team. As an SRE at Lovelace AI, you... 
    Full time

    Lovelace Ai

    Pittsburgh, PA
    1 day ago
  • $65 - $80 per hour

    Sr. SRE / DevOps Engineer (Local to Bay Area only) This range is provided by Compunnel Inc.. Your actual pay will be based on your skills...  ...message the job poster from Compunnel Inc. Position Senior Site Reliability Engineer / DevOps Engineer — Local to Pleasanton area... 
    Senior
    Full time
    Local area

    Compunnel Inc.

    Pleasanton, CA
    1 day ago
  •  ..., please send me a copy of your updated resumes Title: Sr. SRE / DevOps Engineer Location: Sunnyvale, CA (Only Local candidate) Client...  .../ DevOps Engineer at Sunnyvale, California location. As Site Reliability Engineer, the individual will work closely with multi-functional... 
    Senior
    Local area
    Immediate start

    Donato Technologies Inc

    Sunnyvale, CA
    4 days ago
  • $175k - $215k

     ...experiences — and we’re constantly looking for new ways to enhance these exciting experiences. Sr. Manager, Site Reliability Engineer provides strategic leadership across multiple SRE teams and their managers, ensuring alignment with organizational priorities and functional... 
    Senior

    Disney Experiences

    Orlando, FL
    12 hours ago
  • $120k - $200k

     ...PermContact: Kunal DaveContact Email: ****@*****.*** Reliability Engineer(SRE) ResponsibilitiesGlobal Architecture & Disaster Recovery Participate...  ...(e.g., Chaos Engineering, resilience testing, automated recovery)SkillsBilingual Mandarin Site Reliability Engineer(SRE)
    Overseas

    Comrise

    New York, NY
    5 days ago
  •  ...Build and operate one or more bounded contexts of the NeoCloud SRE platform — the multi-region substrate that observes, protects, and...  ..., collection-monitor. Alert, Correlation & SLO: alert-engine-framework, alert-correlation, slo-framework, default M-series alert... 
    Senior
    Full time
    Contract work
    Local area

    Bitdeer Technologies Group

    San Jose, CA
    more than 2 months ago
  •  ...preferred) Employment Type: W2, Contract to Hire, Direct Hire Overview Our client is seeking a highly skilled Edge Site Reliability Engineer (Edge SRE) to lead the design, automation, and operations of Google Distributed Cloud Edge (GDCE) environments. This role combines... 
    Senior
    Full time
    Contract work

    CoSourcing Partners - Enterprise-AI and IT Services Company

    Chicago, IL
    1 day ago
  • $160k - $200k

    Role Overview We are looking for a Control System Engineer/Site Reliability Engineer (SRE) to integrate and maintain the hardware and software systems that enable QuEra’s quantum controls and software stack. You’ll work closely with software engineers, physicists, hardware... 
    Local area
    Remote work

    QuEra Computing

    Boston, MA
    4 days ago
  •  ...shape the future of our communities.This is a Software Engineering position at Director level, which is part of the job family...  ...businesses. This role is for an experienced and driven Site Reliability Engineer (SRE) to join our AI Platform team to help support, scale and... 

    Morgan Stanley

    Alpharetta, GA
    5 days ago
  •  ...troubleshooting staging and production cloud environments . Experienced in architectural design for reliability, scalability, and performance. Practical application of SRE principles : SLIs, SLOs, error budgets, automation, incident management, and postmortems.... 

    Purple Drive

    Atlanta, GA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Sr. Site Reliability Engineer (SRE). Be the first to apply!