Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Sr. Site Reliability Engineer (SRE)

$165k - $225k

Moonlite

Job Description

Job Description

Moonlite delivers high-performance AI infrastructure for organizations running intensive computational research, large-scale model training, and demanding data processing workloads.We provide infrastructure deployed in our facilities or co-located in yours, delivering flexible on-demand or reserved compute that feels like an extension of your existing data center. Our team of AI infrastructure specialists combines bare-metal performance with cloud-native operational simplicity, enabling research teams and enterprises to deploy demanding AI workloads with enterprise-grade reliability and compliance.

Your Role:

You will be instrumental in building and operating production-grade AI infrastructure with deep Kubernetes expertise at its core. Working closely with our systems engineers, network engineers, and platform engineering team, you'll architect and operate the Kubernetes infrastructure that powers our control plane and orchestrates compute, storage, and networking at scale. This role requires deep understanding of Kubernetes internals, custom resource definitions (CRDs), storage and network integrations, and building production-grade clusters from the ground up (not just deploying in managed environments). You'll ensure enterprise-grade reliability while establishing the automation, observability, and operational practices.

Job Responsibilities
  • Kubernetes Infrastructure Engineering: Design, build, and operate production Kubernetes clusters on bare-metal infrastructure – including cluster bootstrapping, control plane architecture, etcd management, and scaling strategies for high-performance compute workloads.
  • Kubernetes Networking & CNIs: Implement and operate custom Kubernetes networking solutions with SR-IOV for high-performance GPU interconnects, multi-tenancy isolation and advanced networking policies. Configure CNI plugins and network segmentation for research workloads.
  • Custom Operators & Controllers: Develop and maintain custom Kubernetes operators and controllers for bare-metal provisioning, infrastructure lifecycle management, and resource orchestration across compute, storage, and networking domains.
  • GPU Infrastructure Integration: Deploy and optimize NVIDIA GPU operators, device plugins, and other custom scheduling logic for GPU workload placement and utilization optimization.
  • Platform Integration & Storage: Build deep integrations between Kubernetes and underlying infrastructure including CSI drivers for storage, custom admission controllers for policy enforcement, and scheduling extensions for specialized hardware placement.
  • Infrastructure Automation: Design and implement automation using Terraform, Ansible, Helm, and custom operators to orchestrate infrastructure workflows and enable deployments across multiple regions.
  • Production Operations & Reliability: Manage production bare-metal infrastructure across multiple regions. Build systems ensuring high availability, fault tolerance, and graceful degradation – establishing SLIs, SLOs, and monitoring to meet enterprise reliability commitments.
  • Observability & Incident Response: Build comprehensive monitoring, logging, and alerting using Prometheus, Grafana, and ELK stack. Lead incident response, conduct postmortems, and implement preventative measures to improve reliability and reduce MTTR.
  • Performance & Capacity Planning: Identify and resolve performance bottlenecks across infrastructure domains. Monitor utilization trends, forecast capacity needs, and optimize resource allocation for various workloads.
Requirements
  • Experience: 5+ years in SRE, DevOps, or infrastructure engineering roles with proven experience operating production infrastructure at scale.
  • Kubernetes Infrastructure Expertise: Deep hands-on experience building and operating production Kubernetes clusters on bare-metal infrastructure – not just deploying workloads in managed clusters. Must understand cluster bootstrapping, control plane architecture, etcd operations, and scaling strategies.
  • Kubernetes Internals & Integration: Strong understanding of Kubernetes internals including custom resource definitions (CRDs), operators, controllers, admission webhooks, and scheduling. Experience integrating storage (CSI drivers), networking (CNI, SR-IOV), and specialized hardware (GPU device plugins) with Kubernetes.
  • Linux Systems Experience: Strong fundamentals in Linux systems administration, performance tuning, troubleshooting, and automation in production environments.
  • Infrastructure Automation: Proficiency with infrastructure-as-code tools (Terraform, Ansible, Helm) and building automation to reduce operational overhead.
  • Networking Fundamentals: Solid understanding of networking concepts including IPAM, DNS, DHCP, VLAN/VXLAN, routing, load balancing, and experience troubleshooting network issues in production.
  • Observability & Monitoring: Experience building and maintaining comprehensive monitoring solutions using tools like Prometheus, Grafana, and centralized logging systems.
  • Reliability Practices: Understanding of SRE principles including SLIs/SLOs/SLAs, error budgets, incident management, and blameless postmortems.
  • Scripting & Automation: Strong scripting skills in Go, Python, or Bash for automation, tooling development, and operational efficiency.
  • Problem-Solving Under Pressure: Demonstrated ability to troubleshoot complex issues under pressure, manage incidents effectively, and communicate clearly during outages.
  • Collaboration & Communication: Excellent communication skills and ability to work across teams including systems engineers, network engineers, and software developers.
Preferred Qualifications
  • Experience building custom Kubernetes operators or controllers for infrastructure orchestration
  • Deep familiarity with Kubernetes networking (Calico, Cilium, Multus), service mesh technologies, and network policy management
  • Experience with GPU workload orchestration including NVIDIA GPU Operator, MIG, time-slicing, and device plugins
  • Background with advanced Kubernetes features including custom schedulers, admission controllers, and API server extensions
  • Experience with Kubernetes cluster federation or multi-cluster management
  • Knowledge of high-performance networking technologies (InfiniBand, RDMA, RoCE) and their integration with Kubernetes
  • Experience with enterprise storage systems (VAST, Lightbits, Ceph, or similar)
  • Familiarity with configuration management at scale and GitOps practices
  • Understanding of security best practices for Kubernetes and bare-metal infrastructure
  • Experience operating infrastructure in regulated industries or co-located data center environments
  • Background supporting research institutions, technical computing environments, or enterprise AI infrastructure
Key Technologies
  • Kubernetes, Linux, Terraform, Ansible, Prometheus, Grafana, ELK Stack, Go, Python, Bash, NVIDIA GPU Technologies, High-Performance Networking, Enterprise Storage Systems
Why Moonlite
  • Build Critical Research Infrastructure: Your work will directly enable quantitative research teams and AI practitioners to push the boundaries of what's possible in financial modeling and AI research.
  • Enterprise Impact: Build and operate infrastructure that supports mission-critical research and AI workloads for leading financial institutions and research organizations.
  • Technical Excellence: Join an infrastructure team focused on delivering enterprise-grade reliability while pushing the boundaries of high-performance computing capabilities.
  • Hands-On Ownership: As part of our growing infrastructure team, you'll have significant ownership over critical systems and the autonomy to influence our operational practices and technology choices.
  • Industry Leadership: Work alongside experienced infrastructure professionals who have built and operated systems for the most demanding computing environments.

We offer a competitive total compensation package combining a competitive base salary, startup equity, and industry-leading benefits. The total compensation range for this role is $165,000 – $225,000, which includes both base salary and equity. Actual compensation will be determined based on experience, skills, and market alignment. We provide generous benefits, including a 6% 401(k) match, fully covered health insurance premiums, and other comprehensive offerings to support your well-being and success as we grow together.

#li-remote

Vacancy posted 14 days ago
Similar jobs that could be interesting for youBased on the Sr. Site Reliability Engineer (SRE) in Chicago, IL vacancy
  •  ...Overview: Senior Site Reliability Engineer (SRE) Location: Chicago, IL (Onsite) Type: Contract Role Overview: We are seeking a Senior Site Reliability Engineer (SRE) with strong expertise in AWS infrastructure, automation, observability, and production... 
    Senior
    Contract work

    Purple Drive

    Chicago, IL
    1 day ago
  • $140k - $170k

     ...We are looking for a   Senior   Site Reliability Engineer to work as part of a lean,   product ‑ focused   engineering organization. This role...  ...and work experience   ~8+ years of experience in DevOps, SRE, platform engineering, or similar roles supporting application... 
    Senior
    Work experience placement
    Flexible hours

    SEI

    Chicago, IL
    9 days ago
  •  ...leader in fast food is seeking a Senior Manager for Edge Operations/SRE in Chicago. This pivotal role involves leading edge...  ...operations, collaborating across teams to ensure high availability and reliability of the platform. Candidates should have 10+ years in... 
    Senior

    McDonald's Corporation

    Chicago, IL
    17 hours ago
  • $190.8k - $267.1k

     ...information. For more information, visit redditinc.com. Reddit SRE is rapidly innovating and our teams are working to meet the...  ...and trafficked corners of the internet. As a Senior Site Reliability Engineer on Reddit’s Infrastructure SRE team, you’ll use your knowledge... 
    Senior
    Work experience placement
    Home office
    Flexible hours

    Alien Blue

    Chicago, IL
    3 days ago
  • $140k - $205k

     ...Senior Technology Site Reliability Engineer Cooley is seeking a Senior Site Reliability Engineer to join the Infrastructure & Development Operationsteam...  ...summary: The Senior Technology Site Reliability Engineer("SRE") is responsible for ensuring the reliability, scalability,... 
    Senior
    Full time
    Temporary work
    Work at office
    Flexible hours
    Weekend work

    Cooley

    Berwyn, IL
    3 days ago
  • $125.04k - $187.56k

     ...services, including Finance, Legal, Sustainability, Commercial, Digital and E-commerce, Technology and more. Overview The Site Reliability Engineer (SRE) III is responsible for ensuring the scalability, reliability, and performance of production systems through automation... 
    Senior
    Full time
    Work at office
    Remote work
    Flexible hours

    ViziRecruiter

    Chicago, IL
    1 day ago
  • $130k - $170k

     ...Senior Site Reliability Engineer About Us Founded in 2014, we offer the industry’s first and only cloud‑based, fully‑customisable, end‑to‑end software...  ...and other regulatory standards Qualifications 5–8 years in SRE, operations, or performance engineering roles Bachelor’s... 
    Senior
    Full time
    Flexible hours
    Shift work

    Supernova Technology™

    Chicago, IL
    4 days ago
  • $250k - $350k

     ...where quantitative researchers, engineers, traders, and operational...  ...technically skilled and proactive SRE to join the strategy...  ...boost stability, throughput, and reliability Qualifications Minimum of 3 years...  ...in production support, site reliability, or infrastructure... 
    Full time

    Engtal Inc

    Chicago, IL
    3 days ago
  • $164.6k - $288k

     ...Overview The SRE Community of Practice (CoP) Senior Implementation Lead is responsible for driving the adoption, standardization, and maturity of Site Reliability Engineering (SRE) practices across the organization. This role serves as a key enabler in scaling SRE principles... 
    Visa sponsorship
    Work visa

    Koitecc Solutions

    Chicago, IL
    17 hours ago
  •  ...running systems that must perform reliably under real-time market...  ...culture is highly collaborative, engineering-driven, and focused on continuous...  ...role for someone with an SRE, systems, or production engineering...  ...3+ years of experience in site reliability, systems engineering... 

    Fintal Partners

    Chicago, IL
    4 days ago
  •  ...A leading quantitative trading firm is seeking a Head of Site Reliability Engineering to help scale one of its most critical infrastructure organisations...  ..., engineering organisation, and reliability strategy. The SRE function sits at the heart of the firm's trading platform.... 
    Immediate start

    Acquire Me

    Chicago, IL
    17 hours ago
  • $150k - $200k

     ...Consultant @ Selby Jennings | Financial Technology We are seeking a Site Reliability Engineer to join our team and assist with the design, development,...  ...and Experience: 3+ years of experience in Systems, Network, SRE, or DevOps engineering Strong programming experience in... 
    Full time
    Work at office

    Selby Jennings

    Chicago, IL
    4 days ago
  • $150k - $155k

     ...Site Reliability Engineer Hybrid (3 days onsite, 2 days remote) full‑time. No visa sponsorship. Base pay: $150,000 – $155,000 per year, subject to...  ...leveraging large language models (LLMs) to automate and optimize SRE workflows, including scripting, incident report... 
    Full time
    Work experience placement
    Remote work
    Visa sponsorship

    Request Technology

    Chicago, IL
    4 days ago
  • $114k - $155k

     ...world's most complex and mission-critical systems. As a Site Reliability Engineer III at JPMorgan Chase within the Commercial and Investment Banking...  ...AI capabilities within the work environment to support SRE workflows with strong validation habits and awareness of data... 
    Local area

    JPMorgan Chase Bank, N.A.

    Chicago, IL
    1 day ago
  • $128.5k - $214.1k

     ...Join to apply for the Staff Site Reliability Engineer role at CME Group . We're looking for a Staff Site Reliability Engineer to join our...  ...manual toil and enhancing operational excellence. Integrate SRE principles directly into the software development lifecycle,... 
    Full time
    Work at office
    Worldwide
    2 days per week

    CME Group

    Chicago, IL
    2 days ago
  • $112.5k - $187.5k

     ...TransUnion, this role will report to a DevOps Director. The Site Reliability Engineering team drives reliability strategy, elevates engineering...  ...serve as a senior technical leader and force multiplier on the SRE team. Operating with full autonomy, you will drive reliability... 
    Full time
    Temporary work
    Work experience placement
    Work at office
    Flexible hours
    2 days per week

    TransUnion

    Chicago, IL
    2 days ago
  • $102.6k - $193.43k

     ...Software Engineer Chamberlain Group (CG) is a global leader in intelligent access and Blackstone portfolio company. Powered by our myQ...  ...Provide leadership, assistance and technical guidance to the SRE Team Facilitate and manage discussions with TECH/IT Executives... 
    Temporary work
    Second job
    Worldwide

    Chamberlain Group

    Oak Brook, IL
    1 day ago
  •  ...Direct message the job poster from Algo Capital Group Senior Site Reliability Engineer - Observability and Automation A leading high-frequency...  ...week ago DevOps Engineer (Mid, Senior, or Principal Level) Sr. Site Reliability Engineer - Observability Chicago, IL $138,... 
    Senior
    Full time
    Work at office
    Flexible hours

    Algo Capital Group

    Chicago, IL
    3 days ago
  •  ...upon to keep lives moving forward when it matters most. Learn more about CCC at **The Role**We are seeking a talented Sr. Site Reliability Engineering Developer to be part of the fast moving, innovative CCC Site Reliability Team. We build enterprise class, hosted... 
    Senior
    Night shift

    CCC Information Services

    Chicago, IL
    17 hours ago
  • $130k - $165k

     ...Job Title: Senior Software Engineer Company: Snapsheet Job Location: USA, Remote...  ...Job Department: Technology  Team : Site Reliability Engineering   About Snapsheet: Snapsheet...  ...As a Senior Site Reliability Engineer (SRE) at Snapsheet, you will play a critical... 
    Senior
    Full time
    Temporary work
    Local area
    Remote work
    Visa sponsorship
    Work visa
    Flexible hours

    Snapsheet

    Chicago, IL
    a month ago
  • $130k - $180k

     ...SRE is part of a global organization that leverages the latest technology to communicate with our colleagues across the globe...  ..., collaboration, and accomplishment. Being a Senior Site Reliability Engineer at iManage Means…  You are an engineer, a builder, and a systems... 
    Senior
    Full time
    Work at office
    Local area
    Remote work
    Worldwide
    Monday to Friday
    Flexible hours

    iManage

    Chicago, IL
    6 days ago
  • $145k - $175k

     ...to help you gain your full potential. Job Overview The Site Reliability Engineer supports deployments, cloud infrastructure, and monitoring...  ...infrastructure improvements. You'll be joining a small, senior SRE team with broad ownership of the platforms and... 
    Senior
    Full time
    Temporary work
    Work at office
    Local area
    Flexible hours
    3 days per week

    Rewards Network

    Chicago, IL
    26 days ago
  • $117.96k - $123.01k

     ...software systems, ensuring system performance, reliability, and maintainability (10%). Provide technical leadership across multiple engineering teams, guiding day-to-day development...  ...Application Architects, Systems Engineers, DevOps/SRE, and business stakeholders to ensure... 
    Full time
    Temporary work
    Work at office
    Remote work
    Flexible hours

    Morningstar

    Chicago, IL
    1 day ago
  • $200k - $250k

     ...We're looking for a senior level DevOps Engineer to join a small, high-impact team that builds...  ...: ~5–8 years of experience in DevOps, SRE, or Platform Engineering roles. ~...  ...trading firm — understanding the urgency and reliability requirements of systems that support P&... 
    Senior
    Full time
    Worldwide
    Flexible hours

    DV Trading

    Chicago, IL
    7 hours ago
  • $155k - $222.6k

     ...maintain automation solutions that improve the reliability, scalability, and operational efficiency...  ...cloud environments. Partner with other engineering teams, product management, and business...  ...~2+ years of experience in Site Reliability Engineering, DevOps, Infrastructure... 
    Full time
    Temporary work
    Local area
    Worldwide
    Flexible hours

    Webex Events (formerly Socio)

    Chicago, IL
    1 day ago
  • $86k - $105k

     ...generation of application infrastructure and to be responsible for reliability, automation and scalability using and the latest best...  ...certifications. Minimum of 2 years prior DevOps, software engineering or related experience. Must be able to work different schedules... 
    Hourly pay
    Work at office
    Immediate start
    Visa sponsorship
    Work visa
    Flexible hours

    Early Warning Services

    Chicago, IL
    1 day ago
  • $120k - $160k

     ...Role We are looking for a Senior DevOps Engineer who specializes in AWS and containerized...  ...~3+ years of experience in a DevOps or SRE role. ~ Infrastructure as Code expertise...  ...Continuous Delivery pipelines to enable rapid and reliable software releases. Automate... 
    Senior
    Full time

    Beyond Finance

    Chicago, IL
    9 days ago
  • $118.6k - $148.25k

     ...manufacturers, and consumers. Our Websites Engineering Team drives the Dealer Inspire platform,...  ...scalable systems. As a DevOps Engineer (Sr. Software Engineer), you'll help...  ...+ years of combined developer and DevOps/SRE experience in CI/CD, scripting, and automation... 
    Senior
    Full time
    Local area
    Remote work
    Home office
    Visa sponsorship
    Work visa

    Cars Commerce

    Chicago, IL
    3 days ago
  • $118.6k - $148.25k

     ...simplifies everything about buying and selling cars. Our Websites Engineering Team drives our Dealer Inspire platform. We are embarking on a...  ...and optimize current systems, and build scalable, secure, and reliable API-based backend services that power our platform. In this... 
    Senior
    Full time
    Local area
    Remote work
    Home office
    Visa sponsorship
    Work visa

    Cars.com

    Chicago, IL
    9 days ago
  •  ...clients to succeed in an evolving digital landscape. Role Overview We are seeking an experienced Observability / Site Reliability Engineer (SRE) to design, scale, and maintain our enterprise monitoring and alerting ecosystems. In this role, you will bridge the gap... 
    Remote job
    Contract work

    Ontrac Solutions

    Chicago, IL
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Sr. Site Reliability Engineer (SRE). Be the first to apply!