Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Sr. Site Reliability Engineer (SRE)

$165k - $225k

Moonlite

Moonlite delivers high-performance AI infrastructure for organizations running intensive computational research, large-scale model training, and demanding data processing workloads. We provide infrastructure deployed in our facilities or co-located in yours, delivering flexible on‑demand or reserved compute that feels like an extension of your existing data center. Our team of AI infrastructure specialists blends bare‑metal performance with cloud‑native operational simplicity, enabling research teams and enterprises to deploy demanding AI workloads with enterprise‑grade reliability and compliance.

Your Role

You will be instrumental in building and operating production‑grade AI infrastructure with deep Kubernetes expertise at its core. Working closely with our systems engineers, network engineers, and platform engineering team, you’ll architect and operate the Kubernetes infrastructure that powers our control plane and orchestrates compute, storage, and networking at scale. This role requires deep understanding of Kubernetes internals, custom resource definitions (CRDs), storage and network integrations, and building production‑grade clusters from the ground up (not just deploying in managed environments). You'll ensure enterprise‑grade reliability while establishing the automation, observability, and operational practices.

Job Responsibilities

  • Kubernetes Infrastructure Engineering: Design, build, and operate production Kubernetes clusters on bare‑metal infrastructure – including cluster bootstrapping, control plane architecture, etcd management, and scaling strategies for high‑performance compute workloads.
  • Kubernetes Networking & CNIs: Implement and operate custom Kubernetes networking solutions with SR‑IOV for high‑performance GPU interconnects, multi‑tenancy isolation and advanced networking policies. Configure CNI plugins and network segmentation for research workloads.
  • Custom Operators & Controllers: Develop and maintain custom Kubernetes operators and controllers for bare‑metal provisioning, infrastructure lifecycle management, and resource orchestration across compute, storage, and networking domains.
  • GPU Infrastructure Integration: Deploy and optimize NVIDIA GPU operators, device plugins, and other custom scheduling logic for GPU workload placement and utilization optimization.
  • Platform Integration & Storage: Build deep integrations between Kubernetes and underlying infrastructure including CSI drivers for storage, custom admission controllers for policy enforcement, and scheduling extensions for specialized hardware placement.
  • Infrastructure Automation: Design and implement automation using Terraform, Ansible, Helm, and custom operators to orchestrate infrastructure workflows and enable deployments across multiple regions.
  • Production Operations & Reliability: Manage production bare‑metal infrastructure across multiple regions. Build systems ensuring high availability, fault tolerance, and graceful degradation – establishing SLIs, SLOs, and monitoring to meet enterprise reliability commitments.
  • Observability & Incident Response: Build comprehensive monitoring, logging, and alerting using Prometheus, Grafana, and ELK stack. Lead incident response, conduct postmortems, and implement preventative measures to improve reliability and reduce MTTR.
  • Performance & Capacity Planning: Identify and resolve performance bottlenecks across infrastructure domains. Monitor utilization trends, forecast capacity needs, and optimize resource allocation for various workloads.

Requirements

  • Experience: 5+ years in SRE, DevOps, or infrastructure engineering roles with proven experience operating production infrastructure at scale.
  • Kubernetes Infrastructure Expertise: Deep hands‑on experience building and operating production Kubernetes clusters on bare‑metal infrastructure – not just deploying workloads in managed clusters. Must understand cluster bootstrapping, control plane architecture, etcd operations, and scaling strategies.
  • Kubernetes Internals & Integration: Strong understanding of Kubernetes internals including custom resource definitions (CRDs), operators, controllers, admission webhooks, and scheduling. Experience integrating storage (CSI drivers), networking (CNI, SR‑IOV), and specialized hardware (GPU device plugins) with Kubernetes.
  • Linux Systems Experience: Strong fundamentals in Linux systems administration, performance tuning, troubleshooting, and automation in production environments.
  • Infrastructure Automation: Proficiency with infrastructure‑as‑code tools (Terraform, Ansible, Helm) and building automation to reduce operational overhead.
  • Networking Fundamentals: Solid understanding of networking concepts including IPAM, DNS, DHCP, VLAN/VXLAN, routing, load balancing, and experience troubleshooting network issues in production.
  • Observability & Monitoring: Experience building and maintaining comprehensive monitoring solutions using tools such as Prometheus, Grafana, and centralized logging systems.
  • Reliability Practices: Understanding of SRE principles including SLIs/SLOs/SLAs, error budgets, incident management, and blameless postmortems.
  • Scripting & Automation: Strong scripting skills in Go, Python, or Bash for automation, tooling development, and operational efficiency.
  • Problem‑Solving Under Pressure: Demonstrated ability to troubleshoot complex issues under pressure, manage incidents effectively, and communicate clearly during outages.
  • Collaboration & Communication: Excellent communication skills and ability to work across teams including systems engineers, network engineers, and software developers.

Preferred Qualifications

  • Experience building custom Kubernetes operators or controllers for infrastructure orchestration.
  • Deep familiarity with Kubernetes networking (Calico, Cilium, Multus), service mesh technologies, and network policy management.
  • Experience with GPU workload orchestration including NVIDIA GPU Operator, MIG, time‑slicing, and device plugins.
  • Background with advanced Kubernetes features including custom schedulers, admission controllers, and API server extensions.
  • Experience with Kubernetes cluster federation or multi‑cluster management.
  • Knowledge of high‑performance networking technologies (InfiniBand, RDMA, RoCE) and their integration with Kubernetes.
  • Experience with enterprise storage systems (VAST, Lightbits, Ceph, or similar).
  • Familiarity with configuration management at scale and GitOps practices.
  • Understanding of security best practices for Kubernetes and bare‑metal infrastructure.
  • Experience operating infrastructure in regulated industries or co‑located data center environments.
  • Background supporting research institutions, technical computing environments, or enterprise AI infrastructure.

Why Moonlite

  • Build Critical Research Infrastructure: Your work will directly enable quantitative research teams and AI practitioners to push the boundaries of what’s possible in financial modeling and AI research.
  • Enterprise Impact: Build and operate infrastructure that supports mission‑critical research and AI workloads for leading financial institutions and research organizations.
  • Technical Excellence: Join an infrastructure team focused on delivering enterprise‑grade reliability while pushing the boundaries of high‑performance computing capabilities.
  • Hands‑On Ownership: As part of our growing infrastructure team, you’ll have significant ownership over critical systems and the autonomy to influence our operational practices and technology choices.
  • Industry Leadership: Work alongside experienced infrastructure professionals who have built and operated systems for the most demanding computing environments.

We offer a competitive total compensation package combining a base salary, startup equity, and industry‑leading benefits. The total compensation range for this role is $165,000 – $225,000, which includes both base salary and equity. Actual compensation will be determined based on experience, skills, and market alignment. We provide generous benefits, including a 6% 401(k) match, fully covered health insurance premiums, and other comprehensive offerings to support your well‑being and success as we grow together.

Equal Employment Opportunity

As set forth in Moonlite’s Equal Employment Opportunity policy, we do not discriminate on the basis of any protected group status under any applicable law. For government reporting purposes, we ask candidates to respond to the voluntary self‑identification survey. Completion of the form is entirely voluntary. Whatever your decision, it will not be considered in the hiring process or thereafter. Any information that you do provide will be recorded and maintained in a confidential file.

Voluntary Self‑Identification

We are a federal contractor or subcontractor. The law requires us to provide equal employment opportunity to qualified people with disabilities. We have a goal of having at least 7% of our workers as people with disabilities. The law says we must measure our progress towards this goal. To do this, we must ask applicants and employees if they have a disability or have ever had one. People can become disabled, so we need to ask this question at least every five years. Completing this form is voluntary, and we hope that you will choose to do so. Your answer is confidential. No one who makes hiring decisions will see it. Your decision to complete the form and your answer will not harm you in any way. If you want to learn more about the law or this form, visit the U.S. Department of Labor’s Office of Federal Contract Compliance Programs (OFCCP) website at

How do you know if you have a disability? A disability is a condition that substantially limits one or more of your “major life activities.” If you have or have ever had such a condition, you are a person with a disability. Disabilities include, but are not limited to: Alcohol or other substance use disorder; Autoimmune disorder such as lupus, fibromyalgia, rheumatoid arthritis, HIV/AIDS; Blind or low vision; Cancer; Cardiovascular or heart disease; Celiac disease; Cerebral palsy; Deaf or serious difficulty hearing; Diabetes; Disfigurement; Epilepsy or other seizure disorder; Gastrointestinal disorders such as Crohn’s disease; Intellectual or developmental disability; Mental health conditions such as depression, bipolar disorder, anxiety disorder, schizophrenia, PTSD; Missing limbs; Mobility impairment; Nervous system condition such as migraine, Parkinson’s disease, multiple sclerosis; Neurodivergence such as ADHD, autism spectrum disorder, dyslexia; Partial or complete paralysis; Pulmonary or respiratory conditions; Short stature; Traumatic brain injury.

Public burden statement: According to the Paperwork Reduction Act of 1995, no persons are required to respond to a collection of information unless such collection displays a valid OMB control number. This survey should take about 5 minutes to complete.

#J-18808-Ljbffr
Vacancy posted 12 hours ago
Similar jobs that could be interesting for youBased on the Sr. Site Reliability Engineer (SRE) in Chicago, IL vacancy
  • $165k - $225k

     ...workloads with enterprise-grade reliability and compliance. Your Role:...  ...closely with our systems engineers, network engineers, and platform...  ...Kubernetes networking solutions with SR-IOV for high-performance GPU...  ...Experience: 5+ years in SRE, DevOps, or infrastructure engineering... 
    Senior
    Remote work
    Flexible hours

    Moonlite

    Chicago, IL
    more than 2 months ago
  •  ...preferred) Employment Type: W2, Contract to Hire, Direct Hire Overview Our client is seeking a highly skilled Edge Site Reliability Engineer (Edge SRE) to lead the design, automation, and operations of Google Distributed Cloud Edge (GDCE) environments. This role combines... 
    Senior
    Full time
    Contract work

    CoSourcing Partners - Enterprise-AI and IT Services Company

    Chicago, IL
    5 days ago
  •  ...and companies, alikeKlover’s engineering team powers one of the fastest...  ...systems that prioritize reliability, security, and performance, and...  ...candidateAbout the RoleAs a Senior/Staff Site Reliability Engineer, you...  ...modern AI as a first-class SRE skill, along with cloud... 
    Senior
    Work at office
    Immediate start
    Remote work

    Attain Data

    Chicago, IL
    2 days ago
  • $140k - $170k

    We are looking for a Senior Site Reliability Engineer to work as part of a lean, product‑focused engineering organization. This role is about building...  ...and work experience 8+ years of experience in DevOps, SRE, platform engineering, or similar roles supporting application... 
    Senior
    Full time
    Work experience placement
    Flexible hours

    SEI Investments Developments

    Chicago, IL
    3 days ago
  • $117.63k - $176.44k

     ...seeking an experienced Data SRE to join the FreeWheel Data SRE...  ...responsible for ensuring the reliability, scalability, and performance...  ...systems. Working closely with data engineers and other operation sub-teams...  ...summary on our careers site for more details.EducationBachelor... 
    Senior
    Full time

    Comcast

    Chicago, IL
    2 days ago
  • $127k - $249k

    The TeamPlatform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that support...  ...alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and... 
    Senior
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    Chicago, IL
    4 days ago
  • $130k - $180k

    SRE is part of a global organization that leverages the latest technology to communicate with our colleagues across the globe...  ...belonging, collaboration, and accomplishment.Being a Senior Site Reliability Engineer at iManage Means… You are an engineer, a builder, and a... 
    Senior
    Work at office
    Local area
    Remote work
    Worldwide
    Monday to Friday
    Flexible hours

    Imanage

    Chicago, IL
    4 days ago
  • $127k - $249k

     ...Central time zones. We are looking for an experienced Senior Engineer for our SRE, Atlas team to support, maintain and grow the Atlas...  ...crucial workloads. Role OverviewWe are seeking a talented Site Reliability Engineer (SRE) with a strong infrastructure background. This... 
    Senior
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    Chicago, IL
    8 hours ago
  • $130k - $165k

     ...Job Title: Senior Software Engineer Company: Snapsheet Job Location: USA, Remote...  ...Job Department: Technology  Team : Site Reliability Engineering   About Snapsheet: Snapsheet...  ...As a Senior Site Reliability Engineer (SRE) at Snapsheet, you will play a critical... 
    Senior
    Full time
    Temporary work
    Local area
    Remote work
    Visa sponsorship
    Work visa
    Flexible hours

    Snapsheet

    Chicago, IL
    more than 2 months ago
  • $125.04k - $187.56k

     ...services, including Finance, Legal, Sustainability, Commercial, Digital and E-commerce, Technology and more. Overview The Site Reliability Engineer (SRE) III is responsible for ensuring the scalability, reliability, and performance of production systems through automation... 
    Senior
    Full time
    Work at office
    Remote work
    Flexible hours

    ViziRecruiter,LLC.

    Chicago, IL
    12 hours ago
  • $145k - $175k

     ...to help you gain your full potential. Job Overview The Site Reliability Engineer supports deployments, cloud infrastructure, and monitoring...  ...infrastructure improvements. You'll be joining a small, senior SRE team with broad ownership of the platforms and... 
    Senior
    Full time
    Temporary work
    Work at office
    Local area
    Flexible hours
    3 days per week

    Rewards Network

    Chicago, IL
    more than 2 months ago
  • $130k - $170k

     ...Senior Site Reliability Engineer About Us Founded in 2014, we offer the industry’s first and only cloud‑based, fully‑customisable, end‑to‑end...  ...regulatory standards Qualifications ~5–8 years in SRE, operations, or performance engineering roles ~ Bachelor’s... 
    Senior
    Full time
    Flexible hours
    Shift work

    Supernova Technology™

    Chicago, IL
    12 hours ago
  • $190.8k - $267.1k

     ...information. For more information, visit redditinc.com. Reddit SRE is rapidly innovating and our teams are working to meet the...  ...and trafficked corners of the internet. As a Senior Site Reliability Engineer on Reddit’s Infrastructure SRE team, you’ll use your knowledge... 
    Senior
    Work experience placement
    Home office
    Flexible hours

    Alien Blue

    Chicago, IL
    5 days ago
  • Play a key role in ensuring system reliability at one of the world’s most iconic and largest financial institutions.As a Site Reliability Engineer II at JPMorgan Chase within the Commercial...  ...the work environment to support SRE workflows (e.g., troubleshooting support... 

    JP Morgan Chase

    Chicago, IL
    3 days ago
  • $100k - $120k

    OverviewThe Site Reliability Engineer is a key force behind improving Origami’s time to resolution and advancing overall site reliability and scalability...  ...on our platform.Partners with the larger Cloud Operations, SRE, Engineering teams, and the business-at-large to advance our... 
    Full time
    Temporary work
    Work experience placement
    Flexible hours

    Origami Risk

    Chicago, IL
    1 day ago
  • $130k - $225k

     ...integrity, innovation and a willingness to challenge consensus.The Algorithmic Trading Team is looking for a Site Reliability Engineer for our Chicago office. The SRE team is critical to the success of our trading - ensuring that our production trading systems, test... 
    Temporary work
    Work at office
    Flexible hours

    DRW

    Chicago, IL
    2 days ago
  •  ...right treatments for the right patients, at the right time.The Site Reliability Engineering team works with all departments and business units to...  ...ISOResponsibilities for the Position:Works as a member of an SRE teamReview application requirements and recommend solutions... 
    Full time

    Tempus

    Chicago, IL
    4 days ago
  • $130k - $150k

     ...and hybrid infrastructure, meaning experience with cloud technologies is essential for this role. Position OverviewThe Site Reliability Engineer (SRE) helps ensure CRA’s critical business services are reliable, scalable, and performant across on-premises and cloud environments... 
    Work at office
    Work from home
    3 days per week

    CRA International

    Chicago, IL
    1 day ago
  • $100.7k - $167.8k

    Job SummaryThe Site Reliability Engineer III is a pivotal architect of stability for CME Clearing & Risk. You will engineer secure, scalable, and...  ...experience managing data layers like Oracle, Postgres, or BigQuery.SRE DNA: A profound understanding of SRE principles,... 
    Full time
    Worldwide

    CME- Group

    Chicago, IL
    2 days ago
  •  ...the world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Commercial and Investment Banking...  ...AI capabilities within the work environment to support SRE workflows with strong validation habits and awareness of data... 

    JP Morgan Chase

    Chicago, IL
    3 days ago
  •  ...the world's most complex and mission-critical systems. As a Site Reliability Engineer III at JPMorgan Chase within the Corporate Technology team,...  ...experienceExposure to or hands-on experience in supporting SRE practices for Data management/migration platforms and products... 

    JP Morgan Chase

    Chicago, IL
    3 days ago
  • $160k - $210k

     ...you'll do:Join our Platform Engineering team, where you'll ensure the...  ...as the technical lead for the SRE function, setting technical direction...  ...mentoring engineers across reliability initiativesAnalyze,...  ...years of experience in DevOps, Site Reliability Engineering, or Platform... 
    Work at office
    Worldwide
    Monday to Friday
    Flexible hours

    NinjaTrader Group

    Chicago, IL
    23 hours ago
  • $194k - $267k

     ...If you are too, let's talk.The TeamThe Site Reliability team is dedicated to architecting and owning...  ...CI/CD platforms that support Okta’s SRE ecosystem. In this development-focused...  ...that maximize platform reliability and engineering velocity.The ideal candidate is someone... 
    Local area
    Worldwide
    Flexible hours

    Okta

    Chicago, IL
    8 hours ago
  • $112.5k - $187.5k

     ...TransUnion, this role will report to a DevOps Director. The Site Reliability Engineering team drives reliability strategy, elevates engineering...  ...serve as a senior technical leader and force multiplier on the SRE team. Operating with full autonomy, you will drive reliability... 
    Full time
    Temporary work
    Work experience placement
    Work at office
    Flexible hours
    2 days per week

    TransUnion

    Chicago, IL
    8 hours ago
  • $132.1k - $220.1k

    We're looking for a Staff Site Reliability Engineer to join our team, focusing on the core systems that power global financial markets. This isn't...  ...manual toil and enhancing operational excellence.Integrate SRE principles directly into the software development lifecycle,... 
    Full time
    Work at office
    Worldwide
    2 days per week

    CME- Group

    Chicago, IL
    2 days ago
  • $194k - $267k

     ...something more than once, automate it” and who can rapidly self-educate on new concepts and tools.Position Overview:The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and services... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta

    Chicago, IL
    4 days ago
  • $194k - $267k

     ...Overview:We are seeking a highly technical StaffObservabilitySite Reliability Engineer with a specialty in Splunk to own and evolve our Splunk...  ...comprehensive, scalable Observability Platform that enables our SRE teams and business partners. You will treat infrastructure as... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta

    Chicago, IL
    2 days ago
  • $118.3k - $219.8k

    Are you excited to lead Site Reliability Engineering teams that keep mission-critical, 24/7 services running reliably and securely?Do you enjoy building...  ...LexisNexis Risk at .About our TeamOur globally distributed SRE team operates across the FCC market, supporting a broad set... 
    Full time
    Local area

    RELX Group

    Chicago, IL
    4 days ago
  • $158.5k - $172k

     ...exceptional value they deserve.About The OpportunityAs a Senior Engineer on the Runtime Automation team, you will design, automate, and...  .... This is a high-impact position driving continuous reliability, deep system optimization, and automation across our entire technology... 
    Senior
    Full time
    Temporary work
    Work at office
    Flexible hours
    3 days per week

    GrubHub

    Chicago, IL
    2 days ago
  • $150k - $200k

     ...@ Selby Jennings | Financial Technology We are seeking a Site Reliability Engineer to join our team and assist with the design, development, and...  ...Experience: ~3+ years of experience in Systems, Network, SRE, or DevOps engineering ~ Strong programming experience in... 
    Full time
    Work at office

    Selby Jennings

    Chicago, IL
    12 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Sr. Site Reliability Engineer (SRE). Be the first to apply!