Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Lead Site Reliability Engineer

Intellum, Inc.

Job Description

Job Description

About us  

Intellum is the leader in corporate education technology and powers the largest, most successful customer, partner, and employee learning programs in the world. Large brands and fast-moving companies like Google, Meta, Amazon, Walmart, Xero, Atlassian, Mailchimp, Airbnb, Stripe, and TikTok rely on Intellum to engage and educate the audiences they touch. 

We have always been a "remote first" company and are proud to have team members located all over the world. We value Curiosity, Creativity, Perseverance, and Kindness and strive to demonstrate these core values every day. Our culture is very important to us. We invest in our people in fun and exciting ways, including personal development budgets and an annual all-company retreat that is focused less on work and more on human connections. We are in growth mode, and our "smart growth" approach ensures that we will continue to scale our company effectively. 

 

The Lead Systems Engineer is a senior individual contributor responsible for the reliability, scalability, and modernization of Intellum's platform infrastructure. Intellum serves large enterprise customers with demanding availability expectations, and this role will help shape the technical direction for how our platform runs, deploys, and scales.

This is a highly hands-on role with significant ownership across infrastructure architecture, cloud environments, deployment systems, observability, and platform reliability. The Lead Systems Engineer will also provide technical leadership across the Systems Engineering function through architecture guidance, mentorship, knowledge sharing, and strong operational standards.

A key focus of this role is continuing to modernize the platform toward portable, container-orchestrated infrastructure, improving deployment and observability capabilities, and maintaining an architecture that can operate effectively across multiple cloud providers.

Responsibilities
  • Own and drive key infrastructure modernization initiatives, including the continued evolution from legacy compute environments toward modern, container-orchestrated infrastructure while maintaining reliable service for enterprise customers.
  • Design and maintain infrastructure as code across multiple cloud providers, ensuring infrastructure decisions support portability, maintainability, and long-term scalability.
  • Improve the reliability and maturity of Intellum's CI/CD systems and deployment tooling so releases are efficient, observable, and recoverable.
  • Provide technical leadership across the Systems Engineering team through mentorship, architecture guidance, knowledge sharing, and support for strong engineering practices.
  • Establish and evolve SLI and SLO practices, along with the monitoring, alerting, and load-testing capabilities needed to support platform reliability.
  • Participate in and provide leadership during platform incidents, including troubleshooting, root cause analysis, and follow-through on corrective actions.
  • Drive visibility into cloud infrastructure costs and incorporate cost considerations into architecture and infrastructure decisions.
  • Improve developer experience by evolving the infrastructure and tooling engineers depend on, including development environments, deployment workflows, and production feedback loops.
  • Partner closely with Security and Engineering teams on access controls, infrastructure hardening, compliance requirements, and secure infrastructure practices.
  • Identify operational and infrastructure risks early, recommend priorities, and help drive the technical roadmap for the Systems Engineering function.
  • Contribute to the continued development of the Systems Engineering team and function, including mentoring engineers and helping build strong technical practices as the organization evolves.
  • Perform other duties as assigned.

 

Required Skills
  • 8+ years of hands-on experience in infrastructure, DevOps, platform engineering, site reliability engineering, or a related discipline, including experience building and operating production systems.
  • Deep hands-on experience designing, operating, and troubleshooting highly available production infrastructure.
  • Production experience across more than one major cloud provider, with depth in at least one of AWS or Google Cloud and working fluency in the other.
  • Significant experience with container orchestration and Kubernetes in production environments, including cluster operations, workload configuration, reliability, and troubleshooting.
  • Experience modernizing production infrastructure, including migrations from VM-based or legacy environments toward containerized or cloud-native architectures.
  • Strong infrastructure-as-code experience using Terraform or comparable tooling, with an emphasis on repeatability and automation.
  • Experience building, operating, or significantly improving CI/CD systems and deployment infrastructure.
  • Strong incident response and troubleshooting capabilities, including experience diagnosing complex distributed-system failures and contributing to effective post-incident review.
  • Strong Linux administration skills and scripting or programming ability in Ruby, Python, or a comparable language.
  • Experience working in a SaaS environment where reliability, availability, and production stability are critical.
  • Ability to collaborate effectively with distributed teams across US and European time zones and participate in an on-call rotation.
  • Strong communication skills and the ability to provide technical direction, mentor other engineers, and influence infrastructure decisions across teams.
Preferred Qualifications
  • Experience operating production infrastructure across both AWS and Google Cloud simultaneously.
  • Prior experience leading or managing engineers, whether through formal people management, technical leadership, or mentorship.
  • Experience developing engineers and helping build strong, high-performing technical teams.
  • Experience with cloud cost management or FinOps practices at meaningful scale.
  • Experience managing deployment platforms such as Spinnaker, Jenkins, or comparable tooling.
  • Experience operating a Ruby on Rails enterprise application or comparable production codebase.
  • Familiarity with SOC 2 or similar compliance frameworks and customer-facing security requirements.
  • Working knowledge of AI-assisted development tooling and its infrastructure implications.
  • Prior people leadership or management experience in a player-coach capacity, balancing hands-on technical contribution with mentorship, team guidance, and development of engineers.
  • Background in learning management systems, learning technologies, or adult education platforms.
Education
  • Bachelor's degree in a related field or equivalent practical experience. Equivalent experience is genuinely accepted for this role.

BENEFITS

  • Medical - 100% of employee premiums for selected individual plans
  • Dental - 100% of employee premiums covered
  • Vision - 100% of employee premiums covered
  • LinkedIn Learning
  • 401(k) plus matching (US Based Only)
  • Flexible PTO
  • Calm subscription
  • Annual Company Retreat

 

Intellum is an equal-opportunity employer. We're committed to building an inclusive team that celebrates diversity in people, perspectives, and backgrounds regardless of race, color, national origin, gender, sexual orientation, age, religion, disability, citizenship, veteran status, or any other protected status. We encourage you to apply for an open position and if you have questions about whether or not your job experience and skill set meet the requirements for a specific role, reach out to us directly at View email address on ziprecruiter.com.  


If you are an individual applying from CA, NY, CO, CT, MD, NV, or RI, please reach out to View email address on ziprecruiter.com to inquire about specific pay ranges.

 

Vacancy posted 11 days ago
Similar jobs that could be interesting for youBased on the Lead Site Reliability Engineer in Atlanta, GA vacancy
  •  ...Fluency: English (Required)Work Shift:1st shift (United States of America)Please review the following job description:The Site Reliability Engineering Lead role focuses on enhancing the reliability and operational excellence of enterprise platforms across hybrid cloud and... 
    Suggested
    Permanent employment
    Full time
    Part time
    H1b
    Work at office
    Local area
    Immediate start
    Work visa
    Monday to Friday
    Shift work
    Day shift

    Truist

    Atlanta, GA
    4 days ago
  • $35 - $44 per hour

    DescriptionKforce has a client seeking a remote Site Reliability Engineer to join their team. We are seeking a Site Reliability Engineer (SRE) to support a large-scale system modernization and legacy platform retirement initiative. This role will focus on maintaining and... 
    Suggested
    Remote work

    KForce

    Atlanta, GA
    3 days ago
  • $98k - $148.5k

     ...Service, Security, and other teams across the organization.As a Site Reliability Engineer I on the Core Infrastructure team in our Atlanta office,you...  ...has four key dimensions that define our Leadership Impact: Lead Self, Lead the Team, Lead the Business, and Lead the Future... 
    Suggested
    Work at office
    Local area
    Flexible hours

    PagerDuty

    Atlanta, GA
    4 days ago
  • $100k - $120k

    OverviewThe Site Reliability Engineer is a key force behind improving Origami’s time to resolution and advancing overall site reliability and scalability. This person participates in efforts to identify root causes during post-incident investigations, while also identifying... 
    Suggested
    Full time
    Temporary work
    Work experience placement
    Flexible hours

    Origami Risk

    Atlanta, GA
    3 days ago
  • $70 - $85 per hour

     ...there.What You’ll Do* Define and establish enterprise reliability standards, Site Reliability Engineering (SRE) practices, SLIs/SLOs, operational governance...  ...that enable scalable, resilient technology platforms.* Lead the design and implementation of observability, monitoring... 
    Suggested
    Temporary work
    Local area
    Flexible hours
    3 days per week

    Slalom

    Atlanta, GA
    4 days ago
  • $60 - $68 per hour

     ...Site Reliability Engineer Immediate need for a talented Site Reliability Engineer. This is a 12+ months contract opportunity with long-term...  ...SRE) activities, Monitoring & Alerting Our client is a leading Airlines organization and we are currently interviewing to... 
    Contract work
    Local area
    Immediate start

    Pyramid Consulting

    Atlanta, GA
    18 hours ago
  • $127k - $249k

     ...Central time zones. We are looking for an experienced Senior Engineer for our SRE, Atlas team to support, maintain and grow the Atlas...  ...workloads. Role Overview We are seeking a talented Site Reliability Engineer (SRE) with a strong infrastructure background. This... 
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    Atlanta, GA
    1 day ago
  • $55 - $60 per hour

     ...Description Our client is looking for an SRE that will Lead the reliability, scalability, security, and operational excellence of customer...  ...to improve operational efficiency. • Collaborate with Engineering, Product, Security, and Infrastructure teams to enhance customer... 

    Insight Global

    Atlanta, GA
    3 days ago
  • $120k - $175k

     ...Senior Site Reliability Engineer (SRE) Atlanta, GA preferred, Remote At PrizePicks, we are the fastest-growing sports company in North America, as recognized by Inc. 5000. As the leading platform for Daily Fantasy Sports, we cover a diverse range of sports leagues... 
    Full time
    Remote work
    Work visa
    Flexible hours

    PrizePicks

    Atlanta, GA
    5 days ago
  •  ...Purple Drive Site Reliability Engineer (SRE) Contractual Atlanta, GA Key Highlights: Proven expertise in Google Cloud Platform (GCP) services, including BigQuery, Cloud Logging, IAM, and Service Accounts. Strong background in provisioning, monitoring, and... 

    Purple Drive

    Atlanta, GA
    5 days ago
  •  ...rate for C2C/1099/W2. Job Description: Job Title : Sr. Site Reliability Engineer Location : Atlanta, GA - Hybrid Duration : 6+ Months...  ...) Job Description - Key Responsibilities: Lead and mentor a team of SREs, fostering a culture of collaboration... 
    Contract work
    Local area
    Immediate start

    Navtech

    Atlanta, GA
    3 days ago
  •  ...OpenShift - Site Reliability Engineer Atlanta , GA / Onsite Qualifications: This position is 60 % SRE and 40% SDE....  ...with teams to create SLI/SLO's . • Actively monitor and lead troubleshooting of degraded performance and hard to define... 
    Work experience placement

    Fisec Global

    Atlanta, GA
    1 day ago
  • $167.7k - $245.2k

     ...Cisco Meraki, we are responsible for building and growing the cloud that supports these customers and their networks. As a Site Reliability Engineer, you will be focused on supporting a specific, highly available, and very secure production environment. You will analyze... 
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    Atlanta, GA
    3 days ago
  • $111.61k - $131.3k

    At U.S. Bank, we’re on a journey to do our best. Helping the customers and businesses we serve to make better and smarter financial decisions and enabling the communities we support to grow and succeed. We believe it takes all of us to bring our shared ambition to life,...
    Full time
    Work experience placement
    Local area
    3 days per week

    US Bank

    Atlanta, GA
    2 days ago
  •  ...Role: Site Reliability Engineering (SRE) Architect Location: Atlanta, GA (Hybrid on-site) Contract Role Summary: As an...  ...into system health and behaviour With overall maturity lead the definition and implementation strategy for Service Level... 
    Contract work
    Early shift

    AceStack LLC

    Atlanta, GA
    3 days ago
  • Purchasing Power is a leading employee purchase program that helps people buy the products...  ...of customers. You will work in a modern engineering environment that embraces automation, AI...  ...evolve.WE OFFER:- Hybrid work model (on-site and remote flexibility)- Comprehensive... 
    Full time
    Work at office
    Local area
    Remote work

    Purchasing Power

    Atlanta, GA
    6 days ago
  • $151k - $297k

     ..., you will partner with SRE leaders and engineers to scale the platform that underpins all...  ...program execution, strengthen production reliability practices, and coordinate cross-...  ...Strengthen Production Reliability - Lead change management and launch readiness programs... 
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    Atlanta, GA
    4 days ago
  • #CareersJC 1483593Qualifications· Strong experience supporting production systems hosted on AWS, including EC2, VPC, ALB/NLB, RDS, Lambda, and EKS.· Hands-on experience with incident management and 24/7 production support models.· Proficiency with monitoring and observability...

    LTM

    Atlanta, GA
    6 days ago
  •  ...Engineer Lead (FDE) - Platform & Automation FDE - Forward Deployment Engineer Location: This role requires associates to be in-office 1 - 2 days per week, fostering collaboration and connectivity, while providing flexibility to support productivity and work-life... 
    Full time
    Temporary work
    Work at office
    Local area
    2 days per week
    1 day per week

    Elevance Health

    Atlanta, GA
    3 days ago
  •  ...development opportunities, a world-class training facility, and leading market tools, we help our people continue to grow both professionally...  ...benefits can be found towards the bottom of our KPMG US Careers site at Benefits & How We Work.Follow this link to obtain salary... 
    Full time
    H1b
    Local area

    KPMG

    Atlanta, GA
    2 days ago
  • $60k - $135k

     ...Job Title: Lead Cloud Architect City: Atlanta State/Province: Georgia Posting Start Date: 1/6/26...  ...our holistic portfolio of capabilities in consulting, design, engineering, and operations, we help clients realize their boldest ambitions... 
    Minimum wage
    Local area

    Wipro

    Atlanta, GA
    1 day ago
  •  ...formsKnowledge of formal database architecture and design.Must have experience with full life cycle projects.Should be able to play a technical lead and mentor junior ETL developer and ramp up the teamQualificationsBachelor’s degree or foreign equivalent required from an... 

    Sonsoft

    Atlanta, GA
    6 days ago
  •  ...at Emory UniversityEmory University is a leading research university that fosters...  ...RESPONSIBILITIES: The Lead Network Systems Engineer, Science and Research Networks is a central...  ...to ensure appropriate availability and reliability. Develops, implements, and monitors policies... 
    Work at office
    Remote work
    Work from home
    Flexible hours

    Emory University

    Atlanta, GA
    4 days ago
  • PrimeFlight Aviation Services Inc. is seeking a Ramp Supervisor in Atlanta, Georgia. The successful candidate will oversee a team of ramp agents, ensuring efficient and safe handling of aircraft on the ground. Responsibilities include coordinating activities such as baggage...

    PrimeFlight Aviation Services Inc.

    Atlanta, GA
    3 days ago
  • $101.5k - $169.1k

     ...program.Job DescriptionThe Release Train Engineer (RTE) has a primary purpose of...  ...influence. Other The RTE may be called upon to lead projects that impact the ART but may not...  ...policies and standards. Monitor AI tool reliability across teams. Create backup plans for system... 
    Full time
    Work at office
    Remote work
    Visa sponsorship
    Flexible hours

    Cox Enterprises

    Atlanta, GA
    4 days ago
  •  ...evolving the foundational systems and practices that ensure the reliability, scalability, performance, and efficiency of our critical...  ...highly resilient systems. Leverage deep expertise in software engineering, distributed systems, cloud infrastructure, and SRE principles... 

    Merican Inc

    Atlanta, GA
    3 days ago
  •  ...demeanor upon arrival and departure. Conduct payroll for your site as well as revenue reconciliation. Hire and train new Valet staff...  ...level of customer service. Provide strong leadership and lead by example to enhance the team's performance. Monitor and maintain... 
    Immediate start

    EVOLUTION PARKING & GUEST SERVICES

    Atlanta, GA
    4 days ago
  •  ...I have an opportunity for "Senior Lead Cloud Security Architect ___ Atlanta, GA - Onsite" and I am looking for a candidate who can...  ...group based on standards that can be adopted and implemented by engineering teams. • Contribute to the development of non-cyber architecture... 
    Contract work
    Immediate start

    Navtech

    Atlanta, GA
    5 days ago
  •  ...ResponsibilitiesThe Trenchless Geotechnical Team Lead will report to the Trenchless Services Engineering Lead, having primary...  ...geotechnical investigations, site visits, and attend client meetings...  ...motivated, accountable, responsible, reliable, and have strong abilities to... 
    For contractors
    Local area
    Remote work

    HDR

    Atlanta, GA
    6 days ago
  • $100,000 - $120,000 per week

    Job Type:PermanentBuild a brilliant future with HiscoxBilling Excellence Lead - Hiscox USOverview:The Billing Excellence Lead is responsible for driving continuous improvement and process excellence within the billing function. This role focuses on transformation, innovation... 
    Full time
    Temporary work
    Work at office

    Hiscox AG

    Atlanta, GA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Lead Site Reliability Engineer. Be the first to apply!