Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Lead Site Reliability Engineer

Intellum, Inc.

Job Description

Job Description

About us  

Intellum is the leader in corporate education technology and powers the largest, most successful customer, partner, and employee learning programs in the world. Large brands and fast-moving companies like Google, Meta, Amazon, Walmart, Xero, Atlassian, Mailchimp, Airbnb, Stripe, and TikTok rely on Intellum to engage and educate the audiences they touch. 

We have always been a "remote first" company and are proud to have team members located all over the world. We value Curiosity, Creativity, Perseverance, and Kindness and strive to demonstrate these core values every day. Our culture is very important to us. We invest in our people in fun and exciting ways, including personal development budgets and an annual all-company retreat that is focused less on work and more on human connections. We are in growth mode, and our "smart growth" approach ensures that we will continue to scale our company effectively. 

 

The Lead Systems Engineer is a senior individual contributor responsible for the reliability, scalability, and modernization of Intellum's platform infrastructure. Intellum serves large enterprise customers with demanding availability expectations, and this role will help shape the technical direction for how our platform runs, deploys, and scales.

This is a highly hands-on role with significant ownership across infrastructure architecture, cloud environments, deployment systems, observability, and platform reliability. The Lead Systems Engineer will also provide technical leadership across the Systems Engineering function through architecture guidance, mentorship, knowledge sharing, and strong operational standards.

A key focus of this role is continuing to modernize the platform toward portable, container-orchestrated infrastructure, improving deployment and observability capabilities, and maintaining an architecture that can operate effectively across multiple cloud providers.

Responsibilities
  • Own and drive key infrastructure modernization initiatives, including the continued evolution from legacy compute environments toward modern, container-orchestrated infrastructure while maintaining reliable service for enterprise customers.
  • Design and maintain infrastructure as code across multiple cloud providers, ensuring infrastructure decisions support portability, maintainability, and long-term scalability.
  • Improve the reliability and maturity of Intellum's CI/CD systems and deployment tooling so releases are efficient, observable, and recoverable.
  • Provide technical leadership across the Systems Engineering team through mentorship, architecture guidance, knowledge sharing, and support for strong engineering practices.
  • Establish and evolve SLI and SLO practices, along with the monitoring, alerting, and load-testing capabilities needed to support platform reliability.
  • Participate in and provide leadership during platform incidents, including troubleshooting, root cause analysis, and follow-through on corrective actions.
  • Drive visibility into cloud infrastructure costs and incorporate cost considerations into architecture and infrastructure decisions.
  • Improve developer experience by evolving the infrastructure and tooling engineers depend on, including development environments, deployment workflows, and production feedback loops.
  • Partner closely with Security and Engineering teams on access controls, infrastructure hardening, compliance requirements, and secure infrastructure practices.
  • Identify operational and infrastructure risks early, recommend priorities, and help drive the technical roadmap for the Systems Engineering function.
  • Contribute to the continued development of the Systems Engineering team and function, including mentoring engineers and helping build strong technical practices as the organization evolves.
  • Perform other duties as assigned.

 

Required Skills
  • 8+ years of hands-on experience in infrastructure, DevOps, platform engineering, site reliability engineering, or a related discipline, including experience building and operating production systems.
  • Deep hands-on experience designing, operating, and troubleshooting highly available production infrastructure.
  • Production experience across more than one major cloud provider, with depth in at least one of AWS or Google Cloud and working fluency in the other.
  • Significant experience with container orchestration and Kubernetes in production environments, including cluster operations, workload configuration, reliability, and troubleshooting.
  • Experience modernizing production infrastructure, including migrations from VM-based or legacy environments toward containerized or cloud-native architectures.
  • Strong infrastructure-as-code experience using Terraform or comparable tooling, with an emphasis on repeatability and automation.
  • Experience building, operating, or significantly improving CI/CD systems and deployment infrastructure.
  • Strong incident response and troubleshooting capabilities, including experience diagnosing complex distributed-system failures and contributing to effective post-incident review.
  • Strong Linux administration skills and scripting or programming ability in Ruby, Python, or a comparable language.
  • Experience working in a SaaS environment where reliability, availability, and production stability are critical.
  • Ability to collaborate effectively with distributed teams across US and European time zones and participate in an on-call rotation.
  • Strong communication skills and the ability to provide technical direction, mentor other engineers, and influence infrastructure decisions across teams.
Preferred Qualifications
  • Experience operating production infrastructure across both AWS and Google Cloud simultaneously.
  • Prior experience leading or managing engineers, whether through formal people management, technical leadership, or mentorship.
  • Experience developing engineers and helping build strong, high-performing technical teams.
  • Experience with cloud cost management or FinOps practices at meaningful scale.
  • Experience managing deployment platforms such as Spinnaker, Jenkins, or comparable tooling.
  • Experience operating a Ruby on Rails enterprise application or comparable production codebase.
  • Familiarity with SOC 2 or similar compliance frameworks and customer-facing security requirements.
  • Working knowledge of AI-assisted development tooling and its infrastructure implications.
  • Prior people leadership or management experience in a player-coach capacity, balancing hands-on technical contribution with mentorship, team guidance, and development of engineers.
  • Background in learning management systems, learning technologies, or adult education platforms.
Education
  • Bachelor's degree in a related field or equivalent practical experience. Equivalent experience is genuinely accepted for this role.

BENEFITS

  • Medical - 100% of employee premiums for selected individual plans
  • Dental - 100% of employee premiums covered
  • Vision - 100% of employee premiums covered
  • LinkedIn Learning
  • 401(k) plus matching (US Based Only)
  • Flexible PTO
  • Calm subscription
  • Annual Company Retreat

 

Intellum is an equal-opportunity employer. We're committed to building an inclusive team that celebrates diversity in people, perspectives, and backgrounds regardless of race, color, national origin, gender, sexual orientation, age, religion, disability, citizenship, veteran status, or any other protected status. We encourage you to apply for an open position and if you have questions about whether or not your job experience and skill set meet the requirements for a specific role, reach out to us directly at View email address on ziprecruiter.com.  


If you are an individual applying from CA, NY, CO, CT, MD, NV, or RI, please reach out to View email address on ziprecruiter.com to inquire about specific pay ranges.

 

Vacancy posted 8 days ago
Similar jobs that could be interesting for youBased on the Lead Site Reliability Engineer in Atlanta, GA vacancy
  •  ...in Atlanta, Georgia, and serves customers in more than 35 countries worldwide.Site Reliability EngineerOnsite: Atlanta, GAJob SummaryAt NCR Voyix, we're looking for a Site Reliability Engineer II to help build, support, and scale the cloud platforms that power our... 
    Suggested
    Full time
    Worldwide
    Flexible hours

    NCR

    Atlanta, GA
    2 days ago
  •  ...We are seeking an experienced Site Reliability Engineer (SRE) to support and maintain production systems hosted on AWS. The role focuses on production support, incident management, monitoring, observability, troubleshooting, and improving system reliability and availability... 
    Suggested
    Contract work

    2T Consulting

    Atlanta, GA
    a month ago
  •  ...solving and decision-making abilities and the highest degree of professionalism. We are seeking an experienced AWS solution design engineer/architect to join our infrastructure cloud team. The infrastructure cloud team is responsible for internal services that provide... 
    Suggested

    Black Knight Financial Services

    Atlanta, GA
    2 days ago
  •  ...Intercontinental Exchange (NYSE:ICE), we engineer technology, exchanges and clearing...  ...capital and derivative markets. With a leading-edge approach to developing technology...  ...people to join our team.We are seeking a Site Reliability Engineer to bring 3+ years of hands-on... 
    Suggested

    Black Knight Financial Services

    Atlanta, GA
    2 days ago
  • $100k - $120k

    OverviewThe Site Reliability Engineer is a key force behind improving Origami’s time to resolution and advancing overall site reliability and scalability. This person participates in efforts to identify root causes during post-incident investigations, while also identifying... 
    Suggested
    Full time
    Temporary work
    Work experience placement
    Flexible hours

    Origami Risk

    Atlanta, GA
    1 day ago
  •  ...of America)Please review the following job description:The Site Reliability Engineer role focuses on enhancing the reliability and operational excellence...  ...business and technology teams.Responsibilities include leading major incident responses, driving problem management, and... 
    Permanent employment
    Full time
    Part time
    H1b
    Work at office
    Local area
    Immediate start
    Work visa
    Monday to Friday
    Shift work
    Day shift

    Truist

    Atlanta, GA
    2 days ago
  •  ...Join to apply for the Site Reliability Engineer role at Motion Recruitment Join to apply for the Site Reliability Engineer role at Motion...  ...with teams to create SLI/SLO’s Actively monitor and lead troubleshooting of degraded performance and hard to define... 
    Contract work
    Worldwide

    Motion Recruitment

    Atlanta, GA
    14 hours ago
  • $123.4k - $222.53k

     ...Responsibilities Enhance system reliability and resilience by identifying issues and implementing preventive measures to reduce downtime...  ...) ~ Acceptable areas of study include Computer Science, Engineering or related field (Required) ~4-7 years Working in operations... 
    Full time
    Temporary work
    Part time
    Work experience placement
    Local area
    Flexible hours

    T-Mobile

    Atlanta, GA
    4 days ago
  • $130k - $150k

     ...recruiter to learn more. Base pay range $130,000.00/yr - $150,000.00/yr Overview: We are seeking a highly skilled Site Reliability Engineer (SRE) to join our team and help build and maintain scalable, reliable, and efficient systems. The ideal candidate will... 
    Full time
    Remote work

    Prestige Staffing

    Atlanta, GA
    1 day ago
  • $178.13k - $205.4k

     ...customers, and the world around us. As a Fortune 500 company and a leading AI platform for managing people, money, and agents, we’re...  ...~​Bachelor’s degree or foreign degree equivalent in Computer Engineering, Computer Science, Engineering, or related field plus five (5)... 
    Work at office
    Remote work
    Flexible hours

    Workday

    Atlanta, GA
    1 day ago
  • $141.8k - $195k

     .... We're one of the fastest-growing private companies and a leading player in a massive, fast-moving market. With a global workforce...  ....Why You'll Love This RoleCribl Inc is seeking a Senior Site Reliability Engineer to join our mission where you will unlock the value of all... 
    Remote work

    Cribl

    Atlanta, GA
    2 days ago
  • $149.8k - $241.5k

     ...efficiency, accelerate time-to-value, and deliver better customer experiences. About The Role We're looking for a Senior Site Reliability Engineer who's passionate about building reliable, scalable infrastructure that helps developers ship better software faster. You'... 
    Work at office
    Local area
    Remote work
    Work from home
    Worldwide
    Home office
    Flexible hours

    Camunda

    Atlanta, GA
    4 days ago
  •  ...Job Title :- Site Reliability Engineer (SRE) Employment Type :- W2 Duration :- Long Term Visa Type :- All Visa applicable which are ready for W2 Location :- Atlanta, GA (Onsite) Job Description We are seeking a highly skilled Site Reliability Engineer (SRE... 

    Highbrow

    Atlanta, GA
    3 days ago
  •  ...Inspire Brands is hiring two Senior Site Reliability Engineers to help build and scale reliable, resilient, and observable systems supporting high...  ...infrastructure metrics Incident Management Lead technical response for high-severity incidents Drive blameless... 
    Worldwide

    Inspire Brands Inc

    Atlanta, GA
    4 days ago
  •  ...Overview: About Us: Atlanta-based Incident IQ is the leading workflow management platform built exclusively for K-12...  ...thrive, and districts to shape the future of education. Site Reliability Engineer (SRE) Overview: We are looking for a Site Reliability Engineer... 
    Full time
    Live in
    Work at office

    Incident IQ LLC

    Atlanta, GA
    1 day ago
  • $136.2k - $214.01k

     ...outcomes Visionary in future focused problem-solving Exceptional in execution and impact The Role As a Senior Site Reliability Engineer at Proofpoint you will develop a deep understanding of the various services and applications that come together to... 
    Full time
    Flexible hours

    Proofpoint

    Atlanta, GA
    2 days ago
  • $81.75k - $138.98k

     ...Job Schedule Full time Job Description As the Senior Site Reliability Engineer, you will serve as a trusted technical resource responsible...  ...Wesco, we build, connect, power and protect the world. As a leading provider of business‑to‑business distribution, logistics... 
    Full time
    Work at office
    Immediate start
    Worldwide
    Shift work

    Anixter

    Atlanta, GA
    1 day ago
  • $120k - $175k

     ...company in North America, as recognized by Inc. 5000. As the leading platform for Daily Fantasy Sports, we cover a diverse range...  ...We are seeking a highly skilled and experienced Senior Site Reliability Engineer to join our team. We are passionate about delivering cutting... 
    Full time
    Remote work
    Work visa
    Flexible hours

    AEG Presents

    Atlanta, GA
    14 hours ago
  •  ...Site Reliability Engineer At Acuity, you will join an Agile team focused on building and supporting advanced platforms and applications that...  ...cross-geo team providing operational & escalation coverage, leading incident response and recovery for critical services.... 

    Acuity

    Atlanta, GA
    3 days ago
  •  ...’re Looking For We’re looking for a proactive, hands‑on Site Reliability Engineer who thrives in building and scaling cloud infrastructure in...  ...improve system performance, reliability, and scalability Leading incident response efforts, conducting postmortems, and... 
    Work experience placement
    Flexible hours

    Rainforest

    Atlanta, GA
    1 day ago
  •  ...expertise. We deliver faster, smarter, more reliable insights to insurance carriers and...  ...the right place. The Role As a Site Reliability Engineer, you'll be responsible for the...  ...our AWS-hosted infrastructure. You'll lead incident response, build the automation... 
    Flexible hours

    Seek Now

    Atlanta, GA
    4 days ago
  •  ...Purple Drive Site Reliability Engineer (SRE) Contractual Atlanta, GA Key Highlights: Proven expertise in Google Cloud Platform (GCP) services, including BigQuery, Cloud Logging, IAM, and Service Accounts. Strong background in provisioning, monitoring, and... 

    Purple Drive

    Atlanta, GA
    1 day ago
  • $61.09k - $104.36k

     ...Site Reliability Engineer Choosing Capgemini means choosing a company where you will be empowered to shape your career in the way you’d like...  ...to reimagine what’s possible. Join us and help the world’s leading organizations unlock the value of technology and build a more... 
    Permanent employment
    Full time
    Contract work
    Local area

    Capgemini

    Atlanta, GA
    14 hours ago
  •  ...rate for C2C/1099/W2. Job Description: Job Title : Sr. Site Reliability Engineer Location : Atlanta, GA - Hybrid Duration : 6+ Months...  ...) Job Description - Key Responsibilities: Lead and mentor a team of SREs, fostering a culture of collaboration... 
    Contract work
    Local area
    Immediate start

    Navtech

    Atlanta, GA
    14 hours ago
  •  ...OpenShift - Site Reliability Engineer Atlanta , GA / Onsite Qualifications: This position is 60 % SRE and 40% SDE....  ...with teams to create SLI/SLO's . • Actively monitor and lead troubleshooting of degraded performance and hard to define... 
    Work experience placement

    Fisec Global

    Atlanta, GA
    3 days ago
  • $127k - $249k

     ...Central time zones. We are looking for an experienced Senior Engineer for our SRE, Atlas team to support, maintain and grow the Atlas...  ...workloads. Role Overview We are seeking a talented Site Reliability Engineer (SRE) with a strong infrastructure background. This... 
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    Atlanta, GA
    3 days ago
  •  ...Site Reliability Engineer We're looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You'll partner with engineers and data scientists to build, automate, and... 

    Alembic Technologies

    Atlanta, GA
    1 day ago
  • $75.7k - $136.3k

     ...and building systems that scale? Join our highly skilled Site Reliability Engineering team! Our team designs, develops, and manages...  ...Employee Stock Purchase Plan (ESPP). Akamai provides industry-leading benefits including healthcare, 401K savings plan, company... 
    Work experience placement
    Work at office

    Akamai

    Atlanta, GA
    4 days ago
  •  ...Site Reliability Engineer We are looking for a Site Reliability Engineer to ensure the reliability, security, and continuous operation of a multi‑cloud application security platform. This role combines platform engineering and security automation, focusing on Kubernetes... 
    Work at office
    Remote work
    Visa sponsorship
    Work visa
    Flexible hours

    AgileEngine

    Atlanta, GA
    1 day ago
  • $70 - $75 per hour

     ...Site Reliability Engineering (SRE) Architect Get AI-powered advice on this job and more exclusive features. This range is provided by STAFFWORXS...  ...Site Reliability Engineering (SRE) Architect to lead the strategic design, development, and maturity of our reliability... 
    Contract work

    Staffworxs Inc

    Atlanta, GA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Lead Site Reliability Engineer. Be the first to apply!