Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer (SRE)

Amplifire Elearning

Site Reliability Engineer

Amplifire is seeking a Site Reliability Engineer to improve the reliability, scalability, performance, and operational efficiency of our cloud-based platform. Working alongside DevOps engineers within the Platform Operations team, this role combines software engineering and systems operations with a focus on observability, automation, incident reduction, and operational excellence.

You will partner closely with software engineering, QA, security, and DevOps to establish reliability practices, improve production visibility, strengthen incident response, automate operational workflows, and help engineering teams deliver changes safely and confidently.

This role supports systems operating under regulatory and compliance requirements, including FedRAMP and SOC 2. The ideal candidate understands that reliability, traceability, security, and change management enable sustainable development velocity rather than compete with it.

Amplifire expects all technical team members to leverage AI-assisted tools and workflows as force multipliers for productivity, learning, automation, and problem-solving while maintaining strong engineering skills, sound judgment, security standards, and operational accountability.

Reliability, Observability & Performance

· Establish and maintain service-level indicators (SLIs), service-level objectives (SLOs), error budgets, and other measures of system health.

· Build and continuously improve monitoring, logging, tracing, dashboards, and alerting that provide actionable visibility into application and infrastructure health.

· Analyze system behavior, performance, capacity, and reliability trends to proactively identify risks and improvement opportunities.

· Partner with engineering teams to define reliability requirements and improve the availability, scalability, and performance of production systems.

Incident Response & Operational Readiness

· Participate in the on-call rotation and respond to production incidents with urgency and sound technical judgment.

· Improve incident detection, triage, escalation, communication, mitigation, and recovery processes.

· Lead or contribute to blameless post-incident reviews, root cause analysis, and corrective actions.

· Develop and maintain runbooks, recovery procedures, and operational practices that improve service resilience and reduce mean time to detect and restore service.

Automation, Infrastructure & Delivery

· Identify and eliminate operational toil through automation, self-service tooling, and continuous improvement.

· Build and maintain cloud infrastructure using Infrastructure as Code practices and tools such as Terraform or AWS CDK.

· Improve CI/CD pipelines, deployment safeguards, rollback capabilities, and progressive delivery practices.

· Develop internal tools and automation that improve reliability, resilience, and engineering productivity.

· Support scalable, secure, and cost-effective production environments.

Collaboration & Continuous Improvement

· Partner with software engineers to embed reliability, operability, and observability throughout the development lifecycle while reducing operational friction and helping teams safely own their services in production.

· Help engineering teams diagnose complex production issues across application and infrastructure layers.

· Contribute to security, compliance, capacity-planning, cloud cost optimization, and operational standards.

· Use AI-assisted tools responsibly to improve troubleshooting, automation, documentation, and operational efficiency while sharing knowledge and continuously improving team practices.

Requirements

· 4+ years of experience in Site Reliability Engineering, DevOps, cloud infrastructure, systems engineering, software engineering, or a related production-focused role.

· Demonstrated experience supporting reliable, customer-facing applications in a production cloud environment.

· Hands-on experience designing, deploying, and operating production workloads on AWS.

· Experience building or maintaining Infrastructure as Code, observability solutions, incident response processes, and CI/CD automation.

· Experience identifying and reducing operational toil through automation and continuous improvement.

· Experience participating in on-call rotations, troubleshooting production issues, performing root cause analysis, and implementing preventive improvements.

Technical Skills

· Proficiency with at least one scripting or programming language, such as Python, Bash, JavaScript/TypeScript, Go, or Java.

· Working knowledge of Linux systems, networking, DNS, load balancing, and common cloud architecture patterns.

· Experience with containers and container-based deployment practices; Docker experience is required, while Kubernetes or similar orchestration experience is beneficial.

· Ability to analyze logs, metrics, traces, and system behavior to troubleshoot issues across application and infrastructure layers.

· Understanding of reliability concepts such as SLIs, SLOs, error budgets, availability, latency, capacity planning, and graceful degradation.

· Familiarity with secure configuration, secrets management, access controls, vulnerability remediation, and other operational security fundamentals.

AI & Modern Tooling

· Experience using AI-assisted development or operational tools to improve productivity, automation, troubleshooting, documentation, or engineering workflows.

· Ability to critically evaluate AI-generated outputs and apply appropriate validation before production use.

· Interest in adopting emerging tools and practices that improve engineering effectiveness while maintaining operational excellence.

Collaboration & Problem-Solving Skills

· Strong debugging and systems-thinking skills, including the ability to work through ambiguous, cross-service production issues.

· Ability to communicate clearly during incidents and translate technical findings for engineering and business stakeholders.

· Experience collaborating with software engineering, QA, security, support, and product teams.

· Ability to independently own reliability improvements while seeking input and alignment when appropriate.

· A proactive approach to problem-solving and a willingness to challenge existing practices constructively.

What Success Looks Like

Within your first year:

· Service-level indicators and objectives are established for critical systems.

· Monitoring, alerting, and observability provide actionable visibility across the platform.

· Incident response processes are more structured, repeatable, and measurable.

· Mean time to detect and restore service trends improve through better tooling and operational practices.

· Engineering teams have greater self-service access to operational insights and reliability tooling.

· Manual operational work is reduced through automation and platform improvements.

· Reliability, performance, and scalability risks are identified proactively rather than reactively.

· Demonstrates a proactive approach to problem solving and does not accept existing processes simply because they have historically been done that way.

Amplifire is the leading AI Learning Platform built on brain science that delivers the proven results of 1:1 expert instruction at enterprise scale. We deliver one-on-one AI instruction at scale, detect where people are confidently wrong, and fix it before it becomes a mistake—reducing training time by 50–80% while driving measurable performance outcomes. Trusted by leading organizations in healthcare, accounting, life sciences and other high-stakes industries, Amplifire enables teams to achieve verified competency faster, reduce risk, and perform at the highest level when it matters most.

Salary Description 130,000 - 155,000

Vacancy posted 9 hours ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer (SRE) in United States vacancy
  •  ...unwavering security to responsibly propel the global lottery industry ever forward.Position SummaryWe are looking for a skilled Site Reliability Engineer (SRE) to enhance the stability, performance, and reliability of our production systems. The SRE will work closely with... 
    Suggested
    Permanent employment
    Full time
    Work experience placement
    Local area

    Scientific Games Corporation

    Alpharetta, GA
    5 days ago
  • We are looking for an experienced Site Reliability Engineer (SRE) to strengthen observability and operational resilience across a Microsoft Azure environment. This long-term Contract role will work closely with DevOps and engineering teams to establish monitoring standards... 
    Suggested
    Long term contract

    Robert Half

    Maumee, OH
    5 days ago
  • $175k - $215k

     ...experiences — and we’re constantly looking for new ways to enhance these exciting experiences.Sr. Manager, Site Reliability Engineer provides strategic leadership across multiple SRE teams and their managers, ensuring alignment with organizational priorities and functional... 
    Suggested

    Disney Interactive

    Orlando, FL
    6 days ago
  •  ...Lovelace is the only provider of enterprise-scale context engines capable of analyzing trillions of real-time data points...  ...~ Lovelace AI is seeking a highly skilled and motivated Site Reliability Engineer (SRE) to join our growing team. As an SRE at Lovelace AI, you... 
    Suggested
    Full time

    Lovelace Ai

    Pittsburgh, PA
    1 day ago
  •  ...About The Role: We're looking for a Senior Site Reliability Engineer to help us mature and scale the infrastructure behind our multi-cloud SaaS...  ...across the team What You Bring: ~​​6+ years in SRE, DevOps, or infrastructure engineering roles, with deep,... 
    Suggested
    Remote work
    Flexible hours

    Dental Intelligence

    United States
    4 days ago
  •  ...About the Role: We are looking for a Senior Site Reliability Engineer (SRE) to help modernize large-scale infrastructure and improve the reliability, scalability, and operational excellence of critical production systems. In this role, you will lead OS modernization... 
    Remote work

    Halo Media

    United States
    1 day ago
  •  ...Site Reliability Engineer (SRE) Location: Remote Shift Timings: 5:30 PM to 3:00 AM IST to ensure support for global operations. Job Description: We are seeking a skilled Site Reliability Engineer (SRE) to join our dynamic team. The ideal candidate will have... 
    Remote work
    Shift work

    InOrg Global

    United States
    9 hours ago
  •  ...Senior Site Reliability Engineer At Swile, we believe that good products can help reduce friction in daily professional life and boost employee...  ...Brazil. Your role as a Senior Site Reliability Engineer (SRE) centers around creatively solving problems, ensuring a balance... 
    Remote work

    Swile

    United States
    10 hours ago
  • $100k - $180k

     ...Site Reliability Engineer (SRE) - Remote Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States. This is a fantastic opportunity to join an established... 
    Full time
    H1b
    Local area
    Immediate start
    Remote work
    Visa sponsorship

    Bright Vision Technologies

    United States
    1 day ago
  • $165k - $225k

     ...demanding AI workloads with enterprise‑grade reliability and compliance. Your Role You will...  ...core. Working closely with our systems engineers, network engineers, and platform...  ...Requirements Experience: 5+ years in SRE, DevOps, or infrastructure engineering roles... 
    Contract work
    For contractors
    For subcontractor
    Work at office
    Flexible hours

    Moon Lite Inc

    Chicago, IL
    3 days ago
  •  ...Site Reliability Engineer (SRE) Location: Remote (Secaucus, NJ) Duration: Contract Experience: 7+ Years Job Description 4+ years of experience with multiple APM tools and extensive experience with Dynatrace 4+ years of experience executing software load and... 
    Contract work
    Work experience placement
    Remote work

    Syntricate Technologies

    United States
    3 days ago
  •  ...in Computer Science, Information Technology, Engineering, or equivalent field ~3-5 years of experience in Site Reliability Engineering, Production Support, Platform Engineering...  ...application health ~ Understanding of SRE principles, including observability,... 
    Remote work

    Anveta

    United States
    2 days ago
  •  ...Site Reliability Engineer (SRE) FLUIX is building the AI operating system that plans, designs, and optimizes AI infrastructure. We are based in Silicon Valley. We specialize in providing AI-driven solutions for data centers and power providers, leveraging cutting-edge... 
    Work at office
    Weekend work

    Fluix AI

    San Francisco, CA
    3 days ago
  •  ...SRE Key Responsibilities • Design and manage multi-account AWS infrastructure (VPC, Route Tables, EC2, ECS, EKS 1.33, RDS, DynamoDB...  ...• Develop automation in Bash, Python, Go, C#/.NET (Unity Game Engine) • Maintain developer experience (Backstage, Click Up, Miro,... 
    Remote work

    ACI Infotech

    United States
    5 days ago
  •  ...Site Reliability Engineer (Sre) Our client, an IT Services and Consulting company, is looking for a Site Reliability Engineer (SRE) for their Plano, TX/Atlanta, GA/Middletown, NJ location. Responsibilities: Design and execute performance, load, stress, failover... 

    ICONMA

    Plano, TX
    2 days ago
  •  ...Job Title:  Site Reliability Engineer (Azure Government & Infrastructure) Pay Type : SALARIED EXEMPT  Location:  Remote Citizenship Requirement...  ...Role/Responsibilities The Site Reliability Engineer (SRE) for Azure Government & Infrastructure plays a critical role... 
    Full time
    Remote work
    Monday to Friday

    Quzara LLC

    United States
    2 days ago
  •  ...Site Reliability Engineer (SRE) Location: North Little Rock AR (onsite) Duration: Contract Required/Desired Skills: • Strong web development skills with a strong focus in C#/.NET • Someone who currently works in a hybrid skillset of BOTH.Net development AND... 
    Contract work

    Software Technology Inc

    North Little Rock, AR
    3 days ago
  • $170k - $230k

     ...Site Reliability Engineer (SRE) Palo Alto / San Francisco Bay Area About Mithril Mithril is an AI infrastructure platform built to make GPU compute more accessible and affordable for the world's leading enterprises, AI startups, and the AI research community,... 
    Work at office
    Local area
    1 day per week

    Mithril

    Palo Alto, CA
    5 days ago
  •  ...automate, deploy, and operate highly reliable cloud systems supporting mission-critical...  ...role is centered on DevSecOps and site reliability engineering, with a strong emphasis on deployment...  ...years of professional experience as an SRE, DevOps, reliability, infrastructure,... 
    Permanent employment
    Remote work

    Quindar

    United States
    3 days ago
  • $92.7k - $203.94k

     ...one community at a time. Position Summary The Senior Site Reliability Engineer is pivotal in ensuring the reliability, scalability, and performance...  ...engineering, infrastructure, and operations teams to embed SRE best practices, improve application resiliency, and optimize... 
    Hourly pay
    Full time
    Temporary work
    Local area

    CVS Health

    Scottsdale, AZ
    2 days ago
  •  ...Senior Site Reliability Engineer (SRE) We are looking for a highly experienced and driven Senior Site Reliability Engineer to join our forward-thinking cloud development and operations team. In this role, you will contribute to the design, development, and operation... 
    Remote work

    Mirantis

    United States
    5 days ago
  •  ..., TX** Our Opportunity: We are looking for a skilled engineer with disciplines that incorporate aspects of software...  ...including AI/ML-driven approaches to observability and reliability. What you’ll do: • Evangelize SRE mindset and solve problems through systematization. •... 

    Mindlance

    Austin, TX
    4 days ago
  •  ...Site Reliability Engineer (SRE) New York (Remote) We are seeking an experienced Site Reliability Engineer (SRE) with strong expertise in Dynatrace to join our growing engineering team. The ideal candidate will be responsible for ensuring the reliability, scalability... 
    Remote work

    Staffing the Universe

    United States
    1 day ago
  • $101k - $161k

     ...several prestigious awards, such as Best Engineering Team, Best Company for Diversity,...  ...DescriptionWho You'll Work WithWe’re looking for Site Reliability Engineers to join our growing Arista’s...  ...-as-a-Service (CVaaS) global SRE team. SREs at Arista combine strong software... 

    Arista Networks

    Santa Clara, CA
    2 days ago
  • OB SUMMARYThe Systems Engineer - Site Reliability Engineering (SRE) is responsible for the reliability, scalability, and performance of mission-critical cloud and on-prem services that support millions of Marriot customers globally. This role involves overseeing incident... 
    Full time
    For contractors
    Work at office
    Remote work
    Flexible hours
    Shift work

    Marriott International

    Bethesda, MD
    3 days ago
  • $160k - $200k

    Role Overview We are looking for a Control System Engineer/Site Reliability Engineer (SRE) to integrate and maintain the hardware and software systems that enable QuEra’s quantum controls and software stack. You’ll work closely with software engineers, physicists, hardware... 
    Local area
    Remote work

    QuEra Computing

    Boston, MA
    5 days ago
  • $172k - $300k

    Job DescriptionGM Vehicle Autonomy is forming a centralized Site Reliability Engineering team to make reliability a measurable, engineered property...  ...reliably without depending on heroics.If you are an expert in SRE practices who loves building the machine that builds the... 
    Full time
    Work at office
    Local area
    Remote work
    Work from home
    Relocation
    Relocation package
    Flexible hours

    General Motors

    Austin, TX
    1 day ago
  • $160k - $185k

     ...fitness journey and revolutionized the industry along the way. And we’re just getting started!OverviewThe Sr. Manager, Site Reliability Engineering (SRE) leads the strategy, execution, and continuous improvement of reliability, availability, and performance across Planet... 
    Work at office
    Local area
    Remote work
    Work from home

    Planet Fitness

    Hampton, NH
    1 day ago
  • $73.41 per hour

     ...Job Description Job Description Job Title: Site Reliability Engineer (SRE) Location: Pennington, NJ / Charlotte, NC Duration: Contract - 1 months Pay Range: $73.41/hr (W2) Job ID: 409373 About BCforward BCforward is a leading global IT consulting... 
    Contract work

    BC Forward

    Pennington, NJ
    5 days ago
  • $149.4k - $202k

     ...Senior Software Engineer- Site Reliability Engineering (SRE) DC, MD, VA, CA The Site Reliability Engineering discipline at Noctua Technology, LLC is a strategic force driving digital transformation. We treat operations as a software engineering challenge, focusing... 
    Remote work

    Noctua Technology

    Virginia, MN
    5 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer (SRE). Be the first to apply!