Site Reliability Engineer (SRE)
Amplifire Elearning
Site Reliability Engineer
Amplifire is seeking a Site Reliability Engineer to improve the reliability, scalability, performance, and operational efficiency of our cloud-based platform. Working alongside DevOps engineers within the Platform Operations team, this role combines software engineering and systems operations with a focus on observability, automation, incident reduction, and operational excellence.
You will partner closely with software engineering, QA, security, and DevOps to establish reliability practices, improve production visibility, strengthen incident response, automate operational workflows, and help engineering teams deliver changes safely and confidently.
This role supports systems operating under regulatory and compliance requirements, including FedRAMP and SOC 2. The ideal candidate understands that reliability, traceability, security, and change management enable sustainable development velocity rather than compete with it.
Amplifire expects all technical team members to leverage AI-assisted tools and workflows as force multipliers for productivity, learning, automation, and problem-solving while maintaining strong engineering skills, sound judgment, security standards, and operational accountability.
Reliability, Observability & Performance
· Establish and maintain service-level indicators (SLIs), service-level objectives (SLOs), error budgets, and other measures of system health.
· Build and continuously improve monitoring, logging, tracing, dashboards, and alerting that provide actionable visibility into application and infrastructure health.
· Analyze system behavior, performance, capacity, and reliability trends to proactively identify risks and improvement opportunities.
· Partner with engineering teams to define reliability requirements and improve the availability, scalability, and performance of production systems.
Incident Response & Operational Readiness
· Participate in the on-call rotation and respond to production incidents with urgency and sound technical judgment.
· Improve incident detection, triage, escalation, communication, mitigation, and recovery processes.
· Lead or contribute to blameless post-incident reviews, root cause analysis, and corrective actions.
· Develop and maintain runbooks, recovery procedures, and operational practices that improve service resilience and reduce mean time to detect and restore service.
Automation, Infrastructure & Delivery
· Identify and eliminate operational toil through automation, self-service tooling, and continuous improvement.
· Build and maintain cloud infrastructure using Infrastructure as Code practices and tools such as Terraform or AWS CDK.
· Improve CI/CD pipelines, deployment safeguards, rollback capabilities, and progressive delivery practices.
· Develop internal tools and automation that improve reliability, resilience, and engineering productivity.
· Support scalable, secure, and cost-effective production environments.
Collaboration & Continuous Improvement
· Partner with software engineers to embed reliability, operability, and observability throughout the development lifecycle while reducing operational friction and helping teams safely own their services in production.
· Help engineering teams diagnose complex production issues across application and infrastructure layers.
· Contribute to security, compliance, capacity-planning, cloud cost optimization, and operational standards.
· Use AI-assisted tools responsibly to improve troubleshooting, automation, documentation, and operational efficiency while sharing knowledge and continuously improving team practices.
Requirements
· 4+ years of experience in Site Reliability Engineering, DevOps, cloud infrastructure, systems engineering, software engineering, or a related production-focused role.
· Demonstrated experience supporting reliable, customer-facing applications in a production cloud environment.
· Hands-on experience designing, deploying, and operating production workloads on AWS.
· Experience building or maintaining Infrastructure as Code, observability solutions, incident response processes, and CI/CD automation.
· Experience identifying and reducing operational toil through automation and continuous improvement.
· Experience participating in on-call rotations, troubleshooting production issues, performing root cause analysis, and implementing preventive improvements.
Technical Skills
· Proficiency with at least one scripting or programming language, such as Python, Bash, JavaScript/TypeScript, Go, or Java.
· Working knowledge of Linux systems, networking, DNS, load balancing, and common cloud architecture patterns.
· Experience with containers and container-based deployment practices; Docker experience is required, while Kubernetes or similar orchestration experience is beneficial.
· Ability to analyze logs, metrics, traces, and system behavior to troubleshoot issues across application and infrastructure layers.
· Understanding of reliability concepts such as SLIs, SLOs, error budgets, availability, latency, capacity planning, and graceful degradation.
· Familiarity with secure configuration, secrets management, access controls, vulnerability remediation, and other operational security fundamentals.
AI & Modern Tooling
· Experience using AI-assisted development or operational tools to improve productivity, automation, troubleshooting, documentation, or engineering workflows.
· Ability to critically evaluate AI-generated outputs and apply appropriate validation before production use.
· Interest in adopting emerging tools and practices that improve engineering effectiveness while maintaining operational excellence.
Collaboration & Problem-Solving Skills
· Strong debugging and systems-thinking skills, including the ability to work through ambiguous, cross-service production issues.
· Ability to communicate clearly during incidents and translate technical findings for engineering and business stakeholders.
· Experience collaborating with software engineering, QA, security, support, and product teams.
· Ability to independently own reliability improvements while seeking input and alignment when appropriate.
· A proactive approach to problem-solving and a willingness to challenge existing practices constructively.
What Success Looks Like
Within your first year:
· Service-level indicators and objectives are established for critical systems.
· Monitoring, alerting, and observability provide actionable visibility across the platform.
· Incident response processes are more structured, repeatable, and measurable.
· Mean time to detect and restore service trends improve through better tooling and operational practices.
· Engineering teams have greater self-service access to operational insights and reliability tooling.
· Manual operational work is reduced through automation and platform improvements.
· Reliability, performance, and scalability risks are identified proactively rather than reactively.
· Demonstrates a proactive approach to problem solving and does not accept existing processes simply because they have historically been done that way.
Amplifire is the leading AI Learning Platform built on brain science that delivers the proven results of 1:1 expert instruction at enterprise scale. We deliver one-on-one AI instruction at scale, detect where people are confidently wrong, and fix it before it becomes a mistake—reducing training time by 50–80% while driving measurable performance outcomes. Trusted by leading organizations in healthcare, accounting, life sciences and other high-stakes industries, Amplifire enables teams to achieve verified competency faster, reduce risk, and perform at the highest level when it matters most.
Salary Description 130,000 - 155,000
- ...candidate for this role to work on site in the specified location(s).... ...by delivering innovative and reliable technology solutions that... ...Bank Platform Operations and Engineering organization, you will help ensure... ...Site Reliability Engineering (SRE) practices that enhance client...SuggestedFull timeWork at office
- ...Lovelace is the only provider of enterprise-scale context engines capable of analyzing trillions of real-time data points... ...~ Lovelace AI is seeking a highly skilled and motivated Site Reliability Engineer (SRE) to join our growing team. As an SRE at Lovelace AI, you...SuggestedFull time
$101k - $161k
...several prestigious awards, such as Best Engineering Team, Best Company for Diversity,... ...DescriptionWho You'll Work WithWe’re looking for Site Reliability Engineers to join our growing Arista’s... ...-as-a-Service (CVaaS) global SRE team. SREs at Arista combine strong software...Suggested$120k - $200k
...PermContact: Kunal DaveContact Email: ****@*****.*** Reliability Engineer(SRE) ResponsibilitiesGlobal Architecture & Disaster Recovery Participate... ...(e.g., Chaos Engineering, resilience testing, automated recovery)SkillsBilingual Mandarin Site Reliability Engineer(SRE)SuggestedOverseas- ...Overview We are seeking an experienced Site Reliability Engineer (SRE) – Microsoft Hyper-V & Private Cloud to operate highly available private cloud and Virtual Desktop Infrastructure (VDI) platforms based on Microsoft Hyper-V. This role combines deep Hyper-V expertise...Suggested
- ...ITIL-based processes. Define and monitor SRE metrics including SLIs, SLOs, and error... ...~ Bachelor’s degree in Computer Science, Engineering, or a related technical field. ~3+ years of experience in Site Reliability Engineering, DevOps, Cloud Engineering, or...
- ...Site Reliability Engineer (SRE) Location: Remote Shift Timings: 5:30 PM to 3:00 AM IST to ensure support for global operations. Job Description: We are seeking a skilled Site Reliability Engineer (SRE) to join our dynamic team. The ideal candidate will have...Remote workShift work
$175k - $185k
...Senior Site Reliability Engineer (SRE) Remote, US Branch is on a mission to empower workers with financial freedom. We do this by helping companies accelerate payments and providing working Americans with accessible, free financial services. We're committed to building...Daily paidRemote workHome officeFlexible hours- ...risk—the leading cause of cybersecurity breaches—and build safer, more resilient organizations. The Role: As a Senior Site Reliability Engineer (SRE) at Dune Security, you will play a critical role in ensuring our platform's stability, scalability, and security. You will...Full timeWork at office
- ...Senior Site Reliability Engineer At Swile, we believe that good products can help reduce friction in daily professional life and boost employee... ...Brazil. Your role as a Senior Site Reliability Engineer (SRE) centers around creatively solving problems, ensuring a balance...Remote work
- ...in Computer Science, Information Technology, Engineering, or equivalent field ~3-5 years of experience in Site Reliability Engineering, Production Support, Platform Engineering... ...application health ~ Understanding of SRE principles, including observability,...Remote work
$160k - $200k
Role Overview We are looking for a Control System Engineer/Site Reliability Engineer (SRE) to integrate and maintain the hardware and software systems that enable QuEra’s quantum controls and software stack. You’ll work closely with software engineers, physicists, hardware...Local areaRemote work$160k - $185k
...fitness journey and revolutionized the industry along the way. And we’re just getting started!OverviewThe Sr. Manager, Site Reliability Engineering (SRE) leads the strategy, execution, and continuous improvement of reliability, availability, and performance across Planet...Work at officeLocal areaRemote workWork from home- ...shape the future of our communities.This is a Software Engineering position at Director level, which is part of the job family... ...businesses. This role is for an experienced and driven Site Reliability Engineer (SRE) to join our AI Platform team to help support, scale and...
- ...Tenable cloud products and ensuring they’re reliable and highly available in cloud... ...complex projects Collaboration with cloud engineers in understanding new cloud technologies,... ...citizen required ~2+ years of related SRE experience ~ Apply core software engineering...Full timeWork experience placementRemote work
$100k - $180k
...Site Reliability Engineer (SRE) - Remote Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States. This is a fantastic opportunity to join an established...Full timeH1bLocal areaImmediate startRemote workVisa sponsorship- ...security to responsibly propel the global lottery industry ever forward. Position Summary We are looking for a skilled Site Reliability Engineer (SRE) to enhance the stability, performance, and reliability of our production systems. The SRE will work closely with...Permanent employmentWork experience placementLocal area
- ...more than the needs of businesses today; we are building the nervous system for a borderless global economy. As a Site Reliability Engineer (SRE) at Unlimit, you will help ensure the reliability, scalability, and performance of our core platform and services. You’...Full timeLocal area
- ...Job Description:- Our client is seeking a Senior Site Reliability Engineer (SRE) with 10 15 years of experience to support front-office trading systems in a production environment. This role focuses on troubleshooting complex trading infrastructure, managing observability...
$80k - $95k
...join our dynamic team supporting the company’s users, applications, and web-based product offerings. In this role, the Site Reliability Engineer (SRE) will play a key role in maintaining resources at peak efficiency to guarantee staff are able to perform their...Remote workVisa sponsorshipWork visa- JOB SUMMARY: The SRE Service Availability Manager plays a key role in ensuring the peak performance and availability of our... ...and services. This position combines proactive site reliability engineering with adept incident command to lead our efforts in minimizing...Full timeRemote workFlexible hoursShift work
$150k - $160k
Front-End & AdTech Site Reliability Engineer (SRE)Haymarket Media, Inc. is seeking a Front-End & AdTech Site Reliability Engineer (SRE) to join the Engineering team. This position is located in our New York, NY office; three (3) days in office depending on business needs...Work at officeLocal area- Site Reliability Engineer - Vice PresidentSite Reliability Engineering (SRE) is an engineering discipline that combines software and systems engineering to build and run scalable, massively distributed, fault-tolerant systems. At Goldman Sachs, SRE is responsible for improving...
$149.4k - $202k
...Senior Software Engineer- Site Reliability Engineering (SRE) DC, MD, VA, CA The Site Reliability Engineering discipline at Noctua Technology, LLC is a strategic force driving digital transformation. We treat operations as a software engineering challenge, focusing...Remote work- ...startups across the US. We’re building a pool of world-class Site Reliability Engineers for current roles and for upcoming opportunities. You will... ...into one of our partner startups or added to our vetted SRE network for future projects. This role is ideal for engineers...Local area
$100k - $200k
OPPO US Research Center is seeking a skilled and proactive Site Reliability Engineer (SRE) to join our team. In this role, you will be responsible for ensuring the stability, scalability, and performance of our application systems. The ideal candidate is passionate about...Full time- Compliance EngineeringWe are Compliance Engineering, a global team of more than 500 engineers... ...the Compliance application portfolio.SRE at Goldman Sachs combines software and systems... ...for changes that improve capacity and reliability.Practicing sustainable incident...
- ...Information Technology group delivers secure, reliable technology solutions that enable... ...RoleAs a Senior Application Support Engineer, you will help power DTCC's global... ...and settlement.Leveraging Site Reliability Engineering (SRE) principles, you will support a portfolio...Remote workFlexible hours
$106.5k - $177.5k
Role Description The Site Reliability Engineering discipline at Noctua Technology, LLC is a strategic force driving digital transformation. We treat... .... We are seeking a motivated Site Reliability Engineer (SRE) to join our dynamic team. As a key contributor, you will apply...Full timeRemote work- DescriptionJob Description SummaryThe Digital Site Reliability Engineer (SRE) - GCP Cloud Adoption Engineer is responsible for facilitating the migration, adoption, and optimization of Google Cloud Platform (GCP) services within the organization.Job DescriptionSummary:...Full timeH1bWork at officeRemote workWork from homeFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer (SRE). Be the first to apply!
- site reliability engineer United States
- lead site reliability engineer United States
- site reliability engineering manager United States
- site reliability engineer remote United States
- site reliability engineer sre United States
- website content developer United States
- after school site coordinator United States
- site leader United States
- site merchandiser United States
- on-site clinical research associate (traveling/remote) United States




