Site Reliability Engineer II
$103.5k - $150kMedallia
Medallia is the pioneer and market leader in Experience Management. Our award-winning SaaS platform, Medallia Experience Cloud, leads the market in the management of experiences, insights, and actions for candidates, customers, employees, patients, and residents alike.
We believe that every experience is a memory that can last a lifetime. Experiences shape the way people feel about a company. And they greatly influence how likely people are to advocate, contribute, and stay. At Medallia, we are committed to creating a world where organizations are loved by their customers and their employees.
We empower exceptional people to create extraordinary experiences together.
Bring your whole self.
The Role and Team
The Site Reliability Engineering organization at Medallia brings together the infrastructure and applications that power a highly reliable global SaaS platform.
As an SRE II, you will help operate and improve the reliability, scalability, and performance of services running across Kubernetes-based environments in cloud and hybrid infrastructure. You will work closely with software engineering teams to build automation, improve operational excellence, and support production services used globally by Medallia customers.
We are looking for engineers who enjoy solving complex technical problems, automating repetitive tasks, improving system reliability, and learning modern cloud-native technologies in a fast-paced environment.
We value engineers who actively seek opportunities to improve scalability and operational efficiency through automation, AI-assisted engineering workflows, and continuous process improvement.
Please note this role participates in a rotating on-call schedule supporting production systems and services.
Engineering Leverage
At Medallia, we hire engineers who scale systems, teams, and outcomes through automation, platform thinking, and AI-assisted engineering.
We value engineers who challenge manual processes, reduce operational toil, and create reusable solutions that improve reliability and productivity for the broader engineering organization.
Successful engineers do not simply solve problems-they eliminate recurring problems through automation, simplification, and self-service capabilities.
Responsibilities- Collaborate with software engineering teams to improve application reliability, scalability, and operational maturity.
- Operate and support production services running in Kubernetes environments.
- Troubleshoot and resolve infrastructure and application issues across the full technology stack.
- Build automation and tooling to reduce operational overhead and eliminate manual work.
- Leverage AI-assisted engineering tools and automation platforms to accelerate troubleshooting, improve productivity, and reduce operational toil.
- Identify opportunities to streamline operational processes through automation, AI-enabled workflows, and self-service solutions.
- Create reusable solutions, tooling, and operational improvements that increase engineering leverage across the team.
- Support CI/CD and GitOps-based deployment workflows.
- Develop and maintain infrastructure-as-code configurations and operational tooling.
- Monitor system health, availability, and performance using observability and alerting platforms.
- Participate in incident response, root cause analysis, and operational improvements.
- Continuously improve reliability, deployment processes, and operational standards.
Candidates based in the Tysons vicinity will be prioritized as this role is Hybrid, 3 days per week onsite.
QualificationsMinimum Qualifications
- 2+ years of experience in Site Reliability Engineering, DevOps, Systems Engineering, Cloud Operations, or related roles.
- Demonstrated experience supporting production environments running on Kubernetes or other containerized platforms.
- Demonstrated experience with cloud infrastructure platforms such as AWS, OCI, or GCP.
- Demonstrated experience with Linux systems administration and troubleshooting.
- Demonstrated experience with scripting or programming languages such as Python, Bash, or Go.
- Familiarity with CI/CD pipelines and Git-based workflows.
- Demonstrated understanding of networking fundamentals including DNS, load balancing, TLS/SSL, and routing concepts.
- Demonstrated experience troubleshooting distributed systems and production incidents.
- Ability to participate in an on-call rotation supporting production systems.
- Fluency in English, both oral and written.
Preferred Qualifications
- Experience with GitOps and tools such as ArgoCD.
- Experience with infrastructure-as-code tools such as Terraform.
- Familiarity with observability platforms such as Prometheus, Grafana, Loki, or OpenTelemetry.
- Experience operating services in hybrid-cloud or multi-region environments.
- Understanding of release strategies such as rolling deployments, canary releases, or blue/green deployments.
- Familiarity with incident management and operational best practices.
- Exposure to security and compliance concepts in production environments.
- Experience using AI-assisted development, automation, or operational tooling to improve engineering productivity and service reliability.
- Demonstrated passion for automation, process improvement, and operational efficiency.
- Strong communication and collaboration skills.
Medallia is committed to equal pay and transparency. The annual base salary range for this position is $103,500 - $150,000. Please note that the salary range information provided is a general guideline and combines all of the distinct labor markets within the US. It is uncommon for an individual to be hired at or near the top of the range for their role and compensation decisions are dependent on a variety of factors. Medallia considers factors such as (but not limited to) scope and responsibilities of the position, candidate's work experience, candidate's work location, education/training, key skills, internal peer equity, external market data, as well as, market and business considerations when making compensation decisions.
Medallia also offers competitive health and wellness benefits, including but not limited to medical, dental, vision, 401(k), short-term and long-term disability, life and AD&D insurance, statutory leaves, paid parental leave, and paid holidays. Benefits and eligibility may vary by location and role.
At Medallia, we celebrate diversity and recognize the value it brings to our customers and employees. Medallia is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age (40 and over), disability, genetic information, veteran status or military service, or any other status protected by state or local law. Individuals with a disability who need an accommodation to apply please contact us at View email address on click.appcast.io. For information regarding how Medallia collects and uses personal information, please review our Privacy Policies. Applications will be accepted for 30 days from the date this role was posted or until the role has been filled.
- ...Site Reliability Engineer II Join the leader in providing smarter solutions for a safer world. The property technology space is growing rapidly, and Kastle Systems is leading the way. Kastle Systems is the leader in managed security, with a track record of introducing...SuggestedRemote work
$95k - $171k
...Opportunities exist to focus on GPU infrastructure, Kubernetes, and ensuring reliability for AI workloads within Akamai's serverless inference platform. As an Site Reliability Engineer II, you will be responsible for: Building and maintaining dashboards, alerts,...SuggestedPermanent employmentWork experience placementWork at officeRemote workWork from homeWorldwideFlexible hours- ...properties, we focus on enhancing property operations and improving guest experiences with cutting‑edge technology. As a Software Engineer II, you will play a critical role in our dynamic team, driving the development and delivery of our world‑class smart property...SuggestedWork experience placementCasual workWork at officeImmediate startFlexible hours
- Exiger is seeking a Platform Engineer II to design and build AWS platform features and contribute to CI/CD pipelines and Kubernetes operations... ...with Terraform, Helm, and observability tools to improve reliability and security across environments. The role requires 0-4 years...Suggested
- A technology solutions provider is seeking an experienced individual for the position of Infrastructure Management - Level II in Arlington, VA. The role involves designing, configuring, and maintaining servers, as well as troubleshooting complex infrastructure issues....Suggested
$92.5k - $146.3k
Platform - Engineering Productivity - Software Engineer II Elastic Cloud 21 July 2025 Elastic, the Search AI Company, enables everyone to find the answers... ...efficient and independent by providing realistic and reliable developer environments. We’re continuously evolving...Local areaWorldwideFlexible hours- ...ears, and hands on the ground at a government customer site, ensuring the reliability and performance of Twenty's mission-critical platform running... ...of deep technical ownership and customer-facing engineering: you'll define how we measure reliability, lead incident...Full timeWork at officeRemote workFlexible hours
- Overview Description : As the Software Engineer II, you will be responsible for engineering and developing stellar software solutions. This position will work with a team of engineers, product and QA to build secure and scalable platforms and applications that will be released...Work at officeFlexible hours
$106k - $169k
...potential. Title and Summary Software Engineer II - Backend/Platform Agentic AI Who is Mastercard... ...ensuring correctness, performance, and reliability in a multi-tenant distributed... ...eligible roles; fitness reimbursement or on‑site fitness facilities; eligibility for tuition...Full timePart timeWorldwideFlexible hours- ...Site Reliability Engineer Mc Lean, VA Long Term Client's Enterprise Data Machine Learning (EDML) employs innovative minds like yourself to design and develop software-systems that can meet the demand of our ever-growing customer base. Like a...Immediate start
$165k - $230k
...with the ultimate goal of enabling human life on Mars. SR. SITE RELIABILITY ENGINEER (STARSHIELD) Starshield leverages SpaceX’s Starlink... ...regulations, applicant must be a (i) U.S. citizen or national, (ii) U.S. lawful, permanent resident (aka green card holder), (iii...Permanent employmentTemporary workImmediate startWeekend work- ...with tech SME and PM Period of performance: Up to 2 years in duration MUST HAVES: •Minimum of 8 years of experience as a Site Reliability Engineerwith a strong understanding of SRE principles for highly scalable and reliable systems •Possess a bachelor's degree...Local areaRelocation package3 days per week
$88.3k - $122.55k
Software Engineer II Are you passionate about building technology that connects the physical and digital worlds? Do you get excited about improving how smart home devices communicate and perform? Alarm.com is looking for a Software Engineer II to join our Protocols...Full timeCasual workWork at officeImmediate startWorldwide$86.8k - $198k
...Job Number: R0238722 Site Reliability Engineer The Opportunity: Engineering to make a system more resilient and efficient frees up time and... ...~ HS diploma or GED ~ DOD 8570 or 8140 IAT Level II Certification, such as Security+ CE, SSCP, CySA+, CCNA-Security...Full timeContract workPart timeWork at officeLocal areaRemote work$150k - $225k
...Site Reliability Engineer Location: Herndon, VA Work Type: Full-Time / Onsite Remote Work: No Job Description Engineering to... ...polygraph ~ HS diploma or GED ~ DOD 8570 or 8140 IAT Level II Certification, such as Security+ CE, SSCP, CySA+, CCNA-...Full timePart timeWork experience placementRemote work$125k - $135k
...Site Reliability Engineer Job number: 880 This is a remote position. Ad Hoc is a technology company that empowers organizations to deliver scalable, impactful digital services. Using modern, agile methods, our team creates products that meet people's needs...Remote workFlexible hours- ...Detail Description: The AWS Site Reliability Engineer (SRE) is responsible for the operational health, availability, and performance of the AWS and Databricks environments built by the Platform Engineering team. You prepare and take ownership of "day two" operations...
- A leading technology firm is looking for a Software Engineer II in McLean, VA to develop and maintain software applications across various platforms including web and cloud-based solutions. The ideal candidate will have a Bachelor's degree in Computer Science or a related...Flexible hours
$106.3k - $221.1k
...more. Join us to drive positive, lasting change that moves missions and the government forward! Job Description The Site Reliability Engineer will ensure the reliability, performance, and scalability of the Client System. The engineer will define and track Key...Live inWork at officeLocal area- Aretum, Llc is seeking a Power Platform Developer II to collaborate with project teams to develop functional applications that meet client needs. You will leverage your expertise in Power Apps and Dynamics 365 to create effective solutions. This remote position involves...Remote job
- Integral Federal, Inc. seeks a Test Engineer II in McLean, Virginia, to provide test engineering support for biometrics projects. This role involves logistics management, coordinating biometric device deployments, and maintaining vital inventory records. The ideal candidate...
- Integral Federal, Inc. is seeking a Cybersecurity Engineer II in McLean, Virginia. The role involves providing comprehensive cybersecurity support for ensuring systems comply with military Risk Management Framework (RMF) requirements and supporting the Authority to Operate...
$115k - $120k
Amentum is a global leader in advanced engineering and innovative technology solutions, trusted by the United States and its allies to address... ...Center (TNOC). Job Description TNCC personnel provide Tier II customer service support to US ARMY INSCOM customers....Hourly payContract workFor contractorsCasual workWork at officeLocal areaShift workNight shiftWeekend work- ...Exiger is a recognized, award‑winning leader in supply chain AI and a FedRAMP authorized provider to the federal government. Site Reliability Engineer Location: U.S. (Hybrid) This role requires U.S. citizenship and eligibility for a U.S. security clearance. Role Summary...Work at officeWork from homeFlexible hours
$100.2k - $203.4k
...training and more. Join us to drive positive, lasting change that moves missions and the government forward! The work As a Site Reliability Engineer, you will play a pivotal role in advancing operational AI adoption within a cutting‑edge Hub-and-Spoke architecture. Your...Live inWork at officeLocal area$116.9k - $234.1k
Site Reliability Engineer The Site Reliability Engineer will ensure the reliability, performance, and scalability of the Client System. The engineer will define and track Key Performance Indicators and Service Level Objectives, identify and resolve performance bottlenecks...Local area$175k - $250k
Senior Cloud Infrastructure Engineer Location: San Francisco, CA (On‑site only) — must live within commuting distance or be willing to relocate. Compensation... ...vision while ensuring scalability, performance, and reliability across environments. What You’ll Do Design, build,...Full timeRelocation- Twenty Technologies is looking for a Site Reliability Engineer to ensure the reliability of its mission-critical platform at a government customer site in Arlington, Virginia. The role involves defining reliability metrics, leading incident response in a secure environment...
- ...(SME) for the Intelligence and Security Command (INSCOM) G3’s Army Trojan Management Office (ATMO). The role involves providing Tier II customer service support to the US ARMY INSCOM customers and troubleshooting issues related to various network hardware including CISCO...Work at office
- Ad Hoc LLC is seeking a Software Engineer II - Front End to contribute to federal civilian digital services projects in a fully remote role. You will build and maintain user-facing apps using Angular, HTML/CSS, and JavaScript, translating UI designs from Figma into accessible...Remote job
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer II. Be the first to apply!
- IT site lead McLean, VA
- junior website developer McLean, VA
- site safety McLean, VA
- site services specialist McLean, VA
- site recruiter McLean, VA
- website content developer McLean, VA
- site leader McLean, VA
- on-site clinical research associate (traveling/remote) McLean, VA
- site reliability engineer
- junior site reliability engineer


