Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer

GrabJobs

SRE Support Engineer - Observability While this position is not currently open, we are interviewing strong candidates for upcoming opportunities on this team. Location: Remote | Time Zone: (US, Canada, Brazil, Chile, Colombia, Mexico)(8AM–5PM Pacific) Freedom to grow. Power to deliver. Virtasant is a global technology services company delivering large-scale cloud, data, and engineering solutions across 130+ countries. We partner with some of the world’s largest organizations to help them build, operate, and scale internal platforms used by tens of thousands of engineers. For this role, you will be supporting one of the most advanced internal developer platforms in the world, powering products used by hundreds of millions of people. The problems you will solve are deep, complex, and essential to keeping a global-scale organization moving. Role Overview The Observability & Tools Support Engineer provides high-impact technical support for customers of a large technology company’s internal IaaS platform, with a focus on monitoring, alerting, telemetry, and operational tooling . This role spans a wide range of support—from white-glove onboarding and end-to-end customer enablement, to deep technical troubleshooting across Linux, networking, and observability systems (especially Prometheus and AlertManager ). You will also contribute to improving the support function itself: strengthening tooling, documentation, workflows, and feedback loops so the service scales. Success depends on excellent troubleshooting, strong written communication, comfort working with highly technical customers, and the maturity to identify patterns and drive operational improvements beyond individual ticket resolution. Business Outcome Become a trusted frontline expert for the customer’s observability ecosystem and operational tooling - delivering fast, accurate support across Slack and tickets, improving monitoring reliability, and reducing incident impact through better triage, troubleshooting, onboarding, and knowledge capture. Success Measures Healthy volume of threads and tickets handled with high-quality outcomes Consistent achievement of time-based SLAs High customer satisfaction through surveys Accurate classification of issue type, severity, and recurring patterns Reduced repeat issues through better docs, tooling, and scalable onboarding What Will Be True When You Succeed Customers can onboard smoothly to monitoring/alerting with minimal friction Monitoring and alerting issues are resolved quickly, with fewer escalations Linux and networking-related incidents reach resolution faster due to strong troubleshooting and clean handoffs Engineering and SRE teams receive clear, actionable feedback based on real customer trends Knowledge base content prevents tickets and accelerates self-service Core Work Units 1) Frontline Support for Observability & Tooling Manage Slack threads and tickets (roughly 50/50) Handle a broad range of customer support: simple issue resolution through end-to-end onboarding Provide clear, structured guidance to highly technical customers Maintain strong attention to detail while managing multiple interactions in parallel 2) Deep-Dive Troubleshooting & Incident Support Troubleshoot, isolate, and resolve monitoring and alerting issues (especially Prometheus + AlertManager ) Troubleshoot complex Linux and networking issues (TCP/IP fundamentals required) Support OpenTelemetry, tracing, and telemetry pipelines , including investigation of gaps in signals and instrumentation Drive incidents to resolution in partnership with Engineering/SRE teams 3) Documentation & Knowledge Development Build and maintain customer-facing and internal knowledge base articles Create informational posts for the community support platform Turn repeated issues into reusable guides, checklists, and onboarding playbooks 4) Trend Analysis & Feedback to Engineering Analyze and categorize customer interaction trends Provide accurate, meaningful feedback to Engineering and SRE orgs to improve product/tooling Identify “top offenders” and propose practical fixes (tooling, docs, process, product) 5) Operational Excellence & Continuous Improvement Participate in post-mortem reviews and drive follow-through on improvements Contribute meaningfully to team objectives and goals (process, tooling, and service scaling) Bring creativity and discretion to resolve highly complex issues “outside the box” High-Quality Work - what top performance looks like Frontline Support Moves smoothly from triage to deeper analysis without losing the customer Communicates clearly and confidently with technical users Maintains clean follow-ups and thread hygiene even with high context switching Troubleshooting Rapidly isolates issues across monitoring/alerting configs, Linux runtime behavior, and network connectivity Uses structured approaches to incident handling: hypothesis → test → evidence → resolution Produces high-signal writeups that accelerate downstream resolution Documentation & Enablement Documentation is clear enough that customers avoid opening tickets Onboarding flows reduce time-to-value and prevent common misconfigurations Captures “tribal knowledge” quickly and makes it reusable Operational Excellence Obsessing over details: correct severity, accurate tagging, clean timelines, strong handoffs Spots patterns early and proactively proposes improvements that scale support Typical Day / Work Patterns ~50% Slack support, ~50% ticket handling Deep-dive investigations during lower ticket volume periods Documentation writing and lightweight tooling/process improvements when patterns emerge Weekly team review of escalations, themes, and operational improvements High rate of context switching and parallel issue management Required Skills & Experience (Non-Negotiable) Several years supporting highly scalable applications and web services Hands-on experience with open-source observability and cloud-native tooling, including: Kubernetes (and container fundamentals) Prometheus and AlertManager troubleshooting OpenTelemetry and distributed tracing concepts Strong understanding of the Linux operating system (command line, process/network debugging, logs) Good understanding of infrastructure observability principles (signals, alerting strategy, SLO thinking, noise reduction) Good understanding of the TCP/IP suite and practical networking troubleshooting Strong experience troubleshooting ambiguous, multi-layer issues Excellent analytical capability and strong attention to detail Strong written and verbal communication (clear, structured, customer-friendly) Comfortable working with a very technical customer base Passion for Technical Support and a service mindset Nice-to-Haves Experience improving or supporting internal support tooling or workflows (automation, templates, runbooks) Experience operating at scale in a services environment (pattern detection, KPI/SLA awareness, operational process maturity) Familiarity with Grafana, log aggregation, incident tooling, and production support practices Prior SRE or platform support experience Minimum Qualifications 3–7+ years in Technical Support Engineering, SRE support, DevOps, Platform Support, or similar Demonstrated experience supporting distributed systems, IaaS, or cloud platforms Strong Linux, troubleshooting, and customer-facing communication background Evidence of documentation, knowledge-base contributions, and process improvement mindset Disqualifiers: weak Linux fundamentals, inability to troubleshoot systematically, poor written communication, or discomfort supporting highly technical users. What You’ll Love Real technical problem solving with tangible customer impact A role that blends deep troubleshooting with scaling support via docs, tooling, and process High autonomy in a remote-first environment What May Be Challenging High context switching and managing multiple threads in parallel Repeated patterns that require discipline to convert pain into scalable improvements Supporting high-visibility systems where speed and accuracy matter Differentiation Industry: Remote-first, trust-based culture; global team; autonomy; modern systems; meaningful technical challenges Internal: High-impact, customer-facing observability support; direct influence on tooling and process maturity; opportunity to shape scalable support practices

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in Washington DC vacancy
  • $230k - $250k

    GovCIO is hiring a Site Reliability Engineer with an active Secret clearance to ensure reliability, scalability, performance, and availability of mission-critical systems by combining software engineering practices with infrastructure operations expertise. This role is... 
    Suggested
    Remote work

    Govcio

    Arlington, VA
    19 hours ago
  • $166k - $220k

     ...requirements and customer expectations. Our systems integration engineers internalize the nuances of each deployment, ensuring the...  ...-to-end solutions we ship.ABOUT THE JOBWe are looking for a Site Reliability Engineer (SRE) to join AGD, our rapidly growing team in Irvine... 
    Suggested
    Full time
    Work experience placement
    Immediate start

    Anduril Industries

    Washington DC
    4 days ago
  • $115.5k - $164.8k

     ...mission that matters at a company where you matter.Your ImpactAs an engineer on the APX SRE CloudOps team, you will spend a significant...  ...that replace what previously required human intervention with reliable, tested automation. You will also participate in on-call rotations... 
    Suggested
    Work experience placement
    Work at office
    Remote work

    Axon

    Washington DC
    2 days ago
  • $112k - $179k

     ...delivery of system, network, software, and security solutions.About The RolePeraton is seeking a self-driven and resourceful Site Reliability Engineer to join our dynamic of Network and UC engineers in Washington, DC. This position combines software engineering and systems... 
    Suggested
    Contract work
    Worldwide
    Shift work

    Peraton Corporation

    Washington DC
    5 days ago
  • $150k - $180k

     ...Umbra.About the JobWe are seeking an experienced SeniorSite Reliability Engineer to help design, build, operate, and scale the mission- and business...  ...impact across the organization.This position is based on-site in either our Arlington, VA office, Reston, VA office or... 
    Suggested
    Permanent employment
    Full time
    Work at office
    Local area
    Remote work
    Worldwide

    Umbra

    Arlington, VA
    4 days ago
  • $165k - $230k

     ...is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SR. SITE RELIABILITY ENGINEER (STARSHIELD)Starshield leverages SpaceX’s Starlink technology and launch capability to support national security efforts.... 
    Permanent employment
    Temporary work
    Immediate start
    Weekend work

    SpaceX

    Washington DC
    2 days ago
  • $185k - $230k

    As a Sr. Site Reliability Engineer (SRE) III, you’ll work as part of a collaborative and high-performing team providing your expertise to deliver technical solutions within the highest levels of the federal government.We know that you can’t have great technology services... 
    Full time
    Local area
    Immediate start

    MetroStar Systems

    Washington DC
    5 days ago
  • $125k - $185k

     ...lifesaving drugs, forecast supply chain disruptions, locate missing children, and more.The RoleWe’re looking for Forward Deployed Site Reliability Engineers who can help us build, operate, and maintain high-performance, scalable, and reliable services for our production... 
    Full time
    Work experience placement
    Work at office
    Remote work
    Work from home
    Relocation package

    Palantir Technologies

    Washington DC
    5 days ago
  • $125k - $185k

    Washington, D.C.Engineering /Full-time /HybridA World-Changing CompanyPalantir builds the world’s leading software for data-driven decisions...  ...locate missing children, and more.The RoleWe’re looking for Site Reliability Engineers who can help us build, operate, and maintain high-... 
    Full time
    Work experience placement
    Work at office
    Remote work
    Work from home
    Relocation package

    Palantir Technologies

    Washington DC
    2 days ago
  •  ...support Pension plan Paid maternity leave 401(k) Get notified when a new job is posted. Sign in to set job alerts for “Senior Site Reliability Engineer” roles. Bellevue, WA $204,000.00-$259,000.00 1 day ago Seattle, WA $115,000.00-$175,000.00 5 months ago Senior ServiceNow... 
    Contract work
    Remote work

    Signature IT World Inc

    Washington DC
    4 days ago
  • $175k - $250k

     ...Senior Cloud Infrastructure Engineer Location: San Francisco, CA. Remote unavailable. Modality: On‑Site only. Must live within commuting distance of San Francisco or...  ...while ensuring scalability, performance, and reliability across environments. What You’ll Do Design, build... 
    Full time
    Remote work
    Relocation
    Relocation package

    The Recruiting Guy

    Washington DC
    4 days ago
  •  ...ears, and hands on the ground at a government customer site, ensuring the reliability and performance of Twenty's mission-critical platform running...  ...of deep technical ownership and customer-facing engineering: you'll define how we measure reliability, lead incident... 
    Full time
    Contract work
    Remote work
    Flexible hours

    Twenty Technologies

    Arlington, VA
    2 days ago
  • $140k - $205k

     ...Senior Technology Site Reliability Engineer Cooley is seeking a Senior Site Reliability Engineer to join the Infrastructure & Development Operationsteam. Position summary: The Senior Technology Site Reliability Engineer("SRE") is responsible for ensuring the reliability... 
    Full time
    Temporary work
    Work at office
    Flexible hours
    Weekend work

    Cooley

    Washington DC
    3 days ago
  • $106.3k - $221.1k

     ...Senior Site Reliability Engineer At Accenture Federal Services, nothing matters more than helping the US federal government make the nation stronger and safer and life better for people. Our 13,000+ people are united in a shared purpose to pursue the limitless potential... 
    Live in
    Work at office
    Local area

    Accenture Federal Services

    Arlington, VA
    4 days ago
  • $82.3k - $228.8k

     ...inclusive environment, empowering our employees to be their authentic selves. We are seeking a highly experienced Senior Site Reliability Engineer – Compute Platforms to design, implement, and support Kubernetes on baremetal and hypervisor platforms in a private cloud... 
    Temporary work
    Work at office
    Remote work
    Worldwide
    3 days per week

    GrabJobs

    Washington DC
    22 hours ago
  •  ...Job Title: Site Reliability Engineer (SRE) Location: Washington, DC (Onsite) Clearance: TS/SCI Position Overview Seeking a highly motivated Site Reliability Engineer (SRE) to support mission-critical enterprise applications and infrastructure in... 

    Input Technology Solutions

    Washington DC
    4 days ago
  •  ...Site Reliability Engineer Location- Wilmington De, Washington DC, Dallas, TX (Onsite Position) Full time position Minimum Qualifications Bachelor’s degree in computer science, Engineering, or a related technical field. Minimum of 5 years of experience... 
    Full time

    Yochana

    Washington DC
    1 day ago
  •  ...Site Reliability Engineer (SRE) Randstad is seeking a skilled and proactive Site Reliability Engineer (SRE) to join our client in the Washington D.C. area, focusing on optimizing the availability, performance, and scalability of critical production services. The ideal... 

    Software Technology Inc

    Washington DC
    4 days ago
  •  ...Site Reliability Engineer (SRE) Dexian is seeking a savvy Site Reliability Engineer (SRE) who will play a key role in building a sustainable platform by developing systems for analyzing environments, predicting, and resolving issues, and supporting the production environment... 
    Work experience placement

    Samprasoft

    Washington DC
    22 hours ago
  • $75.7k - $136.3k

     ...solve complex challenges? Do you have a passion for automation and building systems that scale? Join our highly skilled Site Reliability Engineering team! Our team designs, develops, and manages applications and infrastructure that support Akamai Cloud's products and... 
    Work experience placement
    Work at office

    Akamai

    Washington DC
    3 days ago
  • $100k - $120k

     ...Site Reliability Engineer This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Site Reliability Engineer based in United States. The Site Reliability Engineer will play a critical role in... 
    Temporary work
    Remote work
    Flexible hours

    Jobgether

    Washington DC
    4 days ago
  • $100k - $110k

     ...for new hire onboarding and occasional in-person team meetings and company events. We are seeking an operational-focused Site Reliability Engineer (SRE) to maximize the availability, performance, and resilience of our production healthcare systems. In this role, you will... 
    Permanent employment
    Remote work
    Flexible hours

    GrabJobs

    Washington DC
    1 day ago
  •  ...A leading security infrastructure firm in Washington, D.C. is seeking a hands-on Site Reliability Engineer (SRE) with expertise in Kubernetes and cloud infrastructure. The role emphasizes total ownership of security infrastructure while defending against advanced threats... 

    Cyrad Solutions LLC

    Washington DC
    22 hours ago
  •  ...Role Overview We are seeking a high-caliber Site Reliability Engineer (SRE) to join our Forward Engineering team. You will be the guardian of our production ecosystems, ensuring that our complex, data-driven AI platforms remain resilient, scalable, and highly performant... 
    Local area

    Tiger Analytics

    Washington DC
    2 days ago
  • $174k - $239k

     ...From core infrastructure to enterprise platforms, we partner across functions to drive scale, reliability, and innovation through technology.The Staff Site Reliability Engineer OpportunityOkta Federal, Inc. is looking for an experienced Staff TDI Site Reliability... 
    Local area
    Worldwide
    Flexible hours

    Okta

    Washington DC
    4 days ago
  • $207k - $284.9k

     ...on this mission. If you are too, let's talk.Senior Manager, Site Reliability EngineeringSecure Every Identity, from AI to HumanIdentity is...  ...mission. If you are too, let's talk.The Federal Operations Engineering GroupOkta's Federal Operations team supports government customers... 
    Permanent employment
    Local area
    Worldwide
    Flexible hours
    Day shift

    Okta

    Washington DC
    2 days ago
  • $174k - $238k

     ...work. We're all in on this mission. If you are too, let's talk.The Federal SRE TeamWe are looking for an experienced Staff Site Reliability Engineer to join Okta's Federal SRE team for the Emerging Products Group (EPG). Our mission is to build highly reliable, scalable,... 
    Local area
    Worldwide
    Flexible hours

    Okta

    Washington DC
    2 days ago
  • $166k - $220k

     ...failure. As such, it is critical that Anduril services are reliable and maintainable. This means that all services &...  ...ground systems & Kubernetes infrastructure.ABOUT THE JOBAs a Site Reliability Engineer on the Observability team, you will build & operate Anduril... 
    Full time
    Work experience placement
    Immediate start

    Anduril Industries

    Washington DC
    1 day ago
  • $130k - $180k

     ...alongside some of the most experienced and innovative leaders and engineers in the field. Where we work Headquartered in Amsterdam and...  ...an in-house AI R&D team. The role Nebius is looking for a Site Reliability Engineer in Hardware Infrastructure team. You’re welcome to... 
    Temporary work
    Work at office
    Immediate start
    Remote work
    Flexible hours

    GrabJobs

    Washington DC
    2 days ago
  • $149.4k - $202k

     ...Senior Software Engineer- Site Reliability Engineering (SRE) DC, MD, VA, CA The Site Reliability Engineering discipline at Noctua Technology, LLC is a strategic force driving digital transformation. We treat operations as a software engineering challenge, focusing... 
    Remote work

    Noctua Technology

    Washington DC
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!