Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer

GrabJobs

SRE Support Engineer - Observability While this position is not currently open, we are interviewing strong candidates for upcoming opportunities on this team. Location: Remote | Time Zone: (US, Canada, Brazil, Chile, Colombia, Mexico)(8AM–5PM Pacific) Freedom to grow. Power to deliver. Virtasant is a global technology services company delivering large-scale cloud, data, and engineering solutions across 130+ countries. We partner with some of the world’s largest organizations to help them build, operate, and scale internal platforms used by tens of thousands of engineers. For this role, you will be supporting one of the most advanced internal developer platforms in the world, powering products used by hundreds of millions of people. The problems you will solve are deep, complex, and essential to keeping a global-scale organization moving. Role Overview The Observability & Tools Support Engineer provides high-impact technical support for customers of a large technology company’s internal IaaS platform, with a focus on monitoring, alerting, telemetry, and operational tooling . This role spans a wide range of support—from white-glove onboarding and end-to-end customer enablement, to deep technical troubleshooting across Linux, networking, and observability systems (especially Prometheus and AlertManager ). You will also contribute to improving the support function itself: strengthening tooling, documentation, workflows, and feedback loops so the service scales. Success depends on excellent troubleshooting, strong written communication, comfort working with highly technical customers, and the maturity to identify patterns and drive operational improvements beyond individual ticket resolution. Business Outcome Become a trusted frontline expert for the customer’s observability ecosystem and operational tooling - delivering fast, accurate support across Slack and tickets, improving monitoring reliability, and reducing incident impact through better triage, troubleshooting, onboarding, and knowledge capture. Success Measures Healthy volume of threads and tickets handled with high-quality outcomes Consistent achievement of time-based SLAs High customer satisfaction through surveys Accurate classification of issue type, severity, and recurring patterns Reduced repeat issues through better docs, tooling, and scalable onboarding What Will Be True When You Succeed Customers can onboard smoothly to monitoring/alerting with minimal friction Monitoring and alerting issues are resolved quickly, with fewer escalations Linux and networking-related incidents reach resolution faster due to strong troubleshooting and clean handoffs Engineering and SRE teams receive clear, actionable feedback based on real customer trends Knowledge base content prevents tickets and accelerates self-service Core Work Units 1) Frontline Support for Observability & Tooling Manage Slack threads and tickets (roughly 50/50) Handle a broad range of customer support: simple issue resolution through end-to-end onboarding Provide clear, structured guidance to highly technical customers Maintain strong attention to detail while managing multiple interactions in parallel 2) Deep-Dive Troubleshooting & Incident Support Troubleshoot, isolate, and resolve monitoring and alerting issues (especially Prometheus + AlertManager ) Troubleshoot complex Linux and networking issues (TCP/IP fundamentals required) Support OpenTelemetry, tracing, and telemetry pipelines , including investigation of gaps in signals and instrumentation Drive incidents to resolution in partnership with Engineering/SRE teams 3) Documentation & Knowledge Development Build and maintain customer-facing and internal knowledge base articles Create informational posts for the community support platform Turn repeated issues into reusable guides, checklists, and onboarding playbooks 4) Trend Analysis & Feedback to Engineering Analyze and categorize customer interaction trends Provide accurate, meaningful feedback to Engineering and SRE orgs to improve product/tooling Identify “top offenders” and propose practical fixes (tooling, docs, process, product) 5) Operational Excellence & Continuous Improvement Participate in post-mortem reviews and drive follow-through on improvements Contribute meaningfully to team objectives and goals (process, tooling, and service scaling) Bring creativity and discretion to resolve highly complex issues “outside the box” High-Quality Work - what top performance looks like Frontline Support Moves smoothly from triage to deeper analysis without losing the customer Communicates clearly and confidently with technical users Maintains clean follow-ups and thread hygiene even with high context switching Troubleshooting Rapidly isolates issues across monitoring/alerting configs, Linux runtime behavior, and network connectivity Uses structured approaches to incident handling: hypothesis → test → evidence → resolution Produces high-signal writeups that accelerate downstream resolution Documentation & Enablement Documentation is clear enough that customers avoid opening tickets Onboarding flows reduce time-to-value and prevent common misconfigurations Captures “tribal knowledge” quickly and makes it reusable Operational Excellence Obsessing over details: correct severity, accurate tagging, clean timelines, strong handoffs Spots patterns early and proactively proposes improvements that scale support Typical Day / Work Patterns ~50% Slack support, ~50% ticket handling Deep-dive investigations during lower ticket volume periods Documentation writing and lightweight tooling/process improvements when patterns emerge Weekly team review of escalations, themes, and operational improvements High rate of context switching and parallel issue management Required Skills & Experience (Non-Negotiable) Several years supporting highly scalable applications and web services Hands-on experience with open-source observability and cloud-native tooling, including: Kubernetes (and container fundamentals) Prometheus and AlertManager troubleshooting OpenTelemetry and distributed tracing concepts Strong understanding of the Linux operating system (command line, process/network debugging, logs) Good understanding of infrastructure observability principles (signals, alerting strategy, SLO thinking, noise reduction) Good understanding of the TCP/IP suite and practical networking troubleshooting Strong experience troubleshooting ambiguous, multi-layer issues Excellent analytical capability and strong attention to detail Strong written and verbal communication (clear, structured, customer-friendly) Comfortable working with a very technical customer base Passion for Technical Support and a service mindset Nice-to-Haves Experience improving or supporting internal support tooling or workflows (automation, templates, runbooks) Experience operating at scale in a services environment (pattern detection, KPI/SLA awareness, operational process maturity) Familiarity with Grafana, log aggregation, incident tooling, and production support practices Prior SRE or platform support experience Minimum Qualifications 3–7+ years in Technical Support Engineering, SRE support, DevOps, Platform Support, or similar Demonstrated experience supporting distributed systems, IaaS, or cloud platforms Strong Linux, troubleshooting, and customer-facing communication background Evidence of documentation, knowledge-base contributions, and process improvement mindset Disqualifiers: weak Linux fundamentals, inability to troubleshoot systematically, poor written communication, or discomfort supporting highly technical users. What You’ll Love Real technical problem solving with tangible customer impact A role that blends deep troubleshooting with scaling support via docs, tooling, and process High autonomy in a remote-first environment What May Be Challenging High context switching and managing multiple threads in parallel Repeated patterns that require discipline to convert pain into scalable improvements Supporting high-visibility systems where speed and accuracy matter Differentiation Industry: Remote-first, trust-based culture; global team; autonomy; modern systems; meaningful technical challenges Internal: High-impact, customer-facing observability support; direct influence on tooling and process maturity; opportunity to shape scalable support practices

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in Chesapeake, VA vacancy
  •  ...scale, we invite you to bring your talents to Zscaler to help shape the future of cybersecurity. Role We are looking for a Site Reliability Engineer-SkillBridge Intern (San JosA Ca or Bellevue WA) to join our Zero Trust Exchange team. This is a remote role based in San... 
    Suggested
    Internship
    Work at office
    Local area
    Remote work
    Worldwide

    GrabJobs

    Chesapeake, VA
    4 days ago
  • $68.6k - $85.6k

    A Brief Overview Reporting to the Senior Project Manager(s), the Project Engineer will work with various divisions to support Hunt Companies' construction department needs. The Project Engineer is responsible for assisting the Senior Project Manager(s) with planning... 
    Suggested
    Full time
    Contract work
    For contractors
    For subcontractor
    Local area

    Hunt Military Communities

    Norfolk, VA
    1 day ago
  •  ...networks for USVs. You will configure, troubleshoot, and maintain mission systems in a tactical environment. Ideal candidates have a engineering/technical degree and 7+ years’ experience with software-driven or unmanned platforms, plus strong problem-solving and radio/... 
    Suggested

    W R Systems

    Norfolk, VA
    1 day ago
  • $85.7k - $107.12k

     ...Company: MCA - Alpolic Job Description: The Mechanical / Reliability Engineer is responsible for improving the reliability, performance, and lifecycle of manufacturing equipment across ALPOLIC's production operations. This role focuses on root cause analysis... 
    Suggested
    Work experience placement
    Weekend work

    Mitsubishi Chemical

    Chesapeake, VA
    3 days ago
  • Centurum is seeking a highly experienced technician to install, repair, and operate Global Broadcast Service (GBS) across its product line. The role includes responding to ship and field issues (CASREPS) and providing technical leadership to other technicians. The candidate...
    Suggested

    Centurum

    Chesapeake, VA
    1 day ago
  • $69.3k - $158k

     ...solutions by using out-of-the-box features to create custom pages, site templates, features, and Web parts. Write custom code to create...  ...with Agile methodology, extreme programming, software engineering, product management, and software products Experience with Java... 
    Full time
    Contract work
    Part time
    Work at office
    Local area
    Remote work

    Booz Allen Hamilton

    Norfolk, VA
    4 days ago
  • Syms Strategic Group (SSG)  is seeking a talented Software Developer Location: Remote Department: Veterans Affairs Type:  Full Time Min. Experience:  Experienced Security Clearance Level:  Public Trust (MBI)  Military Veterans are highly encouraged...
    Full time
    Remote work

    Ssg

    Norfolk, VA
    1 day ago
  • $116.36k - $155.15k

     ...demonstrated knowledge and experience in system architecture and engineering disciplines. Specific technical knowledge of enterprise level...  ...Amazon Web Services. Supports due diligence activities including site surveys, design, design review, bill of materials creation,... 
    Full time
    Temporary work
    Remote work
    1 day per week

    Lumen

    Norfolk, VA
    4 days ago
  •  ...Join to apply for the Technical Solutions Engineer role at Epic Join to apply for the Technical Solutions Engineer role at Epic Get AI-powered advice on this job and more exclusive features. Please note that this position is based on our campus in Madison, WI, and requires... 
    Full time
    Work at office
    Relocation
    Visa sponsorship
    Relocation package

    Epic

    Chesapeake, VA
    5 days ago
  • $103.7k - $162.1k

    We invite you to explore a future with us at PRA Group, a diverse and growing company that has a tangible impact on the global economy. Position Summary Designs and monitors internal and external access controls, security safeguards and associated processes to protect ...
    Work experience placement

    PVH (Tommy Hilfiger/Calvin Klein)

    Norfolk, VA
    1 day ago
  • Moffatt & Nichol in Norfolk, Virginia is seeking a Systems Solutions & Adoption Lead. This role involves technical business enablement focusing on AI tools and system capabilities expansion. The ideal candidate must have in-depth knowledge of enterprise technology ecosystems...

    Moffatt & Nichol

    Norfolk, VA
    8 hours ago
  •  ...develop models of possible future configurations Create daily test metrics and reporting Occasionally perform other IT systems engineering activities such as requirements, design, installation, operation, sustainment, and support Apply technical principles,... 
    Full time
    Work experience placement
    Remote work

    Ssg

    Norfolk, VA
    1 day ago
  •  ...MASSACHUSETTS MARITIME ACADEMY is seeking a First Assistant Engineer (1AE) to support the safe and efficient operation of the vessel’s Engineering Department. You will assist the Chief Engineer in maintaining, operating, and repairing propulsion and auxiliary systems.... 

    Massachusetts Maritime Academy

    Norfolk, VA
    8 hours ago
  • About Keycard At Keycard , we’re building identity & access infrastructure for the agent-native era —where software isn’t static, but a dynamic, constantly changing system of AI agents working on behalf of people and businesses to complete dynamic tasks at runtime....
    Remote work
    Shift work

    GrabJobs

    Chesapeake, VA
    5 days ago
  •  ..., configure, upgrade, monitor, and troubleshoot web and application platforms to ensure security, performance, availability, and reliability. Document system configurations, changes, and procedures, and coordinate with vendors as needed to resolve system issues or outages... 
    Full time
    Part time
    Work from home

    Virginia Department of Human Resource Management

    Norfolk, VA
    2 days ago
  •  ...application issues and make recommendations to improve performance and reliability. They are also able to review requirements and estimate...  ...various business development efforts The senior Software Engineer is expected to train and oversee the work of less experienced... 
    Temporary work
    Remote work
    Flexible hours

    VSolvit

    Chesapeake, VA
    7 hours ago
  • $110k - $130k

    At StratasCorp, our mission strives to put employees first while still being recognized as a leader in the Department of Defense Information Technology sector. We believe in a continuing pursuit of customer satisfaction and operational excellence while exceling in service...
    Full time
    Immediate start

    STRATASCORP

    Norfolk, VA
    3 days ago
  • Our Benefits - Designed with You in Mind Comprehensive Health & Well-being Coverage From your very first day, you’ll have access to medical, dental, vision, and prescription drug coverage - ensuring you and your family stay healthy and protected. Generous Paid Time...
    Full time
    Immediate start

    Stellantis

    Chesapeake, VA
    6 hours ago
  • Old Dominion University in Norfolk, VA seeks a Research Systems Administrator to support Sponsored Programs Administration's computing and enterprise systems. The role involves maintaining IT infrastructure, assisting staff with system issues, and ensuring security and...

    Old Dominion University

    Norfolk, VA
    7 hours ago
  • $80k - $110k

     ...Lead Engineer - Electrical Components Veolia Group is a global leader in environmental services, operating across all five continents with nearly 218,000 employees. Specializing in water, energy, and waste management, Veolia Group designs and implements innovative... 
    Contract work
    Work experience placement
    Live in
    Flexible hours

    Sarpi Thinktech

    Norfolk, VA
    5 days ago
  • $50 - $60 per hour

     ...Software Engineer Location: Norfolk, VA (hybrid) Employment Type: Contract to Hire Employment Length: 6 month (conversion based on performance) Pay: $50.00 - $60.00/ hr. Description: We are seeking a highly skilled Software Engineer / Systems Engineer... 
    Contract work

    Apex Systems

    Norfolk, VA
    5 days ago
  • $86.8k - $198k

    Undersea Warfare Wargame Adjudicator The Opportunity: As an expert in defense missions, your unique background inspires you to think bigger, push further, and ask questions others don’t. We need your extensive industry knowledge and advisory skills to solve some of our...
    Full time
    Part time
    Local area

    Booz Allen Hamilton

    Norfolk, VA
    1 day ago
  •  ...Job Description The Project Systems Engineer leverages a strong foundation in systems engineering to guide the division's engineering...  ...the rigorous demands of the complex environments in which they operate, delivering results without compromising safety or reliability.

    Oceaneering

    Chesapeake, VA
    5 days ago
  •  ...The Project Systems Engineer leverages a strong foundation in systems engineering to guide the division's engineering and design activities under the supervision of the Design Engineering Manager. This role requires effective collaboration with Oceaneering personnel and... 

    Oceaneering

    Chesapeake, VA
    4 days ago
  •  ...Systems Engineer III SEACORP is seeking a well-qualified Systems Engineer III. Primary Duties and Responsibilities: A local Defense...  ...and training exercises for multiple platforms at various labs, sites, and shipboard locations. Stay current with advancements in... 
    Full time
    Temporary work
    For contractors
    Work at office
    Local area

    Seacorp Inc

    Norfolk, VA
    4 days ago
  • $99k - $111.5k

     ...General Summary Family Dollar is seeking a highly capable Systems Engineer, to design, deliver, and operate our hybrid infrastructure...  ...a forward-looking mindset focused on automation, scalability, reliability, and cost optimization. Principal Duties & Responsibilities Hybrid... 
    Full time
    Visa sponsorship

    Family Dollar

    Chesapeake, VA
    4 days ago
  •  ...experienced Project Manager specialized in naval navigation systems. The role involves managing a multidisciplinary team of developers and engineers, ensuring project delivery within timelines and budget constraints. Candidates must have a minimum of 7 years' experience in the... 

    Connect Talent Solutions

    Norfolk, VA
    7 hours ago
  • $69.1k - $141.5k

     ...Job Title: Systems Engineer Job Category: Information Technology Time Type: Full time Minimum Clearance Required to Start: Secret Employee Type: Regular Percentage of Travel Required: Up to 10% Type of Travel: Local The Opportunity We are seeking a skilled Software Engineer... 
    Full time
    Local area
    Worldwide

    CACI International

    Norfolk, VA
    1 day ago
  •  ...Serco seeks a Simulation Systems and Test Engineer for its Combat Air Force Distributed Mission Operations (CAF DMO) 3.0 program in Hampton, VA. The CAF DMO 3.0 program, via its Distributed Mission Operations Network (DMON), provides world‑class integration training in... 
    Contract work
    Local area
    Flexible hours

    Serco

    Norfolk, VA
    1 day ago
  • $50 - $60 per hour

     ...Job#: 3040305 Job Description: Software Engineer Location: Norfolk, VA (hybrid) Employment Type: Contract to Hire Employment Length: 6 month (conversion based on performance) Pay: $50.00 - $60.00/ hr. Description: We are seeking a highly skilled Software Engineer / Systems... 
    Contract work

    Apex Systems

    Norfolk, VA
    23 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!