Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer

GrabJobs

SRE Support Engineer - Observability While this position is not currently open, we are interviewing strong candidates for upcoming opportunities on this team. Location: Remote | Time Zone: (US, Canada, Brazil, Chile, Colombia, Mexico)(8AM–5PM Pacific) Freedom to grow. Power to deliver. Virtasant is a global technology services company delivering large-scale cloud, data, and engineering solutions across 130+ countries. We partner with some of the world’s largest organizations to help them build, operate, and scale internal platforms used by tens of thousands of engineers. For this role, you will be supporting one of the most advanced internal developer platforms in the world, powering products used by hundreds of millions of people. The problems you will solve are deep, complex, and essential to keeping a global-scale organization moving. Role Overview The Observability & Tools Support Engineer provides high-impact technical support for customers of a large technology company’s internal IaaS platform, with a focus on monitoring, alerting, telemetry, and operational tooling . This role spans a wide range of support—from white-glove onboarding and end-to-end customer enablement, to deep technical troubleshooting across Linux, networking, and observability systems (especially Prometheus and AlertManager ). You will also contribute to improving the support function itself: strengthening tooling, documentation, workflows, and feedback loops so the service scales. Success depends on excellent troubleshooting, strong written communication, comfort working with highly technical customers, and the maturity to identify patterns and drive operational improvements beyond individual ticket resolution. Business Outcome Become a trusted frontline expert for the customer’s observability ecosystem and operational tooling - delivering fast, accurate support across Slack and tickets, improving monitoring reliability, and reducing incident impact through better triage, troubleshooting, onboarding, and knowledge capture. Success Measures Healthy volume of threads and tickets handled with high-quality outcomes Consistent achievement of time-based SLAs High customer satisfaction through surveys Accurate classification of issue type, severity, and recurring patterns Reduced repeat issues through better docs, tooling, and scalable onboarding What Will Be True When You Succeed Customers can onboard smoothly to monitoring/alerting with minimal friction Monitoring and alerting issues are resolved quickly, with fewer escalations Linux and networking-related incidents reach resolution faster due to strong troubleshooting and clean handoffs Engineering and SRE teams receive clear, actionable feedback based on real customer trends Knowledge base content prevents tickets and accelerates self-service Core Work Units 1) Frontline Support for Observability & Tooling Manage Slack threads and tickets (roughly 50/50) Handle a broad range of customer support: simple issue resolution through end-to-end onboarding Provide clear, structured guidance to highly technical customers Maintain strong attention to detail while managing multiple interactions in parallel 2) Deep-Dive Troubleshooting & Incident Support Troubleshoot, isolate, and resolve monitoring and alerting issues (especially Prometheus + AlertManager ) Troubleshoot complex Linux and networking issues (TCP/IP fundamentals required) Support OpenTelemetry, tracing, and telemetry pipelines , including investigation of gaps in signals and instrumentation Drive incidents to resolution in partnership with Engineering/SRE teams 3) Documentation & Knowledge Development Build and maintain customer-facing and internal knowledge base articles Create informational posts for the community support platform Turn repeated issues into reusable guides, checklists, and onboarding playbooks 4) Trend Analysis & Feedback to Engineering Analyze and categorize customer interaction trends Provide accurate, meaningful feedback to Engineering and SRE orgs to improve product/tooling Identify “top offenders” and propose practical fixes (tooling, docs, process, product) 5) Operational Excellence & Continuous Improvement Participate in post-mortem reviews and drive follow-through on improvements Contribute meaningfully to team objectives and goals (process, tooling, and service scaling) Bring creativity and discretion to resolve highly complex issues “outside the box” High-Quality Work - what top performance looks like Frontline Support Moves smoothly from triage to deeper analysis without losing the customer Communicates clearly and confidently with technical users Maintains clean follow-ups and thread hygiene even with high context switching Troubleshooting Rapidly isolates issues across monitoring/alerting configs, Linux runtime behavior, and network connectivity Uses structured approaches to incident handling: hypothesis → test → evidence → resolution Produces high-signal writeups that accelerate downstream resolution Documentation & Enablement Documentation is clear enough that customers avoid opening tickets Onboarding flows reduce time-to-value and prevent common misconfigurations Captures “tribal knowledge” quickly and makes it reusable Operational Excellence Obsessing over details: correct severity, accurate tagging, clean timelines, strong handoffs Spots patterns early and proactively proposes improvements that scale support Typical Day / Work Patterns ~50% Slack support, ~50% ticket handling Deep-dive investigations during lower ticket volume periods Documentation writing and lightweight tooling/process improvements when patterns emerge Weekly team review of escalations, themes, and operational improvements High rate of context switching and parallel issue management Required Skills & Experience (Non-Negotiable) Several years supporting highly scalable applications and web services Hands-on experience with open-source observability and cloud-native tooling, including: Kubernetes (and container fundamentals) Prometheus and AlertManager troubleshooting OpenTelemetry and distributed tracing concepts Strong understanding of the Linux operating system (command line, process/network debugging, logs) Good understanding of infrastructure observability principles (signals, alerting strategy, SLO thinking, noise reduction) Good understanding of the TCP/IP suite and practical networking troubleshooting Strong experience troubleshooting ambiguous, multi-layer issues Excellent analytical capability and strong attention to detail Strong written and verbal communication (clear, structured, customer-friendly) Comfortable working with a very technical customer base Passion for Technical Support and a service mindset Nice-to-Haves Experience improving or supporting internal support tooling or workflows (automation, templates, runbooks) Experience operating at scale in a services environment (pattern detection, KPI/SLA awareness, operational process maturity) Familiarity with Grafana, log aggregation, incident tooling, and production support practices Prior SRE or platform support experience Minimum Qualifications 3–7+ years in Technical Support Engineering, SRE support, DevOps, Platform Support, or similar Demonstrated experience supporting distributed systems, IaaS, or cloud platforms Strong Linux, troubleshooting, and customer-facing communication background Evidence of documentation, knowledge-base contributions, and process improvement mindset Disqualifiers: weak Linux fundamentals, inability to troubleshoot systematically, poor written communication, or discomfort supporting highly technical users. What You’ll Love Real technical problem solving with tangible customer impact A role that blends deep troubleshooting with scaling support via docs, tooling, and process High autonomy in a remote-first environment What May Be Challenging High context switching and managing multiple threads in parallel Repeated patterns that require discipline to convert pain into scalable improvements Supporting high-visibility systems where speed and accuracy matter Differentiation Industry: Remote-first, trust-based culture; global team; autonomy; modern systems; meaningful technical challenges Internal: High-impact, customer-facing observability support; direct influence on tooling and process maturity; opportunity to shape scalable support practices

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in Boulder, CO vacancy
  •  ...Job Description Job Description Description: Amplifire is seeking a Site Reliability Engineer to improve the reliability, scalability, performance, and operational efficiency of our cloud-based platform. Working alongside DevOps engineers within the Platform Operations... 
    Suggested

    Amplifire

    Boulder, CO
    14 days ago
  • $156k - $193k

     ...keep a security clearance. Applicants that do not meet these requirements will not be considered. SciTec is seeking a Systems Engineer with experience in software-defined sensor data processing pipelines and supporting computaional systems to support the Ground‑Based... 
    Suggested
    Temporary work
    For contractors
    Work experience placement
    Remote work
    Flexible hours

    SciTec

    Boulder, CO
    28 days ago
  • $152k - $241.5k

     ...how you can make a lasting impact on the world.We are now actively looking for energetic, enthusiastic and technologically savvy engineers to join the Chips System Software team. Here we architect and develop the boot firmware for the flagship CPUs which are the core components... 
    Suggested
    Full time

    Nvidia

    Boulder, CO
    4 days ago
  • $124k - $195.5k

     ...looking for great people like you to help us accelerate the next wave of artificial intelligence.We are now hiring a System Software Engineer for the UEFI Firmware team! You will be heavily involved with the early modeling and simulation required to produce our world-... 
    Suggested
    Full time
    Immediate start

    Nvidia

    Boulder, CO
    1 day ago
  • $110k

     ...joining us at PickNik as a Robotics Software Engineer. PickNik Robotics is the company behind...  .... Requirements Driven to ship reliable software used in production settings to solve...  ...every other month to customer sites & conferences. Less than 20% of the time.... 
    Suggested
    Full time
    Work experience placement
    Live in
    Work at office
    3 days per week

    Picknik Inc.

    Boulder, CO
    1 day ago
  • $140k - $200k

     ...– Speechify has no office. These include frontend and backend engineers, AI research scientists, and others from Amazon, Microsoft, and...  ...→ testing → release → maintenance. Ensure quality, reliability, and consistency across releases. Identify, diagnose, and resolve... 
    Full time
    Work at office

    Speechify

    Boulder, CO
    1 day ago
  • $103.1k - $171.8k

     ...next generation missile warning processing system for the US Space Force. The MDPAP Visualization team leverages web based, Unreal Engine, and backend data ingest and manipulation technologies to seamlessly present in-depth real-time sensor and event data visualization... 
    Full time
    Contract work
    Work experience placement
    Interim role
    H1b
    Remote work

    Smx

    Boulder, CO
    1 day ago
  • $140k - $170k

     ...the Role: We are seeking a highly skilled Senior Software Engineer with deep expertise in FreeBSD and low-level systems programming...  ...systems. This role is critical to advancing our platform’s reliability, performance, and hardware compatibility.     Key Responsibilities... 
    Full time
    Local area
    Flexible hours

    Spectra Logic

    Boulder, CO
    1 day ago
  •  ...in working at NetApp? Search our open jobs -   Job Description NetApp is looking for an experienced Software Sustaining Engineer. Critical technical skills and experience would include : * Linux kernel level crash dump analysis * Linux driver troubleshooting... 
    Full time
    Local area

    Netapp

    Boulder, CO
    1 day ago
  • $165k - $218k

    ABOUT THE TEAM The Anduril Imaging team develops state-of-the-art imaging systems, deployed to tackle the most significant security challenges of America and its allies. We are interested in candidates that are excited to build Anduril's  next generation of imaging...
    Full time
    Work experience placement
    Immediate start

    Anduril Industries

    Boulder, CO
    1 day ago
  • $106k - $145k

     ...while never losing sight of the critical importance of systems engineering process and attention to detail. We take bold, calculated...  ...be assigned at any time.   Design, develop, and deploy reliable, maintainable, scalable, and fault-tolerant backend services... 
    Full time
    Temporary work
    Work experience placement
    Flexible hours

    Infleqtion

    Boulder, CO
    1 day ago
  •  ...heavily in a new product that we believe will help hundreds of thousands of independent websites stay independent. As a Software Engineer within Sovrn Labs you will be tasked with architecting, coding and researching a solution to a serious publisher problem. We have... 
    Full time

    Sovrn

    Boulder, CO
    1 day ago
  • $160k - $185k

     ...Job Description Job Description Senior Platform Engineer – US Platform Team NOTE:  We are unable to sponsor or take over sponsorship...  ...of manual processes and improving deployment speed and reliability across the US engineering organization Act as the technical... 
    Full time
    Work at office
    Worldwide
    Work visa

    SumUp

    Boulder, CO
    10 days ago
  • Job SummaryGeneral Atomics Aeronautical Systems, Inc. (GA-ASI), an affiliate of General Atomics, is a world leader in proven, reliable remotely piloted aircraft and tactical reconnaissance radars, as well as advanced high-resolution surveillance systems.DUTIES AND RESPONSIBILITIES... 
    Part time
    Remote work
    Relocation package

    General Atomics

    Boulder, CO
    2 days ago
  •  ...nation’s most critical problems?   Do you want to work alongside engineers and scientists that are experts in their fields? Are you...  ...registered trademark of The MITRE Corporation. Material on this site may be copied and distributed with permission only. Job SummaryJob... 
    For contractors
    Work experience placement
    Internship
    Local area
    Immediate start

    Mitre

    Boulder, CO
    2 days ago
  • $174k - $252k

     ...and algorithms.1 year of experience in a technical leadership role.Experience developing accessible technologies.Google's software engineers develop the next-generation technologies that change how billions of users connect, explore, and interact with information and one... 

    Google

    Boulder, CO
    1 day ago
  • $130.4k - $195.6k

     ...research to solve real user challenges.As a Software Development Engineer in IDX Search, you will:Write and maintain robust, efficient,...  ...through websites that are not Workday Careers. Please be aware of sites that may ask for you to input your data in connection with a... 
    Full time
    Work at office
    Remote work
    Home office
    Flexible hours

    Workday

    Boulder, CO
    2 hours ago
  • $147k - $210k

     ...development.2 years of experience with data structures and algorithms.Experience developing accessible technologies.Google's software engineers develop the next-generation technologies that change how billions of users connect, explore, and interact with information and one... 

    Google

    Boulder, CO
    1 day ago
  • Major League Baseball is looking for a Software Engineer to join a cross-platform engineering group that ships our Mobile MLB App and MLB...  ...focus based on business priorities. Success means shipping reliable, high-quality user experiences while keeping platform parity with... 
    Temporary work
    Work experience placement
    Work at office
    Remote work

    Major League Baseball

    Boulder, CO
    2 hours ago
  • $143.7k - $194.4k

     ...memberships daily and orchestrate customer ML workloads with strict reliability, privacy, and latency requirements.The AMC Custom Models and...  ...drive measurable advertising outcomes. We're looking for an engineer who writes production-quality code, cares about operational... 
    Internship
    Flexible hours

    Amazon

    Boulder, CO
    2 days ago
  • $130.4k - $195.6k

     ...office location.About the RoleAs a Software Engineer on Workday Everywhere, you'll help build...  ..., and you'll help keep our systems reliable, performant, and scalable as we grow.We foster...  ...not Workday Careers. Please be aware of sites that may ask for you to input your data in... 
    Full time
    Work experience placement
    Work at office
    Remote work
    Home office
    Flexible hours

    Workday

    Boulder, CO
    1 day ago
  • $139.3k - $203.6k

     ...anywhere in the USA.Meet the TeamOur software engineering team develops software using Cisco’s...  ...the quality, scalability, security, and reliability of our software.ResponsibilitiesImprove...  ...insurance. Please see the Cisco careers site to discover more benefits and perks.... 
    Full time
    Temporary work
    Local area
    Remote work
    Flexible hours

    CISCO Systems

    Boulder, CO
    2 hours ago
  • $128.6k - $184.9k

     ...Processor Team serves as the foundational engine of the Splunk Integrated Data Platform,...  ...pipelines with minimal latency and maximum reliability. Optimize existing ingestion pipelines to...  ...insurance. Please see the Cisco careers site to discover more benefits and perks. Employees... 
    Full time
    Temporary work
    Local area
    Remote work
    Flexible hours

    CISCO Systems

    Boulder, CO
    3 days ago
  • $176k - $264k

     ...millions of users of Workday’s products.About the RoleEvisort Engineering is growing! As a Fullstack Software Engineer, you’ll join a collaborative...  ...websites that are not Workday Careers. Please be aware of sites that may ask for you to input your data in connection with a... 
    Full time
    Contract work
    Work at office
    Remote work
    Home office
    Flexible hours

    Workday

    Boulder, CO
    4 days ago
  • $195k - $220k

     ...computational breakthroughs. Join a world-class team of scientists, engineers, and business professionals to advance the state-of-the-art in...  ...and cloud platforms, driving key trade-off decisions across reliability, security, scalability, and cost. This position requires solid... 
    Temporary work
    Work at office
    3 days per week

    Atom Computing

    Boulder, CO
    6 days ago
  • $85k - $115k

    Platform Engineer (Azure), Boulder, CO (Hybrid or Remote)The Platform Engineer is responsible...  ..., and client-facing teams to deliver reliable hosted environments, support the ongoing...  ...networking, including VNets and subnets, site-to-site and point-to-site VPN, private endpoints... 
    Work at office
    Remote work
    Flexible hours

    Ascend Analytics

    Boulder, CO
    4 days ago
  • $223.1k - $301.1k

     ...builds the internal systems that enable more than 3,000 Splunk engineers to build, test, and release world-class products at scale. Our...  ...coverage, and basic life insurance. Please see the Cisco careers site to discover more benefits and perks. Employees may be eligible... 
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    Boulder, CO
    2 hours ago
  • $200k - $240k

     ...processing hundreds of billions of ad requests daily across a global, high-throughput exchange. We're looking for a Principal Software Engineer with deep roots in adtech infrastructure and a genuine conviction about what AI-native engineering looks like in practice. This... 
    Full time
    Temporary work
    Work at office
    Remote work
    Flexible hours

    Sovrn

    Boulder, CO
    1 day ago
  •  ...services. Where needed, you'll help align systems with NIST 800-171/CMMC requirements, collaborating closely with the Principal Security Engineer, AWS infra team, dev tooling team, chief software engineer, and cybersecurity/GRC group. You'll work in a lean, impact-focused... 
    Full time
    Work at office
    Shift work
    3 days per week

    Spire

    Boulder, CO
    1 day ago
  • $90k

     ...techniques from universities and moving them to the production floor. We need talented, motivated and forward-thinking mechanical engineers to help make it happen. Responsibilities ~ Develop and apply thermal-hydraulic models for new production processes Create... 
    Permanent employment
    Full time

    Synthio Chemicals Inc

    Boulder, CO
    8 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!