Site Reliability Engineer
GrabJobs
SRE Support Engineer - Observability While this position is not currently open, we are interviewing strong candidates for upcoming opportunities on this team. Location: Remote | Time Zone: (US, Canada, Brazil, Chile, Colombia, Mexico)(8AM–5PM Pacific) Freedom to grow. Power to deliver. Virtasant is a global technology services company delivering large-scale cloud, data, and engineering solutions across 130+ countries. We partner with some of the world’s largest organizations to help them build, operate, and scale internal platforms used by tens of thousands of engineers. For this role, you will be supporting one of the most advanced internal developer platforms in the world, powering products used by hundreds of millions of people. The problems you will solve are deep, complex, and essential to keeping a global-scale organization moving. Role Overview The Observability & Tools Support Engineer provides high-impact technical support for customers of a large technology company’s internal IaaS platform, with a focus on monitoring, alerting, telemetry, and operational tooling . This role spans a wide range of support—from white-glove onboarding and end-to-end customer enablement, to deep technical troubleshooting across Linux, networking, and observability systems (especially Prometheus and AlertManager ). You will also contribute to improving the support function itself: strengthening tooling, documentation, workflows, and feedback loops so the service scales. Success depends on excellent troubleshooting, strong written communication, comfort working with highly technical customers, and the maturity to identify patterns and drive operational improvements beyond individual ticket resolution. Business Outcome Become a trusted frontline expert for the customer’s observability ecosystem and operational tooling - delivering fast, accurate support across Slack and tickets, improving monitoring reliability, and reducing incident impact through better triage, troubleshooting, onboarding, and knowledge capture. Success Measures Healthy volume of threads and tickets handled with high-quality outcomes Consistent achievement of time-based SLAs High customer satisfaction through surveys Accurate classification of issue type, severity, and recurring patterns Reduced repeat issues through better docs, tooling, and scalable onboarding What Will Be True When You Succeed Customers can onboard smoothly to monitoring/alerting with minimal friction Monitoring and alerting issues are resolved quickly, with fewer escalations Linux and networking-related incidents reach resolution faster due to strong troubleshooting and clean handoffs Engineering and SRE teams receive clear, actionable feedback based on real customer trends Knowledge base content prevents tickets and accelerates self-service Core Work Units 1) Frontline Support for Observability & Tooling Manage Slack threads and tickets (roughly 50/50) Handle a broad range of customer support: simple issue resolution through end-to-end onboarding Provide clear, structured guidance to highly technical customers Maintain strong attention to detail while managing multiple interactions in parallel 2) Deep-Dive Troubleshooting & Incident Support Troubleshoot, isolate, and resolve monitoring and alerting issues (especially Prometheus + AlertManager ) Troubleshoot complex Linux and networking issues (TCP/IP fundamentals required) Support OpenTelemetry, tracing, and telemetry pipelines , including investigation of gaps in signals and instrumentation Drive incidents to resolution in partnership with Engineering/SRE teams 3) Documentation & Knowledge Development Build and maintain customer-facing and internal knowledge base articles Create informational posts for the community support platform Turn repeated issues into reusable guides, checklists, and onboarding playbooks 4) Trend Analysis & Feedback to Engineering Analyze and categorize customer interaction trends Provide accurate, meaningful feedback to Engineering and SRE orgs to improve product/tooling Identify “top offenders” and propose practical fixes (tooling, docs, process, product) 5) Operational Excellence & Continuous Improvement Participate in post-mortem reviews and drive follow-through on improvements Contribute meaningfully to team objectives and goals (process, tooling, and service scaling) Bring creativity and discretion to resolve highly complex issues “outside the box” High-Quality Work - what top performance looks like Frontline Support Moves smoothly from triage to deeper analysis without losing the customer Communicates clearly and confidently with technical users Maintains clean follow-ups and thread hygiene even with high context switching Troubleshooting Rapidly isolates issues across monitoring/alerting configs, Linux runtime behavior, and network connectivity Uses structured approaches to incident handling: hypothesis → test → evidence → resolution Produces high-signal writeups that accelerate downstream resolution Documentation & Enablement Documentation is clear enough that customers avoid opening tickets Onboarding flows reduce time-to-value and prevent common misconfigurations Captures “tribal knowledge” quickly and makes it reusable Operational Excellence Obsessing over details: correct severity, accurate tagging, clean timelines, strong handoffs Spots patterns early and proactively proposes improvements that scale support Typical Day / Work Patterns ~50% Slack support, ~50% ticket handling Deep-dive investigations during lower ticket volume periods Documentation writing and lightweight tooling/process improvements when patterns emerge Weekly team review of escalations, themes, and operational improvements High rate of context switching and parallel issue management Required Skills & Experience (Non-Negotiable) Several years supporting highly scalable applications and web services Hands-on experience with open-source observability and cloud-native tooling, including: Kubernetes (and container fundamentals) Prometheus and AlertManager troubleshooting OpenTelemetry and distributed tracing concepts Strong understanding of the Linux operating system (command line, process/network debugging, logs) Good understanding of infrastructure observability principles (signals, alerting strategy, SLO thinking, noise reduction) Good understanding of the TCP/IP suite and practical networking troubleshooting Strong experience troubleshooting ambiguous, multi-layer issues Excellent analytical capability and strong attention to detail Strong written and verbal communication (clear, structured, customer-friendly) Comfortable working with a very technical customer base Passion for Technical Support and a service mindset Nice-to-Haves Experience improving or supporting internal support tooling or workflows (automation, templates, runbooks) Experience operating at scale in a services environment (pattern detection, KPI/SLA awareness, operational process maturity) Familiarity with Grafana, log aggregation, incident tooling, and production support practices Prior SRE or platform support experience Minimum Qualifications 3–7+ years in Technical Support Engineering, SRE support, DevOps, Platform Support, or similar Demonstrated experience supporting distributed systems, IaaS, or cloud platforms Strong Linux, troubleshooting, and customer-facing communication background Evidence of documentation, knowledge-base contributions, and process improvement mindset Disqualifiers: weak Linux fundamentals, inability to troubleshoot systematically, poor written communication, or discomfort supporting highly technical users. What You’ll Love Real technical problem solving with tangible customer impact A role that blends deep troubleshooting with scaling support via docs, tooling, and process High autonomy in a remote-first environment What May Be Challenging High context switching and managing multiple threads in parallel Repeated patterns that require discipline to convert pain into scalable improvements Supporting high-visibility systems where speed and accuracy matter Differentiation Industry: Remote-first, trust-based culture; global team; autonomy; modern systems; meaningful technical challenges Internal: High-impact, customer-facing observability support; direct influence on tooling and process maturity; opportunity to shape scalable support practices
$113.3k - $205.52k
...important to maintain our strong culture, achieve our goals, and thrive as #OneJamf. What you'll do at Jamf: As a Senior Site Reliability Engineer, you'll help us balance development velocity with the reliability our customers depend on. You'll partner with engineering...SuggestedWork at officeRemote workWorldwideFlexible hours- ...role and responsibilities IBM is seeking a motivated and detail-oriented IT Administrator intern with an interest in Site Reliability Engineering (SRE) and/or Networking Reliabillity Engineering (NRE) to join our team. These roles offer hands-on experience in enterprise...SuggestedFull timeContract workPart timeFixed term contractInternshipShift work
$132.23k - $176.31k
...future of AI‑ready connectivity, join us today. The Role We are seeking a highly skilled and proactive Senior Lead Site Reliability Engineer (SRE) to join our team, focusing on production support and performance optimization across our portal ecosystem. This role...SuggestedFull timeTemporary workRemote work- ...define the new space era by continuously pushing the boundaries of engineering models services and technology development. Visit us at... ...learner who is always ready to gain depth of knowledge ~ A reliable worker who knows the importance of showing up when it counts...SuggestedFull timeFor subcontractor
- ...of a Mayo Clinic campus for occasional on-site expectations based on business needs. The... ...Services Unit is seeking a Software Engineer to develop and support Human Capital Management... ...develop, implement, and support secure, reliable, and scalable integrations between Oracle...SuggestedPermanent employmentInternshipFlexible hoursShift work
$110.61k
Distributed Systems Software Engineer, Python / Go Join to apply for the Distributed Systems Software Engineer, Python / Go role at Canonical... ...testing approaches and infrastructure for validating reliability, performance, and resilience of cloud orchestration tools and...Full timeLocal areaRemote workWorldwide- ...Epic in Rochester, MN, is seeking a Technical Solutions Engineer to work on software that impacts millions of patients globally. The role involves diagnosing problems, identifying solutions, and managing implementations across various locations. Qualified candidates should...RelocationRelocation package
- This is an on-site position located in Rochester, MN. The Senior Media Systems Engineer is accountable for project results and goals set by Unit Head and division level... ...production staff, ensuring the delivery of reliable, scalable, and cutting-edge event technology solutions...Permanent employmentFlexible hours
$105k - $140k
...Today, nearly 200 people around the globe work on Speechify in a 100% distributed setting. These include frontend and backend engineers, AI research scientists, and others from Amazon, Microsoft, and Google, leading PhD programs like Stanford, high growth startups like...- ...teams to deliver scalable, secure, and reliable solutions that help clients modernize and... ...software solutions using modern software engineering practices. • Participate in the full... ...observability, monitoring, logging, and site reliability engineering (SRE) practices....
$75k - $125k
Applications Engineer & Technical Account Manager - Industrial Water Treatment$75,000 - $... ...food/beverage facilities, and industrial sites managing high/low-pressure water systems... ...career history showing commitment, reliability, and growth.Local Proximity: Residing within...Local area$111.07k - $148.1k
...demonstrated knowledge and experience in system architecture and engineering disciplines. •Recommends optimized solutions to support current... ...capabilities. •Supports due diligence activities including site surveys, design, design review, bill of materials creation, statement...Full timeTemporary workRemote work1 day per week- ...developers, data analysts/data scientists, and machine learning engineers.Who Should ApplyRecent computer science/engineering/mathematics... ...work with hands-on experience building projects at the client site are the only way a candidate can be picked by clients. If not having...H1bRemote work
- ...relocation to the area. We recruit nationally and provide financial relocation assistance. Responsibilities As a Technical Solutions Engineer at Epic, you’ll work on software that impacts 305 million patients around the world. Together with customer counterparts, you’ll...RelocationVisa sponsorshipRelocation package
- ...thinking.Define and champion comprehensive standards and best practices for CI/CD, observability, infrastructure as code, and reliability engineering across the Platform Technology team, as well as our vendor partners.Work closely with developers, Product Managers, Infosec...Flexible hours
- Our Benefits - Designed with You in Mind Comprehensive Health & Well-being Coverage From your very first day, you’ll have access to medical, dental, vision, and prescription drug coverage - ensuring you and your family stay healthy and protected. Generous Paid Time...Full timeImmediate start
- ...Construction Engineer Job Locations US-MN-Rochester Tracking Code 2026-13228 Overview TheConstructionEngineer... ...customer goalsand needs.This role will sit primarily on the job site, with occasional travel to offices for training, meetings,etc....Contract workFor contractors
$18 - $25 per hour
...October 2026. If you’re interested in Solutions, Consulting and Engineering this is the place for you! Why Solutions Consulting &... ...skills Ability to collaborate well with others Must provide reliable transportation & housing Certain states and localities...Remote jobHourly paySummer workInternshipWork at officeImmediate startShift work- ...Introduction IBM Hardware and Systems engineers design and deliver the compute platforms that power the world's most advanced AI, cloud... ...-up. Conduct experiments, analyze data, and contribute to reliability, process, and materials development projects. Write scripts...Full timeContract workPart timeFixed term contractInternshipShift work
$41.6k - $49.92k
...Braun Intertec is seeking students pursuing degrees in engineering, construction management, or related field; and other interested candidates... .... Braun Intertec strives to ensure that its careers web site is accessible to all. If you need assistance completing your online...Hourly payFull timeWork experience placementWork at officeWeekend work- ...Construction Engineer I The Construction Engineer I supports the project team by processing, maintaining, and updating logs, instructions... ...goals and needs. This role will sit primarily on the job site, with occasional travel to offices for training, meetings, etc....Contract workFor contractors
- ...teams to deliver scalable, secure, and reliable solutions that help clients modernize and... ...software solutions using modern software engineering practices. Participate in the full... ...observability, monitoring, logging, and site reliability engineering (SRE) practices....Full timeContract workPart timeFixed term contractInternshipShift work
- ...background aligns with future opportunities, we’ll reach out directly when formal applications become available. About Software Engineering Roles at Danaher Are you passionate about building real-world applications, writing clean code, and solving meaningful...Remote jobInternship
$40 per hour
...The Software Engineer Intern implements developer tools or product features on a rapid-release cycle. You will work in an agile development... ...Travel Requirements & Working Conditions Minimal travel Reliable internet access for any period of time working remotely and not...Remote jobSummer workInternshipSummer internshipWork at office- ...markets. Our platform, Ubuntu, is very widely used in breakthrough enterprise initiatives such as public cloud, data science, AI, engineering innovation, and IoT. Our customers include the world's leading public cloud and silicon providers, and industry leaders in many...Full timeWork at officeRemote workWork from home
- ...This is an excellent opportunity for Interns to contribute to enterprise-scale software development projects, learn from experienced engineers, and grow within a global technology leader. You'll help build and maintain backend systems that power IBM's internal and external...Full timeContract workPart timeFixed term contractInternshipShift work
$75 - $85 per hour
About The Role The FormativGroup is adding the role of the Salesforce Developer is a short-term, project-based opportunity focused on supporting Salesforce development, integrations, production support, and platform maintenance across core Salesforce technologies. This...Temporary workWork at officeRemote workWork visa- Net Developer A Few Words About Us Integrated Resources, Inc is a premier staffing firm recognized as one of the tri-states most well-respected professional specialty firms. IRI has built its reputation on excellent service and integrity since its inception in 1996....Long term contractContract work
$18 - $50 per hour
...0 GPA Perks: Employee discounts at our top customer sites Networking with our global leaders Mentorship from senior... ...the program Position Overview: As an Electronics Reliability Product Engineering Intern , you will support the development of Simcenter...Remote jobHourly payFull timeInternshipWork at officeLocal area- AI/ML Engineers at Mayo Clinic play a pivotal role in the union of data, systems, and computer sciences. They work closely with a multidisciplinary team, including clinicians, user experience designers, product managers, IT professionals, and external partners, to develop...Flexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
- on-site clinical research associate (traveling/remote) Rochester, MN
- site safety Rochester, MN
- junior website developer Rochester, MN
- construction site safety Rochester, MN
- IT site lead Rochester, MN
- historic site Rochester, MN
- site services specialist Rochester, MN
- official site Rochester, MN
- site leader Rochester, MN
- site reliability engineer



