Site Reliability Engineer
GrabJobs
SRE Support Engineer - Observability While this position is not currently open, we are interviewing strong candidates for upcoming opportunities on this team. Location: Remote | Time Zone: (US, Canada, Brazil, Chile, Colombia, Mexico)(8AM–5PM Pacific) Freedom to grow. Power to deliver. Virtasant is a global technology services company delivering large-scale cloud, data, and engineering solutions across 130+ countries. We partner with some of the world’s largest organizations to help them build, operate, and scale internal platforms used by tens of thousands of engineers. For this role, you will be supporting one of the most advanced internal developer platforms in the world, powering products used by hundreds of millions of people. The problems you will solve are deep, complex, and essential to keeping a global-scale organization moving. Role Overview The Observability & Tools Support Engineer provides high-impact technical support for customers of a large technology company’s internal IaaS platform, with a focus on monitoring, alerting, telemetry, and operational tooling . This role spans a wide range of support—from white-glove onboarding and end-to-end customer enablement, to deep technical troubleshooting across Linux, networking, and observability systems (especially Prometheus and AlertManager ). You will also contribute to improving the support function itself: strengthening tooling, documentation, workflows, and feedback loops so the service scales. Success depends on excellent troubleshooting, strong written communication, comfort working with highly technical customers, and the maturity to identify patterns and drive operational improvements beyond individual ticket resolution. Business Outcome Become a trusted frontline expert for the customer’s observability ecosystem and operational tooling - delivering fast, accurate support across Slack and tickets, improving monitoring reliability, and reducing incident impact through better triage, troubleshooting, onboarding, and knowledge capture. Success Measures Healthy volume of threads and tickets handled with high-quality outcomes Consistent achievement of time-based SLAs High customer satisfaction through surveys Accurate classification of issue type, severity, and recurring patterns Reduced repeat issues through better docs, tooling, and scalable onboarding What Will Be True When You Succeed Customers can onboard smoothly to monitoring/alerting with minimal friction Monitoring and alerting issues are resolved quickly, with fewer escalations Linux and networking-related incidents reach resolution faster due to strong troubleshooting and clean handoffs Engineering and SRE teams receive clear, actionable feedback based on real customer trends Knowledge base content prevents tickets and accelerates self-service Core Work Units 1) Frontline Support for Observability & Tooling Manage Slack threads and tickets (roughly 50/50) Handle a broad range of customer support: simple issue resolution through end-to-end onboarding Provide clear, structured guidance to highly technical customers Maintain strong attention to detail while managing multiple interactions in parallel 2) Deep-Dive Troubleshooting & Incident Support Troubleshoot, isolate, and resolve monitoring and alerting issues (especially Prometheus + AlertManager ) Troubleshoot complex Linux and networking issues (TCP/IP fundamentals required) Support OpenTelemetry, tracing, and telemetry pipelines , including investigation of gaps in signals and instrumentation Drive incidents to resolution in partnership with Engineering/SRE teams 3) Documentation & Knowledge Development Build and maintain customer-facing and internal knowledge base articles Create informational posts for the community support platform Turn repeated issues into reusable guides, checklists, and onboarding playbooks 4) Trend Analysis & Feedback to Engineering Analyze and categorize customer interaction trends Provide accurate, meaningful feedback to Engineering and SRE orgs to improve product/tooling Identify “top offenders” and propose practical fixes (tooling, docs, process, product) 5) Operational Excellence & Continuous Improvement Participate in post-mortem reviews and drive follow-through on improvements Contribute meaningfully to team objectives and goals (process, tooling, and service scaling) Bring creativity and discretion to resolve highly complex issues “outside the box” High-Quality Work - what top performance looks like Frontline Support Moves smoothly from triage to deeper analysis without losing the customer Communicates clearly and confidently with technical users Maintains clean follow-ups and thread hygiene even with high context switching Troubleshooting Rapidly isolates issues across monitoring/alerting configs, Linux runtime behavior, and network connectivity Uses structured approaches to incident handling: hypothesis → test → evidence → resolution Produces high-signal writeups that accelerate downstream resolution Documentation & Enablement Documentation is clear enough that customers avoid opening tickets Onboarding flows reduce time-to-value and prevent common misconfigurations Captures “tribal knowledge” quickly and makes it reusable Operational Excellence Obsessing over details: correct severity, accurate tagging, clean timelines, strong handoffs Spots patterns early and proactively proposes improvements that scale support Typical Day / Work Patterns ~50% Slack support, ~50% ticket handling Deep-dive investigations during lower ticket volume periods Documentation writing and lightweight tooling/process improvements when patterns emerge Weekly team review of escalations, themes, and operational improvements High rate of context switching and parallel issue management Required Skills & Experience (Non-Negotiable) Several years supporting highly scalable applications and web services Hands-on experience with open-source observability and cloud-native tooling, including: Kubernetes (and container fundamentals) Prometheus and AlertManager troubleshooting OpenTelemetry and distributed tracing concepts Strong understanding of the Linux operating system (command line, process/network debugging, logs) Good understanding of infrastructure observability principles (signals, alerting strategy, SLO thinking, noise reduction) Good understanding of the TCP/IP suite and practical networking troubleshooting Strong experience troubleshooting ambiguous, multi-layer issues Excellent analytical capability and strong attention to detail Strong written and verbal communication (clear, structured, customer-friendly) Comfortable working with a very technical customer base Passion for Technical Support and a service mindset Nice-to-Haves Experience improving or supporting internal support tooling or workflows (automation, templates, runbooks) Experience operating at scale in a services environment (pattern detection, KPI/SLA awareness, operational process maturity) Familiarity with Grafana, log aggregation, incident tooling, and production support practices Prior SRE or platform support experience Minimum Qualifications 3–7+ years in Technical Support Engineering, SRE support, DevOps, Platform Support, or similar Demonstrated experience supporting distributed systems, IaaS, or cloud platforms Strong Linux, troubleshooting, and customer-facing communication background Evidence of documentation, knowledge-base contributions, and process improvement mindset Disqualifiers: weak Linux fundamentals, inability to troubleshoot systematically, poor written communication, or discomfort supporting highly technical users. What You’ll Love Real technical problem solving with tangible customer impact A role that blends deep troubleshooting with scaling support via docs, tooling, and process High autonomy in a remote-first environment What May Be Challenging High context switching and managing multiple threads in parallel Repeated patterns that require discipline to convert pain into scalable improvements Supporting high-visibility systems where speed and accuracy matter Differentiation Industry: Remote-first, trust-based culture; global team; autonomy; modern systems; meaningful technical challenges Internal: High-impact, customer-facing observability support; direct influence on tooling and process maturity; opportunity to shape scalable support practices
$158.5k - $172k
...exceptional value they deserve.About The OpportunityAs a Senior Engineer on the Runtime Automation team, you will design, automate, and... .... This is a high-impact position driving continuous reliability, deep system optimization, and automation across our entire technology...SuggestedFull timeTemporary workWork at officeFlexible hours3 days per week$167.7k - $245.2k
...requiring approximately 2 days per week on-site at Cisco offices in either San Francisco... ...AI agents behave as intended, improving reliability and reducing risks. This unified... ...and control.As a Senior Site Reliability Engineer (SRE), you will build, operate, and continuously...SuggestedFull timeTemporary workLocal areaFlexible hours2 days per week$165k - $241.4k
...very effective.We’re looking for talented engineers with a software or operations background... ...development teams to ensure the reliability, performance and security of our infrastructure... ...insurance. Please see the Cisco careers site to discover more benefits and perks....SuggestedFull timeTemporary workWork at officeLocal areaFlexible hours1 day per week$141k - $216.6k
...—it means helping shape the future of emergency response and building a safer, more connected world.Position OverviewAs a Site Reliability Engineer, you'll own the reliability, observability, and operational excellence of our Unified Call (UC) platform—the mission-critical...SuggestedWork experience placementWork at office$123k - $165k
Job Summary:Department/Group OverviewOur engineering fleet is a horizontal set of teams... ...organization. Our specific team provides reliability engineering and operational support to backend... ...products and brands.We are seeking a Site Reliability Engineer who will contribute...Suggested$139k - $257.55k
The ChallengeThe Adobe Creative Community CCM organization is seeking an outstanding Senior Site Reliability Engineer (SRE) to support innovation through machine learning, autonomous AI workflows, and cloud-native infrastructure. Adobe Stock gives designers and businesses...Full timeTemporary workLocal areaRemote workWorldwide$138.1k - $198.2k
...more intuitive with technology that simply works. The SRE Engineering Enablement Team supports our CI Platforms, Developer Environments... ...Our customers are all engineers at Cisco. Your Impact As a Site Reliability Engineer, you will be at the epicenter of our engineering...Permanent employmentFull timeTemporary workWork experience placementLocal areaRemote workFlexible hours$120k - $200k
...PermContact: Kunal DaveContact Email: ****@*****.*** Reliability Engineer(SRE) ResponsibilitiesGlobal Architecture & Disaster Recovery... ...practices (e.g., Chaos Engineering, resilience testing, automated recovery)SkillsBilingual Mandarin Site Reliability Engineer(SRE)Overseas- ...and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Commercial & Investment Bank, Production management team, you will solve complex and broad...Shift work
- ...also has offices in New York, NY, Miami, FL, Zurich, Switzerland and Lisbon, Portugal. About the Role At Luma, our Site Reliability Engineer (SRE) team keeps our platform reliable, secure, and lightning fast. They own everything from AWS infrastructure and Kubernetes...
$115k - $160k
...with Barclays to connect them with exceptional professionals for this role. Embark on a transformative journey as a Senior Site Reliability Engineer - AVP - Credit Trade Floor. At Barclays, our vision is clear –to redefine the future of banking and help craft innovative...Hourly payWork at office$160k - $180k
...Socure is seeking a Site Reliability Engineer in New York to enhance our identity trust infrastructure. In this role, you will take full ownership of AWS and Kubernetes platforms, ensuring high reliability and operability. The ideal candidate will possess extensive experience...- ...Job Description:- Our client is seeking a Senior Site Reliability Engineer (SRE) with 10 15 years of experience to support front-office trading systems in a production environment. This role focuses on troubleshooting complex trading infrastructure, managing observability...
- ...Job Description Job Description Location: New York, NY, USA Exp: 8-12 Years Client: Amex Job Description: SRE Engineer (This is not a Devops role, strictly need an SRE Engineer, who has great analytical skills and is a good incident manager as well)...
- ...We are seeking an experienced Site Reliability Engineer (SRE) – Microsoft Hyper-V & Private Cloud to operate highly available private cloud and Virtual Desktop Infrastructure (VDI) platforms based on Microsoft Hyper-V. This role combines deep Hyper-V expertise with modern...Local area
$80k - $95k
...join our dynamic team supporting the company’s users, applications, and web-based product offerings. In this role, the Site Reliability Engineer (SRE) will play a key role in maintaining resources at peak efficiency to guarantee staff are able to perform their functions...Remote workVisa sponsorshipWork visa- ...Site Reliability Engineer I, Abhishek, would like to share a job opportunity as Site Reliability Engineer in Jacksonville, FL, Cary, NC or New York, NY (Onsite) location for a Fulltime position. In case, if you are not comfortable with this location, please share your...Full timeWork visa
$110k - $120k
...largest companies to small and mid-market firms, rely on SS&C for expertise, scale, and technology.Job DescriptionJob Title: Site Reliability Engineer (SRE) / L3 Support EngineerGetting to know us:As a leading financial services and healthcare technology company based on...Ongoing contractFull timeCasual workRemote workFlexible hours$200k - $250k
Hudson River Trading (HRT) is seeking a Senior Site Reliability Engineer focused on storage to join our growing Enterprise SRE team. This team is responsible for developing and maintaining productivity service infrastructure for the entire firm, both on-prem and in the...Work at officeLocal areaImmediate start- ...researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: At... ...new cloud infrastructure company, we seek to improve our reliability dramatically while scaling the size of our platform and customer...
- ...Karsun Solutions in the DMV area is seeking a Site Reliability Manager to ensure reliability, scalability, and performance of our systems. You will lead a team focusing on Application Reliability, DevSecOps, and Platform Lifecycle Management. The ideal candidate has 1...
$111k - $218k
...The Site Reliability Engineering team designs and builds the global infrastructure on which we deploy our services, focusing on the above mentioned flagship MongoDB Atlas platform. As our customers grow and globalize, our services must satisfy demands for low-latency...Full timeLocal areaWorldwideFlexible hours- ...A dynamic fintech company in New York is seeking a Product & Platform Monitoring Manager to ensure the reliability of its fintech products. The role focuses on end-to-end monitoring of customer journeys, incident management, and collaboration with various teams. Candidates...
- ...Karsun Solutions, LLC is seeking a Site Reliability Manager to lead a multi-disciplinary team responsible for reliability, security, and platform lifecycle across AWS-based services. The role emphasizes collaboration, observability, and continuous improvement in a client...
- ...Overview We are seeking an experienced Site Reliability Engineer (SRE) – Microsoft Hyper-V & Private Cloud to operate highly available private cloud and Virtual Desktop Infrastructure (VDI) platforms based on Microsoft Hyper-V. This role combines deep Hyper-V expertise...
- ...The Office of Technology and Innovation (OTI) in Brooklyn, NY, seeks a Principal Automation Engineer to provide Site Reliability engineering for automation services and lead DevOps initiatives. You will write/update infrastructure code, manage CI/CD pipelines, and be...Work at office
$500 per month
...accounts. Our global team is a diverse group of experienced engineers, traders, and brokerage professionals who are working to... ...significant impact, we encourage you to apply. Your Role: As a Site Reliability Engineer at Alpaca, you'll help keep our brokerage platform...Home office- ...Site Reliability Engineer (SRE) Job Title Site Reliability Engineer (SRE) Job Summary We are seeking a skilled Site Reliability Engineer (SRE) to build, automate, and maintain highly available, scalable, and reliable infrastructure and applications...Flexible hours
- ...About the job Senior Site Reliability Engineer About the Company Stellar is a decentralized, public blockchain that gives developers the tools to create experiences that are more like cash than crypto. The network is faster, cheaper, and far more energy-efficient...
- ...A financial technology company based in New York is seeking a Product & Platform Monitoring Manager to ensure the reliability and health of their fintech products. The role involves monitoring customer journeys, APIs, and incident management, requiring 5+ years of experience...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
- site reliability engineer remote New York, NY
- site reliability engineer sre New York, NY
- site reliability engineering manager New York, NY
- site reliability engineer New York, NY
- on-site clinical research associate (traveling/remote) New York, NY
- site merchandiser New York, NY
- website coordinator New York, NY
- junior website developer New York, NY
- site leader New York, NY
- historic site New York, NY


