Future Openings - SRE Support Engineer - Observability
Virtasant
Job Description
Job Description
SRE Support Engineer - Observability
While this position is not currently open, we are interviewing strong candidates for upcoming opportunities on this team.
Location: Remote | Time Zone: (US, Canada, Brazil, Chile, Colombia, Mexico) (8AM–5PM Pacific)
Freedom to grow. Power to deliver.
Virtasant is a global technology services company delivering large-scale cloud, data, and engineering solutions across 130+ countries. We partner with some of the world’s largest organizations to help them build, operate, and scale internal platforms used by tens of thousands of engineers.
For this role, you will be supporting one of the most advanced internal developer platforms in the world, powering products used by hundreds of millions of people. The problems you will solve are deep, complex, and essential to keeping a global-scale organization moving.
Role OverviewThe Observability & Tools Support Engineer provides high-impact technical support for customers of a large technology company’s internal IaaS platform, with a focus on monitoring, alerting, telemetry, and operational tooling .
This role spans a wide range of support—from white-glove onboarding and end-to-end customer enablement, to deep technical troubleshooting across Linux, networking, and observability systems (especially Prometheus and AlertManager ). You will also contribute to improving the support function itself: strengthening tooling, documentation, workflows, and feedback loops so the service scales.
Success depends on excellent troubleshooting, strong written communication, comfort working with highly technical customers, and the maturity to identify patterns and drive operational improvements beyond individual ticket resolution.
Business OutcomeBecome a trusted frontline expert for the customer’s observability ecosystem and operational tooling - delivering fast, accurate support across Slack and tickets, improving monitoring reliability, and reducing incident impact through better triage, troubleshooting, onboarding, and knowledge capture.
Success MeasuresHealthy volume of threads and tickets handled with high-quality outcomes
Consistent achievement of time-based SLAs
High customer satisfaction through surveys
Accurate classification of issue type, severity, and recurring patterns
Reduced repeat issues through better docs, tooling, and scalable onboarding
Customers can onboard smoothly to monitoring/alerting with minimal friction
Monitoring and alerting issues are resolved quickly, with fewer escalations
Linux and networking-related incidents reach resolution faster due to strong troubleshooting and clean handoffs
Engineering and SRE teams receive clear, actionable feedback based on real customer trends
Knowledge base content prevents tickets and accelerates self-service
1) Frontline Support for Observability & Tooling
Manage Slack threads and tickets (roughly 50/50)
Handle a broad range of customer support: simple issue resolution through end-to-end onboarding
Provide clear, structured guidance to highly technical customers
Maintain strong attention to detail while managing multiple interactions in parallel
2) Deep-Dive Troubleshooting & Incident Support
Troubleshoot, isolate, and resolve monitoring and alerting issues (especially Prometheus + AlertManager )
Troubleshoot complex Linux and networking issues (TCP/IP fundamentals required)
Support OpenTelemetry, tracing, and telemetry pipelines , including investigation of gaps in signals and instrumentation
Drive incidents to resolution in partnership with Engineering/SRE teams
3) Documentation & Knowledge Development
Build and maintain customer-facing and internal knowledge base articles
Create informational posts for the community support platform
Turn repeated issues into reusable guides, checklists, and onboarding playbooks
4) Trend Analysis & Feedback to Engineering
Analyze and categorize customer interaction trends
Provide accurate, meaningful feedback to Engineering and SRE orgs to improve product/tooling
Identify “top offenders” and propose practical fixes (tooling, docs, process, product)
5) Operational Excellence & Continuous Improvement
Participate in post-mortem reviews and drive follow-through on improvements
Contribute meaningfully to team objectives and goals (process, tooling, and service scaling)
Bring creativity and discretion to resolve highly complex issues “outside the box”
Frontline Support
Moves smoothly from triage to deeper analysis without losing the customer
Communicates clearly and confidently with technical users
Maintains clean follow-ups and thread hygiene even with high context switching
Troubleshooting
Rapidly isolates issues across monitoring/alerting configs, Linux runtime behavior, and network connectivity
Uses structured approaches to incident handling: hypothesis → test → evidence → resolution
Produces high-signal writeups that accelerate downstream resolution
Documentation & Enablement
Documentation is clear enough that customers avoid opening tickets
Onboarding flows reduce time-to-value and prevent common misconfigurations
Captures “tribal knowledge” quickly and makes it reusable
Operational Excellence
Obsessing over details: correct severity, accurate tagging, clean timelines, strong handoffs
Spots patterns early and proactively proposes improvements that scale support
Typical Day / Work Patterns
~50% Slack support, ~50% ticket handling
Deep-dive investigations during lower ticket volume periods
Documentation writing and lightweight tooling/process improvements when patterns emerge
Weekly team review of escalations, themes, and operational improvements
High rate of context switching and parallel issue management
Several years supporting highly scalable applications and web services
Hands-on experience with open-source observability and cloud-native tooling, including:
Kubernetes (and container fundamentals)
Prometheus and AlertManager troubleshooting
OpenTelemetry and distributed tracing concepts
Strong understanding of the Linux operating system (command line, process/network debugging, logs)
Good understanding of infrastructure observability principles (signals, alerting strategy, SLO thinking, noise reduction)
Good understanding of the TCP/IP suite and practical networking troubleshooting
Strong experience troubleshooting ambiguous, multi-layer issues
Excellent analytical capability and strong attention to detail
Strong written and verbal communication (clear, structured, customer-friendly)
Comfortable working with a very technical customer base
Passion for Technical Support and a service mindset
Experience improving or supporting internal support tooling or workflows (automation, templates, runbooks)
Experience operating at scale in a services environment (pattern detection, KPI/SLA awareness, operational process maturity)
Familiarity with Grafana, log aggregation, incident tooling, and production support practices
Prior SRE or platform support experience
3–7+ years in Technical Support Engineering, SRE support, DevOps, Platform Support, or similar
Demonstrated experience supporting distributed systems, IaaS, or cloud platforms
Strong Linux, troubleshooting, and customer-facing communication background
Evidence of documentation, knowledge-base contributions, and process improvement mindset
Disqualifiers: weak Linux fundamentals, inability to troubleshoot systematically, poor written communication, or discomfort supporting highly technical users.
What You’ll LoveReal technical problem solving with tangible customer impact
A role that blends deep troubleshooting with scaling support via docs, tooling, and process
High autonomy in a remote-first environment
High context switching and managing multiple threads in parallel
Repeated patterns that require discipline to convert pain into scalable improvements
Supporting high-visibility systems where speed and accuracy matter
Industry: Remote-first, trust-based culture; global team; autonomy; modern systems; meaningful technical challenges
Internal: High-impact, customer-facing observability support; direct influence on tooling and process maturity; opportunity to shape scalable support practices
- ...Description Job Description Build & Release Support Engineer – CI/CD While this position is not currently open, we are interviewing strong candidates for upcoming... ...Monitoring tools (Prometheus/Grafana) Prior SRE experience Minimum Qualifications ~2–5 years...SuggestedImmediate startRemote work
- ...products that shape the future of business and... ...As a Site Reliability Engineer, you will work in an agile... ...responsibilities include: • 24x7 Observability: Be part of a... .... • Maintenance and Support: Tasks related to... ...always staying curious, open to feedback and learning...SuggestedFull timeContract workPart timeFixed term contractInternshipWorldwideFlexible hoursShift work
- ...in Site Reliability Engineering (SRE) and/or Networking Reliabillity... ...professionals to support infrastructure... ...Exposure to monitoring and observability tools. Familiarity... ...architects of the future. Join us to help... ...always staying curious, open to feedback and learning...SuggestedFull timeContract workPart timeFixed term contractInternshipShift work
$106.9k - $176.5k
...EY, we’re all in to shape your future with confidence. We’ll help... ...We are seeking an AI Systems Engineer to own the delivery, model-serving, routing, and observability layer of EY’s AI-native platform... .... Your key responsibilities Supports DevOps and delivery for AI workloads...SuggestedSummer holidayFlexible hours- ...considered for visa sponsorship now or in the future Bachelor of Engineering or Computer Science required with... ...knowledge of USRP Hardware and Open Source UHD SW for SDR applications, a... ...that enables the success of the Global Support Organization and Top Tier accounts...SuggestedTemporary workWork experience placementFlexible hours
- ...re looking for Software Engineers to help build that... ...ship the first version, observe how it behaves in the real... ..., and quality bar that future engineers will inherit.... ...scale, made meaningful open-source contributions, or... ..., user research, support, data work, or operations...Flexible hours
$144.6k - $198.8k
...the Team Shaping the Future of Industrial Technology... ...architecture and engineering standards that shape how... ...improvement in DevSecOps, observability, CI/CD, and compliance... .... Contributions to open-source projects or... ...You will work with a supportive team that fosters a genuine...Ongoing contractFull timeContract workLocal areaRemote workWorldwide- ...driven solutions that support digital commerce and... ...Collaborate with DevOps and SRE partners on cloud... ...and AI-driven engineering or observability capabilities. Provide... ...sponsorship now or in the future. This includes direct... ...updates about GM, open roles, career...Full timeContract workWork experience placementLocal areaWork from homeRelocation package
$178.42k - $230.5k
...maintaining the tools and services engineers here at GM use every day to... ...delivering impact through observability frameworks and will evolve... ...and SLOs Own or contribute to Open Source projects Passion for... ...your ambitions. Learn how GM supports a rewarding career that rewards...Full timeWork experience placementLocal areaWork from homeRelocation packageFlexible hours$144.2k - $288.4k
...Software Development Engineer to join our Digital Caremark... ...or similar) • Observability tools (Datadog, Prometheus... ...Our people fuel our future. Our teams reflect the... ...package designed to support the physical, emotional... ...application window for this opening will close on: 10/30/2...Hourly payFull timeTemporary workLocal area$105.79k - $141.05k
...impact, and help shape the future of AI-ready... ...Lead Salesforce Software Engineer, you will be responsible... ...deployment, and ongoing support of Salesforce Field... ...processes. Build observability, exception handling, automated... .... All legitimate job openings will be posted on our...Permanent employmentTemporary workRemote work- ...provider of technology services and support. Integritek was formed with... ...: The Senior Systems Engineer leverages technical expertise... ...traveling. Flexibility - Is open to change, enjoys the challenge... ...when things change, can flex to future consequences and trends appropriately...Contract workWork at officeFlexible hours
- ...Description Are you an SRE ready to grow your... ...As a Site Reliability Engineer at Brivo, you will bridge... ...manual "toil". Observability: Apply best practices... ...(SLOs). Deployment Support: Support production readiness... ...enough to support the future of AI-driven security....Worldwide
- ...Job Description Sr. Software Engineer - Site Reliability About... ...product-led company shaping the future of e-commerce logistics. Position... ...practices, and automation to support and improve our complex cloud... ...in AWS Build and maintain observability, monitoring, and logging...Full timeWork at office
- ...Principal Site Reliability Engineer About ShipperHQ:... ...led company shaping the future of e-commerce... ...confidently. As a Principal SRE, you'll own the... ...deployment architecture, observability, and platform reliability... .... ~ Experience supporting high-traffic SaaS applications...Full timeWork at office
- ...AI and shape the future of digitalization.... ...Site Reliability Engineer at TeamViewer, you... ...cloud infrastructure supporting TeamViewer’s... ...experience.5+ years in SRE, DevOps, or software... ...monitoring and observability tools (e.g. Datadog... ...proud to have an open and embracing workplace...Temporary workCasual workWorldwide
- ...using modern software engineering practices.... ..., deployment, and support. Collaborate with... .... Knowledge of observability, monitoring, logging... ...reliability engineering (SRE) practices.... ...architects of the future. Join us to help build... ...staying curious, open to feedback and...Full timeContract workPart timeFixed term contractInternshipShift work
- ...Chronosphere, a Palo Alto Networks company, is the observability platform built for control in the modern, containerized world.... ...remediating issues faster. You will be responsible for supporting the engineering team that owns the distributed systems powering Chronosphere...Remote job
$152k - $241.5k
...services. Our work opens up new universes... ...looking for a Senior SRE to join our... ...experience building and supporting critical services.... ...auto-healing, E2E observability or data-driven... ...Ruby.Mentored other engineers and influenced technical... ...our current and future employees, we do...Full time$135.2k - $306.4k
...Infrastructure's state of the art observability platform, powering... ...platforms used by OCI engineering teams to operate and... ...platforms supporting metrics, logs, traces,... ...Concurrent Programming Open source technologies for... ...promise into a better future for all. Discover your...Temporary workFlexible hours$144k - $329.1k
AI Engineering Consultant - Senior Manager - Consulting - Open Location Join to apply for the AI Engineering Consultant - Senior... ...EY, we’re all in to shape your future with confidence. We'll help you... ...projects that have supported data science, business intelligence...Work experience placementSummer holidayFlexible hours- ...up? If so, being a Software Engineer III at Frost could be the job... ...As a Software Engineer III - Open Banking at Frost, you are our... ...writing, testing, implementing, supporting, and documenting solutions... ...health, your family, and your future and strive to have our benefits...Full time
- ...Sonar is driving the future of agent-centric software... ...member of one of our engineering teams, you'll be a key... ..., providing guidance, support, and mentorship to... ...Reliability Engineering (SRE), Cloud Operations, or... ...infrastructure components. ~ Observability & Resiliency: Proven...RelocationFlexible hours
- ...Celestica is looking for a dynamic software engineer who is passionate about working closely... ...customers feel valued, respected and supported. Special arrangements can be made for candidates... ...imagine, develop and deliver a better future with our customers. Celestica would...Contract workWork at officeRemote work
- ...beyond automation - they observe, reason, adapt, and... ...As a member of our AI engineering team, you'll play a critical... ...to shape the future of enterprise AI from the... ...infrastructure and tools to support multi-agent collaboration... ...agent frameworks (e.g.,Open claw, Calude AP,...Permanent employmentFlexible hours
- ...Are Synopsys is the leader in engineering solutions from silicon to... ...Provide direct technical support, troubleshooting, and workflow... ...leverage AI-enabled simulation, opening new possibilities for design... ...make a measurable impact on the future of electronic design. Rewards...
- EY is seeking an AI Systems Engineer to own the delivery, model-serving, and observability layer of EY’s AI-native platform. You will oversee CI/CD/CV pipelines, secure model execution, governance of AI assets, and cost attribution across cloud and on-prem environments...
$114.6k - $234.6k
...tooling; sets standards for observability, security, and compliance. Guides... ...guidance and coaching to engineers to drive improvements.... ...Collaboration - Software Products Support: Collaborates with... ...long-term solutions (e.g., future enhancements). Practices...Temporary workFlexible hoursShift work- ...power of AI and shape the future of digitalization. The... ...capabilities that support reliable AI-powered products... ...systems are scalable, observable, secure, and ready for... .... Collaborate with AI engineers, software engineers,... ..."All Hands" meetings Open door policy and...Temporary workCasual workFlexible hours
- EY seeks an AI Systems Engineer to own delivery, model-serving, routing, and observability across EY’s AI-native platform. You will operate CI/CD/CV pipelines, govern AI assets, and ensure cost-aware governance while enabling scalable, compliant AI workloads. The role blends...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Future Openings - SRE Support Engineer - Observability. Be the first to apply!
- site reliability engineer sre Austin, TX
- site reliability engineer Austin, TX
- site reliability engineer remote Austin, TX
- senior IT engineer Austin, TX
- IT software developer Austin, TX
- junior application support engineer Austin, TX
- senior support engineer Austin, TX
- IT network engineer Austin, TX
- implementation support engineer Austin, TX
- line support engineer Austin, TX



