Future Openings - SRE Support Engineer - Observability
Virtasant
Job Description
Job Description
SRE Support Engineer - Observability
While this position is not currently open, we are interviewing strong candidates for upcoming opportunities on this team.
Location: Remote | Time Zone: (US, Canada, Brazil, Chile, Colombia, Mexico) (8AM–5PM Pacific)
Freedom to grow. Power to deliver.
Virtasant is a global technology services company delivering large-scale cloud, data, and engineering solutions across 130+ countries. We partner with some of the world’s largest organizations to help them build, operate, and scale internal platforms used by tens of thousands of engineers.
For this role, you will be supporting one of the most advanced internal developer platforms in the world, powering products used by hundreds of millions of people. The problems you will solve are deep, complex, and essential to keeping a global-scale organization moving.
Role OverviewThe Observability & Tools Support Engineer provides high-impact technical support for customers of a large technology company’s internal IaaS platform, with a focus on monitoring, alerting, telemetry, and operational tooling .
This role spans a wide range of support—from white-glove onboarding and end-to-end customer enablement, to deep technical troubleshooting across Linux, networking, and observability systems (especially Prometheus and AlertManager ). You will also contribute to improving the support function itself: strengthening tooling, documentation, workflows, and feedback loops so the service scales.
Success depends on excellent troubleshooting, strong written communication, comfort working with highly technical customers, and the maturity to identify patterns and drive operational improvements beyond individual ticket resolution.
Business OutcomeBecome a trusted frontline expert for the customer’s observability ecosystem and operational tooling - delivering fast, accurate support across Slack and tickets, improving monitoring reliability, and reducing incident impact through better triage, troubleshooting, onboarding, and knowledge capture.
Success MeasuresHealthy volume of threads and tickets handled with high-quality outcomes
Consistent achievement of time-based SLAs
High customer satisfaction through surveys
Accurate classification of issue type, severity, and recurring patterns
Reduced repeat issues through better docs, tooling, and scalable onboarding
Customers can onboard smoothly to monitoring/alerting with minimal friction
Monitoring and alerting issues are resolved quickly, with fewer escalations
Linux and networking-related incidents reach resolution faster due to strong troubleshooting and clean handoffs
Engineering and SRE teams receive clear, actionable feedback based on real customer trends
Knowledge base content prevents tickets and accelerates self-service
1) Frontline Support for Observability & Tooling
Manage Slack threads and tickets (roughly 50/50)
Handle a broad range of customer support: simple issue resolution through end-to-end onboarding
Provide clear, structured guidance to highly technical customers
Maintain strong attention to detail while managing multiple interactions in parallel
2) Deep-Dive Troubleshooting & Incident Support
Troubleshoot, isolate, and resolve monitoring and alerting issues (especially Prometheus + AlertManager )
Troubleshoot complex Linux and networking issues (TCP/IP fundamentals required)
Support OpenTelemetry, tracing, and telemetry pipelines , including investigation of gaps in signals and instrumentation
Drive incidents to resolution in partnership with Engineering/SRE teams
3) Documentation & Knowledge Development
Build and maintain customer-facing and internal knowledge base articles
Create informational posts for the community support platform
Turn repeated issues into reusable guides, checklists, and onboarding playbooks
4) Trend Analysis & Feedback to Engineering
Analyze and categorize customer interaction trends
Provide accurate, meaningful feedback to Engineering and SRE orgs to improve product/tooling
Identify “top offenders” and propose practical fixes (tooling, docs, process, product)
5) Operational Excellence & Continuous Improvement
Participate in post-mortem reviews and drive follow-through on improvements
Contribute meaningfully to team objectives and goals (process, tooling, and service scaling)
Bring creativity and discretion to resolve highly complex issues “outside the box”
Frontline Support
Moves smoothly from triage to deeper analysis without losing the customer
Communicates clearly and confidently with technical users
Maintains clean follow-ups and thread hygiene even with high context switching
Troubleshooting
Rapidly isolates issues across monitoring/alerting configs, Linux runtime behavior, and network connectivity
Uses structured approaches to incident handling: hypothesis → test → evidence → resolution
Produces high-signal writeups that accelerate downstream resolution
Documentation & Enablement
Documentation is clear enough that customers avoid opening tickets
Onboarding flows reduce time-to-value and prevent common misconfigurations
Captures “tribal knowledge” quickly and makes it reusable
Operational Excellence
Obsessing over details: correct severity, accurate tagging, clean timelines, strong handoffs
Spots patterns early and proactively proposes improvements that scale support
Typical Day / Work Patterns
~50% Slack support, ~50% ticket handling
Deep-dive investigations during lower ticket volume periods
Documentation writing and lightweight tooling/process improvements when patterns emerge
Weekly team review of escalations, themes, and operational improvements
High rate of context switching and parallel issue management
Several years supporting highly scalable applications and web services
Hands-on experience with open-source observability and cloud-native tooling, including:
Kubernetes (and container fundamentals)
Prometheus and AlertManager troubleshooting
OpenTelemetry and distributed tracing concepts
Strong understanding of the Linux operating system (command line, process/network debugging, logs)
Good understanding of infrastructure observability principles (signals, alerting strategy, SLO thinking, noise reduction)
Good understanding of the TCP/IP suite and practical networking troubleshooting
Strong experience troubleshooting ambiguous, multi-layer issues
Excellent analytical capability and strong attention to detail
Strong written and verbal communication (clear, structured, customer-friendly)
Comfortable working with a very technical customer base
Passion for Technical Support and a service mindset
Experience improving or supporting internal support tooling or workflows (automation, templates, runbooks)
Experience operating at scale in a services environment (pattern detection, KPI/SLA awareness, operational process maturity)
Familiarity with Grafana, log aggregation, incident tooling, and production support practices
Prior SRE or platform support experience
3–7+ years in Technical Support Engineering, SRE support, DevOps, Platform Support, or similar
Demonstrated experience supporting distributed systems, IaaS, or cloud platforms
Strong Linux, troubleshooting, and customer-facing communication background
Evidence of documentation, knowledge-base contributions, and process improvement mindset
Disqualifiers: weak Linux fundamentals, inability to troubleshoot systematically, poor written communication, or discomfort supporting highly technical users.
What You’ll LoveReal technical problem solving with tangible customer impact
A role that blends deep troubleshooting with scaling support via docs, tooling, and process
High autonomy in a remote-first environment
High context switching and managing multiple threads in parallel
Repeated patterns that require discipline to convert pain into scalable improvements
Supporting high-visibility systems where speed and accuracy matter
Industry: Remote-first, trust-based culture; global team; autonomy; modern systems; meaningful technical challenges
Internal: High-impact, customer-facing observability support; direct influence on tooling and process maturity; opportunity to shape scalable support practices
$184k - $287.5k
...seeking a Senior System Software Engineer to lead the evolution of our next-generation Data & Observability Platform. We serve and... ...such as Apache Spark, Elastic/Open Search, Grafana, Prometheus, and... ...diversity in our current and future employees, we do not discriminate...SuggestedFull time- ...Description Job Description Build & Release Support Engineer – CI/CD While this position is not currently open, we are interviewing strong candidates for upcoming... ...Monitoring tools (Prometheus/Grafana) Prior SRE experience Minimum Qualifications ~2–5 years...SuggestedImmediate startRemote work
- ...Job Title: SRE Engineer: Location: Austin [Hybrid] Job Description... ...skillset to be expertise in Observability as service, Telemetry data... ...Dynatrace APM, SolarWinds, Open-Source tools (Prometheus and... ...improvements to prevent future incidents. • Analyze resource...Suggested
$60k - $135k
...Job Title: SRE ENGINEER City: Austin State... ...ambitions and build future-ready, sustainable... ...designing, managing and supporting distributed systems across... ...to be locals, or be open to relocate.... ...AWS • Monitoring & Observability: Splunk, Grafana, AppDynamics...SuggestedMinimum wageLocal areaRelocation- ...About the Role: We are looking for a Senior SRE to join our Platform Engineering team as the operations owner of our observability platforms. You’ll be responsible for the... ...spent on steady-state operations and platform support, and the other half on engineering projects...SuggestedFull time
- ...AI and shape the future of digitalization.... ...Site Reliability Engineer at TeamViewer, you... ...cloud infrastructure supporting TeamViewer’s... ...experience.5+ years in SRE, DevOps, or software... ...monitoring and observability tools (e.g. Datadog... ...proud to have an open and embracing workplace...Temporary workCasual workWorldwide
$152k - $241.5k
...services. Our work opens up new universes... ...looking for a Senior SRE to join our... ...experience building and supporting critical services.... ...auto-healing, E2E observability or data-driven... ...Ruby.Mentored other engineers and influenced technical... ...our current and future employees, we do...Full time- ...seeking a Kafka Site Reliability Engineer to help build, operate, and... ...streaming services that support critical business capabilities... ...excellence through automation, observability, Infrastructure as Code, and... ...opportunity to influence the future of event streaming at Schwab...Full timeWork at office
- ...up? If so, being a Software Engineer III at Frost could be the job... ...As a Software Engineer III - Open Banking at Frost, you are our... ...writing, testing, implementing, supporting, and documenting solutions... ...health, your family, and your future and strive to have our benefits...Full time
- ...integrated product, engineering, strategy and risk... ...shaping Schwab’s future with AI, to build... ...Engineer you will support reliability efforts... ...comprehensive observability frameworks to minimize... ...to integrate SRE practices early in... ...with proprietary or open-source LLMs (e.g.,...Full time
- ...About the Company At Future Secure AI, we're... ...and accessible, with an open-door policy that means... ...for a Site Reliability Engineer to help design, build,... ...infrastructure supporting AI Co-Workers Own Kubernetes... ...Build and improve observability across monitoring, logging...Flexible hours
$178.42k - $230.5k
...maintaining the tools and services engineers here at GM use every day to... ...delivering impact through observability frameworks and will evolve... ...and SLOsOwn or contribute to Open Source projectsPassion for self... ...your ambitions. Learn how GM supports a rewarding career that rewards...Full timeWork experience placementLocal areaWork from homeRelocation packageFlexible hours- ...passionate employees and is supported by world-renowned investors including... ...If you’re ready to shape the future of AI and grow your career in... ...and systems & integrations (engineering).Full compliance with legal... ...to expand our network.Open communication, regular feedback...Full timeWork at officeLocal areaRemote workWork from homeFlexible hours
$224k - $356.5k
...immersed in a diverse, supportive environment where... ...LLMs out of the box. The open-source community builds... ...interconnect throughput) to observed inference throughput,... ...Science, Computer Engineering, Electrical... ...package. As you plan your future, see what we can offer...Full timeLocal area$47.85 - $57.85 per hour
...made to facilitate the recruiting process are not a guarantee of future or continued accommodations once hired. If you would like to be... ...have accommodation needs such as for a disability or religious observance, please call us toll free at (***) ***-**** or send us an...Hourly payWork experience placementLive inWork at officeLocal areaFlexible hours$114.6k - $234.6k
...Principal Software Engineer focused on... ...Product Management, SRE, and OCI platform... ....Establish robust observability through metrics, logging... ....Contributions to open-source software, security... ...decision making supported by clear metrics,... ...into a better future for all. Discover...Temporary workFlexible hours$144.6k - $198.8k
...the Team Shaping the Future of Industrial Technology... ...architecture and engineering standards that shape how... ...improvement in DevSecOps, observability, CI/CD, and compliance... ....Contributions to open-source projects or involvement... ...with a collaborative, supportive team that shares your...Ongoing contractFull timeContract workLocal areaRemote workWorldwide$135.2k - $306.4k
...Infrastructure's state of the art observability platform, powering... ...platforms used by OCI engineering teams to operate and... ...promise into a better future for all. Discover your... ...benefits that support our people with flexible... ...Concurrent Programming Open source technologies...Temporary workFlexible hours$152k - $241.5k
...immersed in a diverse, supportive environment where... ...for a Senior Software Engineer to join our mission to... ...implementation, testing, rollout, observability, and iterative... ...responsibilities for an open source component used... ...in our current and future employees, we do not discriminate...Full time- ...power of AI and shape the future of digitalization.We... ..., logging, and observability solutions to ensure platform... ...optimization while supporting incident response and... ...call operations.Mentor engineers and partner across Engineering... ...are proud to have an open and embracing...Temporary workCasual workWorldwide
- ...Services enables the future of how clients manage... ...role partners across engineering, infrastructure, and security... ...— automated testing, observability, reliability, and... ...architects, production support, and infrastructure teams... ...patterns, or open-source projects in the...Full timeWork at office
- ...ImpactThe Principal Software Engineer shapes and evolves our architecture... ...management systems, and observability systems like Logging and... ...party components (Commercial or Open Source) that provide operational... ...sponsorship now or in the future. DISCO is not currently sponsoring...Visa sponsorship
$210k - $280k
...identity company of the future. Our mission is to... ...Senior Fullstack Software Engineers to help build the next... ...and team matching (open roles across the three... ...rollout, and operational support.Partner closely with... ...performance, testing, observability, and developer experience...Casual workWork at officeFlexible hours$166.9k - $242.4k
...Area, or New York City (opening Summer 2026), you’ll have the support to work in the way that... ...scale. As a Senior Software Engineer on the Upstart Bank team... ..., performance, observability, and data consistency across... ...to help you plan for the future, including a 401(k) or Group...Summer workBank staffCurrently hiringLocal areaRemote workWork from home- ...Enterprise AI Platform Engineer transforms citizen-... ...comprehensive monitoring and observability for all production AI... ...Enablement and Support: This role also involves... ...sponsorship now or in the future. DISCO is not... ...transfers.Perks of DISCO Open, inclusive, and fun environmentBenefits...H1bVisa sponsorship
$148k - $203.5k
...and coached manager, a supportive team (each with their... ...fits. As a Solutions Engineer , you’ll be the technical... ...their discoveries openly, and help define best... ...federal/statutory holidays observed 4 BetterUp Inner Work... ...be modified in the future. The base salary range...Full timeWork experience placementSummer holidayLive outWork at officeLocal areaFlexible hours2 days per week- ...the team powering the future of vision AI Hello, hola... ...role As a Solutions Engineer , you'll turn complex... ...the AEs and CSMs you support. 9 months: Converting... ...environment. Passion for AI, open source, and helping... ...monitoring, CI/CD, or observability. Familiarity with AWS,...Full timeWork at officeLocal areaWorldwideFlexible hours
$115.57k - $195.44k
..., and thrive with our open, AI-driven commerce ecosystem... ...who shape the future of commerce, this is the... ...for a Senior Software Engineer who can own complex... ...design, infrastructure, support, and other engineering... ...through tests, code review, observability, and simpler designs....Local area$95k - $171k
...you want to build your SRE career on one of the... ...As an Site Reliability Engineer II, you will be responsible... ...Akamai's existing observability platform Writing... ...areas for improvement Supporting CI/CD pipeline... ...for today and in the future. We provide benefits surrounding...Permanent employmentWork experience placementWork at officeRemote workWork from homeWorldwideFlexible hours- ...Site Reliability Engineer (SRE) Location: Austin, TX We’re searching... ...and proactively anticipate future needs. Documentation &... ...strong relationships and provide support to partner teams like... ...performance. Telemetry and Observability: Expertise in implementing and...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Future Openings - SRE Support Engineer - Observability. Be the first to apply!
- site reliability engineer Austin, TX
- site reliability engineer sre Austin, TX
- site reliability engineer remote Austin, TX
- application support engineer Austin, TX
- IT software developer Austin, TX
- customer support engineer Austin, TX
- senior application support engineer Austin, TX
- lab support engineer Austin, TX
- IT developer Austin, TX
- remote support engineer Austin, TX

