Senior Site Reliability Engineer (SRE) - Dynatrace & Azure Observability Expert
RaceTrac
Job Description
Job Description
We are seeking a highly experienced Site Reliability Engineer (SRE) with deep expertise in Dynatrace, observability engineering, and Azure cloud technologies . This role will be exclusively focused on building, enhancing, and managing enterprise observability, telemetry, monitoring, and proactive reliability engineering practices across critical digital platforms.
The ideal candidate must possess advanced hands-on expertise in Dynatrace, especially Dynatrace Query Language (DQL), along with strong knowledge of Azure Monitor, Azure KQL, Application Insights, Azure Functions, APIM, and distributed telemetry concepts. The candidate should have a strong understanding of .NET application architecture and the ability to read and analyze .NET code to support troubleshooting, root cause analysis, and observability implementation within Azure environments. Experience enabling observability for mobile platforms such as iOS and Android is also required.
This is a highly technical, hands-on role requiring a proactive engineering mindset, strong analytical capabilities, and the ability to collaborate across engineering, cloud, mobile, and business teams.
What You'll Do
Dynatrace & Observability Engineering
Serve as the primary Dynatrace SME across the organization.
Design, develop, and optimize enterprise observability solutions using Dynatrace.
Develop advanced Dynatrace DQL queries, dashboards, workflows, alerts, and analytics.
Implement intelligent monitoring strategies for applications, APIs, integrations, Azure services, mobile platforms, and distributed systems.
Continuously improve observability maturity through telemetry standardization, proactive monitoring, and automation.
Configure and tune alerting mechanisms to improve signal-to-noise ratio and reduce alert fatigue.
Leverage Dynatrace Davis AI, anomaly detection, and AI-driven root cause analysis capabilities.
Enable and enhance observability for mobile applications across iOS and Android platforms.
Azure Monitoring & Cloud Operations
Build and maintain monitoring solutions using:
Azure Monitor
Application Insights
Azure Log Analytics
Azure KQL
Monitor and troubleshoot Azure Function Apps, App Services, APIs, integrations, and backend services.
Analyze telemetry, traces, logs, metrics, and distributed transactions to identify root causes and performance bottlenecks.
Troubleshoot cloud-native applications and Azure infrastructure issues.
Develop proactive monitoring for cloud services, integrations, APIs, and backend processing systems.
API & Integration Monitoring
Monitor and troubleshoot Azure API Management (APIM), API Gateways, API endpoints, and integrations.
Understand end-to-end API transaction flows and dependency mapping.
Build observability solutions for APIs, middleware platforms, and integration services.
Diagnose latency issues, transaction failures, authentication issues, and backend service degradation.
Mobile Application Observability
Enable telemetry, monitoring, tracing, and performance analysis for iOS and Android applications.
Analyze mobile-to-backend transaction flows and end-user experience metrics.
Troubleshoot mobile application latency, crash analytics, API failures, and connectivity issues.
Correlate mobile telemetry with backend application and infrastructure monitoring data.
Application Engineering & Troubleshooting
Utilize prior .NET development experience to troubleshoot application behavior, performance, and deployment issues.
Read and understand .NET application code to support root cause analysis and observability implementation.
Work closely with development teams to understand application logic, API flows, dependencies, and exception handling.
Support Azure Function deployments, configuration management, scaling, and runtime troubleshooting.
Collaborate with development teams during architecture reviews and production releases.
Ensure observability and monitoring readiness before deployments go live.
Site Reliability Engineering (SRE)
Perform deep technical analysis across systems by correlating logs, metrics, traces, and application telemetry.
Conduct root cause analysis (RCA) for recurring incidents and systemic issues.
Partner with engineering and operations teams to implement preventive improvements and automation.
Develop KPI-driven reliability improvements focused on system stability, performance, and operational excellence.
Proactively identify risks, bottlenecks, failure patterns, and reliability concerns before business impact occurs.
Continuous Improvement & Automation
Automate operational workflows and monitoring processes wherever possible.
Improve operational efficiency using AI-driven insights and automation capabilities.
Build reusable monitoring frameworks, dashboards, and telemetry standards.
Drive observability best practices across engineering teams.
What We're Looking For
Mandatory Technical Skills
10+ years of overall IT experience.
Expert-level hands-on experience with Dynatrace.
Advanced expertise in Dynatrace Query Language (DQL).
Strong hands-on expertise in Azure Kusto Query Language (KQL).
Deep understanding of telemetry, observability, distributed tracing, metrics, and logging concepts.
Strong Azure cloud experience with emphasis on:
Azure Monitor
Application Insights
Azure Functions
Azure API Management (APIM)
Azure Log Analytics
App Services
Strong understanding of API architectures, API Gateways, and backend integrations.
Prior hands-on experience developing .NET applications.
Strong ability to read, analyze, and understand .NET application code.
Experience troubleshooting and deploying Azure Functions and cloud-native applications.
Experience enabling observability and telemetry for mobile applications on iOS and Android.
Understanding of mobile telemetry, crash analytics, API monitoring, and end-user experience monitoring.
Strong understanding of distributed systems and enterprise application architectures.
Preferred Skills
Experience with OpenTelemetry implementation and instrumentation.
Experience with CI/CD pipelines and DevOps practices.
Knowledge of AI-driven observability and AIOps concepts.
Experience monitoring high-volume enterprise digital platforms.
Familiarity with ServiceNow and incident management workflows.
Experience with Databricks, SQL platforms, and integration technologies.
Core Competencies
Strong analytical and troubleshooting skills.
Excellent communication and stakeholder management abilities.
Ability to work independently and drive proactively.
Strong collaboration skills across engineering, cloud, SRE, mobile, and business teams.
Ability to quickly adapt to new technologies and evolving environments.
Success Criteria
Reduction in recurring incidents through proactive monitoring and RCA.
Improved observability coverage across enterprise systems, APIs, and mobile applications.
Faster incident detection and resolution.
Reduction in monitoring noise and false positives.
Increased automation and operational efficiency.
Improved reliability and performance of critical systems and APIs.
Strong partnership with engineering teams to ensure production readiness and operational excellence.
Fueled by Growth, Driven by You
At RaceTrac, our people make the difference. Whether you’re working in a store, at our corporate office, or on the road, you’ll be part of a team that brings energy, innovation, and a passion for serving others every day. We support each other, celebrate wins big and small, and create opportunities for growth at every level. With four operating divisions RaceTrac, RaceWay, Energy Dispatch, and Gulf - there’s always a new challenge to take on and a new path to pursue. Join us and discover how far your career can go.
To see what #LifeatRaceTrac is like, visit our LinkedIn, Facebook, and Instagram pages.
All qualified applicants will receive consideration for employment with RaceTrac without regard to their race, national origin, religion, age, color, sex, sexual orientation, gender identity, disability, or protected veteran status, or any other characteristic protected by local, state, or federal laws, rules, or regulations.
- ...Job Opening: AWS Site Reliability Engineer (SRE) We’re hiring a Site... ...hiring process. Seniority level Seniority level... ...Engineer with Azure and Dynatrace Atlanta, GA $70,00... ...Reliability Engineer - Observability Atlanta, GA $100,0... ...in a new way. Experts add insights directly...SuggestedContract work
- ...Technology Consultant - Site Reliability Engineer (SRE) Location- - Atlanta... ...expertise in Kubernetes, Observability, Java, and production reliability... ...tools such as Splunk, Dynatrace, Prometheus, Grafana, Datadog... ...Experience with AWS, Azure, or Google Cloud Platform....SuggestedPermanent employment
- OneTrust is seeking a Senior Software Engineer in Atlanta, Georgia. The role involves designing and maintaining a reliable application platform, collaborating with engineering teams... ...customer experiences through observability tools. The ideal candidate will have a...Senior
- ...We are seeking an experienced Site Reliability Engineer (SRE) to support and maintain production systems... ..., incident management, monitoring, observability, troubleshooting, and improving... ...observability dashboards using CloudWatch, Dynatrace, Quantum Metric, and ThousandEyes....Suggested
- Inspire Brands is hiring two Senior Site Reliability Engineers to help build and scale reliable, resilient, and observable systems supporting high-... ...applying and implementing SRE principles — not just supporting... ...one major cloud platform (Azure, AWS, or GCP)Expertise with...SeniorWorldwide
- ...software platforms using SRE and AI-native engineering practices. Own production reliability, monitoring, and... ...cloud platforms (AWS/Azure), Kubernetes, and production... .... Key Skills Site Reliability Engineering... ...Kubernetes Docker CI/CD Vault Observability AI-native Engineering...SeniorTemporary workFlexible hours
- ...Senior Systems Reliability Engineer (SRE)At T-Mobile, we invest in YOU! Our Total Rewards... ...solutions using Power Platform, Azure DevOps Pipelines,... ...reliability, scalability, observability, and continuous improvement... ...software engineers, compliance experts, security teams, and AI-...SeniorFull timeTemporary workPart timeWork experience placementFlexible hours
- ...Job Title: Site Reliability Engineer II (SRE II) Data & Intelligence Location: Atlanta, GA Contract... ...continuous improvement through observability, automation, and reliability engineering... ...and platform services across Azure, AWS, or GCP environments. Work...Contract work
$70k - $110k
...position focused on reliability, scalability,... ...knowledge of SRE principles,... ..., AI-powered observability, or intelligent... ...AI-enabled engineering tools in production... ...such as Dynatrace, Prometheus, Grafana... ...Serve as a senior technical advisor... ...matter expert on reliability...Long term contractFull time- ...& Infrastructure Expert Role Type: Contractor... ..., DevOps, SRE, and platform engineering. You will test AI... ...for accuracy and reliability. Work with AWS, Azure, GCP, Kubernetes,... ...Infrastructure Site Reliability Engineering... ...GitLab / JIRA Observability ChatGPT Work...Remote jobFor contractors
- ...Summary: As an SRE Architect, you... ...ensure the reliability, scalability,... ...expertise in software engineering, distributed... ...much beyond observability pillar. Key... ...: Act as a senior technical... ...subject matter expert on reliability... ...solutions (e.g., Dynatrace, Prometheus,...Early shift
- ...Hands-on experience with incident management and 24/7 production support models.· Proficiency with monitoring and observability tools such as CloudWatch, Dynatrace, and Quantum Metric.· Experience building and maintaining monitoring dashboards.· Strong troubleshooting...
- ...job description:The Site Reliability Engineer role focuses on enhancing... ...environments. This senior technical leader... ...improvements in automation, observability, and incident... ..., mentoring SRE team members, and contributing... ...platforms (Splunk, Dynatrace) and event-driven...Permanent employmentFull timePart timeH1bWork at officeLocal areaImmediate startWork visaMonday to FridayShift workDay shift
$100k - $120k
OverviewThe Site Reliability Engineer is a key force behind improving Origami’s... ...colleagues on how to best leverage observability tools during incident and... ...larger Cloud Operations, SRE, Engineering teams, and the... ...platforms (e.g., AWS, Azure) and architectural patternsExcellent...Full timeTemporary workWork experience placementFlexible hours$87.95k - $162.88k
...currently seeking a Senior AI DevOps Engineer (AI Ops /... ..., Security, SRE, and AI teams... ...improving reliability, observability, and operational... ...across AWS, Azure, or Google... ...CloudWatch, Splunk, Dynatrace, ELK, or... ...Engineering, Site Reliability Engineering... ..., we have experts in more than...SeniorTemporary workWork at officeRemote workFlexible hours3 days per week$86.36k - $101.6k
...Responsibilities Lead Observability Strategy across... ...with business outcomes, reliability goals, and customer... ...with Product, Engineering, SRE, and Operations teams... ...Observability Engineering , Site Reliability... ...Proficiency with Datadog, Dynatrace, Splunk, Grafana, Prometheus...Full timeTemporary workWork experience placementLocal area3 days per week$116.48k - $174.71k
...Challenge We're looking for a Senior Software Engineer that will report to the... ...and implement application observability and platform monitoring... ...budgets to balance system reliability with product feature velocity... ...‑based infrastructure (Azure, AWS, GCP, etc.). Experience...SeniorWork experience placementWork at officeLocal areaWorldwideFlexible hours3 days per week1 day per week$98.18k - $115.5k
...DescriptionResponsibilitiesThe Reliability Observability Engineer 3 is responsible for... ...indicators.As a senior-level Reliability... ...engineering teams, SRE teams, and business... ...Engineering, Site Reliability Engineering... ...Proficiency with Datadog, Dynatrace, Splunk, Grafana,...Full timeWork experience placementLocal area3 days per week- T-Mobile seeks a senior SRE leader to design and operate secure, scalable... .... You will own production reliability, monitoring, and automation,... ...hands-on Terraform, AWS/Azure, Kubernetes, and production... ...with a focus on reliability, observability, and incident response. Leadership...Senior
- ...countries worldwide.Observability Automation... ...an Observability Engineer to design, implement... ...to improve service reliability, accelerate troubleshooting... ...with application, SRE, cloud, and... ...in Observability, Site Reliability Engineering... ...such as Microsoft Azure, AWS, or Google...Full timeWorldwideFlexible hours
- ...to shape the future of education. Site Reliability Engineer (SRE) Overview: We are looking for... ...playbook. You'll work with leading-edge observability and reliability tooling, and the... ...AWS), Google Cloud Platform (GCP), or Azure). Communication: Experienced,...Full timeLive inWork at office
- ...pivotal to our platform engineering and site reliability initiatives, owning... ...build and run. The Senior Engineer, Platform &... ...mesh, and modern observability to keep our systems... ...maintains CI/CD pipelines (Azure DevOps) and... ...standards in cloud, DevOps, SRE, and web development...SeniorWork experience placement
$151k - $297k
...As a TPM for SRE, you will partner with SRE leaders and engineers to scale the platform that underpins... ...strengthen production reliability practices, and... ..., cloud networking, or observability stacks (metrics, logs,... ...Google Cloud, and Microsoft Azure. With offices worldwide...Local areaRemote workWorldwideFlexible hours- ...We are seeking a Senior DevOps Developer... ...drive platform reliability, scalability, security... ...development and engineering teams.... ...strong expertise in Azure cloud technologies... ...health using Dynatrace and other observability tools while proactively... ...Engineering, Site Reliability...Senior
$293.5k
General Information Job Title Expert Senior Manager, AI Engineering Job ID 104335 Work Areas Analytics, Data & Research, Management... ...API development, microservices, CI/CD pipelines, observability, and cloud-native deploymentBuild scalable GenAIOps processes...SeniorPermanent employmentFull timeApprenticeshipWork at officeLocal areaWork from homeHome office3 days per week- ...Are you a seasoned engineer who loves building cloud... ...is looking for a Senior Software Engineer to... ...systems for accuracy, reliability and optimal... ...Cloud Platform, AWS or Azure Build tooling with... ...management with Apigee Observability with Dynatrace and Google Cloud...SeniorWork at officeRemote workMonday to Friday
$98k - $148.5k
...other teams across the organization.As a Site Reliability Engineer I on the Core Infrastructure team in... ...infrastructure (e.g., AWS, GCP, Azure),including networking and compute concepts... ...Experience with monitoring, observability, and logging platforms (e.g., DataDog...Work at officeLocal areaFlexible hours$35 - $44 per hour
DescriptionKforce has a client seeking a remote Site Reliability Engineer to join their team. We are seeking a Site Reliability Engineer (SRE) to support a large-scale system... ...automation, and operational support* Manage observability tools such as Grafana and Prometheus*...Remote work- EY seeks an AI Systems Engineer to own delivery, model-serving, routing, and observability for EY’s AI-native platform. You will oversee CI/CD/CV pipelines, governance of AI assets, cost attribution, and deep observability across cloud, on-prem, edge, and air-gapped environments...Senior
- ...typically as part of a larger core team working on a priority. At a Senior Manager level, it could include the following:Lead client... ...have accommodation needs such as for a disability or religious observance, please call us toll free at (***) ***-**** or send us an email...SeniorFull timeLive inWork at officeLocal area
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Site Reliability Engineer (SRE) - Dynatrace & Azure Observability Expert. Be the first to apply!
- site reliability engineer Atlanta, GA
- site reliability engineer sre Atlanta, GA
- azure data engineer Atlanta, GA
- azure developer Atlanta, GA
- microsoft azure architect Atlanta, GA
- cloud engineer azure Atlanta, GA
- azure specialist Atlanta, GA
- azure infrastructure engineer Atlanta, GA
- subject matter expert Atlanta, GA
- fulfillment expert Atlanta, GA


