Platform Reliability Engineer, Azure
Wellfit Technologies
Job Description
Job Description
Wellfit is the dental industry’s fintech solution , breaking down financial barriers so patients, providers, employers, and payors can all access better care. As a healthcare fintech innovator, we’re transforming the patient journey and redefining what’s possible in dental care.
About Wellfit
Wellfit is the dental industry’s fintech solution, breaking down financial barriers so patients, providers, employers, and payors can access better care. As a healthcare fintech innovator, we are transforming the patient journey and redefining what is possible in dental care. Today, Wellfit supports a growing production platform serving 1,100+ offices and processing $1.5B+ in annual transactions . As we continue to scale, reliability, observability, alerting, and production readiness are critical to how we support our customers and deliver with confidence.
About the Role:We are seeking a hands-on Platform Reliability Engineer, Azure to help strengthen the reliability, visibility, and operational maturity of our Azure-based platforms.
This role is ideal for someone who enjoys working directly in Azure, improving production systems, troubleshooting issues across infrastructure and application layers, and building practical monitoring and alerting solutions that help teams respond faster and operate more confidently.
You do not need to be an expert in every part of the stack on day one. We are looking for someone with strong Azure experience, solid troubleshooting instincts, a DevOps/reliability mindset, and the ability to collaborate closely with engineering teams across systems, services, and applications.
What You’ll Do
• Own and improve monitoring, alerting, and observability across Azure-based production systems.
•Work directly in Azure Monitor, Application Insights, App Services, logs, metrics, traces, and related Azure tooling to troubleshoot reliability and performance issues.
• Build and refine practical alerting workflows, including Sev0/Sev1 alert routing, escalation paths, and runbook integration.
• Create and maintain clear, actionable runbooks that help on-call engineers respond confidently to production incidents.
• Partner with engineering teams to investigate issues across infrastructure, configuration, deployments, services, and application behavior.
• Support release readiness by improving visibility into critical Azure resources before, during, and after production deployments.
• Build and maintain dashboards in tools such as Grafana, Azure Monitor, Application Insights, or similar observability platforms.
• Help configure incident routing integrations, including Slack/webhook-based alert delivery to the appropriate team channels.
• Automate repeatable operational tasks using PowerShell, Logic Apps, Azure tooling, or similar workflow automation methods.
• Contribute to RCA documentation, incident follow-up, reliability improvements, and operational playbook development.
What We’re Looking For
• Hands-on experience supporting production systems in Azure.
• Strong working knowledge of Azure App Services, Azure Monitor, Application Insights, and Azure production troubleshooting.
• Experience with DevOps, cloud operations, site reliability, platform engineering, or production support in a hands-on environment.
• Strong troubleshooting instincts and the ability to work through ambiguous production issues.
• Comfort working across logs, metrics, traces, alerts, configurations, deployments, and service dependencies.
• Ability to collaborate with software engineering teams across the stack, including .NET, Angular, SQL, APIs, and cloud services.
• Experience building or improving dashboards, alerts, runbooks, incident workflows, or operational playbooks.
• Working knowledge of scripting or automation, preferably with PowerShell, Logic Apps, CLI tooling, or similar technologies.
• Clear communication skills with the ability to document findings, explain issues, and drive follow-through after incidents.
• A high-ownership mindset with the ability to create structure, improve processes, and operate effectively in a fast-moving environment.
Preferred Experience
• Azure certifications.
• Grafana, Prometheus, DataDog, Dynatrace, or similar observability/APM tools.
• Slack integrations, webhooks, Logic Apps, or incident routing workflows.
• Azure Front Door, CDN, Function Apps, WebJobs, Service Bus, Event Hub, Event Grid, SQL Pools, App Service Plans, or related Azure services.
• Experience in healthcare, fintech, payments, or other high-availability environments.
• Experience in startup, SMB, or scale-up environments where ownership is broad and hands-on.
What Success Looks Like
• You understand how production systems are monitored, where alerting gaps exist, and how to improve them.
• You can work directly in Azure to investigate issues, improve visibility, and support reliable operations.
• You build practical runbooks, dashboards, and alerting workflows that teams actually use.
• You collaborate well with engineers, ask strong troubleshooting questions, and help drive issues to resolution.
• You bring ownership, curiosity, and a builder mindset to a growing platform environment.
Why Wellfit
• Make an Impact: Your work will directly strengthen the reliability of a fast-growing healthcare fintech platform supporting 1,100+ offices and $1.5B+ in annual transactions.
• Build and Own: This is a high-impact role where you will help shape how we monitor, operate, and scale production systems.
• Work Flexibly: Hybrid model based in Dallas with 3 days per week in office.
• Comprehensive Benefits: Full medical, dental, vision, generous PTO, bonus eligibility, and 401(k) matching.
• Fast-Growth Environment: A rare opportunity to grow with a profitable startup on a national trajectory.
Alongside a competitive annual bonus, we offer a 401(k) with up to a 4% match, generous paid time off, and comprehensive healthcare benefits.
We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.
- Site Reliability Engineer - Vice PresidentSite Reliability Engineering (SRE) is an engineering discipline that combines software and systems... ...the availability and reliability of the firm’s most critical platform services and ensures they meet the requirements of our internal...Suggested
- ...ResponsibilitiesDesigns, implements, and deploys platforms to support and meet business... ..., configure, and support Azure Virtual Desktop host pools,... ...in improving repeatability, reliability, and documentation of... ....Work closely with senior engineers, architects, and cross‑...SuggestedFull timeWork experience placement
$85 - $90 per hour
...Role: Senior SRE Engineer Location: Dallas / Fort Worth, Texas Rate: up to $85-$90 per hour INC Structure... ...work experience. ~ Extensive experience with Azure. ~ Experience with container orchestration platforms such as Kubernetes. ~ Experience using IAC tools...SuggestedHourly payContract workWork experience placement$119k - $224k
Wells Fargo is seeking a Lead Platform Reliability Engineer to join the CTO Platform organization. This role is designed for highly experienced infrastructure engineers who possess deep technical expertise in one core platform discipline (Network, Middleware, Database,...SuggestedFull timeWork experience placement$180.5k - $236.91k
...re Oscar. We're hiring a Senior Software Engineer, Cloud Infrastructure / SRE to join our... ...company built around a full stack technology platform and a relentless focus on serving our... ...technical domains such as DevOps, site reliability, and cloud best practices Lead the...SuggestedFull timeWork at officeRemote work$160k - $210k
...redefining media buying with our Deep Learning Advertising Platform. Since 2015, we have harnessed the power of cutting-edge deep... ...cornerstone of our success. We are looking for a Senior Site Reliability engineer to work on expanding our global footprint of datacenters and...Work at officeLocal areaImmediate startRemote work- ...160 data centers globally, the Zscaler Zero Trust Exchange platform combined with advanced AI combats billions of cyber threats... ...the future of cybersecurity. Role We are looking for a Site Reliability Engineer-SkillBridge Intern (San JosA Ca or Bellevue WA) to join our...InternshipWork at officeLocal areaRemote workWorldwide
$107.48k - $143.31k
...locations across the U.S. and Canada with approximately 9,000 employees. What You'll Be Doing Lead and apply regional reliability engineering strategies to improve equipment performance, uptime, and maintenance cost effectiveness across multiple cement plants....Temporary workRemote workFlexible hours- ...available and apply online.Job SummaryThe Principal Devops Platform Engineer will be responsible for driving development and strategy for... ...AWS Certification is preferredWell versed in Devops or Site reliability engineering (SRE) tenets2+ years in an SRE/Operations/DevOps...Full timeLocal area
$106k - $160k
...There are dynamic career paths awaiting you - rewarding opportunities to impact the lives of others and inspire love. Join us!Platform Engineer - AI, Cloud & eCommerceRemotePOSITION SUMMARY:We are looking for a highly hands-on Platform Engineer who enjoys building, experimenting...Full timeRemote workWorldwide$60 - $72 per hour
...60/hr - $72/hrWe're hiring for a large-scale infrastructure migration moving thousands of servers from Azure to Google Cloud Platform. This is a hands-on engineering role — not a "check the DevOps box" position — ideal for someone with a strong Windows systems administration...Full timeContract workTemporary workLocal areaRemote workRelocationFlexible hours$138.4k - $173k
...infrastructure as well as help improve the reliability, quality of services and overall... ...recovery. You’ll collaborate or embed with engineering teams, helping them to improve the... ...components of the AppFolio Real Estate Platform. You’ll help build the future of reliable...Full timeFlexible hours- ...From generative AI and cloud-native platforms to advanced release engineering practices, our teams are redefining... ..., specifically within Microsoft Azure, but may also include AWS and GCPKnowledge... ...behavior preferred Exposure to reliability engineering concepts such as SLOs/...Work experience placementH1bWork at officeRemote workVisa sponsorshipFlexible hoursShift work2 days per week
- Qualifications: 8+ years of Software Engineering experience, or equivalent demonstrated through one or a combination... ...Contractor will implement and maintain scalable and reliable infrastructure on Google Cloud Platform (“GCP”) for Snowflake data warehousing. Monitor, troubleshoot...Contract workFor contractorsWork experience placement
$100k - $115k
...powered by our industry-leading platform and team of experts, who help leaders... ...(IDP) as a product, treating engineering teams as customers and optimizing for reliability, usability, and delivery... ...platform abstractions across AWS and Azure that standardize security, reliability...Temporary work- Core Responsibilities:Evaluate applications, platforms, and vendors to assess resiliency, reliability, and operational risk.Design and implement processes that... ...tooling.Actively participate in reliability engineering and resilience communities of practice, contributing...Full time
- ...From generative AI and cloud-native platforms to advanced release engineering practices, our teams are redefining... ...automation, and cloud platforms (Azure).Familiarity with AI-assisted development... ...accelerate development and improve reliability. Your work will directly influence...H1bWork at officeRemote workVisa sponsorshipFlexible hours2 days per week
- About LanternLantern is the specialty care platform connecting people with the best care when... ...Lantern is seeking an experienced Senior Site Reliability Engineer to champion the reliability, availability, and performance of our Azure-based healthcare platform. In this pivotal...
$104.9k - $174.7k
...:We are hiring a hands-on Senior Site Reliability Engineer (SRE) to actively build, operate, and... ...pipelines and operational workflows using Azure DevOps and GitHubWork directly with... ...environmentsExperience operating monitoring and uptime platforms such as Grafana, Pingdom, and...Full timeWork at officeLocal areaRemote workWork from home- ...enhance patient care. Originating from a successful internal platform developed within Pacific Dental Services, Wellfit is on a... ...clear product-market fit. Role Overview: We are seeking an Azure DevOps Engineer who embodies our values of curiosity, innovation, and...Remote work
- Summary:The Cloud Platform Engineer is responsible for designing, engineering, and operating enterprise compute platforms across a hybrid environment... ....Hands-on experience with GCP or major cloud platforms (AWS/Azure transferable).Experience with VMware virtualization and...Work at office
$197.3k - $313.7k
...of Salesforce.Job Title: Director, Site Reliability EngineeringLocation: New York, NY; San... ...looking for a Director of Site Reliability Engineering to spearhead the evolution of our... ...closely across Application Engineering, Platform, Architecture, Security, Infrastructure...Full timeImmediate start$138.72k - $208.08k
...38,720.00 - $208,080.00Category: Technology, Digital Software Engineering, ProfessionalCompany: CitiThe Digital S/W Engineer Sr Mgr accomplishes... ...budget approval.Responsibilities: Deliver End to end Advisor Platform integration in target state architecture. Drive integration...Full timeTemporary work- About the role:The Release Train Engineer (RTE) is a servant leader and coach for the Agile Release Train (ART).... ...both verbally and in writingExperience with Microsoft Azure DevOps (work hierarchies) and visual platforms such as Miro, IdeaBoardz, etc.Experience:5-7 years...Work at officeFlexible hours2 days per week
- ...Job Title: Senior Platform Engineer Location: Irving, TX/ Dallas TX Job Type: 6 months Contract Work Arrangement: Onsite Role Interview: Three rounds of Video interview Job Description: Design, deploy, administer, and optimize large-scale Apache...Contract work
- ...Technical Skills • Linux bash, Python, pip, conda, Node/npm, rpm, GNU tools (g++,make,configure)Desirable Technical Skills • Public cloud platform experience• Bash script development• Python development, package management, package building and testing• Source code management...
- ...automation. You will optimize performance, drive reliability, and mentor teammates while aligning with... .... You will work across Java apps, IBM middleware, Azure, scheduling tools, and observability to improve operational efficiency and platform strategy. #J-18808-Ljbffr...
$185k - $227k
...for more details. ROLE AND RESPONSIBILITIES: A Senior Site Reliability Engineer (SRE) is expected to own the operational stability and performance... ...to ensure theplatform is scalable and efficient. Nutanix Platform Management Design, deploy, and maintain enterprise-scale...Remote work$210k - $220k
...important workflows. Our intelligent workflow platform applies AI, automation, and integration... ..., it’s popular with security, IT, engineering, finance, and other security-focused... ...to join us on our journey. Senior Site Reliability Engineer - Government Cloud You'll join...Work at officeRemote work$114k - $148k
...Site Reliability Engineer Location: Remote, United States Employment Type: Full-Time Benefits Offered... ..., you will focus on ensuring the platform and services customers rely on are reliable... ...build integrations between Dynatrace, Azure DevOps and Jira. Solid knowledge in focused...Full timeTemporary workWork experience placementRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Platform Reliability Engineer, Azure. Be the first to apply!
- platform developer Irving, TX
- platform engineer Irving, TX
- reliability maintenance engineering technician Irving, TX
- azure specialist Irving, TX
- azure developer Irving, TX
- platform product manager Irving, TX
- platform manager Irving, TX
- digital platform specialist Irving, TX
- director of digital platform Irving, TX
- power platform Irving, TX

