Technology Engineering - Lead Platform Engineer (SRE) (IC3)
Mindlance
Job-ID29344582Reference26-26649Remote50% Remote The Lead Monitoring and Observability Engineer (IC3) serves as a senior technical contributor responsible for ensuring monitoring reliability, telemetry quality, automation maturity, and operability across OMF’s infrastructure and application ecosystem. This role acts as a technical mentor, monitoring and networking SME, and hands-on engineer who guides monitoring-platform evolution, improves service quality, and collaborates with product and engineering teams to deliver scalable, stable, and observable systems. Key Responsibilities Monitoring Reliability & Performance • Establish platform SLOs, availability goals, latency/error budgets, reliability metrics, and monitoring coverage expectations in partnership with teams. • Continuously measure service and platform health and implement changes to improve reliability, performance, alert quality, and operational stability. • Define and maintain standards for actionable alerting, dashboards, logs, metrics, traces, and service-health reporting. Architecture & Technical Leadership • Lead observability design efforts for Elastic/ELK, telemetry pipelines, monitoring platforms, dashboards, alerting, distributed tracing, synthetic monitoring, microservices platforms, cloud infrastructure, or network monitoring, depending on assignment. • Provide “out-of-the-box” technical solutions that balance velocity, reliability, operational visibility, and cost. • Evaluate monitoring and observability tools, patterns, integrations, and emerging requirements. • Define reusable observability patterns, reference architectures, onboarding guidance, and validation practices for product and platform teams. Hands-On Engineering & Automation • Perform advanced configuration, IaC development (Terraform/Ansible), monitoring-as-code development, CI/CD pipeline engineering, and cloud platform automation. • Build and maintain scalable, resilient monitoring and observability components that support product teams. • Implement and maintain monitoring, alerting, data visualization, logging, metrics, tracing, and telemetry collection capabilities as noted in the IC3 role reference. • Improve automation for monitoring onboarding, dashboard creation, alert configuration, telemetry collection, and operational workflows. Operational Excellence • Reduce manual operations through automation and self-service observability tooling. • Review, optimize, and maintain observability capabilities including logs, metrics, traces, dashboards, alerting, and service-health reporting. • Improve alert signal quality, reduce alert noise, and ensure alerts have clear ownership, escalation paths, and actionable runbook guidance. • Participate in and lead high-severity incident response for monitoring and observability-owned domains. • Support root-cause analysis by correlating telemetry, infrastructure conditions, application behavior, and operational events. Network Monitoring • Define and maintain network-monitoring standards for network availability, reachability, latency, packet loss, interface health, capacity, device health, routing, and dependency-related service impact. • Partner with Networking, Cloud, AppDev, Security, and SRE teams to establish monitoring coverage for critical network devices, services, and customer journeys. • Build and maintain network-monitoring dashboards, alerting patterns, service-health views, and operational workflows. • Establish validation practices for network-device onboarding, telemetry collection, alert quality, dashboard completeness, and operational readiness. • Support the migration of network monitoring from OpsRamp to the selected network-monitoring tool, including requirements definition, technical design, migration planning, testing, validation, and operational handoff. • Evaluate network-monitoring tools, integrations, automation opportunities, and emerging network observability requirements. • Connect network telemetry and alerts to application, infrastructure, and customer-impact signals to improve triage and incident response. Cross-Team Collaboration • Partner with AppDev, Security, Observability, Networking, Cloud, and SRE teams to deliver cohesive monitoring services and shared solutions. • Adjust monitoring-platform design, observability standards, and processes based on evolving product team needs. • Collaborate with service owners to prioritize observability improvements for critical services. Mentorship & Influence • Mentor IC1–IC2 engineers in engineering best practices, TDD, agile ceremonies, monitoring and observability fundamentals, and operational workflows. • Provide code reviews, design feedback, and internal technical enablement. • Guide partner teams in adopting agreed observability standards and reusable implementation patterns. • Provide technical leadership on monitoring architecture, telemetry quality, alerting practices, and operational readiness. Qualifications Required • Demonstrated expertise in monitoring and observability engineering, including Elastic/ELK, dashboards, alerting, logging, metrics, tracing, telemetry pipelines, cloud infrastructure, networking, or equivalent. • Proven ability to define availability and performance objectives and implement monitoring-driven improvements. • Hands-on experience with automation, including Terraform, Ansible, scripts, monitoring-as-code, or cloud-platform automation. • Experience designing, implementing, and maintaining monitoring, alerting, dashboards, and service-health reporting. • Experience with observability practices across logs, metrics, traces, dashboards, alerting, and incident response. • Ability to evaluate alternatives, provide thought leadership, and guide technical direction. • Proficiency with TDD, agile development, and iterative delivery. • Strong communication and mentoring capability. Network Monitoring Requirements • Experience with enterprise network monitoring, network operations, network observability, or a related infrastructure discipline. • Working knowledge of network technologies and protocols, including TCP/IP, DNS, SNMP, ICMP, routing, switching, firewalls, load balancing, VPN, and WAN connectivity. • Experience monitoring network devices, interfaces, availability, latency, packet loss, bandwidth utilization, capacity, and reachability. • Ability to design network-focused dashboards, alerts, escalation workflows, and service-health views. • Experience collaborating with Networking and incident-response teams during complex service-impacting events. • Familiarity with network-monitoring migrations, platform consolidations, tool evaluations, or proof-of-concepts. Preferred • Experience with distributed tracing, observability platforms, monitoring analytics, or OpenTelemetry-aligned telemetry practices. • Familiarity with DevOps/SRE practices, including SLOs, runbooks, deployments, incident response, and post-incident improvement. • Experience with multi-cloud environments or hybrid architectures. • Experience with OpsRamp, SolarWinds, Elastic/ELK, Grafana, ServiceNow, or related monitoring, alerting, and incident-management integrations. • Experience with synthetic monitoring, customer-journey monitoring, or service-level reporting. • Experience leading a network-monitoring platform migration or enterprise monitoring-tool evaluation. Release Comments: Remote candidates will be considered.
$103.71k - $138.28k
...development of defined projects and end user support for our Power Platform and Microsoft 365 productivity services in Lumen’s commercial... ...Required Qualifications: ~ Major/Degree: Computer Science, Engineering, or similar technical degree. ~8+ years of experience with...SuggestedFull timeTemporary workRemote work$132.23k - $176.31k
...functional teams, mentoring junior engineers, and developing innovative... .... Collaborate with technology vendors and stakeholders to implement... ...and maintenance, including platforms such as Cisco, Arista, or... ...Leadership: Demonstrated ability to lead technical teams and mentor...SuggestedFull timeTemporary workLocal areaRemote work$105.79k - $141.05k
...Cloud Service (AEMaaCS). This role combines Full Stack development, cloud-native technologies, and an AI-first approach to software engineering to deliver scalable solutions, accelerate platform innovation, and improve development efficiency. Working closely with...SuggestedFull timeTemporary workRemote work- ...We are looking for an experienced Blockchain Engineer to contribute to a Web3 developer reputation platform currently under active development. The platform creates... ...smart contract security, ZK systems, cross-chain technology, or on-chain data. You do not need experience...SuggestedContract work
- ...We’re building a Web3 developer reputation platform that creates verifiable profiles from on-chain activity, smart contract deployments... ...and other technical evidence. We’re looking for an experienced engineer with strong skills in Blockchain, AI, and Security to help...SuggestedContract work
$152.07k - $202.76k
...individual-contributor role that partners with Technology VPs, Finance, business leaders, and... ...decks, and portfolio reporting. Lead cross-functional follow-through on priorities... ...in Business, Finance, Technology, Engineering, or a related field; advanced degree preferred...Full timeContract workTemporary workWork at officeRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Technology Engineering - Lead Platform Engineer (SRE) (IC3). Be the first to apply!

