Get new jobs by email
- ...systems and notifying them of any potential issues. . Assisting with troubleshooting on call. . Managing and tracking incidents such as outages. . Facilitating analysis meetings to discuss incidents. Lead . Identify the automation opportunities for automation specialists to...SuggestedPermanent employmentFlexible hoursShift workWeekend work
- ...through proactive monitoring, maintenance, and capacity planning. The Site Reliability Engineer will be responsible for resolving system outages, developing tools to automate operational processes, and collaborating with development teams to integrate best practices for...Suggested
- ...that are needed dual control and required security parameters - including focusing on the initial commands each has to avoid outages like a major client incurred last fall. # GDPS Training and Education Provide extensive training to various client teams...SuggestedRemote work
- ...using our established monitoring stack (Prometheus, Grafana) Troubleshoot and resolve cluster-related incidents such as broker outages, replication lag. Develop and enhance automation scripts to streamline operational tasks and reduce manual intervention....SuggestedWork experience placement
- ...Perform root cause analysis for application, infrastructure, and monitoring-related incidents. Provide monitoring expertise during outage investigations. Develop corrective actions and preventative monitoring improvements. ServiceNow & Event Management Integration...SuggestedWork experience placement
- ...solutions, and timelines to appropriate stakeholders Assist with implementation of customer side monitoring tools and lead operation outage events Manage cloud infrastructure including networking, storage, and computer services. Provide timely/responsive technical...Suggested
$60 - $75 per hour
...Perform root cause analysis for application, infrastructure, and monitoring-related incidents. Provide monitoring expertise during outage investigations. Develop corrective actions and preventative monitoring improvements. ServiceNow & Event Management Integration...SuggestedHourly payContract workWork experience placement$85 - $88 per hour
...remotely, with occasional overnight travel (up to 20%) and flexible scheduling, including weekend and evening availability during outages, security incidents, or maintenance events WHAT'S IN IT FOR YOU Opportunity to lead high-impact modernization initiatives...SuggestedHourly payContract workLocal areaRemote workFlexible hoursNight shiftAfternoon shift- ...your time-validation requirement Cloud architecture: SaaS/hybrid deployments, failover design, data sync/reconciliation after outages OT/plant network expertise: segmentation, firewalls, PoE design, connecting security devices in industrial networks Cybersecurity...Suggested
- ...Role Descriptions: Working in production support environment Must be available on call to handle any critical issues / system outages. Triaging of issues and providing resolutions to the business users Identifying recurring issues and suggesting permanent fixes Along...SuggestedPermanent employmentWork experience placement
- ...and long?lead materials. This role serves as the primary escalation point for complex delivery challenges impacting major projects, outages, regulatory commitments, and operational reliability. The Material Expeditor Lead establishes expediting standards, governs...SuggestedContract work
$65 - $74 per hour
...Reliability, Customer Operations, and related business organizations. This role is responsible for analyzing complex utility operations, outage management, customer outage notifications, network and device models, and operational data to provide actionable business insights....SuggestedHourly payContract workMonday to Friday- ...moves, and space reconfigurations. · Participate in emergency response activities related to facility operations, including power outages or equipment failures. Requirements: · An associate degree or equivalent is required; technical certification in MEP systems,...SuggestedFor contractorsRelocation
$69 - $78 per hour
...with engineering teams Provide senior-level support for production IAM applications Lead technical troubleshooting during outages and critical incidents Perform root cause analysis and drive corrective actions Improve application monitoring, alerting,...SuggestedContract work- ...helping teams recover quickly from service disruptions. You will also use operational data to identify weaknesses before they become outages. The team is looking for someone who understands modern monitoring practices and can improve dashboards, application telemetry,...SuggestedContract work
- .../WAN infrastructure including Layer 2/3 technologies Participate in incident response and root cause analysis for critical network outages Evaluate emerging networking and security technologies and recommend improvements Required Skills Cisco ISE (5+ years hands...Long term contractContract work
- ...lifecycle of enterprise IAM platforms-from identifying technical risk and remediating vulnerabilities to restoring services during outages and improving long-term platform resiliency. You will lead efforts around vulnerability assessment and remediation, production...
- ...REST, SOAP-based APIs and files. Performing code reviews for Appian developers Conducting root cause analysis for Appian outages Restoring Appian services in the event of an outage Experience working with Appian plugins Reviewing Application performance...Work at office
- ...Cisco Secure Firewall technologies Participate in major incident response efforts and serve as a technical lead during P1 network outages Conduct root cause analysis and drive long-term remediation efforts Collaborate with infrastructure, security, cloud, and...Hourly payContract work
- ...deny mode. Investigate production incidents, application latency, traffic routing, load balancer interactions, and WAF-related outages. Provide security architecture guidance on OWASP Top 10 , API security, bot mitigation, DDoS protection, and secure web...3 days per week
- ...responsible for designing and governing scalable, secure, resilient, and low-latency technology solutions that support CenterPoint Energy"s Outage to Restoration (OTR) products. This role partners with Product Managers, Engineering Teams, Operations, and Business Stakeholders...Immediate startRemote work1 day per week
- ...IP Required 0 Candidate must have experience working in highly regulated environments and leading technical troubleshooting during outages Required 0 Candidate must have ability to communicate technical issues to technical and executive audiences and an ability to mentor...Work at office
- ...backup policies and disaster recovery plans using Azure Site Recovery and related tools. Support business continuity and emergency outage response activities. Team Collaboration & Leadership Provide technical guidance and mentorship to junior team members. Required...
$80 - $88.3 per hour
...Identifies systemic reliability risks across complex network environments and drives cross-functional remediation before they result in outages. Product Strategy & Roadmap (Product-Centric Focus) Develop and apply complex and long-term strategies on challenging / new...Hourly payLocal area- ...resolve system issues, acting as the primary point of contact. Work closely with analysts and technical teams to resolve system outages or performance issues. Schedule and coordinate SAP system installs, upgrades, patches, backups, and restorations. Perform OS,...For contractorsRemote work
- ...tracking for HPC and CAE environments. Document system configurations, procedures, incidents, and best practices. Track outages, analyze root causes, and implement preventive measures. Follow change management processes for system updates and deployments...Full time
- ...efforts within a global 24x7 IT Operations Center. This role is ideal for a seasoned operational leader who excels in high-pressure outage situations, drives rapid service restoration, and helps maintain the stability of enterprise infrastructure and applications....Night shiftAfternoon shift
$130k - $150k
..., and infrastructure-as-code initiatives. Serve as the primary escalation point for complex production incidents and critical outages. Lead root cause analysis (RCA) efforts and implement permanent corrective actions. Manage platform lifecycle planning, capacity...Permanent employmentFull time- ...Support construction oversight of ancillary buildings and supporting site facilities. Coordinate brownfield tie-ins, utility outages, and operational interfaces with site stakeholders. Ensure infrastructure projects are executed, and long-term site objectives....Temporary workFor contractors