Get new jobs by email
$78.4k - $129.4k
ASM Research, An Accenture Federal Services Company is looking for a Mid-level Root Cause Engineer (RCA) to analyze recurring incidents affecting IT services. The successful candidate will gather data, facilitate discussions, and translate findings into actionable insights...Suggested$70 per hour
...applications using Dynatrace. Configure dashboards, alerts, and observability solutions. Investigate production issues and perform root cause analysis. Support application, API, integration, and infrastructure monitoring. Respond to incidents and participate in an on-...SuggestedContract workImmediate startRemote work$80k - $140k
...Lead incident and problem management - Troubleshoot production issues across all layers, participate in on-call rotation, and drive root cause analysis and corrective actions. Drive continuous improvement - Identify opportunities to simplify, automate, and modernize...SuggestedFlexible hoursShift work$60 - $70 per hour
...deployment patterns, operational procedures, and troubleshooting guidance. Participate in production support, incident response, and root-cause analysis as appropriate. Independently own technical work and drive complex problems through resolution with limited...SuggestedHourly payContract workRemote work- ...improvement, service resilience, and uptime management. Incident Management & RCA - Major Incident Management (MIM), outage tracking, root cause analysis, problem management, and MTTR reduction. Failure Analysis & Risk Assessment - FMEA, risk quantification,...SuggestedFor contractors
$100k - $120k
...Origami's time to resolution and advancing overall site reliability and scalability. This person participates in efforts to identify root causes during post-incident investigations, while also identifying preventative measures to minimize future disruptions. They also...SuggestedFull timeTemporary workWork experience placementRemote workFlexible hours- ...development lifecycle. Support containerized workloads and cloud-native applications. Troubleshoot production issues, perform root cause analysis, and implement long-term reliability improvements. Optimize deployment strategies, release automation, and...Suggested3 days per week
- ...applications, infrastructure, and cloud services interact in production; demonstrated ability to troubleshoot production issues, perform root cause analysis, and drive long-term reliability improvements. ~-Proficiency with Python, Bash, or similar scripting languages;...Suggested
- ...Incident Management & Support - Implement and enhance system monitoring, alerting, and observability. - Lead incident response, root cause analysis, and postmortem reviews. - Drive continuous improvement and oversee break/fix operations. Security & Compliance...Suggested
- ...tasks, incident reduction, and proactive reliability improvements. Monitor platform health, troubleshoot complex issues, and lead root cause analysis efforts to minimize downtime and improve system resiliency. Collaborate closely with engineering, platform, and...SuggestedRemote work
$114.3k - $235.32k
...Building and supporting CI/CD automation and deployment workflows using GitHub Actions Participating in incident response, root cause analysis, and post-incident improvement initiatives Reducing operational toil through scripting, tooling, and process automation...SuggestedWork at officeLocal areaRemote workRelocationRelocation package- ...infrastructure and platform automation ~ Strong scripting skills in Python, Bash, or similar languages ~ Proven troubleshooting and root cause analysis skills in complex distributed systems ~ Excellent written and verbal communication skills ~ Bachelor's degree...SuggestedContract workRemote work
$70 - $80 per hour
...they impact users. Respond to production incidents, participate in on-call rotations, and lead post-incident reviews to drive root cause analysis and reliability improvements. Collaborate with software engineering and security teams to ensure new services and...SuggestedHourly payContract work- ...Implement and enhance monitoring, alerting, and observability frameworks for proactive issue detection. Lead incident response, root cause analysis (RCA), and postmortem reviews. Drive continuous improvement by identifying systemic issues and implementing...Suggested
- ...support for production applications and platforms. Monitor system health, troubleshoot incidents, restore service, and perform root cause analysis. Participate in on-call support and production change activities. Platform Support (Kubernetes & Kafka)...SuggestedRemote work
- ...security to translate business needs into technical designs that align with organizational goals and societal impact. Perform root cause analysis for incidents using observability data and logs from Splunk dynatrace and other tools to reduce recurrence and improve...Permanent employmentContract workWork experience placementRemote work
$20 per hour
...perform performance tuning and implement long-term stability improvements. Respond to and resolve production incidents; perform root cause analysis and drive corrective actions through blameless postmortems. Rotate through the team's on-call schedule to keep...Permanent employmentFull timeImmediate startRelocation packageFlexible hoursWeekend work- ...development productivity, troubleshooting, documentation, and automation. Participate in production support, incident response, root cause analysis, and continuous improvement initiatives. Required Qualifications ~8+ years of experience in Site...Contract workLocal areaRemote work
- ...including metrics, logging, tracing, alerting, and production diagnostics. Experience with production troubleshooting, incident response, root-cause analysis, and operational readiness. Experience with Terraform or another Infrastructure as Code technology. Experience...Long term contractLocal areaImmediate startRelocation
$65.12 per hour
...protocols. Implement and enhance monitoring, alerting, and observability for proactive detection. Lead incident response, root cause analysis, and postmortems with corrective actions. Oversee break/fix operations to ensure timely resolution and low business...Contract work3 days per week- ...and repetitive processes using scripting and Infrastructure as Code (IaC). Lead incident response activities, troubleshooting, root cause analysis (RCA), and post-incident reviews. Collaborate with development, infrastructure, and platform teams to improve system...
- ...~8 Required Experience defining and managing SLIs, SLOs, and error budgets ~8 Required Familiarity with incident management, root cause analysis (RCA), and postmortems ~8 Required Experience integrating security and compliance into operational workflows...Contract workLocal areaRemote work
- ...and operational tools through secure conversational interfaces. Design intelligent remediation workflows for incident detection, root cause analysis, log analysis, and operational troubleshooting. Develop secure Infrastructure-as-Code automation using Terraform...Contract work
- ...Kubernetes clusters (pods, resources, scaling behavior) Linux OS tuning (CPU, memory, I/O, ulimits, networking) Identify root causes and propose clear, actionable engineering solutions. Distributed Systems Design Design, review, and influence high...
- ...resiliency, performance, and operational excellence. Define and drive adoption of SLOs, SLIs, Error Budgets, Incident Management, and Root Cause Analysis. Drive automation initiatives that reduce operational overhead and improve service reliability. Serve as...Contract work
- ...severity production events, coordinating cross-functional teams to restore services and minimize customer impact. Perform and facilitate Root Cause Analysis (RCA) activities, ensuring corrective and preventive actions are identified, prioritized, and implemented. Drive...
- ...workload modeling, benchmarking, forecasting, and scale analysis. Observability & Monitoring metrics, logs, traces, monitoring, and root-cause analysis. Load Testing & Benchmarking performance testing and evidence-based capacity decisions. Messaging Systems...
- ...and peak business events. Implement and improve monitoring, observability, alerting, and service health dashboards. Lead Root Cause Analysis (RCA) and problem management activities. Identify reliability risks and implement proactive solutions to improve...Long term contract3 days per week
- ...platform instability, and production incidents across infrastructure, platform, and application layers Conduct incident response, root cause analysis, and post-incident remediation to drive continuous improvement Maintain the integrity and security of servers,...Remote workFlexible hours
- ...Qualifications ~ Experience conducting maintenance on advanced production machinery. ~ Demonstrated examples of improvements and root cause analysis. ~ Strong hydraulic, pneumatic, mechanical, sensor, automation, and industrial skills. ~2+ years of high-level...Full timeWork at officeImmediate startFlexible hours