Average salary: $96,768 /yearly
More statsGet new jobs by email
$60 - $70 per hour
...deployment patterns, operational procedures, and troubleshooting guidance. Participate in production support, incident response, and root-cause analysis as appropriate. Independently own technical work and drive complex problems through resolution with limited...SuggestedHourly payContract workRemote work$70 per hour
...applications using Dynatrace. Configure dashboards, alerts, and observability solutions. Investigate production issues and perform root cause analysis. Support application, API, integration, and infrastructure monitoring. Respond to incidents and participate in an on-...SuggestedContract workImmediate startRemote work$80k - $140k
...Lead incident and problem management - Troubleshoot production issues across all layers, participate in on-call rotation, and drive root cause analysis and corrective actions. Drive continuous improvement - Identify opportunities to simplify, automate, and modernize...SuggestedFlexible hoursShift work- ...improvement, service resilience, and uptime management. Incident Management & RCA - Major Incident Management (MIM), outage tracking, root cause analysis, problem management, and MTTR reduction. Failure Analysis & Risk Assessment - FMEA, risk quantification,...SuggestedFor contractors
- ...applications, infrastructure, and cloud services interact in production; demonstrated ability to troubleshoot production issues, perform root cause analysis, and drive long-term reliability improvements. ~-Proficiency with Python, Bash, or similar scripting languages;...Suggested
- ...development lifecycle. Support containerized workloads and cloud-native applications. Troubleshoot production issues, perform root cause analysis, and implement long-term reliability improvements. Optimize deployment strategies, release automation, and...Suggested3 days per week
$114.3k - $235.32k
...Building and supporting CI/CD automation and deployment workflows using GitHub Actions Participating in incident response, root cause analysis, and post-incident improvement initiatives Reducing operational toil through scripting, tooling, and process automation...SuggestedWork at officeLocal areaRemote workRelocationRelocation package$100k - $120k
...Origami's time to resolution and advancing overall site reliability and scalability. This person participates in efforts to identify root causes during post-incident investigations, while also identifying preventative measures to minimize future disruptions. They also...SuggestedFull timeTemporary workWork experience placementRemote workFlexible hours- ...and repetitive processes using scripting and Infrastructure as Code (IaC). Lead incident response activities, troubleshooting, root cause analysis (RCA), and post-incident reviews. Collaborate with development, infrastructure, and platform teams to improve system...Suggested
- ...Support Implement and enhance monitoring, alerting, and observability frameworks for proactive issue detection. Lead incident response, root cause analysis (RCA), and postmortem reviews. Drive continuous improvement by identifying systemic issues and implementing...SuggestedContract work3 days per week
- ...Kubernetes clusters (pods, resources, scaling behavior) Linux OS tuning (CPU, memory, I/O, ulimits, networking) Identify root causes and propose clear, actionable engineering solutions. Distributed Systems Design Design, review, and influence high-performance...SuggestedLong term contractContract work
- ...Incident Management & Support - Implement and enhance system monitoring, alerting, and observability. - Lead incident response, root cause analysis, and postmortem reviews. - Drive continuous improvement and oversee break/fix operations. Security & Compliance...Suggested
- ...tasks, incident reduction, and proactive reliability improvements. Monitor platform health, troubleshoot complex issues, and lead root cause analysis efforts to minimize downtime and improve system resiliency. Collaborate closely with engineering, platform, and...SuggestedRemote work
- ...support for production applications and platforms. Monitor system health, troubleshoot incidents, restore service, and perform root cause analysis. Participate in on-call support and production change activities. Platform Support (Kubernetes & Kafka)...SuggestedRemote work
- ...resiliency, performance, and operational excellence. Define and drive adoption of SLOs, SLIs, Error Budgets, Incident Management, and Root Cause Analysis. Drive automation initiatives that reduce operational overhead and improve service reliability. Serve as...SuggestedContract work
- ...security to translate business needs into technical designs that align with organizational goals and societal impact. Perform root cause analysis for incidents using observability data and logs from Splunk dynatrace and other tools to reduce recurrence and improve...Permanent employmentContract workWork experience placementRemote work
- ...development productivity, troubleshooting, documentation, and automation. Participate in production support, incident response, root cause analysis, and continuous improvement initiatives. Required Qualifications ~8+ years of experience in Site...Contract workLocal areaRemote work
$20 per hour
...perform performance tuning and implement long-term stability improvements. Respond to and resolve production incidents; perform root cause analysis and drive corrective actions through blameless postmortems. Rotate through the team's on-call schedule to keep...Permanent employmentFull timeImmediate startRelocation packageFlexible hoursWeekend work- ...including metrics, logging, tracing, alerting, and production diagnostics. Experience with production troubleshooting, incident response, root-cause analysis, and operational readiness. Experience with Terraform or another Infrastructure as Code technology. Experience...Long term contractLocal areaImmediate startRelocation
- ...infrastructure and platform automation ~ Strong scripting skills in Python, Bash, or similar languages ~ Proven troubleshooting and root cause analysis skills in complex distributed systems ~ Excellent written and verbal communication skills ~ Bachelor's degree...Contract workRemote work
$65.12 per hour
...protocols. Implement and enhance monitoring, alerting, and observability for proactive detection. Lead incident response, root cause analysis, and postmortems with corrective actions. Oversee break/fix operations to ensure timely resolution and low business...Contract work3 days per week- ...and operational tools through secure conversational interfaces. Design intelligent remediation workflows for incident detection, root cause analysis, log analysis, and operational troubleshooting. Develop secure Infrastructure-as-Code automation using Terraform...Contract work
$70 - $80 per hour
...they impact users. Respond to production incidents, participate in on-call rotations, and lead post-incident reviews to drive root cause analysis and reliability improvements. Collaborate with software engineering and security teams to ensure new services and...Hourly payContract work- ...and peak business events. Implement and improve monitoring, observability, alerting, and service health dashboards. Lead Root Cause Analysis (RCA) and problem management activities. Identify reliability risks and implement proactive solutions to improve...Long term contract3 days per week
- ...~8 Required Experience defining and managing SLIs, SLOs, and error budgets ~8 Required Familiarity with incident management, root cause analysis (RCA), and postmortems ~8 Required Experience integrating security and compliance into operational workflows...Contract workLocal areaRemote work
- ...workload modeling, benchmarking, forecasting, and scale analysis. Observability & Monitoring metrics, logs, traces, monitoring, and root-cause analysis. Load Testing & Benchmarking performance testing and evidence-based capacity decisions. Messaging Systems...
- ...Splunk AWS CloudWatch Datadog Implement proactive monitoring and advanced alerting for payment flows. Lead incident response, root cause analysis (RCA), and post incident reviews. Drive reduction in MTTR and recurring incidents. Database & Data Layer...
$45 - $80 per hour
...(Geometric Dimensioning & Tolerancing) Experience with FMEA and dimensional tolerance stack-up analysis Ability to perform root cause analysis and corrective actions Experience collaborating with cross-functional engineering teams Strong technical problem...Hourly payRemote work$120k - $145k
...-code, deployment automation, and system observability mechanisms Participate in on-call rotation and drive incident response, root-cause analysis, and reliability improvements. Participate in designs related to developer tools and compute infrastructure....Full timeTemporary workWork experience placementInternshipWork at office3 days per week- ...not compromise the HA guarantees the platform advertises. Work on Customer Escalations with cross-functional teams to help identify root causes and provide assistance with reproductions. Troubleshoot at the linux kernel level. Diagnose failures directly on DNOS/Linux hosts...