SRE
Mondo
DevOps Engineer – Platform Operations
We are seeking an experienced DevOps Engineer to support cloud-based platforms and services within a highly secure enterprise environment. This position is heavily focused on production operations, troubleshooting, CI/CD pipeline support, infrastructure-as-code, Kubernetes, and operational reliability.
The ideal candidate has strong hands-on experience troubleshooting Terraform and Jenkins in production environments and is comfortable independently investigating customer and internal support issues before escalating them.
This is an excellent opportunity for an experienced DevOps professional who enjoys solving complex production problems, working across cloud and container environments, and improving the reliability of enterprise platforms.
Responsibilities
- Troubleshoot customer and internal production tickets, identifying root causes and resolving issues before escalation when possible.
- Monitor platform and application environments and investigate logs using tools such as Grafana, Dynatrace, or similar monitoring/APM platforms.
- Maintain and troubleshoot existing Jenkins CI/CD pipelines, ensuring deployments and automation remain reliable.
- Use Terraform to support infrastructure changes, resolve configuration drift, and maintain consistency across environments.
- Support containerized applications and services running on Kubernetes and AWS EKS.
- Deploy and maintain workloads using Helm.
- Automate configuration management and operational activities using Ansible.
- Support secure configuration and secrets management using technologies such as HashiCorp Vault.
- Diagnose Kubernetes pod, service, deployment, and infrastructure issues.
- Perform root-cause analysis and resolve production defects.
- Participate in an on-call rotation and respond to production incidents within established SLAs.
- Support occasional scheduled after-hours maintenance windows.
- Review technical documentation and identify opportunities to improve processes, procedures, and knowledge sharing.
- Collaborate with engineering, security, cloud operations, data, and application teams.
- Support platforms and services related to SAP Analytics Cloud and other enterprise analytics environments.
Required Qualifications
5 years of DevOps, platform engineering, software engineering, or related experience.
Strong hands-on experience with Terraform, including troubleshooting production infrastructure issues—not simply exposure to the tool.
Strong experience maintaining and troubleshooting Jenkins CI/CD pipelines.
Hands-on experience with Kubernetes and AWS EKS.
Experience with Ansible or comparable configuration-management automation.
Previous experience participating in a production on-call rotation.
Strong AWS cloud knowledge, including familiarity with compute, networking, storage, and IAM.
Experience troubleshooting complex production systems using logs, monitoring tools, and systematic root-cause analysis.
Ability to work independently and take ownership of technical issues before escalating.
Strong written and verbal communication skills.
Strong problem-solving skills and a customer-service mindset.
Bachelor's degree in Computer Science, Information Technology, Computer Engineering, or a related field, or equivalent applicable professional experience.
Preferred Qualifications
SAP Analytics Cloud (SAC) experience.
SAP BTP experience.
Experience with enterprise BI or analytics platforms such as Tableau, Power BI, or SAP BW.
Experience with HashiCorp Vault.
Experience with Dynatrace, Grafana, or comparable monitoring/APM tooling.
Experience with Helm.
Experience with GitLab CI, GitHub Actions, or other CI/CD platforms.
Experience working in Agile delivery environments.
Experience supporting SaaS, analytics, or other large-scale cloud-native platforms.
On-Call Expectations
This position participates in a rotating production support schedule. Engineers may be on call for one week at a time and are expected to respond to incidents within approximately 30 minutes. The team also performs occasional scheduled after-hours maintenance, including quarterly windows that may occur between approximately 12:00 AM and 2:00 AM.
What We're Looking For
We're looking for someone who can hit the ground running in a production DevOps environment. You should be comfortable investigating an issue from the initial ticket through logs, infrastructure, pipelines, and Kubernetes workloads to determine the underlying cause.
Successful candidates are curious, resourceful, collaborative, and comfortable working independently. The environment requires patience during onboarding and a willingness to review existing documentation, learn the platform, and suggest improvements.
Clear and concise communication is especially important. Candidates should be able to explain technical problems and solutions confidently without unnecessarily overcomplicating them.
- ...company, who manages all applications and next steps. Our partner is looking for a Site Reliability Engineering Team Lead (Principal SRE) based in the United States. This Principal-level role owns the reliability, availability, and operational health of a global cloud...SuggestedRemote job
- ...— 5,000+ transactions before go-live — through auto-remediation, capacity planning, and actionable dashboards. Build specialized SRE agents using Cursor AI • Design and ship AI agents for incident triage, log analysis, and root-cause investigation (to name a few...SuggestedRemote work
$53 per hour
...strengthen the reliability and visibility of Shopmonkey's production infrastructure. This is a hands-on contract role for an experienced SRE who has built production-grade observability systems—not simply operated tools configured by others. You will take ownership of...SuggestedHourly payContract workRemote workDay shift- ...SRE Human Power BG is an HR agency that offers consultations and recruitment for some of the best companies in Bulgaria. Our client is an IT services and solutions provider, primarily engaged with Implementation and Maintenance of SAP and other business information...SuggestedRemote workShift work
- ...0 professionals from Eastern Europe, North and South America, Armenia, Georgia and Kazakhstan. We are looking for an experienced SRE Engineer for a long-term project. The customer is an American company. The project is a mobile app and API service that lets users...SuggestedFull timeRemote workFlexible hours
- ...and operational automation. Define operational standards, governance, and guardrails. Drive adoption of AI-enabled DevOps and SRE practices. Improve operational efficiency and reliability through automation and AI. Serve as a technical advisor and escalation...Remote work
- ...SRE Engineer Hi All, We have an immediate requirement on "SRE Engineer" @ Remote till Covid. So please share your suitable resumes to ****@*****.*** (***) ***-**** Ext 205. Role: SRE Engineer Location: Remote till Covid Type: C2C Job Description: ~9+ Yrs....Immediate startRemote work
- ...Lead SRE Engineer This position is a lead SRE role. Carnival Cruise is targeting someone with high level automation skills in at least one scripting or programming language, no preference. This person will be the technical leader in handling issues within production...Ongoing contractFor contractorsRemote work
- ...Engineering Manager for SRE Team Parity is one of the world's most experienced core blockchain infrastructure companies, having built and pioneered some of the most advanced technologies in the blockchain sector. Parity was founded by Dr. Gavin Wood, co-founder and...Remote work
- ...Senior SRE Location: Remote Job Type: Contract Lead SRE efforts to establish and maintain platform reliability, API governance, and scalable infrastructure for high-availability services. Collaborate across engineering teams to implement SRE practices, error budgets...Contract workRemote work
- ...SRE Qualifications • Bachelor degree in Computer Science or related technical field involving coding (e.g., physics or mathematics), or equivalent practical experience, within Security and/or Enterprise Monitoring Context, REQUIRED. • Experience working through...Remote work
- ...RFCs and internal guides covering process improvements, reliability patterns, and best practices Support, and guide engineers on SRE related topics Partner with product-development teams to identify, and eliminate friction Skills you should HODL ~3+...Local areaRemote work
- ...Citizenship Requirement: U.S. Citizen (Required) Summary of Position Role/Responsibilities The Site Reliability Engineer (SRE) for Azure Government & Infrastructure plays a critical role in ensuring the resilience, security, and scalability of the cloud environments...Full timeRemote workMonday to Friday
- ...Site Reliability Engineer (SRE) Location: Remote Contract Length: 12 months w/ high likeliness of conversion or extension to FTE Contract Type: W2 Overview: The selected candidate will be responsible for the administration and engineering work concerning...Contract workCasual workRemote work
- ...SRE / DevOps Engineer Canada / Remote 6+ Months Contract Position Requirements Collaborate closely with Development teams to improve services and operational targets. Implement and maintain CI/CD practices using tools like Sonar. Manage application integration...Contract workRemote work
- ...Lead SRE Position in India This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Lead SRE based in India. This role offers the opportunity to lead and scale a high-performing Site Reliability...Remote work
- ...SRE Engineer – Very Critical SFO CA (100% Remote) End Client: Airbnb Exp : 6 to 8 years only, don’t get profiles beyond 8+ years Must Have: Strong Java coding, Automation Scripting and Python and AWS Consultant who started his career as Java developer...Remote work
$100k - $180k
...Site Reliability Engineer (SRE) - Remote Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States. This is a fantastic opportunity to join an established and...Full timeH1bLocal areaImmediate startRemote workVisa sponsorship- ...About the Role: We are looking for a Senior Site Reliability Engineer (SRE) to help modernize large-scale infrastructure and improve the reliability, scalability, and operational excellence of critical production systems. In this role, you will lead OS modernization...Remote work
- Site Reliability Engineer Amplifire is seeking a Site Reliability Engineer to improve the reliability, scalability, performance, and operational efficiency of our cloud-based platform. Working alongside DevOps engineers within the Platform Operations team, this role...Remote work
- ...Site Reliability Engineer (SRE) Location: Remote Shift Timings: 5:30 PM to 3:00 AM IST to ensure support for global operations. Job Description: We are seeking a skilled Site Reliability Engineer (SRE) to join our dynamic team. The ideal candidate will have...Remote workShift work
- ...Prometheus, Grafana, ELK) • Scripting experience in Python, Bash, or similar • Experience in Linux system administration DevOps / SRE Expertise • Strong understanding of DevOps principles and SRE practices • Experience implementing high-availability, fault-...Remote work
- Site Reliability Engineer At Rocket.net, reliability, performance, and customer experience are at the center of everything we build. We are looking for a Site Reliability Engineer to help maintain the health, stability, and performance of our hosting platform while ...Remote workShift work
- ...layers. ~ Experience developing synthetic monitors and other test scripts to validate application health ~ Understanding of SRE principles, including observability, automation, SLIs/SLOs, availability, and continuous service improvement. ~ Experience participating...Remote work
- ...Site Reliability Engineer (SRE) Location: Remote (Secaucus, NJ) Duration: Contract Experience: 7+ Years Job Description 4+ years of experience with multiple APM tools and extensive experience with Dynatrace 4+ years of experience executing software load and performance...Contract workWork experience placementRemote work
- Sr Application Performance and Observability Engineer At Sequoia Connect, we are a Talent-First Technology Ecosystem that redefines how elite professionals interact with the global digital landscape. We move beyond traditional models to act as a catalyst for the top...Remote workWorldwide
$100k - $180k
...Site Reliability Engineer (SRE) Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States. This is a fantastic opportunity to join an established and well-respected...Full timeH1bImmediate startRemote workVisa sponsorship- ...Benefits to more than 5.5 million users in 85,000 companies in France and Brazil. Your role as a Senior Site Reliability Engineer (SRE) centers around creatively solving problems, ensuring a balance between speed, reliability, pragmatism and excellence. It's less...Remote work
- ...Site Reliability Engineer (SRE) We are seeking an experienced Site Reliability Engineer (SRE) with strong expertise in Dynatrace to join our growing engineering team. The ideal candidate will be responsible for ensuring the reliability, scalability, performance, and...Remote work
- ...SRE Key Responsibilities • Design and manage multi-account AWS infrastructure (VPC, Route Tables, EC2, ECS, EKS 1.33, RDS, DynamoDB, Elastic ache Valley, S3, Transit Gateway, Resource Access Manager, Lambda, CloudFormation, AWS Backup) • Configure load balancing...Remote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to SRE. Be the first to apply!

