Senior Site Reliability Engineer
Zingtree
About Zingtree Zingtree is the next-generation intelligent process automation platform reimagining customer experience operations for the world's top support leaders. With 500+ customers, including Optum, Corpay, Sony, SharkNinja, and Allianz, we transform self-service, surface automation opportunities, and turn every agent into an expert. The Role We're hiring a Senior DevOps / Platform Reliability Engineer to own the platform that powers our agentic CX product. You'll build the CI/CD, infrastructure, and observability backbone that enables us to ship multi-agent systems safely to enterprise customers.If you want to operate a production AI platform and use AI to help operate it, this role is for you.In this role, you will collaborate with development, operations, and infrastructure teams to automate and streamline processes, build and maintain tools for deployment, monitoring, and operations, and troubleshoot issues across development and production environments.\n What You'll DoOwn and evolve CI/CD pipelines using GitHub Actions and OIDC-based authentication for microservices and agentic workloads, with safe, fast, and reversible deployments.Automate infrastructure provisioning using Infrastructure as Code (IaC) tools such as Terraform and CloudFormation.Operate and scale our Kubernetes platform (EKS + Argo CD), including autoscaling, ingress, external-dns, cert-manager, External Secrets Operator, backups, runtime guardrails, and multi-tenant isolation for enterprise customers.Manage the edge and network perimeter, including Cloudflare (CDN, WAF, Bot Management, DDoS protection, Zero Trust / Access), CloudFront, API Gateway, ALB/NLB, Route 53, and network security controls.Operate the data and event tier, including Aurora MySQL, ElastiCache/Redis, S3, and MSK (Kafka), with responsibility for backups, point-in-time recovery (PITR), and multi-AZ disaster recovery aligned to defined RTO/RPO objectives.Build and maintain Lambda workloads where event-driven or serverless architectures are the right fit.Build observability as a product using Prometheus, Grafana, and OpenTelemetry, including telemetry for LLM and agentic systems such as token cost, tool-call latency, evaluation signals, and prompt/version tracking.Strengthen our security and compliance posture for SOC 2 and HIPAA, including least-privilege IAM, SCPs, secrets management, SAST/DAST, dependency and container scanning, image signing, AWS Config, Security Hub, GuardDuty, Inspector, and evidence automation.Drive FinOps initiatives, including tagging standards, Savings Plans and Reserved Instances, per-tenant and per-workload cost attribution, and LLM cost controls.Build and evolve our AI-native DevOps capabilities (see section below).Partner with engineering teams to define platform standards, service templates, deployment best practices, and operational SLOs.Monitor system performance and ensure reliability, scalability, and security across infrastructure and services.Collaborate with software engineering teams to support continuous integration and continuous delivery best practices.Document infrastructure, deployment processes, and operational standards to support knowledge sharing across the team.Agentic AI in DevOpsYou'll help define how Zingtree uses agentic AI to operate and improve our platform using modern AI operational practices. Responsibilities include: Design and operate auto-remediation agents for common production toil such as certificate rotation, noisy pods, infrastructure drift, and flaky CI pipelines, with human-in-the-loop (HITL) controls for any destructive or customer-impacting actions.Use LLMs for incident triage and root cause analysis, including log and trace summarization, signal correlation, and first-draft postmortems that are always reviewed by humans.Connect AI agents to internal systems through the Model Context Protocol (MCP), including GitHub, Jira, PagerDuty, AWS, Kubernetes, Terraform, and related platforms, using scoped credentials, audit logging, and allow-listed access.Apply AI-driven observability techniques, including anomaly detection on metrics, LLM-based log clustering, and alert deduplication and summarization on top of Prometheus and OpenTelemetry.Establish operational guardrails such as prompt/version pinning, evaluation frameworks for agent behavior, cost and rate-limit controls, policy-as-code (OPA/Conftest) for AI-generated infrastructure changes, and clearly defined blast-radius controls.Define best practices for AI coding assistants such as GitHub Copilot, Claude, and Amazon Q in infrastructure repositories, including review workflows, prompt design, and restrictions on auto-merged changes.Treat AI components as production systems with SLOs, observability, on-call readiness, runbooks, and rollback strategies for agents and prompts.About You Required Qualifications 5+ years of experience in DevOps, SRE, or Platform Engineering operating production systems on AWS.Strong experience with CI/CD pipelines and tools such as GitHub Actions, GitLab CI, Jenkins, or CircleCI.Hands-on experience operating production EKS environments, including autoscaling, ingress, secrets management, and cluster upgrades.Strong AWS networking experience, including multi-account VPC design, subnets, routing, security groups, NACLs, Route 53, ACM, and load balancers.Deep experience with Terraform and GitHub Actions, ideally using OIDC-based cloud authentication.Experience with Aurora/RDS MySQL, Redis (ElastiCache), and S3, including backups, PITR, migrations, and lifecycle management.Strong observability experience using Prometheus, Grafana, and OpenTelemetry.Experience operating Argo CD at scale.Experience with Infrastructure as Code tools such as Terraform, CloudFormation, or Ansible.Experience managing Cloudflare services including WAF, Bot Management, Rate Limiting, and Zero Trust / Access, along with CloudFront.Experience operating Kafka/MSK at scale, including topics, consumer groups, and schema registries.Experience with Lambda and event-driven architectures.Comfortable working with Python, Bash, and Linux systems.Strong understanding of security best practices across IAM, KMS, secrets management, networking, and software supply chain security.Familiarity with vulnerability scanning and compliance tooling. Nice to Have Experience operating LLM or ML workloads in production, including LiteLLM, Bedrock, pgvector, prompt caching, or evaluation systems.Experience building or integrating MCP servers or deploying agent frameworks such as LangGraph or CrewAI in production environments.How We WorkWe bias toward automation over toil. If you do it twice, script it. If it pages twice, fix it.We're a small team with high ownership. You'll help define standards, not just follow them.Humans stay in the loop for anything risky. AI accelerates decision-making but does not replace judgment.We value blameless incident reviews, documented decisions, and fast feedback loops.What We OfferCompetitive compensation packagesComprehensive health benefits: 100% of employee premiums covered75%–80% of dependent premiums covered for most health, dental, and vision plans401(k) plans to support retirement planning (no employer matching currently)Paid parental leaveUnlimited PTOFlexible remote work from anywhereUp to $200/month co-working reimbursementHome office stipend: Up to $500 for home office setup$100/month for internet, phone, and related expensesZingtree Values Lead with Action We are doers. We move quickly with purpose, take smart risks, learn fast, and focus on outcomes that benefit our customers and the business. People Really Matter We win as a team. We care deeply about our customers and employees, helping each other achieve professional growth and meaningful impact. Ownership Leads to Results When we commit, we deliver. We operate with integrity, accountability, and high standards. Expertise Creates Value We are learners. We continuously grow our knowledge, share expertise, and apply it to create meaningful results. Transparency Builds Trust We communicate openly, honestly, and respectfully. We share information that matters and build trusted relationships through clarity and empathy.\n
- ...alongside some of the most experienced and innovative leaders and engineers in the field. Where we work Headquartered in Amsterdam and... ...vision, audio, and emerging multimodal architectures — fast, reliable, and effortless to deploy at massive scale. To deliver on that...Senior
$141k - $208k
...be a part of our journey! About the role We are committed to providing our customers with reliable and secure services so we are expanding our central Site Reliability Engineering team. You will be responsible for building and leading processes to ensure the reliability...SeniorLocal areaRemote workHome officeFlexible hours$127k - $249k
...Eastern or Central time zones. We are looking for an experienced Senior Engineer for our SRE, Atlas team to support, maintain and grow the... .... Role Overview We are seeking a talented Site Reliability Engineer (SRE) with a strong infrastructure background. This...SeniorFull timeLocal areaRemote workWorldwideFlexible hours- ...itD is seeking a Site Reliability Engineer to develop and enhance automation solutions that improve the reliability, scalability, and operational efficiency of large-scale cloud infrastructure. The ideal candidate will bring hands-on experience in site reliability engineering...SuggestedWork experience placementRemote work
$180k - $200k
...We are seeking a Site Reliability Engineer (SRE) to play a crucial role in enhancing observability, performance, and reliability, making a significant impact across the federal government. The ideal candidate will drive continuous improvements, ensuring robust and reliable...SuggestedContract workRemote work3 days per week- ...Job Details: Lead Site Reliability Engineer The Lead Site Reliability Engineer is a senior technical leader responsible for the reliability, availability, and operational excellence of a cloud-based infrastructure and distributed platform. This role owns uptime,...
- ...desktop applications to modern web platforms based in the cloud. · Make that user experience world-class by working with process engineers and software quality assurance to ensure low defect rates. · Execute appropriate coding practices that align with team standards...SeniorFull time
$113.89k - $187.1k
...at Comcast. (In most cases, Comcast prefers to have employees on-site collaborating unless the team has been designated as virtual due... ...SummaryThe Comcast Cloud Team is seeking an OpenStack Platform Engineer to help design, test, optimize and deploy our multi-region private...SeniorFull timeWork at officeRemote workWorldwide- Philadelphia, Pennsylvania100% RemoteDirect Hire$180k - $215kThere is an exciting opportunity for a talented Senior C++ Engineer to join a growing company focused on cutting-edge development within the gaming industry. This is a fully remote, full-time position. To excel...SeniorFull timeRemote work
$128k - $187k
...Business Insider. Artera is seeking a highly skilled Senior Software Engineer to join our dedicated team, focusing on federal government... ...ObjectScript, with a commitment to development excellence and reliable client support. As a Senior Software Engineer, you’ll...SeniorFull timeTemporary workSummer workSummer holidayLive inCurrently hiringWork at officeLocal areaRemote workFlexible hours3 days per week- American Heritage Credit Union, a $5.3+ billion credit union, has an immediate opening for a Senior Software Developer in our Information Services Department. This position will analyze business and system requirement needs/documentation to design develop, implement and...SeniorImmediate start
- Philadelphia, PAOnsiteFull Time$100k - $145kA well-established software company is hiring a Senior Software Engineer to join a growing product-focused engineering team. This role functions more like a tech lead, with involvement across the full software development lifecycle...SeniorFull time
$1,000 per month
...edge technology and customer-centric solutions.OverviewAs a Senior Backend Engineer on the Trust Platform team, you'll play a pivotal role in... ...coding practices. Experience with designing resilient and reliable systems that meet high security and regulatory standards.Nice...SeniorTemporary workWork at officeImmediate startRemote workFlexible hours- ...platform architecture with a focus on security, cost-efficiency, and operational excellence. The position requires collaboration with engineering, data, and product teams within an Agile environment to improve platform capabilities. Success is measured by the robustness,...SeniorFull timeTemporary workPart timeWork experience placementLocal areaFlexible hours
- Role SummaryNuuly is hiring a Senior Software Engineer to join our Technology team. Building performant scalable cloud service environments is... ...response and post-mortem processes to improve system reliability. You should be adept at picking up new technologies and patterns...Senior
- ...critically about backlog prioritizationPractical knowledge of software engineering best practices, including Agile development methodologies,... ...well-being for you and your family.WHAT YOU'LL DOAs a Senior Software Engineer with the Life Sciences Data and Analytics Center...SeniorApprenticeshipEasy work
$125.9k - $148.1k
...value your ideas.We are seeking a highly skilled Full-Stack Senior Software Engineer to join our team. In this role, you will design, develop,... ...’s next.Job ResponsibilitiesDevelop and maintain a secure, reliable, and scalable, and efficient platform spanning back-end persistence...SeniorFull timeContract workLocal areaFlexible hours$183.8k - $263.6k
...orchestration, and secure service integration. You will work closely with engineers across control plane, data plane, and platform teams to deliver... ..., and basic life insurance. Please see the Cisco careers site to discover more benefits and perks. Employees may be eligible...SeniorFull timeTemporary workLocal areaRemote workFlexible hours- Role SummaryNuuly is seeking a Senior DevOps Engineer to join our growing Platform Engineering team. In this role, you’ll be instrumental in... ...developers, SREs, and security teams to ensure our systems are reliable, performant, and well-architected for scale. You’ll also...Senior
- Philadelphia, PAOnsiteFull Time$100k - $145kA well-established, purpose-driven financial services company is hiring a Senior Developer to join small backend engineering team. This team builds and maintains a product with a strong quantitative core, working through complex...SeniorFull time
$124.2k - $207k
...educational attainment and/or training. Job Title: Senior Software EngineerLocation: Lake Forest,... ...where you’ll partner with architects and engineers to solve complex technical challenges... ...office, with an expectation of being on-site 50% of your working hours to support...SeniorFull timeWork at officeLocal areaRemote workFlexible hours- ...see the progress at role sits within OIT’s reimagined Software Engineering group. Created in 2019, OIT Software Engineering team is a... ...tooling and are a highly collaborative, productive team. As a Senior Software Engineer, you’ll join the City’s internal development...SeniorWork at office
- Responsibilities:Design, develop, and maintain real-time, fault-tolerant C/C++ applications on Linux platformsMigrate legacy GUIs from Motif/X11 to modern toolkits (GTK, Qt, EFL) using Wayland protocolsWrite Bash scripts for build automation, deployment routines, and system...SeniorContract work
- ...and tomorrow.Section 1: Position SummaryWe are seeking a Senior DevOps Engineer with 10+ years of hands‑on experience designing, building,... ...will lead platform standardization, progressive delivery, reliability engineering, and security‑by‑design to enable high‑quality...SeniorFull time
- We are looking for a Senior DevOps Engineer to join our PISA DevOps team in Alexandria, VA. This is an amazing opportunity to work on several growing products in the Clarivate Intellectual Property Group. We would love to speak with you if you have experience in providing...SeniorPermanent employmentFull timeWork at officeWorldwide2 days per week
$1,000 per month
...to build the future of inclusive finance through cutting-edge technology and customer-centric solutions.We are looking for a Senior iOS Engineer to help architect our Mobile Apps for the next 2-3 years of growth. The ideal candidate should have experience leading the design...SeniorTemporary workWork at officeImmediate startRemote workFlexible hours- ...Senior Level Java Developer Solvepoint is looking to add to its team of extraordinary Senior Level Java Developers qualified individuals ready to be compensated and challenged according to their abilities. Position: Senior Level Java Developer Status: Full Time...SeniorFull timeImmediate start
- ...rapidly transform the way their clients do business. As a Senior AI Engineer, you will understand how AI is positioned to meet business... ...understanding of design for scalability, performance, and reliability in Azure Ability to support go-to-market content with technical...SeniorFull time
$120.7k - $188.73k
ResponsibilitiesNoblis MSD is seeking a Senior Java Developer to support a large-scale software... ...Division (NSWCPD) by providing engineering, software development, systems integration... ...visiting the Benefits page on our Careers site.Compensation at Noblis is determined by various...SeniorFull timeContract workPart timeLocal areaRemote work- ...PennsylvaniaOnsiteDirect Hire$130k - $140kA rapidly growing organization is growing out their Data Center team and looking to hire a Senior Solution Engineer to focus on data center project work. They are looking for someone with extensive VMware experience and who has strong...SeniorFull timeWork at officeFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!
- remote senior business analyst Camden, NJ
- senior leadership Camden, NJ
- senior grant accountant Camden, NJ
- senior storage engineer Camden, NJ
- senior manager automotive Camden, NJ
- senior magento developer Camden, NJ
- senior application administrator Camden, NJ
- senior consulting engineer Camden, NJ
- remote senior project manager Camden, NJ
- senior performance engineer Camden, NJ



