Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer

Zingtree

Senior DevOps / Platform Reliability EngineerZingtree is the next-generation intelligent process automation platform reimagining customer experience operations for the world's top support leaders. With 500+ customers, including Optum, Corpay, Sony, SharkNinja, and Allianz, we transform self-service, surface automation opportunities, and turn every agent into an expert.We're hiring a Senior DevOps / Platform Reliability Engineer to own the platform that powers our agentic CX product. You'll build the CI/CD, infrastructure, and observability backbone that enables us to ship multi-agent systems safely to enterprise customers.If you want to operate a production AI platform and use AI to help operate it, this role is for you.In this role, you will collaborate with development, operations, and infrastructure teams to automate and streamline processes, build and maintain tools for deployment, monitoring, and operations, and troubleshoot issues across development and production environments.What You'll DoOwn and evolve CI/CD pipelines using GitHub Actions and OIDC-based authentication for microservices and agentic workloads, with safe, fast, and reversible deployments.Automate infrastructure provisioning using Infrastructure as Code (IaC) tools such as Terraform and CloudFormation.Operate and scale our Kubernetes platform (EKS + Argo CD), including autoscaling, ingress, external-dns, cert-manager, External Secrets Operator, backups, runtime guardrails, and multi-tenant isolation for enterprise customers.Manage the edge and network perimeter, including Cloudflare (CDN, WAF, Bot Management, DDoS protection, Zero Trust / Access), CloudFront, API Gateway, ALB/NLB, Route 53, and network security controls.Operate the data and event tier, including Aurora MySQL, ElastiCache/Redis, S3, and MSK (Kafka), with responsibility for backups, point-in-time recovery (PITR), and multi-AZ disaster recovery aligned to defined RTO/RPO objectives.Build and maintain Lambda workloads where event-driven or serverless architectures are the right fit.Build observability as a product using Prometheus, Grafana, and OpenTelemetry, including telemetry for LLM and agentic systems such as token cost, tool-call latency, evaluation signals, and prompt/version tracking.Strengthen our security and compliance posture for SOC 2 and HIPAA, including least-privilege IAM, SCPs, secrets management, SAST/DAST, dependency and container scanning, image signing, AWS Config, Security Hub, GuardDuty, Inspector, and evidence automation.Drive FinOps initiatives, including tagging standards, Savings Plans and Reserved Instances, per-tenant and per-workload cost attribution, and LLM cost controls.Build and evolve our AI-native DevOps capabilities (see section below).Partner with engineering teams to define platform standards, service templates, deployment best practices, and operational SLOs.Monitor system performance and ensure reliability, scalability, and security across infrastructure and services.Collaborate with software engineering teams to support continuous integration and continuous delivery best practices.Document infrastructure, deployment processes, and operational standards to support knowledge sharing across the team.Agentic AI in DevOpsYou'll help define how Zingtree uses agentic AI to operate and improve our platform using modern AI operational practices.Responsibilities include:Design and operate auto-remediation agents for common production toil such as certificate rotation, noisy pods, infrastructure drift, and flaky CI pipelines, with human-in-the-loop (HITL) controls for any destructive or customer-impacting actions.Use LLMs for incident triage and root cause analysis, including log and trace summarization, signal correlation, and first-draft postmortems that are always reviewed by humans.Connect AI agents to internal systems through the Model Context Protocol (MCP), including GitHub, Jira, PagerDuty, AWS, Kubernetes, Terraform, and related platforms, using scoped credentials, audit logging, and allow-listed access.Apply AI-driven observability techniques, including anomaly detection on metrics, LLM-based log clustering, and alert deduplication and summarization on top of Prometheus and OpenTelemetry.Establish operational guardrails such as prompt/version pinning, evaluation frameworks for agent behavior, cost and rate-limit controls, policy-as-code (OPA/Conftest) for AI-generated infrastructure changes, and clearly defined blast-radius controls.Define best practices for AI coding assistants such as GitHub Copilot, Claude, and Amazon Q in infrastructure repositories, including review workflows, prompt design, and restrictions on auto-merged changes.Treat AI components as production systems with SLOs, observability, on-call readiness, runbooks, and rollback strategies for agents and prompts.About YouRequired Qualifications5+ years of experience in DevOps, SRE, or Platform Engineering operating production systems on AWS.Strong experience with CI/CD pipelines and tools such as GitHub Actions, GitLab CI, Jenkins, or CircleCI.Hands-on experience operating production EKS environments, including autoscaling, ingress, secrets management, and cluster upgrades.Strong AWS networking experience, including multi-account VPC design, subnets, routing, security groups, NACLs, Route 53, ACM, and load balancers.Deep experience with Terraform and GitHub Actions, ideally using OIDC-based cloud authentication.Experience with Aurora/RDS MySQL, Redis (ElastiCache), and S3, including backups, PITR, migrations, and lifecycle management.Strong observability experience using Prometheus, Grafana, and OpenTelemetry.Experience operating Argo CD at scale.Experience with Infrastructure as Code tools such as Terraform, CloudFormation, or Ansible.Experience managing Cloudflare services including WAF, Bot Management, Rate Limiting, and Zero Trust / Access, along with CloudFront.Experience operating Kafka/MSK at scale, including topics, consumer groups, and schema registries.Experience with Lambda and event-driven architectures.Comfortable working with Python, Bash, and Linux systems.Strong understanding of security best practices across IAM, KMS, secrets management, networking, and software supply chain security.Familiarity with vulnerability scanning and compliance tooling.Nice to HaveExperience operating LLM or ML workloads in production, including LiteLLM, Bedrock, pgvector, prompt caching, or evaluation systems.Experience building or integrating MCP servers or deploying agent frameworks such as LangGraph or CrewAI in production environments.How We WorkWe bias toward automation over toil. If you do it twice, script it. If it pages twice, fix it.We're a small team with high ownership. You'll help define standards, not just follow them.Humans stay in the loop for anything risky. AI accelerates decision-making but does not replace judgment.We value blameless incident reviews, documented decisions, and fast feedback loops.What We OfferCompetitive compensation packagesComprehensive health benefits:100% of employee premiums covered75%–80% of dependent premiums covered for most health, dental, and vision plans401(k) plans to support retirement planning (no employer matching currently)Paid parental leaveUnlimited PTOFlexible remote work from anywhereUp to $200/month co-working reimbursementHome office stipend:Up to $500 for home office setup$100/month for internet, phone, and related expensesZingtree Values

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in Camden, NJ vacancy
  • Senior Site Reliability EngineerLocation: Exton or Philadelphia, PA (Hybrid - 3 times a week in-office)Position SummaryAre you ready to start...  ...looking for you!We are looking for a Senior Site Reliability Engineer to take on the responsibility of automating cloud-based... 
    Suggested
    Casual work
    Work at office
    Worldwide

    Bentley Systems

    Philadelphia, PA
    1 day ago
  • $130k - $150k

     ...Site Reliability Engineer (SRE) Engineer Reliability into the Systems That Move the Nation's Food Supply Who We Are US Cold owns and operates one of the most complex temperature-controlled logistics networks in North America. Every day, our systems coordinate... 
    Suggested

    USCS

    Camden, NJ
    4 days ago
  •  ...relating to any production issues; and guide and mentor junior-level engineers. Position is eligible for 100% remote work.REQUIREMENTS:...  ...everyday life.Please visit the benefits summary on our careers site for more details.Comcast is an equal opportunity workplace. We... 
    Suggested
    Full time
    Remote work

    Comcast

    Philadelphia, PA
    1 day ago
  • $1,000 per month

     ...financial data safe. We're also building the AI infrastructure our engineers use every day. When this team does its job well, engineers...  ....As a Senior DevOps / SRE Engineer on this team, you'll own reliability and deployments across our AWS and Kubernetes environment, take... 
    Suggested
    Temporary work
    Work at office
    Immediate start
    Remote work
    Flexible hours

    Creditly Corp

    Philadelphia, PA
    4 days ago
  •  ...cases, Comcast prefers to have employees on-site collaborating unless the team has been...  ...the team responsible for ensuring the reliability and health of RDK software deployments across...  ...of field issues while partnering across engineering, operations, and vendor teams to deliver... 
    Suggested
    Full time
    Work at office
    Remote work
    Worldwide

    Comcast

    Philadelphia, PA
    4 days ago
  •  ...Central time zones. We are looking for an experienced Senior Engineer for our SRE, Atlas team to support, maintain and grow the Atlas...  ...crucial workloads. Role Overview We are seeking a talented Site Reliability Engineer (SRE) with a strong infrastructure background. This... 
    Full time
    Remote work
    Worldwide

    Mongodb

    Philadelphia, PA
    a month ago
  • This position is on-site in PhiladelphiaAbout ProsciaProscia is revolutionizing pathology, the last major frontier in healthcare...  ...what’s possible in medicine.About This PositionAs a Reliability Engineer reporting to the VP, Technical Operations, you will own the reliability... 
    Work at office
    Shift work

    Proscia

    Philadelphia, PA
    2 days ago
  • $51 - $61 per hour

     ...onsite at the project, significantly reducing and/or eliminating the demands to travel. Key Responsibilities:As a Release Train Engineer, you will be responsible for facilitating Agile Release Train events and processes including communicating with stakeholders... 
    Hourly pay
    Live in
    Work at office
    Local area
    Immediate start
    Flexible hours
    Shift work

    Accenture

    Philadelphia, PA
    5 days ago
  • OverviewLutron is seeking a seasoned engineering leader with deep expertise on software and embedded systems who is passionate about teaching the craft of engineering to junior and mid-career engineers as they create world-class, smart lighting and lighting control systems... 
    Apprenticeship
    Worldwide

    Lutron Electronics

    Philadelphia, PA
    4 days ago
  •  ...Philadelphia office are requiredJob DescriptionThe IT Systems Engineer I supports the administration, maintenance, monitoring, and performance...  ...as needed.Success in this role is measured primarily by the reliability, security, and performance of the organization’s network and... 
    Work at office
    Work from home
    Visa sponsorship
    Flexible hours

    IntegriChain

    Philadelphia, PA
    3 days ago
  •  ...Job TitleBe a part of Reliability Growth Team dedicated to supporting the NGHST program and...  ...Reliability ExpertiseResponsibilitiesDeploy engineering and analytical resources across depots....  ...and ensure workload distribution per site needs.Support verification of corrective... 

    Saxon Global

    Philadelphia, PA
    4 days ago
  • $1,000 per month

     ...technology and customer-centric solutions.OverviewAs a Senior Backend Engineer on the Trust Platform team, you'll play a pivotal role in...  ...coding practices. Experience with designing resilient and reliable systems that meet high security and regulatory standards.Nice to... 
    Temporary work
    Work at office
    Immediate start
    Remote work
    Flexible hours

    Creditly Corp

    Philadelphia, PA
    3 days ago
  •  ...Time$130k - $150kJoin a growing technology-driven organization supporting mission-critical environments through modern platform engineering and cloud-native infrastructure. This full-time opportunity is ideal for an experienced Senior Platform Engineer who is passionate... 
    Full time
    Flexible hours

    Motion Recruitment

    Philadelphia, PA
    4 days ago
  •  ...at Comcast. (In most cases, Comcast prefers to have employees on-site collaborating unless the team has been designated as virtual due...  ...communications at enterprise scale.We are seeking a software engineer to help accelerate our adoption of Agentic AI and intelligent automation... 
    Full time
    Work at office
    Remote work
    Worldwide

    Comcast

    Philadelphia, PA
    1 hour ago
  • $113.1k - $154.1k

     ...production readiness standards, including backup, recovery, telemetry, runbooks, and escalation proceduresDevelop and maintain cloud engineering standards, governance processes, and onboarding frameworksReduce configuration drift and improve consistency through automation,... 
    Full time
    Contract work
    Local area
    Flexible hours

    Armanino

    Philadelphia, PA
    9 hours ago
  •  ...Senior Software Engineer, Services (this role is US based remote, all candidates must be authorized to work in the United States...  ...architectural thinking and hands-on execution, with a focus on building reliable, extensible services that support multiple use cases over time.... 
    Remote job
    Full time

    Crosslake Technologies

    Philadelphia, PA
    1 day ago
  •  ...Overview: Gopuff’s engineering team is building solutions to dramatically change the way people purchase their daily goods. We provide the modern-day solution to meet customers' immediate everyday needs with products ranging from snacks and ice cream to household... 
    Full time
    Immediate start

    Gopuff

    Philadelphia, PA
    23 hours ago
  •  ...right in.  Job Role We’re building small, highly capable engineering pods (2–3 engineers) that own problems end-to-end and move...  ...quality Iterate rapidly while maintaining a strong bar for reliability and maintainability Requirements ~7+ years of software engineering... 
    Remote job
    Full time

    Crosslake Technologies

    Philadelphia, PA
    1 day ago
  • $128k - $187k

     ...Insider. Artera is seeking a highly skilled Senior Software Engineer to join our dedicated team, focusing on federal government...  ...ObjectScript, with a commitment to development excellence and reliable client support. As a Senior Software Engineer, you’ll play... 
    Full time
    Temporary work
    Summer work
    Summer holiday
    Live in
    Currently hiring
    Work at office
    Local area
    Remote work
    Flexible hours
    3 days per week

    Artera

    Philadelphia, PA
    1 day ago
  • $120k - $140k

     ...access throughout the software development lifecycle. You will collaborate with engineering, architecture, QA, product, and UX partners to translate business and user needs into reliable, user-centered solutions. The ideal candidate combines strong engineering fundamentals... 
    Full time
    Local area
    Remote work

    Therapynotes.com

    Philadelphia, PA
    1 day ago
  • Lamps.com is actively looking for experienced developers
    Full time

    Lamps.com

    Philadelphia, PA
    1 day ago
  • Company Description System Canada resources have a broad range of skills in different technologies.  The large skill-set has been made possible by a conscious focus on strengthening our skills base. Every person selected for our team brings something new, something...
    Full time

    System Canada Technologies

    Philadelphia, PA
    1 day ago
  •  ...Full Stack Engineer, Tooling (this role is US based remote, all candidates must be authorized to work in the United States without restriction)   What we believe   In the past few years, private equity investors have invested more than a trillion dollars in... 
    Remote job
    Full time
    Immediate start

    Crosslake Technologies

    Philadelphia, PA
    1 day ago
  •  ...desktop applications to modern web platforms based in the cloud. · Make that user experience world-class by working with process engineers and software quality assurance to ensure low defect rates. · Execute appropriate coding practices that align with team standards... 
    Full time

    Nfinity Inc.

    Philadelphia, PA
    1 day ago
  • $140k - $200k

     ...– Speechify has no office. These include frontend and backend engineers, AI research scientists, and others from Amazon, Microsoft, and...  ...→ testing → release → maintenance. Ensure quality, reliability, and consistency across releases. Identify, diagnose, and resolve... 
    Full time
    Work at office

    Speechify

    Philadelphia, PA
    23 hours ago
  • $116.2k - $229.1k

    Position Summary Join our AI & Engineering team in transforming technology platforms, driving innovation, and helping make a significant impact on our clients' success. You’ll work alongside talented professionals reimagining and re-engineering operations and processes... 
    Local area
    Visa sponsorship

    Deloitte

    Philadelphia, PA
    2 days ago
  • $79.25k - $130.73k

    OverviewAs a GIS subject matter expert, you’re a natural at identifying the right analysis tools for the problem at hand. Not only do you create innovative solutions, you talk about solutions in ways that get others excited about the power of GIS technology. Join an account...

    ESRI

    Philadelphia, PA
    2 days ago
  • $79.25k - $130.73k

    OverviewAs a GIS subject matter expert, you’re a natural at identifying the right analysis tools for the problem at hand. Not only do you create innovative solutions, you talk about solutions in ways that get others excited about the power of GIS technology. Join an account...
    Local area

    ESRI

    Philadelphia, PA
    9 hours ago
  • $145k - $185k

    GFT is seeking a Senior Solutions Engineer, Asset Management  to support our GIS and Asset Management Services team! This role follows...  ...requiring occasional onsite attendance at local offices or client sites. Working on the geospatial team offers a unique opportunity to... 
    Full time
    Local area
    Remote work

    Gannett Fleming

    Philadelphia, PA
    3 days ago
  • $192.3k - $248.1k

     ...9/17/2026This position is location specific and will require in person visits to customer sites on a daily basis in Pennsylvania and New Jersey.Meet the TeamAs a Solution Engineer at Cisco, you will join a high-impact team of technical sales professionals tasked with owning... 
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    Philadelphia, PA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!