Sr DevOps Engineer
Omni Inclusive
Job Summary
We are seeking a highly capable Senior DevOps Engineer / Platform Engineer to build, operationalize, and scale the infrastructure and deployment foundation for a strategic site-builder / network automation platform . This role will focus on creating reliable CI/CD pipelines, production-grade Kubernetes deployment patterns, managed database services, observability, environment reproducibility, secrets management, and Infrastructure as Code across development, testing, staging, and production environments.
This engineer will play a critical role in moving the platform from an early-stage, partially manual operating model into a repeatable, supportable, and production-ready DevOps model. The environment includes Kubernetes-hosted services, AWS managed services, workflow orchestration with Temporal, integration with Nautobot, Argo-based promotion flows, and the supporting tooling required for debugging, snapshotting, local development, and production support.
This is a hands-on engineering role for someone who can design the right platform patterns, implement them directly, and establish a durable operating model between development and DevOps teams.
Key Responsibilities
Platform Deployment & CI/CD
• Design, implement, and maintain CI/CD pipelines for testing, staging, and production environments.
• Build and maintain deployment workflows that support safe and seamless promotion across environments.
• Improve and maintain Argo-based deployment workflows to enable controlled release progression from test to staging to production.
• Establish baseline deployment mechanisms for the site-builder application and related services.
• Standardize Kubernetes application packaging and deployment patterns, with a strong preference toward Helm-based lifecycle management for complex services and third-party components.
• Migrate existing deployments to Helm charts where appropriate.
Kubernetes & Runtime Platform Engineering
• Support the deployment and ongoing operation of services running in Kubernetes.
• Improve runtime reliability, resiliency, and troubleshooting for distributed services operating inside shared Kubernetes clusters.
• Investigate and harden service-to-service connectivity patterns, especially for workflow components such as workers connecting to the Temporal engine.
• Partner with development teams to define production-grade runtime requirements, resource sizing, restart policies, and platform support boundaries.
Infrastructure as Code & Cloud Services
• Design and implement fully declarative Infrastructure as Code for managed cloud services, especially in AWS.
• Provision and maintain managed data services such as RDS/PostgreSQL and MongoDB-compatible document databases across all environments.
• Eliminate manual infrastructure setup where possible and replace it with reproducible, version-controlled deployment patterns.
• Prepare the platform for future scale across multiple environments and regions through repeatable IaC and GitOps-aligned practices.
Data Services, Snapshots & Developer Enablement
• Setup and maintain RDS, MongoDB, Redis/cache services , and related dependencies for all environments.
• Build tooling and operational processes for:
• production and staging database snapshots,
• restoring snapshots into development environments,
• enabling local debugging and development from realistic data states.
• Support creation of local and development environments, including Minikube-based environment-as-code approaches that mirror production behavior as closely as practical.
• Improve platform reproducibility so engineers can quickly stand up close-to-production development environments.
Workflow Orchestration & Temporal Support
• Lead the setup, deployment, and operational support of Temporal for workflow orchestration.
• Support production operations for Temporal, including troubleshooting performance issues, restarts, scaling concerns, and resource shortages.
• Establish maintainable deployment patterns for Temporal using supported packaging and lifecycle management approaches.
• Partner with engineering teams to ensure workflow platform reliability and upgradeability over time.
Observability, Reliability & Incident Readiness
• Design and maintain observability across testing, staging, and production using tools such as Prometheus and Grafana .
• Define and implement monitoring for:
• service and cluster utilization,
• CPU, memory, storage,
• IOPS / throughput metrics,
• database connections and session counts,
• cache hit / miss / coverage metrics,
• RDS and MongoDB utilization,
• service health and alerting.
• Build and maintain logging, tracing, and correlation capabilities, separated appropriately by environment.
• Create tools to support deep debugging and operational inspection, including raw database reads, cleanup of unused volumes, and emergency cache invalidation.
Security, Access & Secrets Management
• Maintain secrets management processes across environments.
• Build tooling for short-lived internal token generation and long-lived secret rotation.
• Support secure access from deployed services to active production devices and southbound systems.
• Help establish credential management patterns for southbound integrations and device-facing access.
• Partner with related teams to define safe operational limits and controls for service integrations.
External Integrations & Platform Support
• Support integration patterns with Nautobot and help define safe client-side behaviors such as rate limiting, retry/backoff, and service protection mechanisms.
• Partner with application teams to understand and mitigate integration issues such as rate limiting or request rejection.
• Support staging and testing by enabling virtual device environments where needed.
• Contribute to end-to-end acceptance testing and production readiness activities.
Operating Model & Cross-Functional Execution
• Help define an effective operating model between Development and DevOps, whether via RACI , embedded Agile delivery, or a hybrid support model.
• Support deployment readiness, incident management, environment ownership boundaries, and lifecycle responsibilities.
• Work closely with software engineering, infrastructure, application owners, and partner teams to drive production readiness and sustainable operations.
Required Qualifications
• Bachelor's degree in Computer Science, Engineering, Information Systems, or equivalent practical experience.
• 7+ years of experience in DevOps, Platform Engineering, SRE, or Infrastructure Engineering roles.
• Strong hands-on experience with Kubernetes in production environments.
• Strong experience building and maintaining CI/CD pipelines for multi-environment software delivery.
• Strong experience with ArgoCD , GitOps workflows, or equivalent deployment tooling.
• Strong experience with Helm and Kubernetes package/deployment lifecycle management.
• Experience with AWS managed services , especially RDS/PostgreSQL , document databases, and related infrastructure.
• Strong experience with Infrastructure as Code , such as Terraform and/or similar declarative tooling.
• Experience with Prometheus, Grafana , and modern observability practices.
• Experience with Redis/cache services , secrets management, and operational debugging.
• Strong Linux, networking, and distributed systems troubleshooting skills.
• Strong scripting and automation skills in one or more languages such as Python, Bash, or Go.
• Proven ability to work cross-functionally and operate effectively in environments where ownership boundaries are still evolving.
Preferred Qualifications
• Experience with Temporal deployment and production operations.
• Experience supporting developer platforms with local environment reproducibility using Minikube, kind, or similar tools.
• Experience with MongoDB / DocumentDB operations and restore workflows.
• Experience integrating with Nautobot , NetBox, or similar infrastructure source-of-truth platforms.
• Experience operating in shared-cluster environments with multi-team tenancy and constrained access models.
• Experience designing platform patterns for internal products that must scale across regions or multiple deployment footprints.
• Familiarity with network automation or infrastructure orchestration platforms is a plus.
What Success Looks Like
• CI/CD pipelines are reliable, repeatable, and support safe promotion across all environments.
• Kubernetes deployments are standardized, maintainable, and production ready.
• Managed infrastructure is defined as code rather than through manual setup.
• Temporal, databases, cache layers, and observability tooling are stable and supportable.
• Development teams can reproduce realistic environments locally for faster debugging and delivery.
• Secrets, access patterns, and operational tooling are mature enough to support production-scale operations.
• The DevOps operating model is clearly defined and enables faster deployments with less operational risk.
Scope Notes
In scope
• CI/CD and deployment foundations
• Kubernetes packaging and release management
• RDS, MongoDB, Redis/cache services
• Temporal platform setup and operational support
• Observability, alerting, and debugging tooling
• Secrets management and access enablement
• Infrastructure as Code and environment reproducibility
• DevOps / Development operational model definition
Candidate Profile
The ideal candidate is a builder-operator: someone who can establish engineering discipline where manual patterns currently exist, create durable automation for platform operations, and raise the overall maturity of the product's deployment and runtime ecosystem. This person should be equally comfortable discussing deployment architecture, writing IaC and Helm code, troubleshooting Kubernetes runtime issues, and defining how DevOps and software engineering teams work together over the full product lifecycle.
We are seeking a highly capable Senior DevOps Engineer / Platform Engineer to build, operationalize, and scale the infrastructure and deployment foundation for a strategic site-builder / network automation platform . This role will focus on creating reliable CI/CD pipelines, production-grade Kubernetes deployment patterns, managed database services, observability, environment reproducibility, secrets management, and Infrastructure as Code across development, testing, staging, and production environments.
This engineer will play a critical role in moving the platform from an early-stage, partially manual operating model into a repeatable, supportable, and production-ready DevOps model. The environment includes Kubernetes-hosted services, AWS managed services, workflow orchestration with Temporal, integration with Nautobot, Argo-based promotion flows, and the supporting tooling required for debugging, snapshotting, local development, and production support.
This is a hands-on engineering role for someone who can design the right platform patterns, implement them directly, and establish a durable operating model between development and DevOps teams.
Key Responsibilities
Platform Deployment & CI/CD
• Design, implement, and maintain CI/CD pipelines for testing, staging, and production environments.
• Build and maintain deployment workflows that support safe and seamless promotion across environments.
• Improve and maintain Argo-based deployment workflows to enable controlled release progression from test to staging to production.
• Establish baseline deployment mechanisms for the site-builder application and related services.
• Standardize Kubernetes application packaging and deployment patterns, with a strong preference toward Helm-based lifecycle management for complex services and third-party components.
• Migrate existing deployments to Helm charts where appropriate.
Kubernetes & Runtime Platform Engineering
• Support the deployment and ongoing operation of services running in Kubernetes.
• Improve runtime reliability, resiliency, and troubleshooting for distributed services operating inside shared Kubernetes clusters.
• Investigate and harden service-to-service connectivity patterns, especially for workflow components such as workers connecting to the Temporal engine.
• Partner with development teams to define production-grade runtime requirements, resource sizing, restart policies, and platform support boundaries.
Infrastructure as Code & Cloud Services
• Design and implement fully declarative Infrastructure as Code for managed cloud services, especially in AWS.
• Provision and maintain managed data services such as RDS/PostgreSQL and MongoDB-compatible document databases across all environments.
• Eliminate manual infrastructure setup where possible and replace it with reproducible, version-controlled deployment patterns.
• Prepare the platform for future scale across multiple environments and regions through repeatable IaC and GitOps-aligned practices.
Data Services, Snapshots & Developer Enablement
• Setup and maintain RDS, MongoDB, Redis/cache services , and related dependencies for all environments.
• Build tooling and operational processes for:
• production and staging database snapshots,
• restoring snapshots into development environments,
• enabling local debugging and development from realistic data states.
• Support creation of local and development environments, including Minikube-based environment-as-code approaches that mirror production behavior as closely as practical.
• Improve platform reproducibility so engineers can quickly stand up close-to-production development environments.
Workflow Orchestration & Temporal Support
• Lead the setup, deployment, and operational support of Temporal for workflow orchestration.
• Support production operations for Temporal, including troubleshooting performance issues, restarts, scaling concerns, and resource shortages.
• Establish maintainable deployment patterns for Temporal using supported packaging and lifecycle management approaches.
• Partner with engineering teams to ensure workflow platform reliability and upgradeability over time.
Observability, Reliability & Incident Readiness
• Design and maintain observability across testing, staging, and production using tools such as Prometheus and Grafana .
• Define and implement monitoring for:
• service and cluster utilization,
• CPU, memory, storage,
• IOPS / throughput metrics,
• database connections and session counts,
• cache hit / miss / coverage metrics,
• RDS and MongoDB utilization,
• service health and alerting.
• Build and maintain logging, tracing, and correlation capabilities, separated appropriately by environment.
• Create tools to support deep debugging and operational inspection, including raw database reads, cleanup of unused volumes, and emergency cache invalidation.
Security, Access & Secrets Management
• Maintain secrets management processes across environments.
• Build tooling for short-lived internal token generation and long-lived secret rotation.
• Support secure access from deployed services to active production devices and southbound systems.
• Help establish credential management patterns for southbound integrations and device-facing access.
• Partner with related teams to define safe operational limits and controls for service integrations.
External Integrations & Platform Support
• Support integration patterns with Nautobot and help define safe client-side behaviors such as rate limiting, retry/backoff, and service protection mechanisms.
• Partner with application teams to understand and mitigate integration issues such as rate limiting or request rejection.
• Support staging and testing by enabling virtual device environments where needed.
• Contribute to end-to-end acceptance testing and production readiness activities.
Operating Model & Cross-Functional Execution
• Help define an effective operating model between Development and DevOps, whether via RACI , embedded Agile delivery, or a hybrid support model.
• Support deployment readiness, incident management, environment ownership boundaries, and lifecycle responsibilities.
• Work closely with software engineering, infrastructure, application owners, and partner teams to drive production readiness and sustainable operations.
Required Qualifications
• Bachelor's degree in Computer Science, Engineering, Information Systems, or equivalent practical experience.
• 7+ years of experience in DevOps, Platform Engineering, SRE, or Infrastructure Engineering roles.
• Strong hands-on experience with Kubernetes in production environments.
• Strong experience building and maintaining CI/CD pipelines for multi-environment software delivery.
• Strong experience with ArgoCD , GitOps workflows, or equivalent deployment tooling.
• Strong experience with Helm and Kubernetes package/deployment lifecycle management.
• Experience with AWS managed services , especially RDS/PostgreSQL , document databases, and related infrastructure.
• Strong experience with Infrastructure as Code , such as Terraform and/or similar declarative tooling.
• Experience with Prometheus, Grafana , and modern observability practices.
• Experience with Redis/cache services , secrets management, and operational debugging.
• Strong Linux, networking, and distributed systems troubleshooting skills.
• Strong scripting and automation skills in one or more languages such as Python, Bash, or Go.
• Proven ability to work cross-functionally and operate effectively in environments where ownership boundaries are still evolving.
Preferred Qualifications
• Experience with Temporal deployment and production operations.
• Experience supporting developer platforms with local environment reproducibility using Minikube, kind, or similar tools.
• Experience with MongoDB / DocumentDB operations and restore workflows.
• Experience integrating with Nautobot , NetBox, or similar infrastructure source-of-truth platforms.
• Experience operating in shared-cluster environments with multi-team tenancy and constrained access models.
• Experience designing platform patterns for internal products that must scale across regions or multiple deployment footprints.
• Familiarity with network automation or infrastructure orchestration platforms is a plus.
What Success Looks Like
• CI/CD pipelines are reliable, repeatable, and support safe promotion across all environments.
• Kubernetes deployments are standardized, maintainable, and production ready.
• Managed infrastructure is defined as code rather than through manual setup.
• Temporal, databases, cache layers, and observability tooling are stable and supportable.
• Development teams can reproduce realistic environments locally for faster debugging and delivery.
• Secrets, access patterns, and operational tooling are mature enough to support production-scale operations.
• The DevOps operating model is clearly defined and enables faster deployments with less operational risk.
Scope Notes
In scope
• CI/CD and deployment foundations
• Kubernetes packaging and release management
• RDS, MongoDB, Redis/cache services
• Temporal platform setup and operational support
• Observability, alerting, and debugging tooling
• Secrets management and access enablement
• Infrastructure as Code and environment reproducibility
• DevOps / Development operational model definition
Candidate Profile
The ideal candidate is a builder-operator: someone who can establish engineering discipline where manual patterns currently exist, create durable automation for platform operations, and raise the overall maturity of the product's deployment and runtime ecosystem. This person should be equally comfortable discussing deployment architecture, writing IaC and Helm code, troubleshooting Kubernetes runtime issues, and defining how DevOps and software engineering teams work together over the full product lifecycle.
Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Sr DevOps Engineer in Santa Clara, CA vacancy
$139k - $257.55k
...startup, backed by the resources and infrastructure of Adobe! How can you participate? In this bold venture, we seek a Senior Devops Engineer to build, redesign, and manage a large, globally distributed infrastructure using modern Reliability Engineering and DevOps...SeniorFull timeTemporary workLocal areaWorldwide- ...Job Title : Sr DevOPS Engineer Location - San Jose, CA FTE Only Job Description Sr DevOPS Engineer (Snowflake, DBT and Qlik) • CI/CD tools: Azure DevOps Pipelines or GitLab CI/CD (hands-on pipeline development) • Infrastructure as...Senior
$140k - $210k
...global scale, come make a difference at Fiserv.Job TitleSr. DevOps / Cloud Engineer, Billing InfrastructureAbout CloverClover is a pioneer in the... ...to support our platform. We are looking for talented Sr Engineers to help us deploy and operate our next generation...SeniorFull timeWorldwide- ...day.Culture, growth and wellbeing underline all aspects of Publicis Global Delivery.#WeArePGDOverviewWe are looking for a Senior DevOps Engineer with 5-7 years of experience in cloud infrastructure, automation, and CI/CD processes. Strong AWS, Terraform, and Kubernetes...Senior
$144k - $198k
...team, ensuring that complex algorithms translate predictably to real-time execution.What you need:BS/Advanced Degree in Aerospace Engineering, Computer Science, Electrical/Computer Engineering, or a related field with at least 8 years of relevant experience Proficiency...SeniorLocal area- ...Senior DevOps Engineer Location: Sunnyvale, CA Onsite position Fulltime position JD: Must Have Skills: AWS, EKS, IAM, S3, Kubernetes, Kustomize, Flux, Crossplane, CRDs, Python, Github, Kafka, Linux, Trino Strong...SeniorFull time
- ...Looking for a Senior DevOps Engineer having good Handson with Linux Infrastructure experience, automating, and supporting enterprise-scale Linux environments across cloud and on-premises infrastructure. Candidate should have proven expertise in Kubernetes, Docker...Senior
- ...Job role--Senior DevOps Engineer Work Location-- Sunnyvale, California, 94085-- 3 days Work from office and 2 days Work from Home Technical Hiring Criteria (Must Haves) :--DevOps Engineer[AWS, Kubernetes, Linux/Python Shell Script, DevOps Tools] ,Top 3 Required...SeniorWork at officeWork from home
$176.36k - $293.94k
...Sr. DevOps Engineer We are Omnissa! Omnissa is the first AI-driven digital work platform, built to support flexible, secure, work-from anywhere experiences. We integrate industry-leading solutions—including Unified Endpoint Management, Virtual Apps and Desktops, Digital...SeniorWork experience placementLocal areaVisa sponsorshipFlexible hours$184k - $287.5k
We are seeking a Senior DevOps / Cloud Simulation Infrastructure Engineer to own the complete end-to-end cloud execution pipeline for SimReady assets! This role is critical to our product strategy, enabling us to transition from local, workstation-driven validation to high...SeniorFull timeLocal area- ...Senior Cloud/DevOps Engineer Location: Sunnyvale, CA Experience: 10 Duration: 6 Months Please mention the current location, DL location, and visa status. Only U.S. Citizens and Green Card Holders. Must have skills: NGINX, Zero Trust Networking, AWS, Kubernetes...Senior
$126k - $203.5k
...cloud-native infrastructure, where reliability, scale, and intelligent automation define the future of operations. As a Senior DevOps Engineer, you will design and operate the platforms that power our applications across GCP, AWS, and global data centers — and you'll push...SeniorFull timeWork at officeVisa sponsorshipWork visa$187.74k - $195.25k
...Product Development Job Sub Function: R&D Software/Systems Engineering Job Category: Scientific/Technology All Job Posting Locations... ...Job Description: Employer: Auris Health, Inc. Job Title: DevOps Engineer Job Code: A011.9081 Job Location: Santa Clara, CA...Full timeLocal areaImmediate startRemote work- ...Job Title: Senior DevOps Engineer Location: Palo Alto (Hybrid) Duration: 6+ months with possibility of extension and FTE conversion Responsibilities Follow-the-Sun Incident Response: Provide SEV-1/SEV-2 incident coverage during PST/PDT hours (JP team off-hours...Senior
$150k - $250k
...signal synchronization, clock tree and cross domain clock designs is a significant plus.Need to work closely with system and test engineers to develop high speed interface, package/board, and system clocks in image sensor and bridge chip products.Responsibilities :• Analog...Senior- ...Principal DevOps Engineer Boston Scientific was recognized by Forbes as one of the Best Workplaces for Engineers in 2026, reflecting a culture where engineers do meaningful work. The Principal DevOps Engineer is a senior, hands-on technical leader responsible for...
- ...immediate opening with my client. If you are looking for a new project, please send me a copy of your updated resumes Title: Sr. SRE / DevOps Engineer Location: Sunnyvale, CA (Only Local candidate) Client Interview In-Person Job Summary For this role, we are...SeniorLocal areaImmediate start
- ...and improve on our cloud and GPU cluster, and create essential devops infrastructure to support our daily research and software... ...such as Python, C++, Rust, Go ~ Strong knowledge of software engineering best practices and design patterns ~ Experience with docker...Senior
$144k - $175k
...are looking for a highly motivated and versatile Senior Platform Engineer to join our platform engineering team. This unique role is... ...Qualifications5+ years of professional experience in Platform Engineering, DevOps, Site Reliability Engineering (SRE), or a related discipline....SeniorLocal area- ...our fast‑growing, highly ambitious team you won’t just drive the future of AI—you’ll help define it. Role Overview Senior DevOps Engineer – architect and maintain the core infrastructure that powers our cutting‑edge AI solutions. Responsibilities Design, build...SeniorFull time
- ...DescriptionSkills & RequirementsVery good communication and presentation skills - should be able to articulate problem statements/solutions & DevOps experienceStrong in scripting languages, such as Bash / Perl, etc. (python is desirable)Deep understanding of version control...Immediate start
$150k
...Senior DevOps Engineer Palo Alto, CA The world is moving towards instant digital payments and TabaPay is leading the way. We help thousands of Fintechs in the US and Canada instantly move money in and out of accounts and we are actively expanding into other countries...SeniorFlexible hours- ...DevOps Engineer Location: Sunnyvale, CA Duration: Long term contract Job requirements: Seeking DevOps engineers for the client's Release Engineering team. Will request candidates to complete a coderpad to be considered for a live skills assessment....Long term contract
- We are looking for an experienced software engineer to lead the development of our distributed edge compute platform. This role demands a deep understanding of hyperscalers like AWS, Azure, or GCP, and the ability to design and operate large-scale systems. The successful...SeniorFull timeTemporary work
- ...over 7,000 consultants. We recruit world-class talent for IT, engineering, and other professional jobs at 70+ Fortune and Global 500 companies... ....Job DescriptionDevOps EngineerSunnyvale CA12 monthsAs a DevOps Engineer, you will specialize in developing tools for a...
- ...Title: DevOps Engineer Location: Austin, TX, or Sunnyvale, CA (Hybrid) Duration: 6 months (possibility of extension) Implementation Partner: Infosys End Client: To be disclosed JD: DevOps and Support Role: Python (Code changes go to GitHub...
$170k - $220k
...we'd love to have you on board. Join us in shaping the future of healthcare. Job Summary We're looking for a Senior DevOps Engineer / Site Reliability Engineer to ensure the reliability, performance, and operational excellence of our production environments powering...SeniorLive inFlexible hours- ...Job Title Engineers: Experience working in public and private cloud environments AWS/OpenStack/Azure and working with their APIs Experience with Ansible in cloud environments Experience with CI/CD specifically using Jenkins and Jenkins Pipeline Experience with Kubernetes...Remote work
- ...Jose, California, United StatesProducts - Engineering /Fulltime /HybridOver 50,000 customers... ...to join the Extreme team.JOB DESCRIPTION Sr. Principal Engineer/ Principal Engineer -... ...the platform. Collaborate with product, DevOps, and security teams to align platform capabilities...SeniorFull timeRemote workFlexible hours
- An innovative AI solutions company is seeking a Senior DevOps Engineer to architect and maintain the core infrastructure supporting cutting-edge AI applications. The role involves designing scalable environments, collaborating with teams for seamless deployments, and championing...SeniorFull timeRemote workFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Sr DevOps Engineer. Be the first to apply!
Related searches
- devops engineer Santa Clara, CA
- devops aws developer (remote) Santa Clara, CA
- senior devops engineer remote Santa Clara, CA
- devops engineer full time Santa Clara, CA
- senior devops cloud engineer Santa Clara, CA
- big data devops engineer Santa Clara, CA
- senior manager tax Santa Clara, CA
- senior devops Santa Clara, CA
- senior associate vice president Santa Clara, CA
- senior director digital marketing Santa Clara, CA

