Sr DevOps Engineer
Omni Inclusive
Job Summary
We are seeking a highly capable Senior DevOps Engineer / Platform Engineer to build, operationalize, and scale the infrastructure and deployment foundation for a strategic site-builder / network automation platform . This role will focus on creating reliable CI/CD pipelines, production-grade Kubernetes deployment patterns, managed database services, observability, environment reproducibility, secrets management, and Infrastructure as Code across development, testing, staging, and production environments.
This engineer will play a critical role in moving the platform from an early-stage, partially manual operating model into a repeatable, supportable, and production-ready DevOps model. The environment includes Kubernetes-hosted services, AWS managed services, workflow orchestration with Temporal, integration with Nautobot, Argo-based promotion flows, and the supporting tooling required for debugging, snapshotting, local development, and production support.
This is a hands-on engineering role for someone who can design the right platform patterns, implement them directly, and establish a durable operating model between development and DevOps teams.
Key Responsibilities
Platform Deployment & CI/CD
• Design, implement, and maintain CI/CD pipelines for testing, staging, and production environments.
• Build and maintain deployment workflows that support safe and seamless promotion across environments.
• Improve and maintain Argo-based deployment workflows to enable controlled release progression from test to staging to production.
• Establish baseline deployment mechanisms for the site-builder application and related services.
• Standardize Kubernetes application packaging and deployment patterns, with a strong preference toward Helm-based lifecycle management for complex services and third-party components.
• Migrate existing deployments to Helm charts where appropriate.
Kubernetes & Runtime Platform Engineering
• Support the deployment and ongoing operation of services running in Kubernetes.
• Improve runtime reliability, resiliency, and troubleshooting for distributed services operating inside shared Kubernetes clusters.
• Investigate and harden service-to-service connectivity patterns, especially for workflow components such as workers connecting to the Temporal engine.
• Partner with development teams to define production-grade runtime requirements, resource sizing, restart policies, and platform support boundaries.
Infrastructure as Code & Cloud Services
• Design and implement fully declarative Infrastructure as Code for managed cloud services, especially in AWS.
• Provision and maintain managed data services such as RDS/PostgreSQL and MongoDB-compatible document databases across all environments.
• Eliminate manual infrastructure setup where possible and replace it with reproducible, version-controlled deployment patterns.
• Prepare the platform for future scale across multiple environments and regions through repeatable IaC and GitOps-aligned practices.
Data Services, Snapshots & Developer Enablement
• Setup and maintain RDS, MongoDB, Redis/cache services , and related dependencies for all environments.
• Build tooling and operational processes for:
• production and staging database snapshots,
• restoring snapshots into development environments,
• enabling local debugging and development from realistic data states.
• Support creation of local and development environments, including Minikube-based environment-as-code approaches that mirror production behavior as closely as practical.
• Improve platform reproducibility so engineers can quickly stand up close-to-production development environments.
Workflow Orchestration & Temporal Support
• Lead the setup, deployment, and operational support of Temporal for workflow orchestration.
• Support production operations for Temporal, including troubleshooting performance issues, restarts, scaling concerns, and resource shortages.
• Establish maintainable deployment patterns for Temporal using supported packaging and lifecycle management approaches.
• Partner with engineering teams to ensure workflow platform reliability and upgradeability over time.
Observability, Reliability & Incident Readiness
• Design and maintain observability across testing, staging, and production using tools such as Prometheus and Grafana .
• Define and implement monitoring for:
• service and cluster utilization,
• CPU, memory, storage,
• IOPS / throughput metrics,
• database connections and session counts,
• cache hit / miss / coverage metrics,
• RDS and MongoDB utilization,
• service health and alerting.
• Build and maintain logging, tracing, and correlation capabilities, separated appropriately by environment.
• Create tools to support deep debugging and operational inspection, including raw database reads, cleanup of unused volumes, and emergency cache invalidation.
Security, Access & Secrets Management
• Maintain secrets management processes across environments.
• Build tooling for short-lived internal token generation and long-lived secret rotation.
• Support secure access from deployed services to active production devices and southbound systems.
• Help establish credential management patterns for southbound integrations and device-facing access.
• Partner with related teams to define safe operational limits and controls for service integrations.
External Integrations & Platform Support
• Support integration patterns with Nautobot and help define safe client-side behaviors such as rate limiting, retry/backoff, and service protection mechanisms.
• Partner with application teams to understand and mitigate integration issues such as rate limiting or request rejection.
• Support staging and testing by enabling virtual device environments where needed.
• Contribute to end-to-end acceptance testing and production readiness activities.
Operating Model & Cross-Functional Execution
• Help define an effective operating model between Development and DevOps, whether via RACI , embedded Agile delivery, or a hybrid support model.
• Support deployment readiness, incident management, environment ownership boundaries, and lifecycle responsibilities.
• Work closely with software engineering, infrastructure, application owners, and partner teams to drive production readiness and sustainable operations.
Required Qualifications
• Bachelor's degree in Computer Science, Engineering, Information Systems, or equivalent practical experience.
• 7+ years of experience in DevOps, Platform Engineering, SRE, or Infrastructure Engineering roles.
• Strong hands-on experience with Kubernetes in production environments.
• Strong experience building and maintaining CI/CD pipelines for multi-environment software delivery.
• Strong experience with ArgoCD , GitOps workflows, or equivalent deployment tooling.
• Strong experience with Helm and Kubernetes package/deployment lifecycle management.
• Experience with AWS managed services , especially RDS/PostgreSQL , document databases, and related infrastructure.
• Strong experience with Infrastructure as Code , such as Terraform and/or similar declarative tooling.
• Experience with Prometheus, Grafana , and modern observability practices.
• Experience with Redis/cache services , secrets management, and operational debugging.
• Strong Linux, networking, and distributed systems troubleshooting skills.
• Strong scripting and automation skills in one or more languages such as Python, Bash, or Go.
• Proven ability to work cross-functionally and operate effectively in environments where ownership boundaries are still evolving.
Preferred Qualifications
• Experience with Temporal deployment and production operations.
• Experience supporting developer platforms with local environment reproducibility using Minikube, kind, or similar tools.
• Experience with MongoDB / DocumentDB operations and restore workflows.
• Experience integrating with Nautobot , NetBox, or similar infrastructure source-of-truth platforms.
• Experience operating in shared-cluster environments with multi-team tenancy and constrained access models.
• Experience designing platform patterns for internal products that must scale across regions or multiple deployment footprints.
• Familiarity with network automation or infrastructure orchestration platforms is a plus.
What Success Looks Like
• CI/CD pipelines are reliable, repeatable, and support safe promotion across all environments.
• Kubernetes deployments are standardized, maintainable, and production ready.
• Managed infrastructure is defined as code rather than through manual setup.
• Temporal, databases, cache layers, and observability tooling are stable and supportable.
• Development teams can reproduce realistic environments locally for faster debugging and delivery.
• Secrets, access patterns, and operational tooling are mature enough to support production-scale operations.
• The DevOps operating model is clearly defined and enables faster deployments with less operational risk.
Scope Notes
In scope
• CI/CD and deployment foundations
• Kubernetes packaging and release management
• RDS, MongoDB, Redis/cache services
• Temporal platform setup and operational support
• Observability, alerting, and debugging tooling
• Secrets management and access enablement
• Infrastructure as Code and environment reproducibility
• DevOps / Development operational model definition
Candidate Profile
The ideal candidate is a builder-operator: someone who can establish engineering discipline where manual patterns currently exist, create durable automation for platform operations, and raise the overall maturity of the product's deployment and runtime ecosystem. This person should be equally comfortable discussing deployment architecture, writing IaC and Helm code, troubleshooting Kubernetes runtime issues, and defining how DevOps and software engineering teams work together over the full product lifecycle.
We are seeking a highly capable Senior DevOps Engineer / Platform Engineer to build, operationalize, and scale the infrastructure and deployment foundation for a strategic site-builder / network automation platform . This role will focus on creating reliable CI/CD pipelines, production-grade Kubernetes deployment patterns, managed database services, observability, environment reproducibility, secrets management, and Infrastructure as Code across development, testing, staging, and production environments.
This engineer will play a critical role in moving the platform from an early-stage, partially manual operating model into a repeatable, supportable, and production-ready DevOps model. The environment includes Kubernetes-hosted services, AWS managed services, workflow orchestration with Temporal, integration with Nautobot, Argo-based promotion flows, and the supporting tooling required for debugging, snapshotting, local development, and production support.
This is a hands-on engineering role for someone who can design the right platform patterns, implement them directly, and establish a durable operating model between development and DevOps teams.
Key Responsibilities
Platform Deployment & CI/CD
• Design, implement, and maintain CI/CD pipelines for testing, staging, and production environments.
• Build and maintain deployment workflows that support safe and seamless promotion across environments.
• Improve and maintain Argo-based deployment workflows to enable controlled release progression from test to staging to production.
• Establish baseline deployment mechanisms for the site-builder application and related services.
• Standardize Kubernetes application packaging and deployment patterns, with a strong preference toward Helm-based lifecycle management for complex services and third-party components.
• Migrate existing deployments to Helm charts where appropriate.
Kubernetes & Runtime Platform Engineering
• Support the deployment and ongoing operation of services running in Kubernetes.
• Improve runtime reliability, resiliency, and troubleshooting for distributed services operating inside shared Kubernetes clusters.
• Investigate and harden service-to-service connectivity patterns, especially for workflow components such as workers connecting to the Temporal engine.
• Partner with development teams to define production-grade runtime requirements, resource sizing, restart policies, and platform support boundaries.
Infrastructure as Code & Cloud Services
• Design and implement fully declarative Infrastructure as Code for managed cloud services, especially in AWS.
• Provision and maintain managed data services such as RDS/PostgreSQL and MongoDB-compatible document databases across all environments.
• Eliminate manual infrastructure setup where possible and replace it with reproducible, version-controlled deployment patterns.
• Prepare the platform for future scale across multiple environments and regions through repeatable IaC and GitOps-aligned practices.
Data Services, Snapshots & Developer Enablement
• Setup and maintain RDS, MongoDB, Redis/cache services , and related dependencies for all environments.
• Build tooling and operational processes for:
• production and staging database snapshots,
• restoring snapshots into development environments,
• enabling local debugging and development from realistic data states.
• Support creation of local and development environments, including Minikube-based environment-as-code approaches that mirror production behavior as closely as practical.
• Improve platform reproducibility so engineers can quickly stand up close-to-production development environments.
Workflow Orchestration & Temporal Support
• Lead the setup, deployment, and operational support of Temporal for workflow orchestration.
• Support production operations for Temporal, including troubleshooting performance issues, restarts, scaling concerns, and resource shortages.
• Establish maintainable deployment patterns for Temporal using supported packaging and lifecycle management approaches.
• Partner with engineering teams to ensure workflow platform reliability and upgradeability over time.
Observability, Reliability & Incident Readiness
• Design and maintain observability across testing, staging, and production using tools such as Prometheus and Grafana .
• Define and implement monitoring for:
• service and cluster utilization,
• CPU, memory, storage,
• IOPS / throughput metrics,
• database connections and session counts,
• cache hit / miss / coverage metrics,
• RDS and MongoDB utilization,
• service health and alerting.
• Build and maintain logging, tracing, and correlation capabilities, separated appropriately by environment.
• Create tools to support deep debugging and operational inspection, including raw database reads, cleanup of unused volumes, and emergency cache invalidation.
Security, Access & Secrets Management
• Maintain secrets management processes across environments.
• Build tooling for short-lived internal token generation and long-lived secret rotation.
• Support secure access from deployed services to active production devices and southbound systems.
• Help establish credential management patterns for southbound integrations and device-facing access.
• Partner with related teams to define safe operational limits and controls for service integrations.
External Integrations & Platform Support
• Support integration patterns with Nautobot and help define safe client-side behaviors such as rate limiting, retry/backoff, and service protection mechanisms.
• Partner with application teams to understand and mitigate integration issues such as rate limiting or request rejection.
• Support staging and testing by enabling virtual device environments where needed.
• Contribute to end-to-end acceptance testing and production readiness activities.
Operating Model & Cross-Functional Execution
• Help define an effective operating model between Development and DevOps, whether via RACI , embedded Agile delivery, or a hybrid support model.
• Support deployment readiness, incident management, environment ownership boundaries, and lifecycle responsibilities.
• Work closely with software engineering, infrastructure, application owners, and partner teams to drive production readiness and sustainable operations.
Required Qualifications
• Bachelor's degree in Computer Science, Engineering, Information Systems, or equivalent practical experience.
• 7+ years of experience in DevOps, Platform Engineering, SRE, or Infrastructure Engineering roles.
• Strong hands-on experience with Kubernetes in production environments.
• Strong experience building and maintaining CI/CD pipelines for multi-environment software delivery.
• Strong experience with ArgoCD , GitOps workflows, or equivalent deployment tooling.
• Strong experience with Helm and Kubernetes package/deployment lifecycle management.
• Experience with AWS managed services , especially RDS/PostgreSQL , document databases, and related infrastructure.
• Strong experience with Infrastructure as Code , such as Terraform and/or similar declarative tooling.
• Experience with Prometheus, Grafana , and modern observability practices.
• Experience with Redis/cache services , secrets management, and operational debugging.
• Strong Linux, networking, and distributed systems troubleshooting skills.
• Strong scripting and automation skills in one or more languages such as Python, Bash, or Go.
• Proven ability to work cross-functionally and operate effectively in environments where ownership boundaries are still evolving.
Preferred Qualifications
• Experience with Temporal deployment and production operations.
• Experience supporting developer platforms with local environment reproducibility using Minikube, kind, or similar tools.
• Experience with MongoDB / DocumentDB operations and restore workflows.
• Experience integrating with Nautobot , NetBox, or similar infrastructure source-of-truth platforms.
• Experience operating in shared-cluster environments with multi-team tenancy and constrained access models.
• Experience designing platform patterns for internal products that must scale across regions or multiple deployment footprints.
• Familiarity with network automation or infrastructure orchestration platforms is a plus.
What Success Looks Like
• CI/CD pipelines are reliable, repeatable, and support safe promotion across all environments.
• Kubernetes deployments are standardized, maintainable, and production ready.
• Managed infrastructure is defined as code rather than through manual setup.
• Temporal, databases, cache layers, and observability tooling are stable and supportable.
• Development teams can reproduce realistic environments locally for faster debugging and delivery.
• Secrets, access patterns, and operational tooling are mature enough to support production-scale operations.
• The DevOps operating model is clearly defined and enables faster deployments with less operational risk.
Scope Notes
In scope
• CI/CD and deployment foundations
• Kubernetes packaging and release management
• RDS, MongoDB, Redis/cache services
• Temporal platform setup and operational support
• Observability, alerting, and debugging tooling
• Secrets management and access enablement
• Infrastructure as Code and environment reproducibility
• DevOps / Development operational model definition
Candidate Profile
The ideal candidate is a builder-operator: someone who can establish engineering discipline where manual patterns currently exist, create durable automation for platform operations, and raise the overall maturity of the product's deployment and runtime ecosystem. This person should be equally comfortable discussing deployment architecture, writing IaC and Helm code, troubleshooting Kubernetes runtime issues, and defining how DevOps and software engineering teams work together over the full product lifecycle.
Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Sr DevOps Engineer in Santa Clara, CA vacancy
- ...Job Title : Sr DevOPS Engineer Location - San Jose, CA FTE Only Job Description Sr DevOPS Engineer (Snowflake, DBT and Qlik) • CI/CD tools: Azure DevOps Pipelines or GitLab CI/CD (hands-on pipeline development) • Infrastructure as...Senior
- ...DevOps Engineer As a DevOps engineer, the candidate will be responsible for the development of tools that allows application engineers to focus just on coding/developing the application while everything else happens automatically in the background, including continuous...Senior
$139k - $257.55k
...startup, backed by the resources and infrastructure of Adobe! How can you participate? In this bold venture, we seek a Senior Devops Engineer to build, redesign, and manage a large, globally distributed infrastructure using modern Reliability Engineering and DevOps...SeniorTemporary workLocal areaRelocation$145k - $170k
...technology community. We seek talented, passionate, and committed engineers, technologists, and business leaders to join us.Job Summary:... ...System Engineering and high-level test automation. A hands-on Linux DevOps Engineer specializing in maintaining Server Infrastructure who...SeniorWork experience placementWorldwide$91k - $147.2k
...Product & Platform Management Job Sub Function: Software Engineering - DevOps Job Category: Scientific/Technology All Job Posting... ...States of America : Johnson & Johnson is hiring for a Sr. DevOps Engineer - Shockwave Medical to join our team...SeniorLocal areaImmediate start- ...this is your chance to be part of something exceptional.Position: Senior DevOps EngineerLocation: Campbell, CA / Remote USA (PST Time Zone)Position OverviewWe are seeking a Senior DevOps Engineer to partner closely with development and engineering teams to build, deploy...SeniorFull timeRemote work
$144k - $198k
...team, ensuring that complex algorithms translate predictably to real-time execution.What you need:BS/Advanced Degree in Aerospace Engineering, Computer Science, Electrical/Computer Engineering, or a related field with at least 8 years of relevant experience Proficiency...SeniorLocal area- ...Job role--Senior DevOps Engineer Work Location-- Sunnyvale, California, 94085-- 3 days Work from office and 2 days Work from Home Technical Hiring Criteria (Must Haves) :--DevOps Engineer[AWS, Kubernetes, Linux/Python Shell Script, DevOps Tools] ,Top 3 Required...SeniorWork at officeWork from home
- ...data warehousing solutions. Prefer experiences of build real time analytics system or machine learning system Education: ~ Must have BS/MS in Computer Science/Engineering with 2 years+ software development experience Required Skills: ~ AWS...SeniorWork experience placement
- ...goals. ~ World-class leadership team: Our Heads of AI, Engineering, and Product bring extensive experience from some of the world... ...ABOUT THE ROLE Kai is seeking a highly skilled Senior DevOps Engineer to design, build, and maintain secure, scalable, and...Senior
$176.36k - $293.94k
...Sr. DevOps Engineer Opportunity At Omnissa We are Omnissa! Omnissa is the first AI-driven digital work platform, built to support flexible, secure, work-from anywhere experiences. We integrate industry-leading solutions—including Unified Endpoint Management, Virtual...SeniorWork experience placementLocal areaVisa sponsorshipFlexible hours$110k - $160k
...applications must be submitted through our official website ( monks.com/careers ). We are seeking a highly motivated and innovative DevOps engineer to join our team. In this position, you will join a team of DevOps engineers focused on automating solutions, securing data and...SeniorWork at office- ...cloud services. The role emphasizes distributed systems, secure engineering, and expert problem solving in a fast‑paced environment. You... ...mentor teams and drive architecture decisions for enterprise‑grade DevOps solutions. Location in Santa Clara, CA with a focus on cutting...Senior
- ...Senior Cloud/DevOps EngineerLocation: Sunnyvale, CAExperience: 10 Duration: 6 MonthsPlease mention the current location, DL location... ...Groovy.Leverage AWS CodeWhisperer to improve code quality and engineering efficiency.Provide production support and lead root-cause...Senior
- BayInfotech in Santa Clara, CA is seeking multiple DevOps Support Engineers to join our customer success team. You will troubleshoot issues with our Cloud Center product suite and serve as a key technical contact for clients to ensure successful deployments and exceptional...Senior
- ...Senior Staff DevOps Engineer Pangyo, South Korea At Sonatus, we're driving the transformation to AI-enabled software-defined vehicles. Traditional automotive software methods can't keep pace with consumer expectations shaped by the mobile industry—where features...SeniorWorldwideShift work
$186k - $282k
...code in sandboxes, spend money per token rather than per request, and fail in ways a 500-rate dashboard never catches. Today, DevOps engineers carry this work alongside the wider fleet. We are making it someone's whole job. You will embed with the Transform and Close...Senior$126k - $203.5k
...cloud-native infrastructure, where reliability, scale, and intelligent automation define the future of operations. As a Senior DevOps Engineer, you will design and operate the platforms that power our applications across GCP, AWS, and global data centers — and you'll push...SeniorFull timeWork at office$102.5k - $187.9k
...the gap between cutting-edge AI technology and actionable audit technology. Your key responsibilities As a Kubernetes DevOps Engineer, you are responsible to design, deploy, and manage containerized applications and orchestration platforms (Kubernetes) to ensure...SeniorSummer holidayWork at officeFlexible hours$180k - $220k
...United States of America the safest country in the world. One Team. One Force. About the Role We are looking for a Senior DevOps Engineer to own the build, deployment, and release infrastructure that keeps Knightscope's autonomous security robot fleet running. You...SeniorFull time- ...Job Description We are seeking a highly skilled Senior Cloud & DevOps Engineer to design, build, and operate scalable cloud-native platforms that support modern data, machine learning, and application workloads. The ideal candidate will have strong expertise in...Senior
- Everpure in Santa Clara is seeking a Senior Software Engineer in Production Engineering to own the CI and test-orchestration platform that enables FlashArray and FlashBlade teams to ship high-quality products at predictable velocity. This is a developer-first role with...Senior
$182k - $260k
jobfuBrowse jobs Zscaler Principal DevOps Engineer San Jose, California, USA USD 182,000-260,000 Skills DevOpsZero TrustDigital TransformationData Centers About The Role About Zscaler is a pioneer and global leader in zero trust security. The world’s largest businesses,...- ...Job Title: Senior DevOps Engineer Location: Palo Alto (Hybrid) Duration: 6+ months with possibility of extension and FTE conversion Responsibilities Follow-the-Sun Incident Response: Provide SEV-1/SEV-2 incident coverage during PST/PDT hours (JP team off-hours...Senior
- ...and improve on our cloud and GPU cluster, and create essential devops infrastructure to support our daily research and software... ...such as Python, C++, Rust, Go ~ Strong knowledge of software engineering best practices and design patterns ~ Experience with docker...Senior
- ...DescriptionSkills & RequirementsVery good communication and presentation skills - should be able to articulate problem statements/solutions & DevOps experienceStrong in scripting languages, such as Bash / Perl, etc. (python is desirable)Deep understanding of version control...Immediate start
- Redolent Infotech Pvt. Ltd. in Sunnyvale, CA seeks a Senior Software Build & Release Engineer to own release pipelines, build automation and deployment across web services. The role requires strong experience with GIT, Jenkins, Python, Perforce and Docker, plus hands-on...Senior
$150k - $250k
...signal synchronization, clock tree and cross domain clock designs is a significant plus.Need to work closely with system and test engineers to develop high speed interface, package/board, and system clocks in image sensor and bridge chip products.Responsibilities :• Analog...Senior- NVIDIA seeks a Senior Board Product Development Engineer to join our Operations Engineering team in Santa Clara, CA. You will drive the release of NVIDIA board and system products to high-volume manufacturing, coordinating cross-functional teams across design, test, and...Senior
$150k - $200k
...team includes leaders with strong expertise in neuroscience and engineering from Stanford University. LVIS has been selected to be a... ...neurology health care industry. LVIS is looking for a senior DevOps engineer who is responsible for commercial cloud-based big-data...SeniorFull timeWork at office
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Sr DevOps Engineer. Be the first to apply!
Related searches
- senior devops cloud engineer Santa Clara, CA
- big data devops engineer Santa Clara, CA
- devops aws developer (remote) Santa Clara, CA
- devops engineer Santa Clara, CA
- senior devops engineer remote Santa Clara, CA
- senior network engineer remote Santa Clara, CA
- senior app developer Santa Clara, CA
- senior manager legal Santa Clara, CA
- sr project manager Santa Clara, CA
- senior account executive Santa Clara, CA


