Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior DevOps Engineer

Namely

Senior DevOps / Cloud Platform Engineer – ML AI Infrastructure Job Summary We are looking for a Senior DevOps / Cloud Platform Engineer with strong experience in AWS, Kubernetes, CI/CD, infrastructure automation, and ML/AI infrastructure to design, deploy, manage, and optimize cloud infrastructure supporting our machine learning and AI services. The ideal candidate will have hands‑on experience with AWS EKS, SageMaker, Bedrock, Docker, Kubernetes, Terraform, Helm, GitHub Actions, Databricks, Elasticsearch, and self-hosted LLM deployments . This role will work closely with Data Engineering, Machine Learning, and Software Engineering teams to build reliable, scalable, secure, and cost‑efficient platforms for ML services across development and production environments. In this role you will...

Key Responsibilities
  • AWS Kubernetes Infrastructure Design, deploy, and administer AWS infrastructure supporting ML and AI workloads.
  • Manage Amazon EKS clusters , including cluster provisioning, upgrades, scaling, networking, and troubleshooting.
  • Work with AWS SageMaker, AWS Bedrock, EKS, ECS, and related AWS services .
  • Configure and manage Kubernetes Ingress controllers such as NGINX and AWS ALB .
  • Manage Cloudflare Tunnels, DNS, Cloudflare configuration, networking, and security .
  • Troubleshoot application, networking, compute, and infrastructure issues across AWS and Kubernetes environments.
  • Implement best practices for security, reliability, availability, and scalability.
  • CI/CD Azure-to-AWS Migration Build and maintain CI/CD pipelines for ML and AI services.
  • Develop and manage GitHub Actions and Azure DevOps pipelines using YAML.
  • Migrate repositories and CI/CD workflows from Azure DevOps to GitHub/AWS .
  • Automate build, test, containerization, deployment, and release processes.
  • Establish deployment strategies across development, staging, and production environments.
  • ML Service Deployment Deploy and manage ML services across AWS EKS/ECS and SageMaker .
  • Build and maintain Docker containers and Kubernetes deployments.
  • Manage environment segregation and configuration across Dev, QA, and Production.
  • Develop and maintain Kubernetes manifests and Helm charts .
  • Troubleshoot ML service deployment, networking, scaling, and runtime issues.
  • App Runner to EKS Migration Lead migration of existing services from AWS App Runner to Amazon EKS .
  • Containerize applications and develop Kubernetes manifests/Helm charts.
  • Design appropriate Kubernetes architecture, networking, ingress, scaling, and deployment strategies.
  • Ensure minimal service disruption during migration and establish operational best practices on EKS.
  • Self-Hosted LLM AI Infrastructure Deploy and manage self-hosted Large Language Models and inference services.
  • Work with model serving frameworks such as vLLM .
  • Design containerized infrastructure for GPU-based model serving and inference.
  • Manage model versions, deployments, configurations, and rollback strategies.
  • Support migration of ML services from managed APIs/services to self-hosted models .
  • Work with engineering teams on API integration and inference infrastructure.
  • Databricks Administration Administer Databricks workspaces, clusters, permissions, and access controls .
  • Manage cluster configuration, policies, and resource utilization.
  • Support LMI Insights and related ML/AI workloads.
  • Troubleshoot Databricks infrastructure and connectivity issues.
  • Implement appropriate security and access-control practices.
  • Elasticsearch Infrastructure Design, deploy, and manage Elasticsearch clusters .
  • Perform cluster sizing, scaling, configuration, and performance optimization.
  • Manage indices, mappings, retention, and data lifecycle requirements.
  • Support Kibana configuration, dashboards, and troubleshooting.
  • Monitor Elasticsearch health, capacity, and performance.
  • Monitoring, Reliability Auto-Scaling Implement monitoring and observability for Kubernetes, AWS, and ML services.
  • Use Prometheus, Grafana, and AWS CloudWatch for monitoring and alerting.
  • Configure Kubernetes HPA/VPA and other auto-scaling mechanisms.
  • Establish proactive alerting for infrastructure and application health.
  • Perform capacity planning and resource optimization.
  • Identify opportunities for AWS infrastructure and compute cost optimization .
  • Infrastructure as Code Automation Build and maintain infrastructure using Terraform .
  • Develop reusable Terraform modules for AWS and Kubernetes infrastructure.
  • Manage Kubernetes deployments using Helm charts .
  • Automate infrastructure provisioning, configuration, deployments, and operational tasks.
  • Maintain infrastructure documentation and deployment standards.
  • Cross-Team Collaboration Partner closely with Data Engineering, ML Engineering, Data Science, and Software Engineering teams.
  • Understand data pipelines, SQL, APIs, and ML service architecture sufficiently to troubleshoot end-to-end workflows.
  • Coordinate infrastructure requirements for new ML models and services.
  • Participate in production incident resolution, root-cause analysis, and continuous improvement.
  • Establish engineering standards around deployment, monitoring, security, and operational ownership.
Required Skills
  • Experience 5+ years of experience in DevOps, Cloud Infrastructure, SRE, or Platform Engineering.
  • Strong hands-on experience with AWS .
  • Strong experience administering Amazon EKS and Kubernetes in production.
  • Hands-on experience with: AWS EKS AWS SageMaker AWS Bedrock AWS ECS AWS App Runner Kubernetes Docker NGINX / AWS ALB Ingress Cloudflare / Cloudflare Tunnels
  • Strong experience with Terraform and Helm .
  • Strong experience developing CI/CD pipelines using GitHub Actions and/or Azure DevOps.
  • Strong YAML scripting and Git experience.
  • Experience migrating CI/CD pipelines and repositories from Azure to AWS/GitHub .
  • Experience deploying and operating ML/AI services.
  • Experience with self-hosted LLM/model serving , preferably vLLM .
  • Experience with GPU-based workloads is highly desirable.
  • Experience with Databricks administration .
  • Experience managing Elasticsearch and Kibana .
  • Experience with Prometheus, Grafana, and CloudWatch .
  • Strong understanding of Kubernetes HPA/VPA, networking, ingress, DNS, and service discovery .
  • Strong understanding of cloud networking fundamentals.
  • Experience with production troubleshooting, monitoring, capacity planning, and cost optimization.
  • Strong understanding of security, IAM, secrets management, and access control.
Preferred / Nice-to-Have Skills
  • Experience supporting Generative AI / LLM platforms .
  • Experience with GPU infrastructure and NVIDIA/CUDA environments.
  • Experience with model lifecycle and model version management.
  • Experience migrating workloads between managed cloud services and Kubernetes.
  • Experience with AWS networking such as VPC, load balancers, security groups, and Route 53.
  • Experience with API gateways and microservice architectures.
  • Experience with Python or shell scripting for infrastructure automation.
  • Experience working with Data Engineering and ML teams in a production environment.
What You'll Own
  • AWS ML/AI infrastructure EKS cluster administration and upgrades
  • ML service deployment and production operations
  • CI/CD automation
  • App Runner → EKS migration
  • Self-hosted LLM infrastructure and vLLM
  • Databricks platform administration
  • Elasticsearch infrastructure
  • Monitoring and auto-scaling
  • Terraform and Helm-based infrastructure automation
  • Cloud cost, reliability, and performance optimization
Ideal Candidate

The ideal candidate is a hands‑on infrastructure engineer who can independently take an ML/AI service from containerization → CI/CD → AWS infrastructure → EKS deployment → monitoring → scaling → production support . They should be comfortable working across both traditional DevOps infrastructure and modern AI/ML infrastructure , and should be able to collaborate closely with Data Engineering and ML teams while taking ownership of the underlying platform. #LI-Onsite

#J-18808-Ljbffr
Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Senior DevOps Engineer in Mountain View, CA vacancy
  •  ...of our fast‑growing, highly ambitious team you won’t just drive the future of AI—you’ll help define it. Role Overview Senior DevOps Engineer – architect and maintain the core infrastructure that powers our cutting‑edge AI solutions. Responsibilities Design,... 
    Senior
    Full time

    New Code Inc

    Palo Alto, CA
    4 days ago
  • $155k - $195k

     ...Overview Senior DevOps Engineer role at Drivemode. This range is provided by Drivemode. Your actual pay will be based on your skills and experience — talk with your recruiter to learn more. Base pay range $155,000.00/yr - $195,000.00/yr Additional compensation... 
    Senior
    Full time

    Drivemode

    Mountain View, CA
    4 days ago
  • $135k - $170k

     ...We are looking for a motivated and enthusiastic Senior DevOps Engineer to join our IT team. The ideal candidate will have a foundational understanding of DevOps principles, experience with cloud platforms, and a desire to grow their skills in a fast-paced and supportive... 
    Senior

    VALID8 Financial

    Sunnyvale, CA
    3 days ago
  • $150k - $209k

     ...readiness. Join us as we define the future of SaaS security! About The Team DevOps focuses on providing an end‑to‑end service to turn software into live services. We work closely with Engineering, QE, and Customer Support teams to continuously improve engineering... 
    Senior
    Work experience placement
    Flexible hours
    Shift work

    Obsidian Security

    Palo Alto, CA
    2 days ago
  •  ...Job Description As a Security Engineer you will be on the front lines, supporting in translating the magic of Google Cloud into mission...  ...: Skills and Requirements ~ Overall, 9 – 10 years of IT DevOps experience ~5 years of experience with large-scale enterprise... 
    Senior

    Insight Global

    Mountain View, CA
    16 hours ago
  • $184k - $287.5k

     ...As a Senior DevOps Engineer, you will help lead the evolution of infrastructure operations within our Networking Software group. Building on a strong Linux systems administration foundation, you will build, automate, and operate scalable platforms that support networking... 
    Senior
    Remote work

    NVIDIA

    Santa Clara, CA
    5 days ago
  • $134.83k - $158.62k

     ...together: Pragmatic Optimism Excellence without Ego Proactive Collaboration Job Overview We are seeking a Senior DevOps Engineer to act as a technical pillar of our platform team. In this role, you will transition from simply maintaining... 
    Senior
    Local area
    Flexible hours

    Ring

    Menlo Park, CA
    9 hours ago
  • $190k - $230k

     ...radius - there's no SRE org or ticket queue between us and the engineers we serve, so when a training run needs a thousand GPUs by Monday...  ...Minimum Requirements Proven track record of 5+ years in a DevOps, SRE, or infrastructure engineering role. Experience with hands... 
    Senior
    Temporary work
    Local area
    Work from home
    Flexible hours

    1x

    San Carlos, CA
    16 hours ago
  • $155k - $165k

     ...infrastructure and scalable cloud platforms to solve real financial challenges for everyday workers. We are seeking a Senior AWS DevOps & Cloud Security Engineer to own and drive the security, production infrastructure, and DevOps automation across our AWS environment. This... 
    Senior
    Work at office
    Local area

    Jobot

    Milpitas, CA
    1 day ago
  • $180k - $220k

     ...United States of America the safest country in the world. One Team. One Force.   About the Role We are looking for a Senior DevOps Engineer to own the build, deployment, and release infrastructure that keeps Knightscope's autonomous security robot fleet running.... 
    Senior
    Full time

    Knightscope

    Sunnyvale, CA
    11 days ago
  •  ...Devops EngineerLocation: Sunnyvale, CACandidate need to relocate to client location from day 1 and can WFH.Mandatory list:LinkedInPhoto...  ...a photo for identification purposes* 7+yrs experienced Devop Engineer* Position is for a hands-on DevOps engineer on cloud.* Experience... 
    Work from home
    Relocation

    Keylent Inc

    Sunnyvale, CA
    3 days ago
  •  ...Contribute to DevOps and SRE initiatives across Horizon Cloud’s global infrastructure Build and maintain CI/CD pipelines...  ...drive improvements to existing systems Report to the Senior Manager in Engineering and collaborate across platform pods in a hybrid agile... 
    Visa sponsorship

    Jobtailor

    Mountain View, CA
    3 days ago
  •  ...Globality is seeking a Sr. AI Software Engineer to build scalable AI platforms and developer tools that empower teams to deploy and iterate intelligent applications reliably. You will design end-to-end AI infrastructure, create prompting abstractions, and integrate... 
    Senior

    Globality Inc

    Palo Alto, CA
    2 days ago
  • $160k - $240k

     ...healthcare. Job Summary: We’re looking for a skilled Platform Engineer to contribute to the development of our Gen AI for Healthcare...  ...) and infrastructure as code (Terraform) Familiarity with DevOps practices, CI/CD pipelines, and monitoring tools Understanding... 
    Senior
    Live in
    Flexible hours
    3 days per week

    Qualified Health PBC

    Palo Alto, CA
    3 days ago
  •  ...one of these locations. We Are Synopsys is the leader in engineering solutions from silicon to systems, enabling customers to...  ...default outcomes. Build and evolve CI/CD pipelines in Azure DevOps, GitHub Actions, or similar platforms to standardize builds, testing... 
    Senior
    Work at office
    Relocation

    Synopsys

    Sunnyvale, CA
    9 hours ago
  •  ...Omnissa, LLC is hiring an AI Application Software Engineer to design, develop, and deploy AI-powered features across our cloud-scale platform. You will own end-to-end engineering initiatives, collaborating with platform, product, and engineering teams to deliver seamless... 
    Senior

    Omnissa, LLC

    Mountain View, CA
    16 hours ago
  • $206k - $238k

     ...Own the Unknown. Samsung currently requires employees to be onsite Monday-Thursday with Friday being a flex day. The Staff DevOps Engineer will design, build and run high quality and non-impactful deployments to power Samsung.com. We do continuous integration/deployment... 
    Monday to Friday
    Flexible hours

    National Black MBA Association

    Mountain View, CA
    4 days ago
  •  ...We Are Synopsys is the leader in engineering solutions from silicon to systems, enabling customers to rapidly innovate AI-powered products. We deliver industry-leading silicon design, IP, simulation and analysis solutions, and design services. We partner closely with... 
    Senior

    Synopsys

    Sunnyvale, CA
    2 days ago
  • $170k - $216k

     ...in an early stage and requires investment for Waymo to grow into a high scale, world-class service. We are building a team of engineers dedicated to further developing our Android Platform team. Our team works on the user interfaces used by customers to interact... 
    Senior
    Full time
    Remote work

    Waymo

    Mountain View, CA
    16 hours ago
  •  ...team builds the API platform and the tooling that every other engineering team at CoreWeave uses to ship and consume APIs. This role sits...  ...make generation the default path. About The Role As a Senior Engineer on API Platform & Client Tooling, you'll design and build... 
    Senior
    Permanent employment
    Full time
    Contract work
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    Sunnyvale, CA
    1 day ago
  • $105.5k - $213.5k

    DevOps Engineer (Sunnyvale, CA)This role has been designed as ‘’Onsite’ with an expectation that you will primarily work from an HPE office. Who We Are: Hewlett Packard Enterprise is the global edge-to-cloud company advancing the way people live and work. We help companies... 
    Full time
    Work experience placement
    Work at office
    Local area
    Immediate start

    Hewlett Packard Enterprise

    Sunnyvale, CA
    1 day ago
  • $102.5k - $187.9k

     ...the gap between cutting-edge AI technology and actionable audit technology. Your key responsibilities As a Kubernetes DevOps Engineer, you are responsible to design, deploy, and manage containerized applications and orchestration platforms (Kubernetes) to ensure... 
    Senior
    Summer holiday
    Work at office
    Flexible hours

    Ernst & Young

    San Jose, CA
    more than 2 months ago
  • $186k - $282k

     ...code in sandboxes, spend money per token rather than per request, and fail in ways a 500-rate dashboard never catches. Today, DevOps engineers carry this work alongside the wider fleet. We are making it someone's whole job. You will embed with the Transform and Close... 
    Senior

    FloQast

    San Jose, CA
    a month ago
  • $193.93k - $291.15k

     ...incident-response practices that let the team catch regressions. About You BS, MS, or PhD in Computer Science, Electrical Engineering, or a closely related field, plus 3+ years of relevant work experience. Willingness to deep-dive into implementation and to... 
    Senior
    Work experience placement
    Immediate start
    Flexible hours

    Nuro

    Mountain View, CA
    20 days ago
  •  ...an agile team that's constantly pushing the envelope to enhance, build, and deliver top-notch technology products. As a Senior Lead Software Engineer at JPMorgan Chase within the Corporate Sector, Infrastructure Platforms team, you are an integral part of an agile team... 
    Senior
    For contractors

    J.P. Morgan

    Palo Alto, CA
    4 days ago
  • $195k - $235k

     ...Here, responsibility is not determined by role, tenure, or seniority. Every team member - no matter their level - has the opportunity...  ...is the place for you. THE ROLE Infrastructure & DevOps Engineer III As a Infrastructure & DevOps Engineer III, you will... 
    Full time
    Local area

    Quince

    Palo Alto, CA
    2 days ago
  •  ...AI Platform Engineer This role is on-site in our San Francisco office (Atherton near Palo Alto) or hybrid in our New York City office...  ...backend engineering, full-stack development, or infrastructure/DevOps: internships and teaching assistantships don't count ~... 
    Senior
    Internship
    Work at office

    Voltai

    Palo Alto, CA
    2 days ago
  • $129.4k - $198.4k

     ...Job Description The Role: We are seeking experienced and motivated candidates for the role of Senior Software Engineer – Virtual Cloud Platform as part of the Virtualization & Embedded Software Development Tools (VEST) team in VSEE. Our mission is to develop... 
    Senior
    Full time
    Work experience placement
    Local area
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Mountain View, CA
    1 day ago
  • $130.6k - $285k

     ...runtime optimization, systems architecture, multi-tenant scheduling, and end-user tooling. About the Role As a Senior Software Engineer on Spark Platform, you will set the technical direction for our in-house Spark deployment and shape the architecture that... 
    Senior
    Hourly pay
    Full time
    Work at office
    Local area
    Remote work
    Relocation
    Flexible hours

    DoorDash USA

    Sunnyvale, CA
    2 days ago
  • $100k - $215k

     ...compliance. What you will do We are seeking a highly skilled Senior Software Test Engineer with expertise in automation, functional, and performance...  ...Collaborate with developers, product managers, and DevOps teams to integrate testing into the CI/CD pipeline... 
    Senior
    Hourly pay
    Full time
    Work experience placement
    Local area

    GEICO

    Palo Alto, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior DevOps Engineer. Be the first to apply!