Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

SRE - Senior AI Platform Reliability Engineer

HTC Global Services Inc

Job Description

Job Description

Are you a Senior Site Reliability Engineer with direct experience operating AI or machine learning platforms in large-scale production environments? We are looking for you  to join our growing team working a hybrid schedule at one of our three locations either in Seattle, Burbak or Orlando. If working on a highly collaborative cutting edge technology team and in a job that is not a traditional infrastructure-only SRE role then this is the position for you! 

The selected engineer will help build, scale, and operate the cloud and Kubernetes infrastructure that enables enterprise AI capabilities across the company and the selected engineer will help design, scale, and operate the cloud and Kubernetes infrastructure supporting enterprise AI workloads.

You do not need to be  an AI model developers, data scientists, or LLM experts. instead your day to day focus and experience with the infrastructure and operational challenges involved in:

  • Deploying AI models or AI services into production
  • Operating AI platforms at enterprise scale
  • Supporting AI workloads across cloud environments
  • Scaling distributed AI services
  • Building reliable infrastructure for model inference, APIs, and AI-enabled applications
  • Monitoring the performance, availability, and capacity of production AI platforms

The ideal candidate combines strong AI platform infrastructure experience with deep expertise in SRE, Kubernetes, multi-cloud engineering, Infrastructure as Code, observability, and production reliability.

Key Responsibilities
  • Lead the design, implementation, and operation of highly available infrastructure supporting enterprise AI platforms and services.
  • Build and operate Kubernetes-based environments used to deploy and scale AI workloads.
  • Support the production deployment of AI models, inference services, AI APIs, agents, and related platform capabilities.
  • Design infrastructure that enables AI workloads across distributed and multi-cloud environments.
  • Partner with AI engineers, platform engineers, architects, and application teams to move AI services from development into reliable production environments.
  • Establish deployment, scaling, capacity, and reliability patterns for AI-powered services.
  • Design and maintain Kubernetes infrastructure using Helm and Terraform.
  • Support AI platform dependencies such as model gateways, vector or operational data stores, messaging platforms, secrets management, and API services.
  • Develop scalable platform solutions capable of maintaining 99.99% availability.
  • Lead capacity planning for AI workloads, including compute, memory, storage, network, and service dependencies.
  • Build automated CI/CD pipelines using Harness or comparable enterprise deployment platforms.
  • Implement blue/green deployments, canary releases, automated rollback, and feature-flag strategies.
  • Build observability for AI platforms using metrics, logs, traces, service-level indicators, and workload-specific health signals.
  • Troubleshoot complex production issues involving AI services, Kubernetes, cloud infrastructure, databases, messaging, networking, and distributed systems.
  • Lead incident response, root-cause analysis, and permanent corrective-action planning.
  • Mentor SRE, DevOps, and platform engineers and establish engineering and operational standards.
  • Ensure AI infrastructure aligns with enterprise security, governance, compliance, and resiliency requirements.
Required Qualifications
  • 7+ years of experience in Site Reliability Engineering, Platform Engineering, DevOps, cloud infrastructure, or a related field.
  • Direct experience supporting AI, machine learning, or model-serving platforms in production .
  • Experience deploying or operating AI models, inference services, AI APIs, agents, or AI-enabled applications in cloud environments.
  • Experience supporting AI workloads at enterprise or high-traffic scale.
  • Strong understanding of the infrastructure required to move AI services from development into production.
  • Expert-level Kubernetes administration and production operations experience.
  • Strong Helm and Terraform experience.
  • Experience designing scalable, highly available, distributed cloud infrastructure.
  • Hands-on experience with Google Cloud Platform, with additional AWS or Azure experience.
  • Experience building automated deployment pipelines using Harness or a comparable enterprise CI/CD platform.
  • Strong scripting and automation skills using Python, Bash, and YAML.
  • Experience supporting production databases, messaging, and platform services such as PostgreSQL, Redis, Kafka, MongoDB, and Vault.
  • Experience implementing observability using OpenTelemetry, Prometheus, Grafana, Splunk, AppDynamics, or similar tools.
  • Strong production troubleshooting and incident-management experience across cloud-native distributed systems.
  • Experience with capacity planning, performance tuning, disaster recovery, and high-availability design.
  • Strong communication and technical leadership skills.
Preferred Qualifications
  • Experience supporting generative AI, LLM, model inference, or AI-agent platforms.
  • Experience operating model gateways, AI APIs, GPU-enabled workloads, or distributed inference services.
  • Experience with AI platform technologies such as LiteLLM, Open WebUI, model-serving frameworks, vector databases, or similar platforms.
  • Experience defining SLOs, SLIs, error budgets, and reliability standards for AI services.
  • Experience with GCP-based AI infrastructure and services.
  • Experience supporting high-volume or globally distributed AI platforms.
  • Experience with progressive delivery, chaos engineering, and automated resilience testing.
  • Cloud or Kubernetes certifications.

What Makes HTC A Great Place To Build Your Future

HTC Global Services wants you to join our team. Come build new things with us and advance your career. At HTC Global, you’ll collaborate with experts, work alongside clients, and be part of high-performing teams driving success together. You’ll have long-term opportunities to grow your career and develop skills in the latest emerging technologies.

At HTC Global Services, our employees have access to a comprehensive benefits package. Benefits can include Group Health (Medical, Dental, and Vision), Paid Time Off, Paid Holidays, 401(k) matching, Group Life and Disability insurance, Professional Development opportunities, Wellness programs, and a variety of other perks.

Our success as a company is built on inclusion and diversity. HTC Global Services is committed to providing a workplace free from discrimination and harassment, where every employee is treated with dignity and respect. We celebrate differences and believe that diverse cultures, perspectives, and skills drive innovation and success. HTC is an Equal Opportunity Employer and a proud National Minority Supplier. We seek to empower each individual, fostering an environment where everyone feels valued, included, and respected.

#LI-NC1 #LI-DT1 #LI-Remote #Hiring #AIJobs #SREJobs #SeattleJobs #BurbankJobs #OrlandoJobs

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the SRE - Senior AI Platform Reliability Engineer in Seattle, WA vacancy
  • $232k - $319k

    Secure Every Identity, from AI to HumanIdentity is the...  ....The Infrastructure Platform and Shared Services...  ...with great people and reliable, cost-effective, and efficient...  ...initiatives across SRE & Infrastructure organization...  ...of SRE and product engineering by developing robust... 
    Senior
    Permanent employment
    Local area
    Worldwide
    Flexible hours

    Okta

    Bellevue, WA
    5 hours ago
  • $170k - $220k

     ...looking for a hands-on, high-agency Site Reliability Engineer to help shape and scale the reliability...  ...and improve deployment confidence. Using AI tools to generate code is encouraged and...  ...re a Great Fit If YouHave 3-6+ years in SRE, DevOps, or infrastructure roles with... 
    Senior

    Supio

    Seattle, WA
    2 days ago
  • $127k - $249k

    The TeamPlatform Engineering is the department within SRE that is responsible for a range of...  ...that ensure cluster reliability and security (e.g., CoreDNS...  ...in cloud infrastructure platforms, including AWS, GCP, or AzureProficiency...  ...data platform for the AI era, enabling builders to... 
    Senior
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    Seattle, WA
    4 days ago
  • $165k - $225.6k

     ...Every Identity, from AI to HumanIdentity is the...  ...to enterprise platforms, we partner across functions...  ...to drive scale, reliability, and innovation through technology.The Senior Site Reliability Engineer OpportunityReporting...  ...champion DevOps and SRE best practices, deliver... 
    Senior
    Permanent employment
    Local area
    Worldwide
    Flexible hours

    Okta

    Bellevue, WA
    2 days ago
  • $134.25k - $214.8k

     ....Your ImpactAre you an engineer who gets excited about...  ...generation observability platform, enabling the entire...  ...team within Axon's Site Reliability organization — a focused...  ...years of experience in SRE, platform engineering,...  ...engineeringExperience with agentic AI tooling or building LLM... 
    Senior
    Work experience placement
    Work at office
    Remote work

    Axon

    Seattle, WA
    3 days ago
  • $167.7k - $245.2k

     ...ThousandEyes is a Digital Experience Assurance platform that empowers organizations to deliver...  ...the ones they don’t own. Powered by AI and an unmatched set of cloud,...  ...offering. Your ImpactAs a FedRAMP Site Reliability Engineer(SRE), you will lead the operations and architecture... 
    Senior
    Full time
    Temporary work
    Work at office
    Local area
    Flexible hours
    2 days per week

    CISCO Systems

    Seattle, WA
    3 days ago
  • $160k - $220k

    Secure Every Identity, from AI to HumanIdentity is the key to unlocking the...  ...mission. If you are too, let's talk.Senior Database Reliability Engineer (DBRE) Experience Level: Mid-Senior...  ...systems. You will work closely with SRE, Platform, and Engineering teams to ensure performance... 
    Senior
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta

    Bellevue, WA
    4 days ago
  • $127k - $249k

    We are looking for an experienced Senior or Staff Engineer for our SRE, InfraSec team, to guide the security...  ...implement controls that reinforce the platform’s security posture.This is an SRE...  ...redefined the data platform for the AI era, enabling builders to create, transform... 
    Senior
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    Seattle, WA
    2 days ago
  •  ...Data Infrastructure SRE team is responsible for the reliability, scalability, and...  ...features, but about engineering the resilience and...  ...of the underlying platform that all product teams...  ...working alongside senior engineers to solve...  ...Pragmatic automation and AI orchestration:... 
    Senior

    TikTok

    Seattle, WA
    4 days ago
  • $78k - $185k

     ...week.ABOUT THE TEAMThe Platform Engineering team is part of...  ...enablement and system reliability in all phases of the Software...  ...etc.ABOUT THE ROLEThe Senior Platform Engineer will...  ...i.e., DevOps, SysOps, SRE) experience with a...  ...toolchainExperience using Gen AI tooling for software... 
    Senior
    Temporary work
    Work at office
    Local area
    Remote work
    Worldwide
    3 days per week

    Morgan Stanley

    Seattle, WA
    4 days ago
  • $160k - $210k

     ...Deep Learning Advertising Platform. Since 2015, we have harnessed...  ...be at the forefront of AI-driven advertising solutions...  ...growing! We are looking for a Senior Site Reliability Engineer to strengthen our AWS...  ...with our datacenter-focused SRE, so while deep, hands-on AWS... 
    Senior
    Work at office
    Immediate start
    Remote work
    Work from home

    Cognitiv

    Bellevue, WA
    7 days ago
  • Are you ready to shape the future of AI-driven reliability engineering? At JPMorganChase, we're building intelligent...  ...technology at enterprise scale.As a Senior Lead Software Engineer at JPMorganChase within the AI/ML Data Platforms organization, you will lead the... 
    Senior

    JP Morgan Chase

    Seattle, WA
    5 hours ago
  • $147k - $202.4k

    Secure Every Identity, from AI to HumanIdentity is the key to unlocking the potential...  ...company is seeking a highly skilled Senior Site Reliability Engineer to join our team. We are a SaaS...  ...for ensuring our systems and data platforms are compliant with industry standards... 
    Senior
    Work at office
    Local area
    Worldwide
    Flexible hours
    Shift work

    Okta

    Bellevue, WA
    3 days ago
  • $126.2k - $264.1k

    What You’ll DoAs a Senior Principal Network Reliability Engineer, you will:Provide technical leadership for the deployment...  ...engineering initiatives supporting AI infrastructure, new OCI Regions,...  ....Guide automation strategy and platform evolution through technical leadership... 
    Senior
    Temporary work
    Flexible hours

    Oracle Corporation

    Seattle, WA
    4 days ago
  • $160k - $250k

     ...DevOps And Systems Engineer Hive is the leading provider of cloud-based AI solutions to understand, search, and generate...  ...to grow our DevOps and Site Reliability team to maintain the reliability...  ...a diverse array of technology platforms, following best practices and procedures... 
    Senior

    Hive

    Seattle, WA
    19 hours ago
  •  ...Infrastructure Experts to evaluate AI-powered workflows across software development...  ..., cloud infrastructure, DevOps, SRE, and platform engineering. You will test AI-generated commands,...  ...workflows for accuracy and reliability. Work with AWS, Azure, GCP, Kubernetes... 
    Remote job
    For contractors

    YO AI Labs

    Seattle, WA
    15 days ago
  •  ...Ai2 in Seattle is seeking a Senior Platform Engineer to design the foundational architecture behind AI research agents. Responsibilities include developing APIs for data access, enhancing system reliability, and collaborating across teams. The ideal candidate will have... 
    Senior

    Ai2

    Seattle, WA
    2 days ago
  •  ...Flyway. Direct experience with the Snowflake platform and large-scale data systems. Excellent...  ...problem-solving skills with a proactive approach. AI/ML experience or a strong interest in applying AI/ML to reliability, security, or operational efficiency is a plus.... 
    Senior
    Full time
    Work at office
    Flexible hours

    Okta

    Bellevue, WA
    16 days ago
  • $148.5k - $223.9k

     ...DetailsAbout SalesforceSalesforce is the #1 AI CRM, where humans with agents drive...  ...for an experienced and hands-on software engineer to join our team to build and scale the next...  ...tools, frameworks, workflows, and validation platforms, applying AI tools, that help Salesforce... 
    Senior
    Full time

    Salesforce

    Bellevue, WA
    4 days ago
  •  ...The Allen Institute for Artificial Intelligence in Seattle is seeking a Senior Engineer to design foundational architecture for AI research agents. The ideal candidate has 8+ years of experience, strong Python skills, and a background in developing production services... 
    Senior

    The Allen Institute for Artificial Intelligence

    Seattle, WA
    2 days ago
  • $92.5k - $209.5k

     ...Infrastructure (OCI) is looking for a Senior Platform Software Engineer to join the PKI team responsible for...  ...deliver new capabilities, improve reliability and operational excellence, and contribute...  ...to life-saving care. And with AI embedded across our products and services... 
    Senior
    Temporary work
    Flexible hours

    Oracle Corporation

    Seattle, WA
    5 hours ago
  •  ...deliver top-notch technology products.As a Senior Lead Software Engineer at JPMorgan Chase within the Corporate Sector, Infrastructure Platforms team, you are an integral part of an...  ..., scalable cloud platforms optimized for AI/ML workloads.Partner with AI teams to translate... 
    Senior
    For contractors

    JP Morgan Chase

    Seattle, WA
    3 days ago
  • Union.ai is seeking a Systems Development Engineer in Seattle for a 50/50 operations and engineering role. You will work across cloud infrastructure, observability...  ...deployments, turning customer signals into durable platform improvements. You will sit in engineering, partner... 

    Linuxconfig

    Seattle, WA
    2 days ago
  • $120k - $170k

     ...Manager/Manager Site Reliability Engineering Join to apply for the...  ...role at Aritzia Get AI-powered advice on this...  ...Delivery team and lead the SRE team responsible for...  ...them with tools and platforms to support their work...  ...of your choosing. Seniority level Seniority level... 
    Senior
    Full time
    Work at office
    Remote work
    Flexible hours

    Aritzia

    Seattle, WA
    2 days ago
  •  ...Intelligence (CDI) is the central data, AI, and machine learning engineering team in ASE Media. We own the...  ...Store and more. The Data Engineering Platform (DEP) is how CDI does that work at...  ...operational tooling. We are looking for a Senior Platform Software Engineer to own... 
    Senior
    Contract work
    Work at office

    Apple

    Seattle, WA
    2 days ago
  • $135k - $155k

     ...making a real-world impact. This role sits at the intersection of platform engineering, data engineering, and data product development within...  ...target-state architecture for custom platform.• Work through AI agents by default. Use agentic tools to design, build, test,... 
    Senior
    Full time

    Valorem Reply

    Seattle, WA
    1 day ago
  • $92.5k - $209.5k

     ...moderately complex components within platform services or SDKs; leads team-...  ...ChatGPT Codex (or comparable AI coding agents) to accelerate...  ...production-quality engineering standards. Ability to effectively...  ..., health, support, and reliability. Core Responsibilities... 
    Senior
    Temporary work
    Immediate start
    Flexible hours
    Shift work

    Oracle

    Seattle, WA
    2 days ago
  •  ...SRE / DevOps Engineer Seattle based client. Seattle-WA (3 days onsite). U.S. Citizens and those authorized to work in the U.S. are encouraged to apply. We are unable to sponsor currently. Must Have Skills F5 load balancer Chef and Terraform - must have... 

    Georgia IT Inc

    Seattle, WA
    19 hours ago
  • $151.2k - $204.6k

     ....About the RolePlatform Core Engineering (PCE) builds the foundational...  ...engineers build on every day. As a Senior Product Manager on PCE, your...  ...their trust. You will own a platform product area and be measured...  ...to be hands-on with modern AI development tools, to... 
    Senior
    Flexible hours

    Twitch

    Seattle, WA
    3 days ago
  • $143k - $191k

     ...powered by Lattice OS, an AI-powered operating...  ...TEAMA Connected Warfare SRE - System Deployment deploys...  .... System Deployment Engineers work in complex environments...  ...data and applications platform and be able to educate...  ...with site reliability engineers to provide and... 
    Full time
    Temporary work
    Work experience placement
    Immediate start

    Anduril Industries

    Seattle, WA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to SRE - Senior AI Platform Reliability Engineer. Be the first to apply!