Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Site Reliability Engineering- AI Infrastructure

$78k
Full-time

HCLTech

HCLTech is looking for a highly talented and self- motivated Senior Site Reliability Engineering– AI Infrastructure

to join it in advancing the technological world through innovation and creativity.

Job Title: Senior Site Reliability Engineering– AI Infrastructure

Job ID: 158107

Position Type: Full-time

Location: Remote

Role/Responsibilities

Engagement summary

The Candidate will provide SRE services for AI platforms and supporting infrastructure with emphasis on reliability engineering, incident response, service health, and operational automation. This role is best suited to a senior hands-on engineer who can improve availability while remaining effective in detailed production troubleshooting.

What this Candidate will be doing

• Operate and improve reliability of AI platform services, cluster dependencies, and shared infrastructure components.

• Lead or support incident triage for service degradation involving Kubernetes, Linux hosts, storage, network, scheduling, job orchestration, or dependency failures.

• Define and refine SLIs, SLOs, alerting thresholds, runbooks, escalation paths, and post-incident actions.

• Analyze recurring failure patterns and convert manual operations into automation and preventive controls.

• Build observability across system, service, workload, and dependency layers using metrics, logs, traces, and event correlation.

• Troubleshoot performance and availability issues affecting training jobs, inference services, internal platforms, and support tooling. • Partner with infrastructure and validation teams to improve production readiness and change safety.

• Drive operational reviews, readiness criteria, and resilience testing.

What we need to see

• 7+ years in SRE, production operations, or reliability-focused infrastructure engineering.

• Strong hands-on troubleshooting across Linux, Kubernetes, networking, and distributed systems.

• Experience building observability, alerting, and response workflows in complex production environments.

• Ability to balance urgent operational response with medium-term reliability engineering improvements.

• Strong scripting and automation skills, with experience reducing toil through tooling.

• Experience participating in incident management, root cause analysis, and post-incident follow through.

• Strong communication skill with the ability to summarize technical issues clearly for cross functional teams.

Preferred experience

• Experience in AI platforms, ML infrastructure, or large-scale HPC-like service environments.

• Familiarity with Prometheus, Grafana, ELK/OpenSearch, Loki, PagerDuty, and incident tooling.

• Experience defining error budgets and applying SRE practices in environments with heavy batch and service traffic.

Pay and Benefits

Pay Range Minimum: $78,000/Annum

Pay Range Maximum: $148,000/Annum

HCLTech is an equal opportunity employer, committed to providing equal employment opportunities to all applicants and employees regardless of race, religion, sex, color, age, national origin, pregnancy, sexual orientation, physical disability or genetic information, military or veteran status, or any other protected classification, in accordance with federal, state, and/or local law. Should any applicant have concerns about discrimination in the hiring process, they should provide a detailed report of those concerns to View email address on click.appcast.io for investigation.

Compensation and Benefits

A candidate’s pay within the range will depend on their work location, skills, experience, education, and other factors permitted by law. This role may also be eligible for performance-based bonuses subject to company policies. In addition, this role is eligible for the following benefits subject to company policies: medical, dental, vision, pharmacy, life, accidental death & dismemberment, and disability insurance; employee assistance program; 401(k) retirement plan; 10 days of paid time off per year (some positions are eligible for need-based leave with no designated number of leave days per year); and 10 paid holidays per year.

How You’ll Grow

At HCLTech, we offer continuous opportunities for you to find your spark and grow with us. We want you to be happy and satisfied with your role and to really learn what type of work sparks your brilliance the best. Throughout your time with us, we offer transparent communication with senior-level employees, learning and career development programs at every level, and opportunities to experiment in different roles or even pivot industries. We believe that you should be in control of your career with unlimited opportunities to find the role that fits you best.

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Senior Site Reliability Engineering- AI Infrastructure in United States vacancy
  • $139.3k - $203.6k

     ...builds and operates secure, reliable cloud services for U.S....  ...with application engineering, security, compliance, and infrastructure teams to support the Webex...  ...protect organizations in the AI era - and beyond. We’ve...  ...see the Cisco careers site to discover more benefits... 
    Senior
    Full time
    Temporary work
    Work at office
    Local area
    Flexible hours
    Shift work

    CISCO Systems

    Boxborough, MA
    3 days ago
  •  ...UsAlembic is the pioneering Causal AI platform. We help the world's...  ...built on Grace Blackwell infrastructure — one of the fastest private supercomputers...  ...under real-world scale, reliability, and security demands — and we're looking for an engineer who wants to own the... 
    Senior

    Alembic

    San Francisco, CA
    1 day ago
  • The Data Infrastructure SRE team is responsible for the reliability, scalability, and efficiency...  ..., but about engineering the resilience and...  ...system stability.As a Site Reliability Engineer...  ...working alongside senior engineers to solve...  ...automation and AI orchestration: Design... 
    Senior

    TikTok

    Seattle, WA
    1 day ago
  •  ...time.Let’s make.Job DescriptionWe're looking for a Senior Site Reliability Engineer, Platform Infrastructure to take hands-on technical ownership of the architecture...  ...seamless 24/7 reliability. It's ideal for an AI-forward engineer with a strong software engineering... 
    Senior
    Work at office
    Remote work
    Relocation
    Relocation package

    Cricut

    Riverton, UT
    4 days ago
  •  ...identity security, delivering an AI-powered platform that...  ...systems. As a Staff Platform Engineer, you will play a critical role...  ...role. You will own reliability for major platform domains,...  ...and maintaining the shared infrastructure services and platforms that... 
    Senior
    Full time

    Saviynt

    Atlanta, TX
    4 days ago
  •  ...THE ROLE This is a senior on-call SRE role at an early-stage AI infrastructure company, where you will...  ...on-call rotation with engineers across multiple time...  ...infrastructure team on long-term reliability improvements and...  .... LOCATION On-site in San Francisco, CA.... 
    Senior
    Immediate start

    Jobleads-US

    San Francisco, CA
    3 days ago
  • $232k - $319k

    Secure Every Identity, from AI to HumanIdentity is the key...  ...building the trusted, neutral infrastructure that enables organizations...  ...service with great people and reliable, cost-effective, and...  ...velocity of SRE and product engineering by developing robust platforms... 
    Senior
    Permanent employment
    Local area
    Worldwide
    Flexible hours

    Okta

    Bellevue, WA
    2 days ago
  •  ...identity security, delivering an AI-powered platform that...  ...systems. As a Staff Platform Engineer, you will play a critical role...  ...role. You will own reliability for major platform domains,...  ...and maintaining the shared infrastructure services and platforms that... 
    Senior

    Saviynt

    Milpitas, CA
    a month ago
  • $55 - $60 per hour

     ...Overview: Our client is looking for an experienced Site Reliability Engineer (SRE) to join the Infrastructure Platform Engineering team. In this role, the...  ...infrastructure management and cutting-edge agentic AI tooling, building robust services, telemetry platforms... 
    Senior
    Temporary work
    Local area

    CYNET SYSTEMS

    Santa Clara, CA
    5 days ago
  •  ...Summary We are seeking a Senior SRE / DevSecOps Engineer with strong experience...  ..., observability, infrastructure automation, and AI-assisted troubleshooting...  ...will focus on platform reliability, incident management, SLO...  ...SLO/SLI governance and site reliability practices.... 
    Senior
    Contract work

    PB consulting

    Charlotte, NC
    a month ago
  • $104.9k - $174.7k

    About the role:A FinOps Site Reliability Engineer (SRE) bridges the gap between engineering, operations...  ...by embedding cost optimization into infrastructure design, automation, monitoring, and...  ...Terraform, observability, automation, AI platforms, and cloud financial... 
    Senior
    Full time
    Local area

    LexisNexis Risk Solutions Group

    Boca Raton, FL
    4 days ago
  •  ...Summary We are seeking a Lead Site Reliability & Environment Monitoring Engineer to establish and evolve our enterprise...  ...: Azure-based SRE tooling or AI-assisted operations Automation...  ...Actions, Runbooks, etc.) Infrastructure as Code (Terraform, ARM, Bicep)... 
    Senior
    Full time
    Shift work

    CRC Group

    Charlotte, NC
    3 days ago
  •  ...Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands...  ...day is currently Tuesday.Engineering at Lambda is responsible...  ...teams to improve service reliability and deployment...  ...5+ years of experience in Site Reliability Engineering, Production... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    1 day ago
  • $165k - $265k

     ...enabling human life on Mars.SR. SITE RELIABILITY ENGINEER (STARSHIELD) At SpaceX we’...  ...'s software and GPU infrastructure, you will design, operate and...  ...and productize solutions for AI clusters (100k+ GPU scale)...  ...train junior engineersAs a senior engineer you must lead the... 
    Senior
    Permanent employment
    Temporary work
    Immediate start
    Weekend work

    SpaceX

    Palo Alto, CA
    4 days ago
  •  ...cybersecurity. We protect how people, data, and AI agents connect across email, cloud,...  ...in execution and impactThe RoleAs a Senior Site Reliability Engineer at Proofpoint you will develop a...  ...team player who cares about the infrastructure, remains calm in crisis,... 
    Senior
    Full time
    Flexible hours

    Proofpoint

    Austin, TX
    5 days ago
  •  ...Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of...  ...work from home day is currently Tuesday.Engineering at Lambda is responsible for building...  ...Kubernetes services, workloads, and platform reliability.You6+ years of experience in a SRE,... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    3 days ago
  • $152.5k - $205k

     ...applications, and programmable blockchain infrastructure. Circle’s platform includes the...  ...What you’ll be responsible for:As a Senior Site Reliability Engineer on Circle’s platform team, you’ll...  ...infrastructure behind critical digital-assets, AI, and application workloads. You will... 
    Senior
    Flexible hours

    Circle

    San Francisco, CA
    3 days ago
  •  ...candidate for this role to work on site in the specified location(s).As a Site Reliability Engineer supporting the Cashiering...  ...application development teams, infrastructure partners, business stakeholders...  ...based applicationsExperience with AI-enabled productivity and... 
    Senior
    Full time
    Work at office

    The Charles Schwab Corporation

    Southlake, TX
    3 days ago
  • $168k - $270.25k

     ...intelligence.Join our team at NVIDIA as a Senior Site reliability engineer focused on HPC storage and play a...  ...with engineering teams to align infrastructure with their evolving needs, document...  ...for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $15k

     ...company that applies state-of-the-art AI and machine learning techniques to...  ...catered lunches, and more.As a Senior Cluster Site Reliability Engineer (SRE), you will help scale our research...  ...support both on-prem and cloud infrastructure, and work to provide the best experience... 
    Senior
    Work at office
    Local area
    Remote work

    The Voleon Group

    Berkeley, CA
    2 days ago
  • $190.8k - $267.1k

     ...grow its business. The reliability of our Ads systems...  ...partners closely with Ads Engineering to improve...  ...We’re looking for a Senior Site Reliability Engineer...  ...build, and maintain infrastructure, tooling, and automation...  ...artificial intelligence (AI). You will have the opportunity... 
    Senior
    For contractors
    Work experience placement

    Reddit

    San Francisco, CA
    4 days ago
  •  ...software solutions harness the power of AI and shape the future of...  ...join our team and make an impact?As a Senior Site Reliability Engineer at TeamViewer, you’ll be a key player...  ...available, secure, and scalable Azure cloud infrastructure supporting TeamViewer’s global SaaS... 
    Senior
    Temporary work
    Casual work
    Worldwide

    TeamViewer

    Austin, TX
    5 days ago
  • $267k - $356k

     ...Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands...  ...Tuesday.Lambda's Storage Engineering team is the backbone behind...  ...the industry, which means reliability and performance aren't just...  ...across new and existing sites using tools such as Ansible... 
    Senior
    Work experience placement
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    2 days ago
  • $91.7k - $163.7k

     ...everyone. From advanced data analytics and AI to cybersecurity, we use innovative...  .... Connecting. Growing together. The Site Reliability Engineer will architect, develop, and maintain...  ...resilient and high performance cloud infrastructure. You'll enjoy the flexibility to work... 
    Senior
    Minimum wage
    Full time
    Work experience placement
    Work at office
    Local area
    Remote work

    UnitedHealth Group

    Eden Prairie, MN
    5 days ago
  • $148k - $235.75k

     ...into the unlimited potential of AI to define the next era of...  ....Join our team of innovative engineers who are building an AI Data Center...  ..., high-volume telemetry into reliable, job-centric insights and...  ...automation.Manage deployment infrastructure and packaging (Helm + Terraform... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $127k - $249k

    The TeamPlatform Engineering sits within SRE and builds the core infrastructure powering MongoDB’s broader...  ...role in engineering the reliable, globally connected,...  ...are seeking a talented Senior Site Reliability Engineer (...  ...data platform for the AI era, enabling builders... 
    Senior
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    San Francisco, CA
    5 days ago
  • $80k - $140k

     ...Management Technology is seeking a Senior Site Reliability Engineer to join its Wealth Management SRE Team...  ...will work closely with development, infrastructure, platform, and support teams to...  ...objectives, and shaping the future of AI-enhanced operations. You will help design... 
    Senior
    Full time
    Flexible hours
    Shift work

    Royal Bank of Canada

    Minneapolis, MN
    5 days ago
  • $160k - $200k

     ...days/per week.Tulip, the leader in AI-native frontline operations, is helping...  ...best practices, SLIs/SLOs, and reliability culture across engineering teams. Contributing to and maintaining...  ..., build, and maintain the core infrastructure & tooling used by all of Tulip’s engineering... 
    Senior
    Temporary work
    Work at office
    Local area
    Flexible hours
    3 days per week

    Tulip Interface

    Somerville, MA
    5 days ago
  •  ...GIPHY is seeking a highly experienced Site Reliability Engineer to join our SRE team. You will help...  ...design, build, operate, and evolve the infrastructure that powers GIPHY, including our...  ...particularly agentic development and AI-assisted engineering, and identify opportunities... 
    Senior
    Full time
    Work experience placement
    Remote work

    Shutterstock

    New York, NY
    1 day ago
  •  ...AirwallexAirwallex is the AI-native financial...  ...in 2015 to build the infrastructure global commerce runs...  ...full.About the teamThe Engineering team at Airwallex is...  ...together to build scalable, reliable, and secure products...  ....What you’ll doAs a Senior Site Reliability Engineer,... 
    Senior
    Temporary work
    Local area

    Airwallex

    San Francisco, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Site Reliability Engineering- AI Infrastructure. Be the first to apply!