Senior Site Reliability Engineering- AI Infrastructure
$78kHCLTech
HCLTech is looking for a highly talented and self- motivated Senior Site Reliability Engineering– AI Infrastructure
to join it in advancing the technological world through innovation and creativity.
Job Title: Senior Site Reliability Engineering– AI Infrastructure
Job ID: 158107
Position Type: Full-time
Location: Remote
Role/Responsibilities
Engagement summary
The Candidate will provide SRE services for AI platforms and supporting infrastructure with emphasis on reliability engineering, incident response, service health, and operational automation. This role is best suited to a senior hands-on engineer who can improve availability while remaining effective in detailed production troubleshooting.
What this Candidate will be doing
• Operate and improve reliability of AI platform services, cluster dependencies, and shared infrastructure components.
• Lead or support incident triage for service degradation involving Kubernetes, Linux hosts, storage, network, scheduling, job orchestration, or dependency failures.
• Define and refine SLIs, SLOs, alerting thresholds, runbooks, escalation paths, and post-incident actions.
• Analyze recurring failure patterns and convert manual operations into automation and preventive controls.
• Build observability across system, service, workload, and dependency layers using metrics, logs, traces, and event correlation.
• Troubleshoot performance and availability issues affecting training jobs, inference services, internal platforms, and support tooling. • Partner with infrastructure and validation teams to improve production readiness and change safety.
• Drive operational reviews, readiness criteria, and resilience testing.
What we need to see
• 7+ years in SRE, production operations, or reliability-focused infrastructure engineering.
• Strong hands-on troubleshooting across Linux, Kubernetes, networking, and distributed systems.
• Experience building observability, alerting, and response workflows in complex production environments.
• Ability to balance urgent operational response with medium-term reliability engineering improvements.
• Strong scripting and automation skills, with experience reducing toil through tooling.
• Experience participating in incident management, root cause analysis, and post-incident follow through.
• Strong communication skill with the ability to summarize technical issues clearly for cross functional teams.
Preferred experience
• Experience in AI platforms, ML infrastructure, or large-scale HPC-like service environments.
• Familiarity with Prometheus, Grafana, ELK/OpenSearch, Loki, PagerDuty, and incident tooling.
• Experience defining error budgets and applying SRE practices in environments with heavy batch and service traffic.
Pay and Benefits
Pay Range Minimum: $78,000/Annum
Pay Range Maximum: $148,000/Annum
HCLTech is an equal opportunity employer, committed to providing equal employment opportunities to all applicants and employees regardless of race, religion, sex, color, age, national origin, pregnancy, sexual orientation, physical disability or genetic information, military or veteran status, or any other protected classification, in accordance with federal, state, and/or local law. Should any applicant have concerns about discrimination in the hiring process, they should provide a detailed report of those concerns to View email address on click.appcast.io for investigation.
Compensation and Benefits
A candidate’s pay within the range will depend on their work location, skills, experience, education, and other factors permitted by law. This role may also be eligible for performance-based bonuses subject to company policies. In addition, this role is eligible for the following benefits subject to company policies: medical, dental, vision, pharmacy, life, accidental death & dismemberment, and disability insurance; employee assistance program; 401(k) retirement plan; 10 days of paid time off per year (some positions are eligible for need-based leave with no designated number of leave days per year); and 10 paid holidays per year.
How You’ll Grow
At HCLTech, we offer continuous opportunities for you to find your spark and grow with us. We want you to be happy and satisfied with your role and to really learn what type of work sparks your brilliance the best. Throughout your time with us, we offer transparent communication with senior-level employees, learning and career development programs at every level, and opportunities to experiment in different roles or even pivot industries. We believe that you should be in control of your career with unlimited opportunities to find the role that fits you best.
$139.3k - $203.6k
...builds and operates secure, reliable cloud services for U.S.... ...with application engineering, security, compliance, and infrastructure teams to support the Webex... ...protect organizations in the AI era - and beyond. We’ve... ...see the Cisco careers site to discover more benefits...SeniorFull timeTemporary workWork at officeLocal areaFlexible hoursShift work- ...UsAlembic is the pioneering Causal AI platform. We help the world's... ...built on Grace Blackwell infrastructure — one of the fastest private supercomputers... ...under real-world scale, reliability, and security demands — and we're looking for an engineer who wants to own the...Senior
- The Data Infrastructure SRE team is responsible for the reliability, scalability, and efficiency... ..., but about engineering the resilience and... ...system stability.As a Site Reliability Engineer... ...working alongside senior engineers to solve... ...automation and AI orchestration: Design...Senior
- ...time.Let’s make.Job DescriptionWe're looking for a Senior Site Reliability Engineer, Platform Infrastructure to take hands-on technical ownership of the architecture... ...seamless 24/7 reliability. It's ideal for an AI-forward engineer with a strong software engineering...SeniorWork at officeRemote workRelocationRelocation package
- ...identity security, delivering an AI-powered platform that... ...systems. As a Staff Platform Engineer, you will play a critical role... ...role. You will own reliability for major platform domains,... ...and maintaining the shared infrastructure services and platforms that...SeniorFull time
- ...THE ROLE This is a senior on-call SRE role at an early-stage AI infrastructure company, where you will... ...on-call rotation with engineers across multiple time... ...infrastructure team on long-term reliability improvements and... .... LOCATION On-site in San Francisco, CA....SeniorImmediate start
$232k - $319k
Secure Every Identity, from AI to HumanIdentity is the key... ...building the trusted, neutral infrastructure that enables organizations... ...service with great people and reliable, cost-effective, and... ...velocity of SRE and product engineering by developing robust platforms...SeniorPermanent employmentLocal areaWorldwideFlexible hours- ...identity security, delivering an AI-powered platform that... ...systems. As a Staff Platform Engineer, you will play a critical role... ...role. You will own reliability for major platform domains,... ...and maintaining the shared infrastructure services and platforms that...Senior
$55 - $60 per hour
...Overview: Our client is looking for an experienced Site Reliability Engineer (SRE) to join the Infrastructure Platform Engineering team. In this role, the... ...infrastructure management and cutting-edge agentic AI tooling, building robust services, telemetry platforms...SeniorTemporary workLocal area- ...Summary We are seeking a Senior SRE / DevSecOps Engineer with strong experience... ..., observability, infrastructure automation, and AI-assisted troubleshooting... ...will focus on platform reliability, incident management, SLO... ...SLO/SLI governance and site reliability practices....SeniorContract work
$104.9k - $174.7k
About the role:A FinOps Site Reliability Engineer (SRE) bridges the gap between engineering, operations... ...by embedding cost optimization into infrastructure design, automation, monitoring, and... ...Terraform, observability, automation, AI platforms, and cloud financial...SeniorFull timeLocal area- ...Summary We are seeking a Lead Site Reliability & Environment Monitoring Engineer to establish and evolve our enterprise... ...: Azure-based SRE tooling or AI-assisted operations Automation... ...Actions, Runbooks, etc.) Infrastructure as Code (Terraform, ARM, Bicep)...SeniorFull timeShift work
- ...Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands... ...day is currently Tuesday.Engineering at Lambda is responsible... ...teams to improve service reliability and deployment... ...5+ years of experience in Site Reliability Engineering, Production...SeniorWork at officeLocal areaWork from homeFlexible hours
$165k - $265k
...enabling human life on Mars.SR. SITE RELIABILITY ENGINEER (STARSHIELD) At SpaceX we’... ...'s software and GPU infrastructure, you will design, operate and... ...and productize solutions for AI clusters (100k+ GPU scale)... ...train junior engineersAs a senior engineer you must lead the...SeniorPermanent employmentTemporary workImmediate startWeekend work- ...cybersecurity. We protect how people, data, and AI agents connect across email, cloud,... ...in execution and impactThe RoleAs a Senior Site Reliability Engineer at Proofpoint you will develop a... ...team player who cares about the infrastructure, remains calm in crisis,...SeniorFull timeFlexible hours
- ...Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of... ...work from home day is currently Tuesday.Engineering at Lambda is responsible for building... ...Kubernetes services, workloads, and platform reliability.You6+ years of experience in a SRE,...SeniorWork at officeLocal areaWork from homeFlexible hours
$152.5k - $205k
...applications, and programmable blockchain infrastructure. Circle’s platform includes the... ...What you’ll be responsible for:As a Senior Site Reliability Engineer on Circle’s platform team, you’ll... ...infrastructure behind critical digital-assets, AI, and application workloads. You will...SeniorFlexible hours- ...candidate for this role to work on site in the specified location(s).As a Site Reliability Engineer supporting the Cashiering... ...application development teams, infrastructure partners, business stakeholders... ...based applicationsExperience with AI-enabled productivity and...SeniorFull timeWork at office
$168k - $270.25k
...intelligence.Join our team at NVIDIA as a Senior Site reliability engineer focused on HPC storage and play a... ...with engineering teams to align infrastructure with their evolving needs, document... ...for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA...SeniorFull time$15k
...company that applies state-of-the-art AI and machine learning techniques to... ...catered lunches, and more.As a Senior Cluster Site Reliability Engineer (SRE), you will help scale our research... ...support both on-prem and cloud infrastructure, and work to provide the best experience...SeniorWork at officeLocal areaRemote work$190.8k - $267.1k
...grow its business. The reliability of our Ads systems... ...partners closely with Ads Engineering to improve... ...We’re looking for a Senior Site Reliability Engineer... ...build, and maintain infrastructure, tooling, and automation... ...artificial intelligence (AI). You will have the opportunity...SeniorFor contractorsWork experience placement- ...software solutions harness the power of AI and shape the future of... ...join our team and make an impact?As a Senior Site Reliability Engineer at TeamViewer, you’ll be a key player... ...available, secure, and scalable Azure cloud infrastructure supporting TeamViewer’s global SaaS...SeniorTemporary workCasual workWorldwide
$267k - $356k
...Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands... ...Tuesday.Lambda's Storage Engineering team is the backbone behind... ...the industry, which means reliability and performance aren't just... ...across new and existing sites using tools such as Ansible...SeniorWork experience placementWork at officeLocal areaWork from homeFlexible hours$91.7k - $163.7k
...everyone. From advanced data analytics and AI to cybersecurity, we use innovative... .... Connecting. Growing together. The Site Reliability Engineer will architect, develop, and maintain... ...resilient and high performance cloud infrastructure. You'll enjoy the flexibility to work...SeniorMinimum wageFull timeWork experience placementWork at officeLocal areaRemote work$148k - $235.75k
...into the unlimited potential of AI to define the next era of... ....Join our team of innovative engineers who are building an AI Data Center... ..., high-volume telemetry into reliable, job-centric insights and... ...automation.Manage deployment infrastructure and packaging (Helm + Terraform...SeniorFull time$127k - $249k
The TeamPlatform Engineering sits within SRE and builds the core infrastructure powering MongoDB’s broader... ...role in engineering the reliable, globally connected,... ...are seeking a talented Senior Site Reliability Engineer (... ...data platform for the AI era, enabling builders...SeniorLocal areaRemote workWorldwideFlexible hours$80k - $140k
...Management Technology is seeking a Senior Site Reliability Engineer to join its Wealth Management SRE Team... ...will work closely with development, infrastructure, platform, and support teams to... ...objectives, and shaping the future of AI-enhanced operations. You will help design...SeniorFull timeFlexible hoursShift work$160k - $200k
...days/per week.Tulip, the leader in AI-native frontline operations, is helping... ...best practices, SLIs/SLOs, and reliability culture across engineering teams. Contributing to and maintaining... ..., build, and maintain the core infrastructure & tooling used by all of Tulip’s engineering...SeniorTemporary workWork at officeLocal areaFlexible hours3 days per week- ...GIPHY is seeking a highly experienced Site Reliability Engineer to join our SRE team. You will help... ...design, build, operate, and evolve the infrastructure that powers GIPHY, including our... ...particularly agentic development and AI-assisted engineering, and identify opportunities...SeniorFull timeWork experience placementRemote work
- ...AirwallexAirwallex is the AI-native financial... ...in 2015 to build the infrastructure global commerce runs... ...full.About the teamThe Engineering team at Airwallex is... ...together to build scalable, reliable, and secure products... ....What you’ll doAs a Senior Site Reliability Engineer,...SeniorTemporary workLocal area
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Site Reliability Engineering- AI Infrastructure. Be the first to apply!
- site reliability engineering manager United States
- site reliability engineer sre United States
- site reliability engineer United States
- site reliability engineer remote United States
- lead site reliability engineer United States
- remote infrastructure engineer United States
- infrastructure engineer United States
- principal infrastructure engineer United States
- senior infrastructure engineer United States
- associate infrastructure engineer United States




