Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer (SRE) Intern — AI Infrastructure

$27.12 - $51.93 per hour

Tencent

About the Hiring Team

What the Role Entails

Role Summary We are seeking a motivated Site Reliability Engineer (SRE) Intern to join our AI Compute team, supporting the daily operations of AI infrastructure. In this role, you will work closely with internal business teams and external engineering partners to build and operate AI infrastructure. This is a hands-on opportunity to gain knowledge and experience in cutting-edge AI infrastructure.

Key Responsibilities Support the deployment, configuration, and maintenance of high-performance AI infrastructure servers, storage servers, networking equipment, and software components in secure environments. Assist with hardware diagnostics, system functionality checks, and firmware updates as required. Collaborate with engineering teams to help deliver tailored customer environments (e.g., bare-metal systems, Kubernetes, Slurm, etc.). Provide first-line engineering support for onsite operational issues, including troubleshooting hardware, network, and software problems, and firmware compliance. Document incident details, resolutions, and lessons learned to improve future problem-solving. Maintain clear, accurate, and up-to-date documentation to support knowledge sharing across the team. Participate in team meetings and knowledge-sharing sessions to foster collaboration and continuous learning. Who We Look For

Qualifications & Requirements

Currently pursuing or recently completed a Bachelor's or Master's degree in computer engineering, computer science, or a related technical field. Basic understanding of server hardware, firmware lifecycle, and Linux environments, with an awareness of physical and system-level security standards. Exposure to scripting languages such as Bash or Python. Familiarity with — or strong interest in — configuration management, CI/CD tools, workload managers, and cluster software (e.g., Slurm, Kubernetes), and observability tools (e.g., Prometheus, Grafana, ELK). Strong problem-solving and analytical skills. Ability to work both independently and as part of a team. Professional fluency in English and Mandarin is highly preferred Preferred Qualifications

Coursework, projects, or hands-on experience related to AI Infrastructure, distributed systems, or cloud infrastructure. Familiarity with networking fundamentals and Linux system administration. Genuine interest in AI/ML infrastructure. Location State(s) US-California-Palo AltoThe expected base pay range for this position in the location(s) listed above is $27.12 to $51.93 per hour. Actual pay may vary depending on job-related knowledge, skills, and experience. This position will be eligible for 1 hour of paid sick leave for every 30 hours worked and up to 13 paid holidays throughout the calendar year. Subject to the terms and conditions of the applicable plans then in effect, full-time interns are also eligible to enroll in the Company-sponsored medical plan.

Equal Employment Opportunity at Tencent

As an equal opportunity employer, we firmly believe that diverse voices fuel our innovation and allow us to better serve our users and the community. We foster an environment where every employee of Tencent feels supported and inspired to achieve individual and common goals.
Vacancy posted 1 hour ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer (SRE) Intern — AI Infrastructure in Palo Alto, CA vacancy
  • $20k

     ...goal of enabling human life on Mars.SITE RELIABILITY ENGINEER (STARSHIELD) At SpaceX we’re leveraging...  ...on Starshield's software and GPU infrastructure, you will design, operate and scale the...  ...validate, and productize solutions for AI clusters (100k+ GPU scale)Develop... 
    Suggested
    Permanent employment
    Temporary work
    Immediate start
    Weekend work

    SpaceX

    Palo Alto, CA
    17 hours ago
  • $195k - $265k

     ...we scale toward launch, our engineering infrastructure has to scale with us. We're hiring a Staff Site Reliability Engineer to own source control...  ...tooling that reduce toil for SRE and product engineering alike...  ...use artificial intelligence (AI) tools to support parts of the... 
    Suggested
    Full time

    Zoox

    Foster, CA
    1 day ago
  • $61k - $101k

     ...or certification in site reliability engineering, along with 5+ years...  ...advanced knowledge of SRE culture and principles...  ...development, automation, and infrastructure as code. We...  ...-authorized AI capabilities in the workplace...  ...reliability community through internal forums, communities... 
    Suggested
    Full time

    J.P. Morgan

    Palo Alto, CA
    19 days ago
  •  ...Description Developer & Infrastructure Expert Role Type:...  ...to evaluate AI-powered workflows across...  ...infrastructure, DevOps, SRE, and platform engineering. You will test AI-...  ...for accuracy and reliability. Work with AWS,...  ...Cloud Infrastructure Site Reliability... 
    Suggested
    Remote job
    For contractors

    YO AI Labs

    Palo Alto, CA
    29 days ago
  • $101k - $161k

     ...prestigious awards, such as Best Engineering Team, Best Company for...  ...Work WithWe’re looking for Site Reliability Engineers to join our...  ...as-a-Service (CVaaS) global SRE team. SREs at Arista combine...  ...microservices stack, monitoring infrastructure, and much more. What You'll... 
    Suggested

    Arista Networks

    Santa Clara, CA
    3 days ago
  •  ...Future Workforce Wand turns AI into labor. It enables humans...  ...the world’s first Agentic Labor Infrastructure enabling governments and...  ...hiring for a hands‑on Head of SRE to establish, lead, and scale our Site Reliability Engineering function. This role combines strategic... 
    Shift work

    Wand AI

    Palo Alto, CA
    8 hours ago
  • $222k - $300.5k

     ...OverviewAbout the TeamIntuit's Infrastructure and Site Reliability organization owns...  ...Platform Systems Engineering team builds and...  ...security, and other SRE/infrastructure leaders...  ...for this role is AI Ops: embedding AI-driven...  ...customer-facing and internal systems.Lead, grow,... 
    Worldwide
    Shift work

    Intuit

    Mountain View, CA
    3 days ago
  • $262k - $364k

     ...ecosystem have reliability and uptime appropriate...  ....Build creative engineering solutions to operations and infrastructure problems, including AI-powered...  ...contribute to the cross-SRE AI Ops program,...  ...a strategic way.Site Reliability...  ...services—both our internally critical and our... 

    Google

    Mountain View, CA
    20 hours ago
  •  ...Site Reliability Engineer There are NO limits to your career: come shape the...  ...Site Reliability Engineering (SRE) is a discipline that...  ...engineering and applies them to infrastructure and operations problems....  ...in Python supported by Gen AI tooling to accelerate development... 
    Immediate start
    Remote work
    Worldwide

    OutSystems

    Menlo Park, CA
    4 days ago
  •  ...Site Reliability Engineer III There's nothing more exciting than being at the...  ...& Analytics Office (CDAO) AI/ML & Data Platforms team, you...  .... Through code and cloud infrastructure, you will configure, maintain...  ...environment to support SRE workflows with strong validation... 
    Work at office

    Chase

    Palo Alto, CA
    4 days ago
  • $152k - $228k

     ...profound opportunity for AI to drive positive...  ...You will own the infrastructure that makes this...  ...much much more.Engineers across the company...  ...quality gates.Platform Reliability & Observability:...  ...: Guide the SRE team through the OS...  ...Introduction), SRE (Site Reliability Engineering... 
    Temporary work
    Immediate start
    Flexible hours

    Nuro

    Mountain View, CA
    3 days ago
  •  ...DESCRIPTION Elevate your engineering prowess to...  ...top echelon in site reliability. As a Senior Lead...  ...Chase within the Infrastructure Platforms and Foundational...  ...-authorized AI capabilities...  ...reliability community via internal forums,...  ...particularly in SRE, observability, or... 

    J.P. Morgan

    Palo Alto, CA
    2 days ago
  • $226k - $369k

     ...approval. LinkedIn’s Reliability Infrastructure team is responsible for...  ...Principal Staff Software Engineer, Reliability...  ...handling are in place.As AI-assisted software development...  ...is not a traditional SRE role focused on...  ...measurable improvements in site stabilityHelp shape how... 
    For contractors
    Work at office
    Remote work
    Work from home
    Flexible hours

    Linkedin

    Mountain View, CA
    20 hours ago
  • $175k - $287k

     ...team. LinkedIn’s Core AI is building the Evaluation...  ...for tracing infrastructure for all LinkedIn AI Agents.As a Staff Engineer, you will own the end-to...  ...continuously improve the quality, reliability, safety, and...  ...environment with other SRE/SWE Engineers, Project... 
    For contractors
    Work experience placement
    Work at office
    Flexible hours

    Linkedin

    Mountain View, CA
    1 day ago
  •  ...Lead Site Reliability Engineer Assume a critical role in defining the future...  ...applications Implements infrastructure, configuration, and network...  ...Uses enterprise-authorized AI capabilities within the work...  ...site reliability engineering (SRE) culture and principles,... 

    Chase

    Palo Alto, CA
    7 hours ago
  •  ...defense tech startups with the tools, infrastructure, and compliance solutions they need...  ...are seeking a highly motivated Systems Reliability Engineer (SRE) to lead the design and...  ...sensitive and cleared workforces. The Site Reliability Engineer (SRE) - SecOps... 
    For contractors
    Work at office
    Flexible hours

    Arkenstone Defense

    Menlo Park, CA
    12 days ago
  • $100k - $200k

    OPPO US Research Center is seeking a skilled and proactive Site Reliability Engineer (SRE) to join our team. In this role, you will be responsible...  ...systems. The ideal candidate is passionate about cloud infrastructure, automation, and building reliable, production-grade... 
    Full time

    OPPO

    Palo Alto, CA
    1 day ago
  • $190.98k - $214.05k

     ...otherwise, all roles are on-site five days per week...  ...thoughtful use of AI tools in our daily...  ...a Senior Software Engineer (Platform & Data Infrastructure) to architect, scale...  ...Platform, Security, SRE, and Product Engineering...  .... Drive Reliability, DR & On-Call Excellence... 
    Full time
    Temporary work
    Work at office
    Remote work
    Flexible hours

    ID.me

    Mountain View, CA
    1 day ago
  • $125k - $195k

     ...Responsibilities Manage GPU and CPU infrastructure deployments to Top Secret...  ..., validate, and productize AI cluster solutions at 100,000...  .... Collaborate with AI engineers to build scalable, operable,...  ...one year of professional site reliability engineering or DevOps... 
    Permanent employment
    Full time
    Temporary work
    Weekend work

    SpaceX

    Palo Alto, CA
    11 hours ago
  •  ...Social, we're building the AI-native social operating...  ...Founded by ex-Meta product and engineering leaders, we've raised over...  ...Senior Software Engineer, Infrastructure to own the reliability, scalability, and...  ...base, and we need a seasoned SRE to help us scale these systems... 
    Work at office
    Remote work
    Flexible hours
    Shift work

    Nectar Social

    Palo Alto, CA
    1 day ago
  • $207k - $300k

     ...team of Software/Systems Engineers on projects for users and...  ...with Large Language Model.Site Reliability Engineering (SRE) combines software and...  ...Google's services—both our internally critical and our externally...  ...systems, building infrastructure and eliminating work through... 

    Google

    Sunnyvale, CA
    1 day ago
  • $160.36k - $240.54k

     ...profound opportunity for AI to drive positive...  ...driving, and the ML Infrastructure team builds and operates...  ...care as much about reliability and operational maturity...  ...Science, Electrical Engineering, or a closely related...  ...distributed training internals, including NCCL and... 
    Work experience placement
    Immediate start
    Flexible hours

    Nuro

    Mountain View, CA
    2 days ago
  • $160k - $260k

     ...RoleWe're looking for a talented, driven Senior Software Engineer to build the critical infrastructure that powers the Aptos blockchain ecosystem.This role...  ...payments, DeFi, or TradFi (a plus)Comfortable leveraging AI tools to enhance productivity and code qualityThe base... 
    Full time
    Work experience placement
    Work at office
    Local area

    Aptos Labs

    Palo Alto, CA
    3 days ago
  • $224k - $356.5k

    NVIDIA is hiring engineers to scale up the introduction...  ...into its EDA Infrastructure. We expect you to have...  ...effective, clear and reliable architecture specificationTranslate...  ..., DevOps and/or SRE practices and/or Platform...  ...crowd:Developing ML/AI infrastructure. Developing... 
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $120k - $195k

     ...part of our world-class software engineering team, you will help build the next-generation infrastructure and platforms that power LinkedIn’s products, business, and AI-first future. This includes...  ...foundational platforms that enable reliable, scalable, and AI-first product... 
    For contractors
    Work experience placement
    Work at office
    Flexible hours

    Linkedin

    Mountain View, CA
    1 day ago
  • $230k - $250k

     ...autonomous networking, giving engineers and AI agents the ability to know...  ....Forward is looking for a Site Reliability EngineerAbout the Role This...  ...not a "keep the lights on" SRE role. As our first or early...  ...closely with engineering, infrastructure, and product to ensure our... 
    Night shift

    Forward Networks

    Santa Clara, CA
    20 hours ago
  • $148k - $235.75k

     ...unlimited potential of AI to define the next era...  ...team of innovative engineers who are building an AI...  ...volume telemetry into reliable, job-centric insights...  ...automation.Manage deployment infrastructure and packaging (Helm +...  ...systems as SRE/DevOps/Platform Ops.Proven... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $152k - $241.5k

     ...looking for a Senior SRE to join our...  ...harness the power of AI to deliver groundbreaking...  ...fabrics.Use IaC(Infrastructure‑as‑Code) and...  ...Service (QoS) for internal customers through...  ...management, fleet reliability/auto-healing, E2E...  ...Ruby.Mentored other engineers and influenced technical... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $180k - $230k

     ...most critical constraint in AI’s growth trajectory:...  ...defining bottleneck in the AI infrastructure race. While leading tech companies...  ...'re looking for a Senior SRE to own the reliability, scalability, and...  ...closely with platform and data engineering to keep high-throughput,... 
    Work at office
    Local area
    Immediate start
    Remote work
    3 days per week

    GridCARE, Inc.

    Redwood City, CA
    3 days ago
  • $145k - $175k

     ...Requirements As a Senior Site Reliability Engineer at Commence, you will own...  ...implement, and own observability infrastructure including metrics, logging,...  ...~7+ years of experience in SRE, platform engineering, or...  ...systems). Exposure to AI/ML infrastructure and the... 
    Full time
    Remote work

    GrabJobs

    Santa Clara, CA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer (SRE) Intern — AI Infrastructure. Be the first to apply!