Site Reliability Engineer (SRE) Intern — AI Infrastructure
$27.12 - $51.93 per hourTencent
About the Hiring Team What the Role Entails Role Summary
We are seeking a motivated Site Reliability Engineer (SRE) Intern to join our AI Compute team, supporting the daily operations of AI infrastructure. In this role, you will work closely with internal business teams and external engineering partners to build and operate AI infrastructure. This is a hands-on opportunity to gain knowledge and experience in cutting-edge AI infrastructure. Key Responsibilities
Support the deployment, configuration, and maintenance of high-performance AI infrastructure servers, storage servers, networking equipment, and software components in secure environments.
Assist with hardware diagnostics, system functionality checks, and firmware updates as required.
Collaborate with engineering teams to help deliver tailored customer environments (e.g., bare-metal systems, Kubernetes, Slurm, etc.).
Provide first-line engineering support for onsite operational issues, including troubleshooting hardware, network, and software problems, and firmware compliance.
Document incident details, resolutions, and lessons learned to improve future problem-solving.
Maintain clear, accurate, and up-to-date documentation to support knowledge sharing across the team.
Participate in team meetings and knowledge-sharing sessions to foster collaboration and continuous learning.
Who We Look For Qualifications & Requirements Currently pursuing or recently completed a Bachelor's or Master's degree in computer engineering, computer science, or a related technical field.
Basic understanding of server hardware, firmware lifecycle, and Linux environments, with an awareness of physical and system-level security standards.
Exposure to scripting languages such as Bash or Python.
Familiarity with — or strong interest in — configuration management, CI/CD tools, workload managers, and cluster software (e.g., Slurm, Kubernetes), and observability tools (e.g., Prometheus, Grafana, ELK).
Strong problem-solving and analytical skills.
Ability to work both independently and as part of a team.
Professional fluency in English and Mandarin is highly preferred
Preferred Qualifications Coursework, projects, or hands-on experience related to AI Infrastructure, distributed systems, or cloud infrastructure.
Familiarity with networking fundamentals and Linux system administration.
Genuine interest in AI/ML infrastructure.
Location State(s)
US-California-Palo AltoThe expected base pay range for this position in the location(s) listed above is $27.12 to $51.93 per hour. Actual pay may vary depending on job-related knowledge, skills, and experience.
This position will be eligible for 1 hour of paid sick leave for every 30 hours worked and up to 13 paid holidays throughout the calendar year. Subject to the terms and conditions of the applicable plans then in effect, full-time interns are also eligible to enroll in the Company-sponsored medical plan. Equal Employment Opportunity at Tencent As an equal opportunity employer, we firmly believe that diverse voices fuel our innovation and allow us to better serve our users and the community. We foster an environment where every employee of Tencent feels supported and inspired to achieve individual and common goals.
Vacancy posted 1 hour ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer (SRE) Intern — AI Infrastructure in Palo Alto, CA vacancy
$20k
...goal of enabling human life on Mars.SITE RELIABILITY ENGINEER (STARSHIELD) At SpaceX we’re leveraging... ...on Starshield's software and GPU infrastructure, you will design, operate and scale the... ...validate, and productize solutions for AI clusters (100k+ GPU scale)Develop...SuggestedPermanent employmentTemporary workImmediate startWeekend work$195k - $265k
...we scale toward launch, our engineering infrastructure has to scale with us. We're hiring a Staff Site Reliability Engineer to own source control... ...tooling that reduce toil for SRE and product engineering alike... ...use artificial intelligence (AI) tools to support parts of the...SuggestedFull time$61k - $101k
...or certification in site reliability engineering, along with 5+ years... ...advanced knowledge of SRE culture and principles... ...development, automation, and infrastructure as code. We... ...-authorized AI capabilities in the workplace... ...reliability community through internal forums, communities...SuggestedFull time- ...Description Developer & Infrastructure Expert Role Type:... ...to evaluate AI-powered workflows across... ...infrastructure, DevOps, SRE, and platform engineering. You will test AI-... ...for accuracy and reliability. Work with AWS,... ...Cloud Infrastructure Site Reliability...SuggestedRemote jobFor contractors
$101k - $161k
...prestigious awards, such as Best Engineering Team, Best Company for... ...Work WithWe’re looking for Site Reliability Engineers to join our... ...as-a-Service (CVaaS) global SRE team. SREs at Arista combine... ...microservices stack, monitoring infrastructure, and much more. What You'll...Suggested- ...Future Workforce Wand turns AI into labor. It enables humans... ...the world’s first Agentic Labor Infrastructure enabling governments and... ...hiring for a hands‑on Head of SRE to establish, lead, and scale our Site Reliability Engineering function. This role combines strategic...Shift work
$222k - $300.5k
...OverviewAbout the TeamIntuit's Infrastructure and Site Reliability organization owns... ...Platform Systems Engineering team builds and... ...security, and other SRE/infrastructure leaders... ...for this role is AI Ops: embedding AI-driven... ...customer-facing and internal systems.Lead, grow,...WorldwideShift work$262k - $364k
...ecosystem have reliability and uptime appropriate... ....Build creative engineering solutions to operations and infrastructure problems, including AI-powered... ...contribute to the cross-SRE AI Ops program,... ...a strategic way.Site Reliability... ...services—both our internally critical and our...- ...Site Reliability Engineer There are NO limits to your career: come shape the... ...Site Reliability Engineering (SRE) is a discipline that... ...engineering and applies them to infrastructure and operations problems.... ...in Python supported by Gen AI tooling to accelerate development...Immediate startRemote workWorldwide
- ...Site Reliability Engineer III There's nothing more exciting than being at the... ...& Analytics Office (CDAO) AI/ML & Data Platforms team, you... .... Through code and cloud infrastructure, you will configure, maintain... ...environment to support SRE workflows with strong validation...Work at office
$152k - $228k
...profound opportunity for AI to drive positive... ...You will own the infrastructure that makes this... ...much much more.Engineers across the company... ...quality gates.Platform Reliability & Observability:... ...: Guide the SRE team through the OS... ...Introduction), SRE (Site Reliability Engineering...Temporary workImmediate startFlexible hours- ...DESCRIPTION Elevate your engineering prowess to... ...top echelon in site reliability. As a Senior Lead... ...Chase within the Infrastructure Platforms and Foundational... ...-authorized AI capabilities... ...reliability community via internal forums,... ...particularly in SRE, observability, or...
$226k - $369k
...approval. LinkedIn’s Reliability Infrastructure team is responsible for... ...Principal Staff Software Engineer, Reliability... ...handling are in place.As AI-assisted software development... ...is not a traditional SRE role focused on... ...measurable improvements in site stabilityHelp shape how...For contractorsWork at officeRemote workWork from homeFlexible hours$175k - $287k
...team. LinkedIn’s Core AI is building the Evaluation... ...for tracing infrastructure for all LinkedIn AI Agents.As a Staff Engineer, you will own the end-to... ...continuously improve the quality, reliability, safety, and... ...environment with other SRE/SWE Engineers, Project...For contractorsWork experience placementWork at officeFlexible hours- ...Lead Site Reliability Engineer Assume a critical role in defining the future... ...applications Implements infrastructure, configuration, and network... ...Uses enterprise-authorized AI capabilities within the work... ...site reliability engineering (SRE) culture and principles,...
- ...defense tech startups with the tools, infrastructure, and compliance solutions they need... ...are seeking a highly motivated Systems Reliability Engineer (SRE) to lead the design and... ...sensitive and cleared workforces. The Site Reliability Engineer (SRE) - SecOps...For contractorsWork at officeFlexible hours
$100k - $200k
OPPO US Research Center is seeking a skilled and proactive Site Reliability Engineer (SRE) to join our team. In this role, you will be responsible... ...systems. The ideal candidate is passionate about cloud infrastructure, automation, and building reliable, production-grade...Full time$190.98k - $214.05k
...otherwise, all roles are on-site five days per week... ...thoughtful use of AI tools in our daily... ...a Senior Software Engineer (Platform & Data Infrastructure) to architect, scale... ...Platform, Security, SRE, and Product Engineering... .... Drive Reliability, DR & On-Call Excellence...Full timeTemporary workWork at officeRemote workFlexible hours$125k - $195k
...Responsibilities Manage GPU and CPU infrastructure deployments to Top Secret... ..., validate, and productize AI cluster solutions at 100,000... .... Collaborate with AI engineers to build scalable, operable,... ...one year of professional site reliability engineering or DevOps...Permanent employmentFull timeTemporary workWeekend work- ...Social, we're building the AI-native social operating... ...Founded by ex-Meta product and engineering leaders, we've raised over... ...Senior Software Engineer, Infrastructure to own the reliability, scalability, and... ...base, and we need a seasoned SRE to help us scale these systems...Work at officeRemote workFlexible hoursShift work
$207k - $300k
...team of Software/Systems Engineers on projects for users and... ...with Large Language Model.Site Reliability Engineering (SRE) combines software and... ...Google's services—both our internally critical and our externally... ...systems, building infrastructure and eliminating work through...$160.36k - $240.54k
...profound opportunity for AI to drive positive... ...driving, and the ML Infrastructure team builds and operates... ...care as much about reliability and operational maturity... ...Science, Electrical Engineering, or a closely related... ...distributed training internals, including NCCL and...Work experience placementImmediate startFlexible hours$160k - $260k
...RoleWe're looking for a talented, driven Senior Software Engineer to build the critical infrastructure that powers the Aptos blockchain ecosystem.This role... ...payments, DeFi, or TradFi (a plus)Comfortable leveraging AI tools to enhance productivity and code qualityThe base...Full timeWork experience placementWork at officeLocal area$224k - $356.5k
NVIDIA is hiring engineers to scale up the introduction... ...into its EDA Infrastructure. We expect you to have... ...effective, clear and reliable architecture specificationTranslate... ..., DevOps and/or SRE practices and/or Platform... ...crowd:Developing ML/AI infrastructure. Developing...Full time$120k - $195k
...part of our world-class software engineering team, you will help build the next-generation infrastructure and platforms that power LinkedIn’s products, business, and AI-first future. This includes... ...foundational platforms that enable reliable, scalable, and AI-first product...For contractorsWork experience placementWork at officeFlexible hours$230k - $250k
...autonomous networking, giving engineers and AI agents the ability to know... ....Forward is looking for a Site Reliability EngineerAbout the Role This... ...not a "keep the lights on" SRE role. As our first or early... ...closely with engineering, infrastructure, and product to ensure our...Night shift$148k - $235.75k
...unlimited potential of AI to define the next era... ...team of innovative engineers who are building an AI... ...volume telemetry into reliable, job-centric insights... ...automation.Manage deployment infrastructure and packaging (Helm +... ...systems as SRE/DevOps/Platform Ops.Proven...Full time$152k - $241.5k
...looking for a Senior SRE to join our... ...harness the power of AI to deliver groundbreaking... ...fabrics.Use IaC(Infrastructure‑as‑Code) and... ...Service (QoS) for internal customers through... ...management, fleet reliability/auto-healing, E2E... ...Ruby.Mentored other engineers and influenced technical...Full time$180k - $230k
...most critical constraint in AI’s growth trajectory:... ...defining bottleneck in the AI infrastructure race. While leading tech companies... ...'re looking for a Senior SRE to own the reliability, scalability, and... ...closely with platform and data engineering to keep high-throughput,...Work at officeLocal areaImmediate startRemote work3 days per week$145k - $175k
...Requirements As a Senior Site Reliability Engineer at Commence, you will own... ...implement, and own observability infrastructure including metrics, logging,... ...~7+ years of experience in SRE, platform engineering, or... ...systems). Exposure to AI/ML infrastructure and the...Full timeRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer (SRE) Intern — AI Infrastructure. Be the first to apply!
Related searches
- site reliability engineer sre Palo Alto, CA
- site reliability engineer Palo Alto, CA
- lead infrastructure engineer Palo Alto, CA
- infrastructure engineer Palo Alto, CA
- infrastructure developer Palo Alto, CA
- remote infrastructure engineer Palo Alto, CA
- senior infrastructure engineer Palo Alto, CA
- IT site lead Palo Alto, CA
- site safety Palo Alto, CA
- website content developer Palo Alto, CA


