Tencent
Palo Alto, CA
Site Reliability Engineer (SRE) Intern — AI Infrastructure
We tailor your resume to this role and apply for you in seconds.
Or apply on Tencent's site yourselfJob details
- Location
- Palo Alto, CA
- Work type
- Onsite
- Compensation
- $56,410 - $108,014/yr
- Visa
- Sponsorship available
- Posted
- 2 weeks ago
- Apply on
- tencent.wd1.myworkdayjobs.com
About this role
Tencent is seeking a Site Reliability Engineer Intern to join its AI Compute team and support the daily operations of AI infrastructure. The role involves deploying, maintaining, troubleshooting, and documenting high-performance infrastructure systems while collaborating with engineering teams and partners.
What you'll do:
- Support the deployment, configuration, and maintenance of high-performance AI infrastructure servers, storage servers, networking equipment, and software components in secure environments
- Assist with hardware diagnostics, system functionality checks, and firmware updates as required
- Collaborate with engineering teams to help deliver tailored customer environments (e.g., bare-metal systems, Kubernetes, Slurm, etc.)
- Provide first-line engineering support for onsite operational issues, including troubleshooting hardware, network, and software problems, and firmware compliance
- Document incident details, resolutions, and lessons learned to improve future problem-solving
- Maintain clear, accurate, and up-to-date documentation to support knowledge sharing across the team
- Participate in team meetings and knowledge-sharing sessions to foster collaboration and continuous learning
What they're looking for:
- Currently pursuing or recently completed a Bachelor's or Master's degree in computer engineering, computer science, or a related technical field
- Basic understanding of server hardware, firmware lifecycle, and Linux environments, with an awareness of physical and system-level security standards
- Exposure to scripting languages such as Bash or Python
- Familiarity with — or strong interest in — configuration management, CI/CD tools, workload managers, and cluster software (e.g., Slurm, Kubernetes), and observability tools (e.g., Prometheus, Grafana, ELK)
- Strong problem-solving and analytical skills
- Ability to work both independently and as part of a team
- Professional fluency in English and Mandarin is highly preferred
- Coursework, projects, or hands-on experience related to AI Infrastructure, distributed systems, or cloud infrastructure
- Familiarity with networking fundamentals and Linux system administration
- Genuine interest in AI/ML infrastructure
Benefits:
- 1 hour of paid sick leave for every 30 hours worked
- Up to 13 paid holidays throughout the calendar year
- Full-time interns are eligible to enroll in the Company-sponsored medical plan, subject to the terms and conditions of the applicable plans then in effect
Ready to apply to Tencent?
ApplyBolt finds matching jobs, tailors your resume, and submits applications for you.