ByteDance
San Jose, CA
Research Intern (Inference Infrastructure) - 2027 Start (PhD)
We tailor your resume to this role and apply for you in seconds.
Apply to Research Intern (Inference Infrastructure) - 2027 Start (PhD) at ByteDanceJob details
- Location
- San Jose, CA
- Work type
- Onsite
- Visa
- Sponsorship available
- Posted
- 2 weeks ago
- Apply on
- joinbytedance.com
About this role
ByteDance is a technology company building cloud and AI computing infrastructure through its DPU team. The Research Intern will design and build scalable cluster management, cloud-native GPU, AI accelerator, and inference infrastructure while collaborating on LLM solutions and developing production-ready systems code.
What you'll do:
- Design and build large-scale, container-based cluster management and orchestration systems with extreme performance, scalability, and resilience
- Architect next-generation cloud-native GPU and AI accelerator infrastructure to deliver cost-efficient and secure ML platforms
- Collaborate across teams to deliver world-class inference solutions using vLLM, SGLang, TensorRT-LLM, and other LLM engines
- Stay current with the latest advances in open source (Kubernetes, Ray, etc.), AI/ML and LLM infrastructure, and systems research; integrate best practices into production systems
- Write high-quality, production-ready code that is maintainable, testable, and scalable
What they're looking for:
- Currently pursuing a PhD in Computer Science, Computer Engineering, Electrical Engineering, or a related technical field
- Able to commit to working for 12 weeks during Summer 2027
- Strong understanding of large model inference, distributed and parallel systems, and/or high-performance networking systems
- Hands-on experience building cloud or ML infrastructure in areas such as resource management, scheduling, request routing, monitoring, or orchestration
- Solid knowledge of container and orchestration technologies (Docker, Kubernetes)
- Proficiency in at least one major programming language (Go, Rust, Python, or C++)
- Experience contributing to or operating large-scale cluster management systems (e.g., Kubernetes, Ray)
- Experience with workload scheduling, GPU orchestration, scaling, and isolation in production environments
- Hands-on experience with GPU programming (CUDA) or inference engines (vLLM, SGLang, TensorRT-LLM)
- Familiarity with public cloud providers (AWS, Azure, GCP) and their ML platforms (SageMaker, Azure ML, Vertex AI)
- Strong knowledge of ML systems (Ray, DeepSpeed, PyTorch) and distributed training/inference platforms
- Excellent communication skills and ability to collaborate across global, cross-functional teams
Benefits:
- Interns have day one access to health insurance, life insurance, wellbeing benefits and more.
- Interns also receive 10 paid holidays per year.
- Interns also receive paid sick time (56 hours if hired in first half of year, 40 if hired in second half of year).
- Interns who are not working 100% remote may also be eligible for housing allowance.
Ready to apply to ByteDance?
ApplyBolt finds matching jobs, tailors your resume, and submits applications for you.