TikTok
San Jose, CA
Machine Learning Engineer Intern (E-Commerce Recommendation Video) - 2027 Start (PhD)
We tailor your resume to this role and apply for you in seconds.
Or apply on TikTok's site yourselfJob details
- Location
- San Jose, CA
- Work type
- Onsite
- Compensation
- $124,800/yr
- Visa
- Sponsorship available
- Posted
- today
- Apply on
- lifeattiktok.com
About this role
TikTok is a short-form mobile video platform with a Global E-Commerce Content Recommendation team responsible for recommendation systems across TikTok Shop. The Machine Learning Engineer Intern will help develop and deploy large-scale recommendation systems for video, livestream, and product scenarios, including generative retrieval, ranking, LLM-based modeling, hardware optimization, and agent-assisted research and development.
What you'll do:
- You will help build — and rewrite — an industrial recommendation system serving a billion-scale user base across short-video, livestream, and product scenarios, covering retrieval, pre-ranking, ranking, and blending end to end. Every iteration ships to production and directly moves user experience and GMV
- Scale recommendation models like LLMs. Push ranking models from hundreds of millions to billions of parameters and chart the scaling laws of recommendation: behavior-corpus pretraining; multi-scenario, multi-task, multi-stage joint training; ultra-long behavior-sequence modeling (10K+ events) with KV caching, sequence compression, user/generation (U-G) disaggregated serving, speculative decoding, and dynamic batching — raising MFU while holding a strict millisecond latency budget
- Build one-stage generative retrieval. Reframe retrieval as generation: tokenize the item space into semantic IDs (RQ-VAE / SID) and train autoregressive models, grounded in MLLM semantics, to generate what a user wants next — collapsing the traditional "multi-channel retrieval + ranking" funnel into a single generative stage. The open problems span the full stack: item tokenizers that balance semantic content against collaborative signal, and SIDs that stay stable while millions of new items arrive daily; post-training the generator directly on live user feedback (preference optimization, GRPO-style RL); and decoding under a millisecond budget — beam search, decoding constrained to the valid item space, and test-time scaling that trades inference compute for better recommendations. The prize is a system freed from its path dependence on ID memorization, where cold-start generalization comes from semantics rather than impression history
- Inject world knowledge. Use large models' real-world knowledge to mine latent user interests and semantic representations beyond what pure ID co-occurrence can express; use reasoning models to run explicit chain-of-thought inference over long-horizon user intent, making the system materially better at discovery and novelty
- Push training and inference to the hardware limit. Custom CUDA / Triton fused kernels, memory and computation-graph optimization, distributed training and inference acceleration, mixed precision and low-bit quantization — engineered for what makes recommendation hard: sparse embeddings, variable-length sequences, and many task heads
- Rewrite R&D with agents. We are embedding coding agents deep into the algorithm-development loop: automated feature mining and pipeline generation, experiment configuration and training orchestration, automated evaluation and online-diagnosis attribution, bad-case mining and patrol. You will be both a user and a builder of this system
- Do original work on open problems. Long-term value modeling, repurchase and retention, transaction attribution, fatigue modeling, new-user recommendation, incremental value modeling, interest exploration, LLM4Rec — problems where industry has no standard answers. We expect, and support, original research: internal papers, patents, and publication at top external venues
What they're looking for:
- Currently pursuing a PhD in Computer Science, Electrical Engineering, Mathematics, Statistics or a related discipline
- Solid ML and engineering fundamentals: you understand the math behind the models, and you write clean, efficient, reproducible code with a strong command of algorithms and data structures
- Deep research or engineering practice in at least one of: LLMs / foundation models, NLP, CV, RL, or recommendation / search / ads — and you can articulate why you made the choices you made, and where they fell short
- Genuine enthusiasm for LLM / LRM techniques: you want frontier methods live in production, not parked at offline metrics
- Strong problem definition and decomposition: faced with an ambiguous problem that has no standard answer, you find your own foothold
- Publications at KDD, SIGIR, RecSys, WWW, ACL, NeurIPS, ICML, ICLR, or comparable venues — or high-quality open-source work
- CUDA / Triton kernel development, source-level deep-learning-framework optimization, large-scale distributed training, or high-performance inference deployment
- Hands-on experience with LLM post-training (SFT / RLHF / DPO / GRPO), agent-system construction, or inference acceleration
- Led or deeply contributed to a key project in search, ads, recommendation, or large models, with a complete problem-to-online-impact loop
- Awards in ACM-ICPC, NOI, Kaggle, or comparable competitions
- Heavy user of AI coding and agentic workflows for building systems and optimizing models
Benefits:
- Interns have day one access to health insurance, life insurance, wellbeing benefits and more.
- Interns also receive 10 paid holidays per year and paid sick time (56 hours if hired in first half of year, 40 if hired in second half of year).
- Interns who are not working 100% remote may also be eligible for housing allowance.
Ready to apply to TikTok?
ApplyBolt finds matching jobs, tailors your resume, and submits applications for you.