TikTok
Seattle, WA
Machine Learning Engineer Intern - E-Commerce Recommendation Video
We tailor your resume to this role and apply for you in seconds.
Or apply on TikTok's site yourselfJob details
- Location
- Seattle, WA
- Work type
- Onsite
- Posted
- today
- Apply on
- lifeattiktok.com
About this role
Global E-Commerce (TikTok Shop) is one of TikTok's fastest-growing businesses and a core driver of the company's revenue growth. The Global E-Commerce Content Recommendation team owns the end-to-end recommendation stack for e-commerce video and image-text content on TikTok worldwide—including retrieval, ranking, and multi-queue blending; supply ecosystem and cold start; and the browsing-to-purchase experience for hundreds of millions of users.
The team believes recommendation is being rewritten in the compute era. ID-based collaborative filtering and supervised learning built today's systems and still run most of the industry, but their returns are diminishing. The team's ambition is to rebuild the stack on LLM foundations and build the most advanced recommendation system in the world.
PhD internships provide students the opportunity to contribute to products, research, future plans, and emerging technologies. The internship experience includes hands-on learning, community-building, professional development events, and collaboration with industry experts. Applications are reviewed on a rolling basis; applicants are encouraged to apply early and clearly state their availability, including start and end dates, in their resume.
## Responsibilities
- Help build and rewrite an industrial recommendation system serving a billion-scale user base across short-video, livestream, and product scenarios, covering retrieval, pre-ranking, ranking, and blending end to end. Iterations ship to production and directly affect user experience and GMV.
- Scale recommendation models like LLMs: grow ranking models from hundreds of millions to billions of parameters; study recommendation scaling laws; and work on behavior-corpus pretraining, multi-scenario and multi-task joint training, ultra-long behavior-sequence modeling (10K+ events), KV caching, sequence compression, user/generation-disaggregated serving, speculative decoding, and dynamic batching while improving MFU within strict millisecond latency budgets.
- Build one-stage generative retrieval by tokenizing the item space into semantic IDs (RQ-VAE / SID) and training autoregressive models grounded in MLLM semantics. Work on item tokenizers balancing semantic content and collaborative signal, stable SIDs as new items arrive, post-training generators on live user feedback, and millisecond-budget decoding techniques such as beam search, valid-item-space constraints, and test-time scaling.
- Use large models' real-world knowledge to mine latent user interests and semantic representations beyond ID co-occurrence, and use reasoning models for explicit chain-of-thought inference over long-horizon user intent to improve discovery and novelty.
- Push training and inference to hardware limits through custom CUDA / Triton fused kernels, memory and computation-graph optimization, distributed training and inference acceleration, mixed precision, and low-bit quantization, accounting for sparse embeddings, variable-length sequences, and many task heads.
- Use and build coding-agent systems in the algorithm-development loop, including automated feature mining and pipeline generation, experiment configuration and training orchestration, automated evaluation and online-diagnosis attribution, and bad-case mining and patrol.
- Conduct original work on open problems such as long-term value modeling, repurchase and retention, transaction attribution, fatigue modeling, new-user recommendation, incremental value modeling, interest exploration, and LLM4Rec. The team supports original research, including internal papers, patents, and publication at top external venues.
## Minimum Qualifications
- Currently pursuing a PhD in Computer Science, Electrical Engineering, Mathematics, Statistics, or a related discipline.
- Solid machine-learning and engineering fundamentals, including understanding the mathematics behind models and writing clean, efficient, reproducible code; strong command of algorithms and data structures.
- Deep research or engineering practice in at least one of LLMs / foundation models, NLP, CV, RL, or recommendation / search / ads, and ability to explain design choices and their limitations.
- Genuine enthusiasm for LLM / LRM techniques and bringing frontier methods into production.
- Strong problem definition and decomposition skills, including finding an approach to ambiguous problems without standard answers.
## Preferred Qualifications
- Publications at KDD, SIGIR, RecSys, WWW, ACL, NeurIPS, ICML, ICLR, or comparable venues, or high-quality open-source work.
- CUDA / Triton kernel development, source-level deep-learning-framework optimization, large-scale distributed training, or high-performance inference deployment.
- Hands-on experience with LLM post-training (SFT / RLHF / DPO / GRPO), agent-system construction, or inference acceleration.
- Led or deeply contributed to a key project in search, ads, recommendation, or large models, with a complete problem-to-online-impact loop.
- Awards in ACM-ICPC, NOI, Kaggle, or comparable competitions.
- Extensive use of AI coding and agentic workflows for building systems and optimizing models.
## Compensation and Benefits
The hourly rate range for this position in the selected city is $57–$57. Interns have day-one access to health insurance, life insurance, wellbeing benefits, and more. Interns also receive 10 paid holidays per year and paid sick time (56 hours if hired in the first half of the year, 40 hours if hired in the second half). Interns who are not working 100% remotely may also be eligible for a housing allowance. Benefits may vary depending on the nature of employment and the country work location.
Ready to apply to TikTok?
ApplyBolt finds matching jobs, tailors your resume, and submits applications for you.