TikTok
Seattle, WA

Machine Learning Engineer Intern - Conversational AI

OnsitePosted todayLikely sponsors

We tailor your resume to this role and apply for you in seconds.

Or apply on TikTok's site yourself

Job details

Location
Seattle, WA
Work type
Onsite
Posted
today
Apply on
lifeattiktok.com

About this role

We build the next-generation unified Agent system for TikTok's global e-commerce customer service, running in 30+ languages across one of the largest e-commerce surfaces on the internet. Our north star is a self-evolving Agent: post-training, harness, memory and context engineering, tools, and evaluation form one closed loop, and every served conversation becomes the next iteration's training, evaluation, retrieval, and skill-induction signal. This loop is already running in production: cases are mined, root-caused, turned into constrained candidates, replayed against frozen regression sets, and shipped behind guardrails. We build the agent runtime itself—not prompts on top of a vendor API—and treat evaluation and experimentation as first-class systems. By combining generative recommendation, large recommendation models, multimodal representation learning, and cross-domain value modeling, the team works on important algorithmic problems in live commerce. Our goal is to improve user experience, optimize ecosystem efficiency, and drive sustainable business growth for TikTok Shop across global markets. We are looking for talented individuals to join us for an internship. PhD internships provide students with the opportunity to actively contribute to our products and research, as well as to the organization's future plans and emerging technologies. The internship experience blends hands-on learning, community-building and professional development events, and collaboration with industry experts. Applications are reviewed on a rolling basis; applicants are encouraged to apply early and clearly state their availability (start date and end date) in their resume. ## Responsibilities - Build the agent runtime (harness and agent loop): orchestrate skills, tools, and context; implement loop control and intervention, progressive disclosure, and behavior-level guardrails. Build the production safety layer, including pre-flight budgets and timeout truncation, serve-time gates, shadow or swap-in answer delivery, and safe fallback paths. - Develop context and memory for long multi-turn agents, including agentic memory, context compaction and summarization, context editing and observation masking, and just-in-time retrieval. Treat context as an evolving, itemized playbook with structured diffs and a deterministic curator. - Develop post-training and the data flywheel: use SFT, DPO, and RL to internalize rules into model weights, distill to smaller serving models, and turn served conversations into training, evaluation, and retrieval signals. - Build tools, skills, and MCP systems, including tools-as-APIs, connectors, skill and tool search for large inventories, and skill-library governance. - Build evaluation systems suitable for launch decisions, including LLM-as-judge with human-agreement calibration; paired comparisons, confidence intervals, repeated sampling, and pass^k; held-out and time-rolling evaluation splits with overfitting alarms; and cascaded scoring and cross-family judge panels. - Build the self-evolving loop: case mining, automatic root-cause analysis, constrained candidate generation, replay verification against frozen regression sets, canary, and flywheel. Make it auditable with a candidate registry, exact runtime read-back, change lineage, and an archive of rejected candidates. - Support online experimentation and causal readout through shadow, canary, and A/B tests; non-inferiority gates; traffic-split health; robust metric definitions; and off-policy counterfactual evaluation where live A/B testing is not possible. ## Minimum Qualifications - Currently pursuing a PhD in Computer Science, Engineering, Operations Research, or a related technical discipline. - Strong Python skills plus one of C++, Go, Rust, or Java. - Solid machine learning, deep learning, and NLP fundamentals, with hands-on experience with LLMs or agents through coursework, research, an internship, competition, open source, or a serious side project. - Basic statistical literacy, including the ability to compute a confidence interval, explain what a p-value does and does not mean, and distinguish an increase in a number from an improvement in a system. - Ability to read a paper or engineering blog and turn it into working code. ## Preferred Qualifications - Hands-on depth in at least one area: post-training (SFT, DPO, RLHF, RLAIF, RLVR, reward modeling, or reward hacking defenses); agent systems (harness, context engineering, MCP or Skills, sub-agents, or tool search); evaluation and experimentation (LLM-as-judge and judge calibration, pass^k, regression suites, A/B and non-inferiority testing, or off-policy evaluation); or self-improving and evolutionary systems. - For PhD candidates, publications; strong competition results such as ACM-ICPC, Kaggle, or ML competitions; or notable open-source contributions. - E-commerce or multilingual experience is a plus, but not required. ## Compensation and Benefits The base salary range for this position in the selected city is $57–$57 annually. Interns have day-one access to health insurance, life insurance, wellbeing benefits, 10 paid holidays per year, and paid sick time (56 hours if hired in the first half of the year or 40 hours if hired in the second half). Benefits may vary depending on employment and work location.
Ready to apply to TikTok?
ApplyBolt finds matching jobs, tailors your resume, and submits applications for you.

About TikTok

TikTok
Seattle, WA