TikTok Data Engineer Interview | TikTok New Grad, Full Process
TikTok Data Engineer new grad interview, start to finish: a resume deep dive (Flink checkpoints, state growth, back pressure), two SQL questions (daily active user growth, top creators by engagement and optimizing big-table joins) and a Python log-pattern question, with SQL and Python reference implementations.
Overview
Usually 60 minutes in three parts: resume discussion + two SQL questions + one Python question.
The interviewer may be in San Jose or in China:
| Interview time | Interviewer |
|---|---|
| After 6 pm PT | Usually based in China, 100% speaks Chinese |
| Before 6 pm PT | Usually based in the US, occasionally speaks Chinese |
Part 1: Resume deep dive
Mostly the candidate talking; no specific test points. As long as your projects relate to data engineering and you know them well, you'll be fine.
A real example: the interviewer asked the candidate to pick the project that best shows their data engineering skills. The candidate described a real-time user behavior analytics system built on Kinesis + Flink, walking through the pipeline from ingestion and real-time computation to loading downstream.
The interviewer's follow-ups:
- What do you do about Flink checkpoint delays or state growth? Answer: tune the RocksDB state backend, control checkpoint frequency, and set a TTL on state to keep it from piling up.
- How do you troubleshoot back pressure in a streaming job?
Back pressure troubleshooting: first check which operator is flagged in the Flink Web UI (back pressure propagates upstream, so find the first operator that can't keep up); then see whether it's CPU-bound, suffering from data skew (some subtasks much slower), or blocked by a slow sink. Respond by increasing parallelism, spreading hot keys, or batching / making sink writes asynchronous.
Part 2: SQL (HackerRank)
Two medium-difficulty SQL questions, with emphasis on clear analysis and optimization ideas.
Q1: Daily active users and day-over-day growth
From a user events table, compute daily active users (DAU) and the growth rate versus the previous day.
WITH daily AS (
SELECT DATE(event_time) AS dt,
COUNT(DISTINCT user_id) AS dau
FROM events
GROUP BY DATE(event_time)
)
SELECT dt,
dau,
ROUND(100.0 * (dau - LAG(dau) OVER (ORDER BY dt))
/ LAG(dau) OVER (ORDER BY dt), 2) AS growth_pct
FROM daily
ORDER BY dt;
The LAG() window function fetches the previous day's DAU; the first day has no previous day, so its growth rate is NULL.
Follow-up: how would you optimize at scale? Partition the table by date and always filter on a date range to avoid full table scans.
Q2: Creators with the highest engagement
Join several tables to find the creators with the highest engagement score over a recent period. Below, score = likes × 1 + comments × 2 + shares × 3 — confirm the scoring rules and time window with your interviewer first (date functions use SQLite syntax; MySQL / Hive differ slightly, e.g. DATE_SUB(CURRENT_DATE, INTERVAL 7 DAY)):
WITH recent AS (
SELECT v.creator_id,
SUM(CASE e.event_type WHEN 'like' THEN 1
WHEN 'comment' THEN 2
WHEN 'share' THEN 3 ELSE 0 END) AS score
FROM engagements e
JOIN videos v ON v.video_id = e.video_id
WHERE e.event_time >= DATE('2026-10-01', '-7 days')
GROUP BY v.creator_id
)
SELECT c.creator_name,
r.score,
RANK() OVER (ORDER BY r.score DESC) AS rnk
FROM recent r
JOIN creators c ON c.creator_id = r.creator_id
ORDER BY rnk, c.creator_name;
Aggregate the big table first, then join the small one (creators) — that saves a lot of computation compared with joining first.
Follow-up: the events table reaches hundreds of millions of rows — how do you optimize the big join?
- Broadcast join the small table (e.g. creators) to avoid shuffling the big one.
- Maintain a daily summary table ahead of time so queries only aggregate summaries.
- In a distributed setting, use window functions and CTEs to simplify the query logic.
Part 3: Python coding (HackerRank)
Easier than the SDE interview. Find users in a set of logs who match a behavior pattern, such as repeating the same action a certain number of times in a row.
Solution: group by user, sort by time, then count consecutive actions in a single pass — O(n log n).
from collections import defaultdict
def users_with_streak(logs, k):
"""logs: [(user_id, timestamp, action)]; returns users who repeat the same action k (or more) times in a row"""
by_user = defaultdict(list)
for user, ts, action in logs:
by_user[user].append((ts, action))
result = set()
for user, events in by_user.items():
events.sort() # sort by time
streak, prev = 0, None
for _, action in events:
streak = streak + 1 if action == prev else 1
prev = action
if streak >= k:
result.add(user)
break
return result
Follow-ups (they felt improvised by the interviewer):
- The logs arrive out of order as a distributed stream — how do you guarantee processing order? Use time buckets + watermarks to bound how long you wait for late logs, then reorder within the window before processing.
- The data is too large to load into memory at once — now what? Use stream processing, or external sorting (shard by user, sort in batches) to compute in batches.
Found this helpful? Let's talk.
Happy to swap interview notes, do mock interviews, or share referral info.

Scan to add me on WeChat
More TikTok notes
View all ›- TikTok 2027 CodeSignal OA: Even Digit Count, Cyclic-Shift Difference Sums, Bouncing Diagonals, Most Points in a RangeTikTok · 2026-10-03›
- TikTok Data Science OA, 4 Questions Done in 28 Minutes: Metrics, Table Merges, Feature Preprocessing, Random Forest Threshold TuningTikTok · 2026-10-03›
- TikTok OA, All 4 Passed: Min in a Range, Sorting by Vowel Gap, Bouncing Diagonals, Subarrays with at Least k Fruit PairsTikTok · 2026-10-03›
- Four Classic TikTok OA Questions: Adjacent Character Changes, Closest Earlier Timestamp, Placing Shapes, Fewest Operations to an Arithmetic SequenceTikTok · 2026-10-03›