TikTok Data Science OA, 4 Questions Done in 28 Minutes: Metrics, Table Merges, Feature Preprocessing, Random Forest Threshold Tuning
TikTok DS OA recap: computing average rating, second-language share and success rate with Pandas; merging multiple tables; train-set-based imputation, categorical encoding and standardization (avoiding data leakage); and tuning a random forest's threshold on a validation set to balance precision and recall.
I took the TikTok DS OA today. The four questions had some substance, but three were familiar, so I finished in about half an hour.
TikTok's OA isn't actually hard — most people just don't prepare enough and get a bit nervous. The questions are fairly standard; here are all four and my approaches.
Q1: Compute three metrics
Compute three metrics from the given data:
| Metric | How to compute |
|---|---|
| Average rating | Mean of the rating column, rounded to two decimals |
| Second-language share | Share of records where second_language isn't no, as a percentage with two decimals |
| Success rate | Combine all rides_*.csv files and compute the share of status equal to success, as a percentage with two decimals |
Return the results in the required format.
Q2: Combine multiple tables
Merge several tables into one. The main task is merging as described and making sure field values follow the required conventions. Do the merge and field extraction as specified, then save the result.
Q3: Driver data preprocessing
Preprocess the driver data — impute missing values, encode categorical variables, standardize numeric features and convert the label:
- Missing values: fill
agewith the training set's mean. - Categorical variables: for
sec_language,car_modeland similar, build mapping dictionaries from the training set; unknown values get a new code. - Numeric features: standardize columns like
net_xxxusing the training set's mean and standard deviation. - Label: convert
driver_classto numbers (A → 0, B → 1).
Avoid data leakage: every transformation must use statistics from the training set only, then be applied unchanged to the validation and test sets.
Q4: Binary classifier + threshold tuning
Train a binary classifier on the cleaned data, with class B as the positive class:
- Fit a random forest on the training set.
- Tune the decision threshold on the validation set: maximize recall while keeping precision high.
- Apply the best threshold from validation to binarize the test set's predicted probabilities.
- Save and return the predictions in the required format.
Good luck, and may all your tests pass!
Found this helpful? Let's talk.
Happy to swap interview notes, do mock interviews, or share referral info.

Scan to add me on WeChat
More TikTok notes
View all ›- TikTok Data Engineer Interview | TikTok New Grad, Full ProcessTikTok · 2026-10-04›
- TikTok 2027 CodeSignal OA: Even Digit Count, Cyclic-Shift Difference Sums, Bouncing Diagonals, Most Points in a RangeTikTok · 2026-10-03›
- TikTok OA, All 4 Passed: Min in a Range, Sorting by Vowel Gap, Bouncing Diagonals, Subarrays with at Least k Fruit PairsTikTok · 2026-10-03›
- Four Classic TikTok OA Questions: Adjacent Character Changes, Closest Earlier Timestamp, Placing Shapes, Fewest Operations to an Arithmetic SequenceTikTok · 2026-10-03›