Vietnamese Sentiment Analysis on TikTok Comments
Crawling TikTok comments, labelling them with a locally hosted LLM, then benchmarking PhoBERT and ViSoBERT against TF-IDF baselines over five emotions.
- PhoBERT
- PyTorch
- Scikit-learn
- Streamlit
- Python
An IE403 coursework project taken end to end: collect Vietnamese TikTok comments, label them, then benchmark Vietnamese language models against TF-IDF baselines on a five-class emotion problem.
No dataset, so I crawled one#
I needed real comments written in TikTok Vietnamese: teen slang, missing diacritics,
emoji wedged between words. The data comes from the api/comment/list/ endpoint — pass
the aweme_id parsed out of the video URL, page through with cursor twenty comments at
a time, sleep 0.4s between requests and retry up to three times on failure.
What I cared about from the start was being able to resume. Each run reads the
post_url column of the existing CSV to see which videos are already done, then appends
instead of overwriting. Nine news and lifestyle videos yielded roughly 42,900 raw comments.
Labelling with a local LLM#
Hand-labelling that many rows was never going to happen. I ran a local LLM through LM Studio and turned it into an annotator: the prompt describes the five classes (happy, angry, sad, fearful, neutral) with teen-slang examples for each, and the output is forced down to a single digit.
# Unambiguous slang goes straight to a dictionary, saving a model call.
teen_code_dict = {"hehe": 0, "hihi": 0, "dm": 1, "dkm": 1, "sml": 2, "hú hồn": 3}
if text.lower().strip() in teen_code_dict:
return teen_code_dict[text.lower().strip()]
payload = {
"messages": [{"role": "user", "content": prompt}],
"temperature": 0.3,
"max_tokens": 1, # capped at one token: no room for the model to explain itself
"stop": ["\n"],
}
response = requests.post("http://localhost:1234/v1/chat/completions", json=payload)max_tokens: 1 is the detail that earned its keep — post-processing collapses to a single
isdigit() check. The loop checkpoints every five rows and remembers which line_number
values are done, so I could stop and resume freely. It labelled 28,811 rows in total.
Vietnamese preprocessing#
def clean_and_tokenize_text(text):
text = re.sub(r'@\S+', '', text) # drop mentions
text = re.sub(emoji_pattern, ' ', text) # drop emoji
text = text.lower()
text = unicodedata.normalize('NFKC', text)
text = text.translate(str.maketrans('', '', string.punctuation))
text = text_normalize(text) # underthesea: diacritic normalisation
return word_tokenize(text, format="text") # "em trai" -> "em_trai"On top of that I collapse repeated characters (hayyy → hay) and drop rows where 80% or
more of the letters fall outside the Latin range. Cleaning left 21,025 rows, skewed hard
towards neutral (6,780) while fear had only 2,756. Downsampling every larger class to the
smallest one gave a balanced set of 13,780 rows, 2,756 per label.
Which model won#
The classical group uses TF-IDF with 5,000 features; the transformers are fine-tuned with
max_length 128, AdamW at lr=5e-5, balanced class weights in the loss and early stopping
on validation accuracy (patience 3). Both groups are evaluated on the same balanced set
with a 90/10 split.
| Model | Accuracy | Macro F1 |
|---|---|---|
| PhoBERT-base | 0.6640 | 0.6648 |
| ViSoBERT | 0.6626 | 0.6628 |
| RoBERTa-base-Vietnamese | 0.6524 | 0.6517 |
| Stacking (LR + SVM + RF) | 0.6466 | 0.6473 |
| viBERT-base-cased | 0.6335 | 0.6298 |
| Logistic Regression | 0.6321 | 0.6332 |
| SVM (linear) | 0.6212 | 0.6229 |
| Random Forest | 0.6154 | 0.6159 |
PhoBERT beats the TF-IDF stack by only 1.7 accuracy points — closer than I expected, and the reason I kept the classical models in the demo. I also tried multilingual BERT with a BiLSTM head, but the BERT encoder stayed frozen and it reached 0.5434: the price of not fine-tuning. Per class, every model reads fear best (F1 0.75–0.82) and happy worst (0.53–0.59) — sarcasm and joking remain the hard part.
Outcome#
- A closed loop: crawl → auto-label → clean → balance → train → demo
- PhoBERT-base came out on top at 0.6640 accuracy and 0.6648 macro F1 across five labels
- A Streamlit app takes a TikTok link, crawls the comments live and runs all eight models so the predictions sit side by side