A Content-Based Anime Recommender
An anime recommender built on self-crawled data, with genre vectors in Spark ML, watch events replayed through Kafka, and a Streamlit demo on top.
- Python
- Streamlit
- Spark
- Kafka
A Big Data coursework project, built by two of us: an anime recommender starting from nothing — crawl the data, clean it, vectorise it with Spark ML, demo it in Streamlit.
Crawling instead of downloading a dataset#
Kaggle has plenty of anime datasets, but the titles are all English or romaji. I wanted
the system to return the names Vietnamese viewers actually use, so I scraped
animevietsub directly. Selenium drives a headless Brave build through two listing types
(list-le for films, list-bo for series) crossed with four seasons, a few dozen pages
each. From every card I pull title, link, poster, episode count, rating, views, quality,
year, genres and synopsis — eight JSON files, merged into 5,389 records.
Most of the work was cleaning#
Nothing came off the site usable as-is. The rating column is a full Vietnamese sentence
along the lines of "9.1 out of 10 based on 76 member ratings", so a regex splits it into a
float rate and an integer nums_of_vote. The quality column carries every variant of
FHD, Full HD, BD, BD Vietsub and 480P — plus one cell that had a date in it — so a mapping
dictionary collapses them into six groups. year contained both "07 - 2021" and
"2010-2011". For episode, every null turned out to be a standalone film, which makes
filling with 1 the honest choice. The result is 5,389 rows across 14 columns, loaded into
MongoDB for the app to read back.
Genre vectors and a hand-rolled cosine#
I built the feature from genres rather than description, because every synopsis on the
site is truncated with an ellipsis — TF-IDF over those would learn opening sentences, not
content. Genres are short, clean and discriminative enough. So Spark ML's
CountVectorizer turns the genre list into a count vector and I compute similarity myself
over the RDD:
data = data.withColumn("genres", split(col("genres"), ","))
# Saved to disk so the consumer and the Streamlit app reload one identical vocabulary
vectorizer = CountVectorizer(inputCol="genres", outputCol="genre_vector")
count_vectorizer_model = vectorizer.fit(data)
def calculate_similarity(row):
target_vector = broadcast_genre_vector.value
row_vector = DenseVector(row.genre_vector.toArray())
dot_product = sum(target_vector[i] * row_vector[i] for i in range(len(target_vector)))
norm_target = sum(x ** 2 for x in target_vector) ** 0.5
norm_row = sum(x ** 2 for x in row_vector) ** 0.5
return row.title, row.item_id, dot_product / (norm_target * norm_row)
# Scan the whole catalogue per query, keep the top N
similar_items = (preprocessed_data.rdd.map(calculate_similarity)
.filter(lambda x: x[1] != item_id)
.takeOrdered(top_n, key=lambda x: -x[2]))Kafka here is a simulation, not a live feed#
The crawl is a static snapshot, so any "real time" had to be manufactured. I built a watch
history file and had a producer push one view at a time into the anime topic, five
seconds apart; the consumer picks each event up, reloads the saved CountVectorizerModel
and prints recommendations right after the watch.
producer = KafkaProducer(
bootstrap_servers='localhost:9092',
value_serializer=lambda v: json.dumps(v).encode('utf-8')
)
for _, row in df.iterrows():
producer.send('anime', value={
'user_id': row['user_id'],
'item_id': row['link'],
'titles_watched': row['titles_watched'],
'genres': row['genres'].split(','),
'rate': row['rate'],
})
time.sleep(5) # pace it so it behaves like a real user streamThe Streamlit demo#
The app has two tabs. The recommendation tab takes a partial title, matches it with
LIKE %...% and shows the top 10 with poster, link, genre list and similarity score — so
a viewer can judge whether a suggestion makes sense instead of trusting a number. The
statistics tab reads straight from Spark: total records, distinct titles, average rating,
and a genre frequency chart.
Outcome#
- A pipeline that runs end to end — Selenium to JSON to CSV to MongoDB to Spark to Streamlit — with every stage a notebook that can be re-run
- Precision@10 came out at 0.75 in the notebook, though I treat it as indicative only: a hit counts whenever a single genre overlaps, which is an easy bar to clear
- The obvious weakness is genre-only vectors: far too many titles tie at a similarity
near 1.0, and splitting on
","without trimming leavesActionandActionas two separate terms in the vocabulary