Skip to main content

YMovies: Creating a Movie Recommendation Web App

00:10:21:33

Hello, friend!

This project started the way many serious engineering projects do.

I was lying on a couch watching Netflix.

Very demanding research.

At some point, I noticed the "Because you watched..." section again.

Nothing unusual there.

Netflix has been recommending things to people for years.

But this time I stopped for a second.

It had recommended something that was annoyingly accurate.

And because I was studying data science at the time, my brain did what it normally does when I see something interesting:

Could I build that?

This question has caused me a surprising amount of unnecessary work over the years.

YMovies was one of those times.

One Small Idea Became a Whole App

Around the same time, I was taking a Machine Learning course on Coursera.

Recommendation systems came up.

Suddenly, the Netflix idea stopped looking like magic and started looking like something I could actually attempt.

Not Netflix itself, obviously.

I did not have several thousand engineers, billions of viewing events, or whatever terrifying infrastructure they have running behind that home page.

I had Python.

A laptop.

The TMDB API.

And unreasonable confidence.

Good enough.

The idea became YMovies: a movie discovery app that learns what you like and recommends things you might actually want to watch.

You can check out the code or play with the live version here:

View on GitHub

Live Demo

The First Problem: I Had No Users

The first recommendation technique I learned about was collaborative filtering.

The idea is pretty intuitive.

Suppose you and I both like Inception.

I also love Interstellar.

You have never seen it.

The system notices that people with similar taste to yours tend to enjoy Interstellar and recommends it to you.

Very powerful.

There is only one small problem.

You need users.

Preferably lots of them.

You need ratings.

Clicks.

Likes.

Watch history.

Patterns between people.

YMovies had approximately zero of those things.

Hard to discover meaningful user behavior when the user base is me refreshing localhost.

So collaborative filtering was not especially useful yet.

That pushed me toward content-based filtering.

Instead of asking:

What do people similar to this user like?

Content-based filtering asks:

What movies are similar to the movies this user already likes?

Much better problem for a new app.

Movies Are Basically Bags of Features

A movie is more than a title.

It has genres.

Actors.

A director.

Keywords.

A synopsis.

Release year.

Runtime.

All of those things give you clues about what kind of movie it is.

TMDB gives you a lot of this metadata.

So I could take a movie like The Dark Knight and describe it using its features.

Action.

Crime.

Drama.

Christian Bale.

Christopher Nolan.

Batman.

Gotham.

Then compare that description against every other movie.

The closer the features are, the more similar the movies probably are.

Simple idea.

The interesting part was turning those words into something a computer could compare.

TF-IDF, But Without Making It Sound Terrifying

At the center of the recommendation engine is TF-IDF vectorization.

Which sounds much worse than it is.

The basic idea is that I combine useful movie information into one piece of text.

Something like:

text
action crime drama christopher_nolan christian_bale batman gotham vigilante ...

Then TF-IDF turns that text into numbers.

The important words get useful weights.

Words that appear everywhere matter less.

Words that help distinguish one movie from another matter more.

Now every movie is represented as a vector.

Once I have vectors, I can compare them using cosine similarity.

The closer two vectors point in the same direction, the more similar the movies are.

So if you like The Dark Knight, movies sharing similar genres, themes, actors, directors, or keywords should end up nearby.

That becomes the foundation of:

Because you liked...

Not Every Feature Should Matter Equally

There was another problem.

If I simply throw every feature into the system once, I am quietly saying:

The director matters exactly as much as one random word in the description.

That did not make much sense.

Genres are usually important.

Directors can be very important.

Major cast members can matter.

Some random word buried in the overview?

Probably less.

So I gave certain features more weight.

Here is part of the feature-building logic:

python
def _combine_features(self, row):
    features = []

    if 'genres' in row and row['genres']:
        genres = row['genres'].split()
        features.extend([g for g in genres for _ in range(3)])

    if 'overview' in row and row['overview']:
        features.append(row['overview'])

    if 'cast' in row and row['cast']:
        cast_list = row['cast'].split()[:5]
        features.extend([c for c in cast_list for _ in range(2)])

    if 'director' in row and row['director']:
        features.extend([row['director']] * 3)

    if 'keywords' in row and row['keywords']:
        features.append(row['keywords'])

    return ' '.join(features)

Genres get repeated three times.

Directors get repeated three times.

Main cast members get repeated twice.

Very sophisticated mathematical explanation:

I wanted them to matter more.

Sometimes machine learning involves complicated equations.

Sometimes it involves writing Christopher Nolan's name three times.

"Because You Liked..." Was the Fun Part

This was the feature that started the whole project.

If a user likes a movie, YMovies can find nearby movies in the TF-IDF feature space.

Like Inception?

Find films with similar genres, themes, cast, directors, and keywords.

Then present them as:

Because you liked Inception...

Simple.

But it makes the recommendation feel much more personal.

There is also an important little detail:

I exclude movies the user has already interacted with.

Because recommending the same movie someone already added to their watchlist is technically a recommendation.

Just not a particularly useful one.

Then I Wanted It to Feel More Personal

Content similarity is useful.

But there is an obvious limitation.

Suppose someone likes Inception.

Does that mean they want twenty Christopher Nolan movies?

Maybe.

Honestly, I would understand.

But recommendation systems get boring very quickly if they only produce increasingly similar versions of the same thing.

So I started experimenting with a hybrid approach.

The idea was to combine several signals:

  • movies the user liked
  • recent interactions
  • preferred genres
  • year preferences
  • runtime preferences
  • content similarity

The recommendation logic started looking something like this:

python
def get_recommendations(self, user_data, n=20):
    recommendations = []

    if liked_movie_ids:
        for movie_id in recent_liked_ids:
            similar_movies = self.get_because_you_liked_recommendations(
                movie_id,
                n=10
            )

            # Filter watched movies and add recommendations

    if quiz_genres:
        quiz_recs = self._get_quiz_based_recommendations(
            quiz_genres,
            quiz_year_range,
            quiz_duration,
            exclude_ids,
            n=20
        )

        # Add quiz-based recommendations

    return recommendations

Nothing magical is happening here.

I am just collecting several clues about what the user might enjoy and combining them instead of trusting one signal completely.

Which is closer to how taste actually works anyway.

The Cold Start Problem

Recommendation systems have an awkward first-date problem.

You open the app.

The app knows absolutely nothing about you.

And somehow it is expected to impress you.

This is called the cold start problem.

Collaborative systems struggle here because there is no interaction history yet.

My first solution was a small onboarding quiz.

What genres do you like?

Recent or classic movies?

Short movies?

Long movies?

Something in the middle?

From those answers, the system could narrow the movie catalog before the user had liked anything.

For example:

  • Recent → movies from roughly the last five years
  • Classic → movies before 2000
  • Short → under 100 minutes

Then I could rank the remaining options using things like popularity and ratings.

It was a nice solution.

There is just one funny detail.

I eventually did not use the quiz in the final experience.

And I think that is worth mentioning.

Building something does not mean you have to keep it.

Sometimes a feature helps you understand the problem and then turns out not to belong in the final product.

That is not wasted work.

That is product development.

TMDB Gave Me Data. It Did Not Give Me Clean Data.

APIs always look beautiful in documentation.

Then you actually use the data.

Genre IDs need mapping.

Some movies have missing overviews.

Some fields are empty.

Cast lists vary.

Keywords can be inconsistent.

Movies have missing values in places you assumed would obviously contain values.

Classic.

So before any recommendation logic could work properly, the data needed cleaning and normalization.

That happens around the movie-loading logic in app.py.

This part of machine learning projects tends to get significantly less attention than the algorithm itself.

But the recommender is only as good as the features you feed it.

A beautiful cosine similarity calculation cannot rescue terrible input data.

Garbage in.

Mathematically sophisticated garbage out.

Then It Got Slow

Another fun discovery:

Comparing thousands of movies against thousands of movies can get expensive surprisingly quickly.

At first everything feels instant.

Then the dataset grows.

Then you click something.

Then the computer begins thinking deeply about its life choices.

TF-IDF helps because its vectors are naturally sparse.

Most movies do not contain most terms.

There is no reason to store enormous dense matrices full of zeros.

Using sparse matrices keeps memory usage much more reasonable and makes the similarity calculations practical.

This was one of those nice moments where something that looks like an implementation detail suddenly matters a lot.

The algorithm can be correct and still be useless if it takes forever.

Recommendation Systems Are Mostly Tradeoffs

One of the harder questions was not technical.

It was:

How personal should the recommendations be?

Imagine someone likes three science fiction movies.

Should the entire home page become science fiction?

Probably not.

But ignoring those likes would also be strange.

You want relevance.

But you also want discovery.

You want familiar things.

But not twenty copies of the same recommendation.

This is where I experimented with things like a diversity_factor.

Essentially:

How much should we trust the obvious matches, and how much should we let something different sneak in?

There is no perfect answer.

Recommendation systems are full of questions like this.

The math can tell you which items are similar.

It cannot tell you exactly how boring your homepage becomes after displaying fourteen of them.

The Machine Learning Was Only Half the Project

This is something I did not fully appreciate when I started.

I thought:

I am building a recommendation system.

What I was actually building was:

a recommendation system,

inside a backend,

connected to an API,

using movie data,

with users,

watchlists,

likes,

a frontend,

movie pages,

search,

deployment,

and all the small pieces required to make the recommendation system useful to an actual human.

The recommendation algorithm is important.

But nobody wants to run:

bash
python content_based_recommender.py

every time they want movie suggestions.

The model needs a product around it.

That was probably one of the more useful lessons from YMovies.

A machine learning model sitting in a notebook is an experiment.

Putting it inside something people can actually use is a completely different problem.

Finishing It Was Harder Than Starting It

I would be lying if I said I was completely sure this project would make it to the end.

The beginning was slow.

There were parts I rewrote.

Features I tried and dropped.

Recommendation logic I adjusted.

Data I had to clean.

UI decisions that looked better in my head.

The usual.

Starting a project is exciting because everything is possibility.

Finishing a project involves discovering what all those possibilities actually cost.

But I am glad I finished it.

Because somewhere along the way, the pieces stopped feeling separate.

content_based_recommender.py finds similar movies.

hybrid_recommender.py adds information about the user.

app.py connects everything to the application.

TMDB provides the movie world around it.

And the interface turns all of that into something someone can actually click.

What I Would Do Differently Now

If I rebuilt YMovies today, I would probably experiment with more advanced recommendation approaches.

Collaborative filtering becomes much more interesting once real user interactions exist.

Embeddings could represent movie meaning better than plain TF-IDF in some cases.

User behavior could influence ranking more dynamically.

And I would like to experiment with modern recommendation libraries and open-source systems like the approaches used around large-scale platforms.

At one point I was also interested in exploring ByteDance's open-source recommendation work.

Maybe that becomes another project.

That is one of the nice things about finishing something.

You stop asking:

Can I build this?

And start asking:

What would I change if I built it again?

Those are much better questions.

The Best Part Was Not the Recommendation Algorithm

The original idea came from Netflix.

Build something that says:

Because you liked this, maybe you will like that.

And yes, I built that.

But the useful part ended up being everything around it.

Choosing an algorithm that made sense with no users.

Working with messy external data.

Deciding which movie features mattered more.

Making similarity calculations fast enough.

Balancing relevance with diversity.

Dropping features that did not improve the product.

Turning a machine learning script into an actual application.

That is the part projects teach you.

Courses give you algorithms.

Documentation gives you APIs.

Tutorials give you examples.

But a project eventually looks at you and says:

Cool.

Now make all of those things work together.

And that is usually where the interesting learning starts.

YMovies began because Netflix recommended me a movie.

Apparently that was enough to create several months of work.

So it goes.

Much love,

Yassine Erradouani