Hello, friend!
I have built some questionable things.
A recommendation engine that seemed determined to recommend the same three movies.
A sentiment model that looked at:
This is terrible.
and occasionally decided:
Positive.
And a dashboard that took so long to load you could probably make Moroccan mint tea before the first chart appeared.
None of these are the projects you proudly pin to the top of your portfolio.
Some of them live in repositories with names like:
dashboard-final
dashboard-final-v2
dashboard-final-ACTUAL
You know.
Professional software engineering.
But after building around a dozen side projects, dashboards, pipelines, ML experiments, and small applications, I noticed something.
The broken ones usually taught me the most.
Courses showed me what the tools could do.
Projects showed me what happens when those tools meet reality.
Reality has significantly worse documentation.
I Used to Think the Model Was the Project
When I started learning data science, the exciting part was obvious.
The model.
Machine learning.
Predictions.
Recommendation systems.
Anything that could reasonably have the words "AI-powered" attached to it.
Then I built YMovies.
I spent a lot of time thinking about recommendation logic.
Similarity.
Movie features.
User preferences.
Ranking.
All the fun stuff.
Then the application started asking less glamorous questions.
Where should user data live?
What happens when the movie API gives me missing values?
How do recommendations stay fast?
What happens when something fails?
How does any of this become something another person can actually use?
That was when I started understanding:
The model is only one part of the product.
A notebook saying:
accuracy = 0.92
is nice.
A system that keeps working when real people touch it is nicer.
Real Data Has Terrible Manners
Academic datasets are suspiciously clean.
Dates are dates.
Columns have sensible names.
Missing values politely appear as NaN.
Then you work with actual data.
One source gives you:
2025-08-12
Another gives:
12/08/2025
Another gives:
Aug 12, 2025
And somewhere, somehow:
12.08.25 FINAL
I ran into the same thing while experimenting with sentiment analysis.
The training dataset behaved beautifully.
Real reviews did not.
French.
English.
Darija.
Sometimes all three in one sentence.
Emoji.
Spelling mistakes.
Reviews containing nothing except:
👍👍👍
Very expressive.
Slightly inconvenient for preprocessing.
At first, the instinct is to blame the model.
Maybe better embeddings.
Maybe a transformer.
Maybe more training.
Sometimes.
But often the problem is much simpler.
The input is a mess.
That is one lesson I had to learn repeatedly:
Data cleaning is not the thing before the real work.
A lot of the time, it is the real work.
The Dashboard That Needed 43 Seconds to Think
At one point I built a small analytics dashboard.
A few charts.
Filters.
A date selector.
Nothing dramatic.
The first version took about 43 seconds to load.
Forty-three.
For around 5,000 rows.
That is not a dashboard.
That is a meditation exercise.
The problem was not mysterious.
I was loading far too much data into Pandas.
Recalculating aggregations every time somebody touched a filter.
Sending more data than the page actually needed.
Basically asking the computer to do unnecessary work and hoping it would be polite about it.
It was not.
The fix was boring.
Better SQL.
Indexes.
Precomputed aggregations where useful.
Caching things that did not change every five seconds.
Only retrieving the data the page actually needed.
Same dashboard.
Same numbers.
Much faster.
That project taught me something I still think about:
Most performance problems do not need bigger infrastructure first.
Sometimes you just need to stop making the computer do stupid things.
I Had an "Everything Needs ML" Phase
I think this is a required stage of learning machine learning.
You discover ML.
Suddenly every problem needs ML.
Classification?
ML.
Prediction?
ML.
Categorization?
Obviously ML.
If my toaster behaved strangely, I was probably three tutorials away from training a random forest.
Then I experimented with a small support-ticket classifier.
Five categories.
Good test scores.
Looked impressive enough.
Then imagine someone writes:
URGENT: BILLING SYSTEM DOWN
and your model confidently throws it into:
General inquiry.
Excellent.
Meanwhile, a boring rule system using keywords and fuzzy matching might handle the important cases more reliably.
Not exciting.
Not AI-powered.
Useful.
That forced me to start asking:
Does this problem actually need machine learning?
Sometimes yes.
Sometimes the answer is SQL.
Sometimes regex.
Sometimes fuzzy matching.
Sometimes a dropdown.
The goal is not to use the smartest-looking technology.
The goal is to solve the problem.
Pipelines Break Quietly
Web development gave me certain expectations.
When a website breaks, it normally tells you.
Page crashes.
API fails.
Button stops working.
Something obviously looks wrong.
Data pipelines can be much sneakier.
I built a small pipeline that collected job listings from several sources.
Extract.
Clean.
Load.
Store.
Display.
Very reasonable.
It worked beautifully.
For about two weeks.
Then one website changed its HTML.
Another started rate limiting.
A field I assumed would never be empty became null.
And because I barely had monitoring, I did not immediately notice.
That was when words like these stopped sounding boring:
Idempotency.
Running the same job twice should not destroy everything.
Validation.
Check that the assumptions you made yesterday are still true today.
Monitoring.
Know when the thing is broken.
Retries.
Because networks occasionally decide they have had enough.
Data engineers can seem slightly paranoid.
After building a few pipelines, I understand why.
Web and Data Stopped Feeling Separate
For a while, I treated web development and data as two different worlds.
Web was interfaces.
Data was SQL, pipelines, models, and analysis.
Then I started building actual data products.
And the line disappeared.
A dashboard needs good SQL.
But it also needs to feel fast.
A model needs to be accurate.
But the API around it needs to work.
A chart can look beautiful.
But if it communicates the wrong thing, it is just a beautiful mistake.
My web background made me ask:
Can someone understand this?
Does this feel fast?
What happens when something fails?
My data work made me ask:
Is the number correct?
Where did it come from?
When was it updated?
Can I reproduce it?
The projects I enjoy most now tend to live somewhere in the middle.
Work Taught Me That Definitions Matter More Than SQL
School assignments tend to have excellent manners.
Here is the dataset.
Here is the question.
Build the thing.
Real work sounds more like:
We need a dashboard for sales.
Cool.
What exactly counts as a sale?
Invoice created?
Payment received?
Order delivered?
What about refunds?
Cancellations?
Partial payments?
Suddenly the difficult part is not the SQL.
It is understanding what the question actually means.
Working in BI taught me this repeatedly.
A technically perfect answer to the wrong business question is still wrong.
That sounds obvious.
It becomes much more obvious after you spend hours calculating the wrong thing perfectly.
Some Projects Deserved to Fail
I have tried things that simply did not work.
Models that lost to basic baselines.
Pipelines with far more technology than the data justified.
Architecture diagrams that looked much more impressive than the actual problem.
There is a stage where you discover:
Airflow.
Kafka.
Spark.
dbt.
Redis.
Docker.
And naturally think:
What if I use all of them?
Then you realize the dataset fits comfortably inside an Excel file.
An important moment.
You can simplify.
Or build distributed infrastructure for 8,000 rows.
Side projects are nice because the stakes are low enough to make these mistakes.
You can overengineer something.
Feel the pain.
Then finally understand why people told you not to.
That lesson sticks.
Boring Patterns Keep Winning
After enough projects, certain things keep coming back.
Validate your inputs.
Do not blindly trust external data.
Process only what changed when possible.
Keep responsibilities separate.
Log important things.
Version your code.
Make failures visible.
None of this is particularly exciting.
Nobody reads about data validation and screams:
THIS CHANGES EVERYTHING.
But boring practices become very attractive after their absence ruins your afternoon a few times.
I think that is how engineering intuition develops.
At first, best practices are rules somebody else wrote.
Later, they become scars.
What I Would Tell the Version of Me Who Started
Do not skip the boring parts.
They become the important parts very quickly.
Build things end to end.
Even if they are small.
A tiny application that actually works will teach you more about systems than a massive architecture you never finish.
Do not use machine learning simply because you know machine learning.
Make it earn its place.
Care about the interface even when the project is "about data."
Someone still has to use the thing.
And most importantly:
Let your side projects break.
Let an API change.
Let your query become slow.
Let an assumption fail.
Let the data become messy.
Then figure out why.
That process is annoying while it is happening.
Later, you realize that was the lesson.
The Unsexy Truth
Most data work is not particularly cinematic.
It is SQL.
Cleaning.
Debugging.
Checking numbers.
Reading logs.
Fixing pipelines.
Discovering that the entire problem was one unexpected space in a CSV header.
Building dashboards that answer simple questions clearly.
And that is probably why these side projects taught me so much.
They slowly changed the questions I ask.
I used to look at a problem and think:
What cool technology can I use here?
Now I am much more interested in:
What is the simplest thing that will work reliably?
A good data system does not need the largest architecture diagram.
It needs to work.
The numbers need to be trustworthy.
Someone else needs to understand it.
And ideally, it should still work tomorrow.
Most of these experiments still live somewhere on my GitHub @yassnemo.
Some are polished.
Some worked once.
Some are probably one dependency update away from becoming archaeology.
That is fine.
They did their job.
They taught me how to build the next one better.
Much love,
Yassine Erradouani
