I wanted to build a data engineering project that did not end with a pipeline diagram.
You know,
API goes in.
Kafka happens somewhere in the middle.
There is a database.
Then a dashboard with four cards telling you everything is "healthy."
Technically correct.
Not especially interesting.
I wanted the final thing to answer a more difficult question:
Can I take messy real-world data, make it reliable, and turn it into something another person can actually use?
That became UrbanPulse (claude suggessted this name).
A real-time city intelligence platform combining weather, air quality and shared-mobility data into one view of what is happening across a city.
The interface is part of the project.
But the interesting bit is everything required to make that interface believable.
The actual problem wasn't getting the data
Public city data is surprisingly available.
Weather can come from Open-Meteo.
Air-quality measurements can come from OpenAQ.
Bike-sharing systems often publish GBFS feeds containing station capacity, available vehicles, docks and status information.
Getting a response from these APIs is easy.
Something like this already gets you surprisingly far:
response = client.get(url)
response.raise_for_status()
data = response.json()
And now you have data.
Well, sort of.
The problem is that none of these sources describe the world in quite the same way.
They have different schemas.
Different refresh frequencies.
Different timestamp formats.
Different ideas of what a location is.
Different failure modes.
And because they are external systems, I don't control any of them.
So UrbanPulse is really a normalization problem disguised as a city dashboard.
The useful part is not collecting three APIs.
It is turning them into one coherent system.
I built the pipeline around one rule
I wanted every number in the interface to have a traceable path back to its source.
That sounds obvious until something breaks.
Suppose the dashboard suddenly shows 312 mobility stations instead of 487.
Did 175 stations actually disappear?
Did the provider stop returning them?
Did ingestion fail halfway through?
Did a transformation filter them out?
Did the frontend query the wrong period?
If all I kept was the final table, debugging this would be unpleasant.
So I kept the raw source data.
The flow looks roughly like this:
Public APIs
↓
Ingestion
↓
Raw snapshots
↓
Validation + normalization
↓
PostgreSQL / PostGIS
↓
dbt models
↓
FastAPI
↓
UrbanPulse
Raw data is the boring part of the architecture that becomes extremely interesting the first time something goes wrong.
If my transformation says one thing and the source says another, I can go back and inspect exactly what arrived.
That gives me a chain of evidence instead of a chain of assumptions.
Not everything needed to be "real-time"
This was one of the more useful decisions in the project.
At first, I was tempted to think:
Real-time platform.
Therefore Kafka everywhere.
That would certainly look impressive in the architecture diagram.
It would also be silly.
Weather does not need millisecond processing.
Some analytical models are naturally scheduled.
Historical aggregates do not suddenly become better because they traveled through a message broker first.
Mobility data is different.
Bike availability changes frequently, and storing those changes over time is useful.
That is where a streaming path makes sense.
So I use Kafka where an event stream gives me something concrete: separation between ingestion and consumers, independently processed events, and the possibility of replay.
For scheduled work, Airflow makes more sense.
It handles jobs where I care about dependencies, retries, execution history and backfills.
So the system deliberately contains both batch and streaming.
Not because having both looks good on a résumé.
Because the workloads are different.
That distinction became important to me while building this.
A technology should enter the architecture because a problem invited it.
Not because the README looked slightly empty.
The database stayed boring on purpose
UrbanPulse uses PostgreSQL with PostGIS.
There are much more exciting combinations I could have chosen.
ClickHouse.
Elasticsearch.
A dedicated time-series database.
Redis.
Some distributed thing that would require a second monitor just to fit the architecture diagram.
I didn't add them.
At least not yet.
The project has relational data.
Cities contain locations.
Stations belong to locations.
Measurements belong to locations and points in time.
And because the application is geographic, I need spatial queries.
PostgreSQL handles those requirements well, and PostGIS gives me proper geospatial capabilities without introducing another system.
This was one of the decisions I wanted the project to demonstrate.
Not adding technology is also an engineering decision.
If historical analytics eventually become slow enough that PostgreSQL is genuinely the bottleneck, I can benchmark it.
Then maybe ClickHouse becomes useful.
But I want the problem before the solution.
Not the other way around.
dbt is where the APIs stop mattering
The frontend should not know that one provider calls something station_id while another nests its coordinates inside some strange object.
That mess belongs near the source.
After normalization, I want the rest of the system to speak one language.
So I model the data into facts, dimensions and marts.
For example:
dim_city
dim_location
dim_station
dim_source
fct_weather_hourly
fct_air_quality_measurements
fct_mobility_station_snapshots
Then I create models closer to the questions the application needs:
mart_city_overview
mart_mobility_hourly
mart_environment_hourly
mart_pipeline_health
This is where dbt earns its place.
The SQL itself is not the interesting part.
The useful part is being able to treat transformations like software.
Dependencies are explicit.
Models are version controlled.
Important fields are documented.
Tests run against the data.
And lineage can be generated from the transformations instead of manually drawing arrows and hoping they remain correct.
Which leads to one of my favourite parts of the project.
I made the pipeline visible inside the product
Most data projects show you the result.
UrbanPulse also shows you the machinery.
There is a Pipeline Monitor for ingestion and transformation jobs.
Instead of just seeing:
Weather: 21.8°C
I can expose information such as:
WEATHER INGESTION
Status Healthy
Last success 42s ago
Duration 1.8s
Rows processed 4,320
Freshness Current
The important detail is that these values come from backend operational metadata.
They are not fake numbers added to make the interface look technical.
There is also a data-quality view.
A pipeline succeeding does not mean the data is good.
A job can run perfectly and still produce stale, incomplete or ridiculous data.
So I test assumptions such as:
station_id_unique
station_capacity_non_negative
station_coordinates_valid
measurement_timestamp_not_null
source_freshness_within_threshold
If the provider sends a bike station somewhere in the middle of the Atlantic, I would quite like the system to notice.
Preferably before I do.
Historical data created the feature I didn't initially expect
Once I started storing mobility and environmental snapshots, the system had something more useful than current state.
It had memory.
That led to City Replay.
Instead of only asking:
What is happening right now?
I can ask:
What was happening at 14:00 yesterday?
And then move through time.
22 September 2026
14:00 ━━━━━●━━━━━━━━ 20:00
▶ Replay
As the selected time changes, mobility availability, environmental readings, weather, map state and KPIs can change with it.
This feature looks like frontend work.
It is actually a data-engineering feature.
The slider is easy.
The reason the slider can exist is that the pipeline kept historical state.
That is the kind of connection I wanted UrbanPulse to show.
Infrastructure decisions eventually become product capabilities.
The frontend gets modeled information, not database leftovers
Another decision I made was to avoid turning the browser into a second data-processing platform.
The API returns information shaped around the product.
For example:
/api/v1/cities/paris/overview
/api/v1/mobility/stations
/api/v1/environment/air-quality
/api/v1/operations/pipelines
/api/v1/replay
The React application should not fetch thousands of raw observations and calculate everything itself.
Large aggregations happen server-side.
The browser gets what it needs.
This keeps the interface lighter, but there is another benefit.
The frontend does not need to understand how the warehouse works.
I can change transformation logic without teaching a React component about my staging models.
That boundary is boring when everything works.
Very useful when things change.
What I deliberately didn't build
This part matters as much as the architecture.
I did not add Kubernetes.
The project does not need it.
I did not add Spark or Flink.
The workload has not earned them.
I did not add Redis simply because applications apparently develop a legal obligation to install Redis after reaching a certain number of Docker containers.
I did not add ClickHouse before proving PostgreSQL was a problem.
And I do not describe the project as "big data" just because Kafka exists somewhere inside it.
UrbanPulse is supposed to show engineering judgment.
Not technology collection.
I would rather explain exactly why eight components exist than show twenty-five logos I barely needed.
What UrbanPulse actually became
At this point, the project is not really about weather.
Or bikes.
Or PM2.5.
Those happen to be useful datasets because they change over time, contain geography and come from independent real-world systems.
The actual project is about this:
fragmented sources
↓
reliable ingestion
↓
preserved raw history
↓
tested transformations
↓
understandable models
↓
useful APIs
↓
a product people can explore
And that last arrow matters.
A data pipeline is not automatically useful because the DAG is green.
Someone eventually needs the information.
UrbanPulse forced me to think about that entire distance.
From the API response all the way to the person looking at the screen.
One thing I would recommend
If you are building a portfolio project, start with the problem you want the final user to solve.
Then work backwards.
Do not begin with:
I want a project using Kafka, Airflow, dbt and Kubernetes.
That almost guarantees you will spend your time inventing reasons for your tools.
Start with:
I want someone to understand what is happening across a city right now, trust the information, and inspect where it came from.
Now the questions become useful.
Do I need history?
Do I need streaming?
What needs orchestration?
What can fail?
What should be tested?
What does the frontend actually need?
And occasionally:
Do I really need another database?
Building UrbanPulse taught me that the architecture gets much easier to defend when every box exists because the product pushed it there.
That is probably the main thing I will carry into the next project.
Build the problem first. Let the stack catch up.
