How we measure delays

In short: public transport runs more punctually than it feels

The median deviation from the timetable, calculated over more than two million readings, is +12 seconds. The impression of constant delays comes from the tail of the distribution — and from the fact that a large part of that tail is not delay at all, but measurement artefacts.

Not every reading means the same thing

The same number in the data can describe completely different situations. That is why every reading is classified, and why the board and the ranking only show the ones that can be treated as a real delay.

classthresholdwhat it means
zwykleeverything elsea reading that can be treated as a real delay
kranieclast stop of the servicedescribes a finished journey, not a passenger waiting
nieprawdopodobnefrom 1800 s (30 min) mid-route, from 7200 s (120 min) at a terminusoutside the plausible range
odrobioneup to −262 s (4 min 22 s)vehicle well ahead of the timetable
rozkladowyno live datatimetable, not measurement

The thresholds are not plucked from the air — they follow from the distribution measured in our own observations. The classification is computed on read, not stored, so a rule can be corrected and the whole history recalculated.

Why the last stop of a service is different

A reading from the last stop describes a journey that has just ended — not the time somebody spends waiting. The numbers show it:

role of the stop in the serviceaverage deviationshare above 30 minutes
last1044 s6.25%
first98 s0.30%
mid-route80 s0.50%

At the last stop the delay is on average thirteen times larger, and extreme values occur more than ten times as often. The same thirty minutes means something different there than mid-route — hence a different threshold.

The role of a stop depends on the line, not on the stop itself: the same pole can be the start for one line, the end for another and an intermediate stop for a third.

Leaving early is a separate problem

A vehicle that leaves a minute ahead of the timetable is worse for a passenger than a late one: you can still catch a late service, but not an early one, because nobody comes to the stop “just in case”. That is why the statistics count it separately, rather than as a “negative delay”.

Forecast drift — can you trust the number on the board

Everything above describes deviation from the timetable. That matters, but it is not the number you look at while standing at a stop — there you see “in 8 minutes”. Forecast drift measures exactly that number: how much the board will still change its mind before the vehicle actually leaves.

We measure it like this: for every service we reconstruct what the board showed at successive moments before departure, and compare it with what it showed at the end. We take the reading at fixed moments — one minute before, two, three, five and so on — not as an average over all the changes. The difference matters: a forecast is only recorded when it changes, so an average over the records would measure how often it changes rather than what the board said. A stable service would leave no record at all and drop out of the statistics.

We keep the two directions of drift apart, because for a passenger they are not symmetric:

you planned;

about that.

It is the same asymmetry as with early arrivals, and for the same reason the result is reported separately for each side.

Why the median, not the mean

The data contains broken records — forecasts hundreds of days away from reality. There are few of them, below one ten-thousandth, but with an arithmetic mean a single such value shifts an entire hour's result by two orders of magnitude. That is why the median is what we put up front, and values above half an hour are discarded as broken — and we count how many there were, rather than quietly dropping them.

What drift does NOT measure

It does not measure the error of the forecast, only its variability. The reference point is the last forecast before the service disappeared from the board, because the real departure time simply is not in this data — the row vanishes and that is all we know. If the board were consistently two minutes wrong and knew it right to the end, the drift would come out at zero.

Only services that disappeared from the board around the announced time enter the statistics. Services closed off by a restart of the service, or vanishing long after their time, are discarded — we do not know what their reference point was, so any figure calculated from them would be invented.

What these numbers do not cover

The statistics describe what is visible in the operators' APIs — that is, forecasts for stops. They are not a measurement taken on board a vehicle, nor an official punctuality indicator of any carrier.