These are scheduled samples, not visitor traffic. A scheduled job asks the live service a spread of questions every five minutes, across whatever is in service. Every prediction below is real and was logged before the arrival happened.
The week nobody looked at
The strongest number here. One week of data was sealed in June and opened once, on 10 September 2026, against a plan written down beforehand.
| Checked against | Sealed week | Live, 2 to 8 Sept | |
|---|---|---|---|
| Average error | 59.7s | 58.3s | 95.9s |
| Median error | 29.0s | 29.1s | 48s |
| Inside the range | 80.2% | 80.0% | 75.1% |
| Predictions | 217,653 | 217,290 | 88,576 |
Why this settles something
On data the model had never seen, the range held 80.0% against 80% claimed. Live coverage was 75.1% over 2 to 8 September 2026. So the model was calibrated when trained, which leaves the railway having changed since.
What does "sealed" mean here?
In June the data was cut three ways: one part to learn from, one part to check every decision against, and one week sealed and never opened. Every choice made since was tested against the middle part, and each of those looks leaks a little into the result. The sealed week had none of that done to it.
Were the misses lopsided?
No. They split 10.4% below the low bound against 9.6% above, where 10 and 10 is perfect. A model can reach 80% with badly lopsided tails. This one does not.
How did it do against a simpler predictor, and further ahead?
It beats the naive predictor that assumes the current delay simply carries forward by 32.1%, and the known weakness reappears unchanged: coverage falls to 73.3% when asked more than an hour ahead.
No operator comparison is possible for that week, because Irish Rail's estimate exists only live and cannot be recovered afterwards.
By line
Blended into one number this looks like a small shortfall. Split by line it is uneven, so each line is listed on its own.
| Line | Scored | Error, all scored | vs operator, matched | Coverage | against 80% claimed |
|---|
How to read this table
The mark on each section of track is the 80% the range claims. Amber is below the 70% at which the retraining policy requires investigation.
Stations off the board sample shows no operator comparison because the poller archives only 30 stations, so there is no operator estimate on record to compare against. Those predictions are still scored against the real arrival.
The rule fired, and nothing was patched
A rule written in advance says a line below 70% for a week needs action. On 8 September 2026 it fired on three groups, the Cork corridor, the Kildare line and other intercity routes, while the DART and the Dublin hubs held. What changed on those lines has not been established.
The ranges could have been widened until the number came back to 80%. That would have hidden the shift instead of explaining it, so the number is published per line, recomputed every night.
Day by day
Average error against Irish Rail's own expected arrival, on matched events only.
RailCast Irish Rail's own estimate
measured coverage the 80% it claims
| Date | Scored | Matched | RailCast | Operator | Better by | Coverage |
|---|
What do Scored, Matched and the dagger mean?
Scored counts every prediction with a real arrival to check against. Matched counts the subset where the operator also published an estimate for the same train, station and moment, with both answers rounded to the minute so neither gets an unearned edge. That is the only fair comparison.
A dagger marks a day with fewer than 100 matched events, which is noise rather than a trend. Those days stay in the table with their sample size and out of the chart.
By how far ahead you ask
The next stop is an easier question than one an hour away. The error roughly triples across the range and coverage decays, which is what an honest model should show.
| Asked | Scored | Matched | RailCast | Operator | Better by | Coverage |
|---|
On real delays over an hour, coverage is zero
Not low. Zero. It was zero for the previous model too.
Why can't it see those coming?
When three Sligo line trains lost 75 to 89 minutes in one stretch on 2 September 2026, each was two to seven minutes late at every moment the service was asked. Nothing in "how late is it now, how far is left to go" can see that coming. A range wide enough to catch it would be wrong about ordinary days to hedge against rare ones.
Didn't the previous model cover some?
It appeared to cover a quarter of them, until the labels behind that figure turned out to belong to other trains.
How often it answers at all
A system that answers rarely and well is a different product from one that answers often and well.
| Outcome | Predictions | Share |
|---|
What do the outcomes mean?
A prediction that could not be scored is reported here, not dropped. No arrival recorded means the operator never logged one for that stop. Arrival identical to the timetable is the signature of a scheduled time echoed back as though it were an observation, which is the data problem that shaped this whole project.