How it works.
Two minutes for the outline. The questions under each section open for the detail, and any of them can be skipped.
1. The question
Station boards say a train is expected at 17:27, with no sign of how sure that is. RailCast answers with a range instead: most likely 17:26, probably between 17:22 and 17:34. The arrival should land inside that range four times in five, and every night that claim is checked.
It only answers about trains already running. Every input describes a journey in progress, so for tomorrow's 08:15 it has nothing to reason from, and it declines instead of guessing.
Why is a range more useful than a single time?
A train that has been steadily two minutes late for five stops and a train that has just picked up an unexplained delay both get the same kind of estimate on a board: one time, stated with the same confidence.
The range is what separates them. With a four minute connection at the other end, a two minute range says you are fine and a twelve minute range says you might not be. A single number cannot tell you that.
What does "four times in five" mean exactly?
The formal name is a prediction interval: a low bound and a high bound, chosen so the true answer falls between them a stated share of the time. This one aims for 80%. That is a claim anyone can check against what actually happened, which is the whole point of the accuracy page.
2. Where the numbers come from
Irish Rail's public realtime feed: no sign up, no key, and provided as is with no support. Four parts of it are in use.
- One train's full journey on one date, with every stop, the scheduled time and the actual arrival. It is the record of what really happened, and what the model learns from.
- A station's board, including Irish Rail's own expected arrival, the baseline being competed against. It exists only live, so it is captured every five minutes.
- Which trains are running right now, so the system knows what it can be asked about.
- The station list, fetched once and cached.
How do you know the feed serves real history, not a replayed timetable?
By asking for the same train on the same date in 2020, 2024, 2025 and 2026. The times all differ, and the 2020 version has three fewer stops and a different scheduled arrival at Cork, which is the 2020 timetable on 2020 infrastructure. It is genuine history, back to at least 2007.
That single check is what makes a missed collection run recoverable. It does not apply to the operator's expected arrival, which exists only live and can never be backfilled, which is why a poller has been running every five minutes since August.
How much data was it trained on?
28,706 gzipped responses covering 1,087 train codes across 34 dates, about half a million stop level records. They were collected once, at two requests per second, overnight, backing off whenever the server complained.
3. The problem that shaped everything
A model learns lateness from true arrival times, its labels. On parts of the network the feed's arrival field holds the timetable instead, presented as though it were observed. Nothing marks it. Train on it and the model learns those lines are perfectly punctual.
Call it an echo. Irish Rail publish a list of ten lines where this happens, and it looked like the answer: those lines echo 21% of the time against under 3% elsewhere. It was the wrong conclusion. Split instead by whether signalling equipment captured the time automatically, and among machine captured records the flagged lines echo less than the rest of the network.
So the rule became: trust the capture flag on each record, never the published line list.
How can an echo be spotted at all?
A real arrival almost never lands on exactly the scheduled second. Across the whole archive, 2.92% of arrivals are identical to the scheduled time.
Why did the line list look so convincing?
Far more of the flagged lines' records were not machine captured: 47% against 1.5% elsewhere, and the non-automatic records are where echoes live. Among machine captured records the flagged lines echo less; among hand-entered records they echo more. The line name was standing in for how the data was captured.
A comparison that reverses when you split it has a name: Simpson's paradox. Three quarters of the suspect records sit on lines the documentation never mentions, so a line filter would have thrown away good data and kept bad.
Why is nothing deleted, even when it looks obviously wrong?
An exact match is suspicious, not proof. About 2.4% of the most trustworthy records genuinely do land on the scheduled second. Part of the reason is that every time in the feed is rounded to six seconds, so there are only ten possible second values in a minute and coincidences are commoner than they look.
So nothing is deleted at ingestion. The flag is carried through, what to exclude is decided when evaluating, and the numbers are reported both ways. A result that agrees with the documentation is exactly the one nobody checks carefully enough, and this one was wrong.
4. The model
One model for the whole network, not one per train, trained three times: at the 10th, 50th and 90th percentile. That gives a low bound, a most likely time and a high bound.
Inputs describe the situation, never the identity. The train code is not an input, so a service launched next March, or renumbered, gets a prediction like any other.
How can a model learn a percentile?
With lopsided penalties. An ordinary model is trained to be wrong by as little as possible in both directions, which gives a middle. To learn the 90th percentile, being too low is punished nine times as heavily as being too high, so the model settles where 90% of outcomes fall below. This is called quantile loss.
The models themselves are gradient boosted trees: many small decision trees, each correcting what the ones before it got wrong.
What does it look at?
Delay accumulated upstream today, stops and minutes remaining, time of day, day of week, and a route proxy. All twelve inputs exist for a train that started running yesterday. The strongest by far is the delay at the previous stop, because lateness carries forward along a journey.
A model that had learned "A218 runs two minutes down" would have nothing to say about a new service, which would arrive as an unknown category, or about the same service after a renumbering.
The three models are trained separately. What stops them disagreeing?
Nothing forces them to agree, so occasionally the 10th percentile comes out above the 90th. That is called quantile crossing. The three values are sorted at prediction time, which is provably no worse than leaving them crossed.
Why is Irish Rail's own estimate not an input?
It is the baseline being compared against, so feeding it in would make the comparison meaningless.
5. How it is measured
Head to head against Irish Rail's own expected arrival, the number a passenger would otherwise be reading. Beating a naive guess would prove nothing.
One week of data, 20 to 26 July, was sealed from the start and opened once, on 10 September 2026: 80.0% of arrivals inside the range against the 80.0% claimed, on 217,290 predictions the model had never seen.
Nothing is reconstructed after the fact. Every prediction is logged the moment it is made, with a timestamp and the model version. The nightly scorer can read that log and has no permission to write to it, enforced by the AWS credentials it runs with.
What makes the comparison fair?
Three rules, each there because ignoring it produces a flattering wrong number.
- Matched events only. A result counts only where both the model and the operator produced a prediction for the same train, at the same station, at the same moment. Comparing all of one against all of the other compares populations, not methods.
- One result per event, not per poll. A single arrival is polled about eighteen times as the train approaches. Counting those as eighteen independent results inflates the sample size roughly eighteenfold and makes noise look significant.
- Matched precision. The operator publishes to the minute and the model to the second, so the model's answer is rounded to the minute before either is scored.
Why can the sealed week only be used once?
It was not looked at through any model change, and the analysis was written down and committed before it was opened. Once opened, every later decision would be made knowing the result. A test set that has been looked at is no longer a test set, so any future held-out check needs new data.
6. What is actually running
Everything is a scheduled function writing to object storage. There is no database and no server that stays up.
Why no database?
The read pattern is "give me one whole day" and the write pattern is append only. Neither needs one, and a database that stays up bills by the hour whether or not anyone asks it anything.
What does it cost to run?
Measured on 23 August 2026, before the site was added, about ten cents a month, almost all of it storage requests.
Why is there a generator asking the service questions?
An accuracy page built on a handful of the author's own test queries would be worthless, and almost nobody visits this site. So a scheduled job asks the live service a spread of questions every five minutes, one per lead band, sampled across whatever is in service rather than favouring any route.
Those are real predictions, made before the outcome existed. They are simply not anyone's, and the accuracy page says so. Predictions requested by this site's boards carry a different source tag, so the two can always be told apart.
7. What it cannot do
The ranges do not cover disruptions. On real delays over an hour, coverage is zero, and it was zero for the previous model too.
- No upstream report, no prediction. The service declines rather than guessing, and the decline rate is published beside the accuracy.
- Fewer answers on a station board. Given a train in service it answers about nine times in ten, but only about four in ten board entries are trains that have already left.
- Some lines are below the promise. In the seven days to 8 September 2026 the range held 67.1% of the time on the Cork corridor, 63.9% on the Kildare line and 56.8% on other intercity routes, against 80% claimed. Published per line.
- It is not a route planner. No map, no accounts, no notifications.
Why can't it see a disruption coming?
On 2 September 2026 three Sligo line trains each lost 75 to 89 minutes in a single stretch. The service had been asked sixteen times across those journeys, and at every one of those moments the train was two to seven minutes late. Nothing in "how late is it now, how far is left to go" can see that coming.
A range wide enough to catch it would be wrong about ordinary days to hedge against rare ones, so the ranges are calibrated for ordinary lateness.
What changed on the lines below the promise?
That has not been established. The retraining rule fired on those three groups, and the ranges were not widened to hide it: coverage is published per line, recomputed every night, rather than blended into one number.
8. The thing that kept happening
Eleven separate failures in this project shared one shape. None raised an error. Every one produced output that looked exactly like a correct, unremarkable result, and every one was caught the same way: by taking a number and asking what it should have been.
| What it looked like | What it was |
|---|---|
| Flagged lines reporting arrival times | Scheduled times echoed back as though observed |
| 420 successful downloads, HTTP 200 | An internet provider's login page, 1,369 identical bytes each time |
| A deploy reporting success on an alarm | A subscription nobody confirmed, deleted 48 hours later. Alerting was dead and looked fine |
| A model with a 22 minute average error | Two ways of computing "late" that disagreed by a day. No model change fixed it |
| Arrival times marked machine verified | Real times, captured by real equipment, filed against the wrong train |
| A model covering a quarter of severe delays | Every covered case was a wrong label matched by a range learned from the same wrong labels |
| A rebuild that "completed, exit code 0" | It had crashed on its first line. The exit code belonged to the command after the pipe |
| A nightly job printing a healthy report | It ran out of memory four minutes later. A frozen page and a working one look identical |
A system which works and a system you can tell is working are different things, and most of the effort here went into the second.
Where did these failures come from?
Four of them were introduced by the author, after the system was working, while writing up the others. The lesson drawn was not "be more careful". It was to attach a denominator and a check to every published number.
How many are still in the project, undetected?
Almost certainly some. Eleven is a count of the ones that were caught, and each was caught only because the number it touched had a denominator attached and a check that asked what it should have been. That is the design, and it is the argument for publishing where the model loses rather than only where it wins.