Four Beyond the Breadcrumbs trails begin in different places in Clyde Street.

Each was developed by returning to Keith Lyons’s original Clyde Street writing and the sources around it. They are not Clyde Street articles, and they are not Keith’s retrospective synthesis. They are reader routes created in Beyond the Breadcrumbs: source-led returns to his writing that follow particular questions while keeping the original Clyde Street posts within reach.

What Can an Observer Know? asks what different kinds of observation allow us to know.

A History with Missing Pieces follows what happens when the surviving past is incomplete.

What Happens to a Prediction After the Future Arrives? looks at what becomes of a forecast once the outcome is known.

What Has to Travel with Knowledge? asks what must accompany a knowledge object if somebody else is to understand and question it later.

You do not need to read those trails first. Each remains open from here, and each leads back to the original Clyde Street posts and sources it follows.

This article exists because, once those four Beyond the Breadcrumbs trails are placed beside one another, another piece of learning becomes visible.

Do not just keep the answer.

Keep enough of the journey to know what the answer can mean.

In October 2019, Keith Lyons ran into a small problem with an apparently solid phrase: official data.

During the Rugby World Cup, figures on the World Rugby website could change after matches had been played. Pass counts were one example. Keith responded by declaring a practical rule for his own dataset: he would use the values as they stood twenty-four hours after each game.

The important part is not that twenty-four hours was the perfect answer. Clyde Street gives us no basis for claiming that. The important part is that Keith made the state he had used visible.

Once an official source can change, “the official data” is no longer enough information. A later reader needs to know which version entered the analysis, and when.

That small practical decision is one example of the larger learning that becomes visible when those four trails are read together.

Knowledge has a state.

That is our shorthand, not Keith’s terminology or a theory we are attributing to him.

An observation has conditions. A historical claim rests on a particular surviving record. A prediction belongs to the information available before an outcome. A dataset, article or analysis can acquire versions, lose context, move platforms or become detached from its maker.

If we keep only the answer, those conditions can disappear.

And when they disappear, the answer can start to look more certain, more complete or more timeless than it ever was.

The practical consequence is therefore not simply to collect more evidence.

It is to keep visible enough of the conditions under which something became knowable that another person can still judge what it means.

First, work out what you are actually looking at

In January 2019, Keith tried an R package for visualising missing data.

His Asian Cup dataset appeared to contain missing values in the fields for red cards and second yellow cards. But Keith knew the coding rule he had used. NA did not necessarily mean that an observation was unknown or had been lost. In those fields, it could mean that the card had not been awarded.

The display was accurate. A careless interpretation of the display would not have been.

This is a small example, but it changes the question a reader should ask.

Before asking, “What does this evidence show?”, ask:

What kind of evidence is this?

A live human observation is not the same object as a television replay. A replay is not the same as a remembered event. A hand-notated dataset is not the same as an official web dataset. A blank field may not mean the same thing as an unknown value. Two visits to the same official website may not encounter the same numbers if the source is being updated.

What Can an Observer Know? follows those distinctions further. Different forms of observation carry different limits: mediation, operational definition, attention, observer drift, memory, recording quality, coding semantics and mutable source states are not interchangeable problems.

That means “get more data” cannot be a universal remedy.

If the recording is poor, a better recording may help. If the operational definition is vague, the problem is definition. If observer reliability has not been tested, another observer or repeated coding might tell us something useful. If NA means “not awarded”, collecting another value would misunderstand the data structure.

The first transferable learning is therefore simple:

Do not begin by asking how much evidence you have.

Begin by establishing what kind of thing you are looking at, how it was produced, and where its uncertainty lives.

Go straight to Clyde Street: Trying visdat.

Second, do not turn the edge of your evidence into the edge of reality

The same discipline becomes historical when the object in front of us is the surviving past.

Keith repeatedly searched for beginnings in performance analysis. The searches mattered to him. He followed references into libraries, contacted people, returned to old papers and made recovered material visible.

But Clyde Street also records occasions when an apparent beginning moved.

His 2015 post on slow-motion film is especially useful. Keith had once regarded a 1939 paper by Roy Priebe and William Burton as the earliest published example he knew of slow-motion pictures being used as a coaching device. He was also aware of a 1935 thesis that might move the date earlier.

Then the post acquired a postscript.

Keith found more material: a 1936 thesis, a 1938 thesis and an example from 1929 in which Earl Hill had used slow-motion film with aeronautical students.

The past had not changed while Keith was writing the page.

The available record had.

That difference is easy to lose when a historical account is later compressed into a neat origin story.

The earliest source we have found is not automatically the first time something happened. Failing to locate a predecessor is not proof that no predecessor existed. A missing archive creates a real evidential limit, but it does not tell us what the missing documents would have shown. A living witness can add detail and correct a representation without becoming a time machine.

A History with Missing Pieces follows those problems in depth.

The transferable learning is not that history is hopelessly subjective. Keith’s searching points in the opposite direction. The missing pieces matter precisely because evidence matters.

The lesson is to preserve the state of the historical evidence.

Sometimes the strongest responsible sentence is not “this was the beginning”.

It is “this is the earliest source I have found”.

That small difference keeps the search open to correction.

It also protects us from a common mistake: mistaking the edge of our search for the edge of the past.

Go straight to Clyde Street: Slow Motion.

Third, keep the claim on the correct side of the event

Predictions reveal another way in which knowledge acquires a state.

Before an event, a forecast is vulnerable. The future can contradict it.

After the event, the result is known. What had been uncertain can start to look obvious. A prediction can be remembered as more confident than it was, or quietly reconstructed around what eventually happened.

Clyde Street contains a useful defence against that hindsight because Keith often wrote before events as well as after them.

The 2017 AFL Grand Final is particularly revealing.

Keith deliberately restricted what he knew. He had recorded quarter-by-quarter scores from the official AFL site, but he had not watched the games. He had avoided newspaper coverage and football programmes. He was asking how far a deliberately narrow indicator might take him.

His median scoring profiles suggested an Adelaide advantage. Before the game, he corrected the data and revised the size of that estimated advantage. He also described conditions under which Richmond might become dangerous.

Then Richmond won by forty-eight points.

The value of the sequence is not that Keith was right. He was not.

Its value is that the before-state survives.

We can see what he used. We can see what he excluded. We can see the correction made while the outcome was still unknown. We can see the expectation. Then we can see the result.

What Happens to a Prediction After the Future Arrives? follows that prospective chronology more fully.

The learning travels beyond sport forecasting.

Whenever we judge a decision, plan, estimate, forecast or proposed intervention after the outcome has become known, we should try to recover the information state that existed beforehand.

What could reasonably be known then?

What had not yet happened?

What was revised while the future was still open?

What only became obvious afterwards?

A correct prediction does not prove an entire model. A failed prediction does not automatically prove the opposite explanation. The result answers the prediction, but it does not explain itself.

The transferable habit is to keep the before-state on the correct side of the event.

Go straight to Clyde Street: Profiling the 2017 #AFLGF Teams.

Fourth, preserve enough of the journey for somebody else to question the destination

In August 2010, Keith was dismantling a conference website.

The IACSS09 site had lived on Ning. When the platform changed, Keith moved what he could: an exported archive, WordPress posts, presentations, proceedings and copies held elsewhere.

Yet some of the environment was already disappearing. The official conference website was unavailable. So was the Twitter hashtag.

The files had travelled.

The whole knowledge environment had not.

This exposes a different problem. A backup can preserve bytes while losing the relationships that made those bytes intelligible.

A later reader may need to know what a file was, where it came from, who created it, whether it was final, what project it belonged to, what version of the data or software produced it, and what conversation once surrounded it.

That is the deeper question in What Has to Travel with Knowledge?

Clyde Street moves through curation, persistent identifiers, reproducible workflows, metadata and mutable sources. Each helps with a different part of the problem.

A DOI can help an object remain findable. It does not make the object good.

A reproducible workflow can make the route to a result inspectable. It does not make the result correct.

Metadata can identify a creator, version, date or relationship. It does not make bad data valid.

An archive can survive and still be incomplete.

The point is not to preserve everything indiscriminately. Nor is it to publish everything. Not everything can be shared publicly.

The learning is narrower and more useful:

If somebody else may inherit the object, preserve enough of its route that they can tell what they have inherited.

Sometimes that means a creator’s name. Sometimes a version. Sometimes a codebook. Sometimes raw data and a script. Sometimes a capture date. Sometimes a note that part of the original environment is already gone.

Go straight to Clyde Street: Migratning IACSS09.

These are not four separate problems

At first, observation, history, prediction and knowledge infrastructure look like four different territories.

But imagine a single chain.

Someone watches a performance.

What they notice becomes notation or data.

Years later, those data become evidence about what happened.

Someone builds an analysis from them.

The analysis contributes to a prediction.

The future arrives.

The result is interpreted.

Later still, another person inherits the dataset, the code, the article, the memory and the conclusion.

At every transition, something can become detached from the conditions that once gave it meaning.

Observation can become detached from the procedure that produced it.

Data can become detached from the coding rule that tells us what a missing value means.

A historical claim can become detached from the gaps in the archive.

A prediction can become detached from what was known before the outcome.

A conclusion can become detached from the version of the source that entered the analysis.

An idea can become detached from its creator.

A file can survive while the environment that made it intelligible disappears.

This is why the four Beyond the Breadcrumbs trails become more powerful when they are read together.

They are all asking, in different ways:

What must remain visible if a later person is to understand not only the answer, but the conditions under which that answer became possible?

“Knowledge has a state” is useful only if it leads to this practical consequence.

States can change.

The source can be updated. The archive can widen. The forecast can be revised. A coding rule can become detached from the data it explains. A platform can disappear. A witness can correct a previous account. A later custodian can add new metadata without becoming the original author.

A useful record does not prevent those movements.

It leaves enough trace that the movements can still be seen.

What this changes in practice

This does not need to become a formal checklist. But the four trails suggest a small set of questions worth carrying into many kinds of inquiry:

1. What am I actually looking at?

2. What state is it in — and when was that state created or captured?

3. What could reasonably be known at that point?

4. What is missing, uncertain, excluded or unavailable?

5. What changed before or after this state?

6. Whose observation, judgement, memory, model or interpretation is this?

Question to carry

7. What needs to survive so that somebody else can question it later?

Those questions do not produce one universal method.

They do something more modest.

They help us match our response to the actual problem.

A poor recording may need a better recording.

A moving source may need a capture rule.

A historical gap may need another search rather than a stronger conclusion.

A forecast may need its earlier version preserved.

A coded dataset may need its semantics explained.

Someone else’s account may need attribution rather than absorption.

An inherited analysis may need a route back through its construction.

The point is not to make every claim cautious.

It is to make the confidence fit the state of the evidence.

Uncertainty becomes useful when you can locate it

There is an easy way to misunderstand all of this.

If observation is mediated, history incomplete, predictions fallible and archives fragile, perhaps the conclusion is that nothing can really be known.

These returns to Clyde Street do not take us there.

The striking thing across these trails is how practical the responses are.

Define the event.

Record the procedure.

Look for the primary source.

Say when the source is missing.

Keep the earlier forecast.

Correct the data before the event and show that you corrected it.

Name the person whose account you are carrying.

Keep the code.

Add the metadata.

Declare the capture rule.

Those are not gestures of resignation.

They are ways of making knowledge more examinable.

The uncertainty becomes useful when it is located.

“We cannot know” is usually too broad.

“We cannot establish this division of labour from the surviving authorship record” is informative.

“The video is not clear enough to classify this detail confidently” is informative.

“This is the earliest source found so far” is informative.

“This forecast was revised before the game” is informative.

“These official figures were captured twenty-four hours after the match” is informative.

Each sentence tells the next person where the boundary is.

That is very different from hiding uncertainty behind a polished endpoint.

What to carry out of Clyde Street

The most important learning from these four trails may therefore be a change in what we think an answer is.

An answer is not always a detachable endpoint.

Sometimes part of its meaning lives in the route that produced it.

The observation matters, but so do the conditions of observing.

The historical claim matters, but so does the state of the archive.

The prediction matters, but so does what could be known before the result.

The dataset matters, but so do its definitions, versions and provenance.

This does not mean every reader needs every scrap of process.

It means that when the process changes what the answer can legitimately mean, the process is no longer disposable background.

That is the handshake between Clyde Street and the reader’s own practice.

You may never analyse a Rugby World Cup, reconstruct the beginnings of performance analysis, forecast an AFL Grand Final or migrate a conference website.

But outside Clyde Street the same kinds of problems recur: claims whose histories have been compressed, inherited files whose makers are no longer present, charts without their codebooks, decisions judged after their outcomes are known, confident accounts built from incomplete archives, data whose source has changed, and work that somebody else may have to understand after you have moved on.

The four deeper Beyond the Breadcrumbs trails remain open: What Can an Observer Know? for the different conditions of observation; A History with Missing Pieces for the limits of reconstructed pasts; What Happens to a Prediction After the Future Arrives? for forecast accountability across time; and What Has to Travel with Knowledge? for the custodial infrastructure that lets an object remain intelligible.

They lead, in turn, back into Clyde Street.

But there is one lesson worth carrying before taking any of those paths.

Do not just keep the answer.

Keep enough of the journey to know what the answer can mean.

Sources followed

Keith Lyons, “Rwc2019 Patterns After 29 Games” — 19/10/09

Keith Lyons, “Trying Visdat” — 19/01/27

Keith Lyons, “Slow Motion” — 15/09/04

Keith Lyons, “Profiling The 2017 Aflgf Teams” — 17/09/29

Keith Lyons, “Migratning IACSS09” — 10/08/17