In August 2010, Keith Lyons was dismantling a conference website.
That is a surprisingly good place to begin a story about knowledge.
The site had been built on Ning for IACSS09. Keith had used the platform because he liked the mixture of social tools it offered around an event. Then Ning changed its account structures. Users began migrating. Keith decided to close the IACSS09 site.
So he moved things.
There was an exported archive on Box.net. There were WordPress posts. SlideShare presentations. An Internet Archive copy of the proceedings. Another copy of the proceedings on Box.net.
It sounds reassuring.
The material had been distributed.
But not everything had survived.
Keith recorded that the official conference website was unavailable. So was the Twitter #iacss09 tag.
His conclusion was not that the migration had solved preservation.
It had taught him about the importance of curating ephemeral content.
That distinction opens a larger Clyde Street question.
Knowledge does not become durable merely because somebody uploads another copy.
If an object is to travel across platforms, projects, people and years, what has to travel with it?
A copy is not the whole object
The IACSS09 migration looks at first like a storage problem.
A platform is changing. Material may disappear. Put copies somewhere else.
Keith did that.
But the post also shows why copying is not enough.
A conference is more than a bundle of files. There was an official site. There were social exchanges around a hashtag. There were presentations, proceedings and blog posts. Some of those things could be moved. Some could not. Some survived in different services with different relationships to one another.
The knowledge environment had begun to come apart.
This matters because later readers do not encounter an object in the same conditions as its original participants.
They need routes.
What was this file?
Where did it come from?
What belonged to the official conference and what was Keith’s own surrounding work?
Were the proceedings complete?
Was this the final version?
What conversation once surrounded it?
A backup can keep bytes alive while losing the relationships that made those bytes intelligible.
Clyde Street does not give us a perfect preservation system in 2010. It gives us a live encounter with loss.
Keith has copied what he can and can already see what he cannot carry forward.
That is more useful than a simple claim that “the internet never forgets”.
Sometimes it forgets remarkably quickly.
Curation becomes a custodial problem
Two years later, Keith gave this kind of work a more explicit name.
In ProCurate, he was thinking about curation, aggregation and communities of practice. He described ProCuration as collecting digital information so that other people could find it and develop it.
But the post does not stop at collection.
Keith follows Deanna Dahlsad’s account of curation, which emphasises selection, arrangement and publication but also context, annotation and proper credit.
Those additions are crucial.
A knowledge object needs more than exposure.
It needs enough surrounding information for another person to know what they are looking at, why it was selected and whose work it is.
Then the post takes an unexpected turn.
A tweet leads Keith into writing about digital estates: what happens to online profiles, cloud-stored files and other digital assets after someone dies. Keith follows Paul Wallbank and Erica Swallow into the practical problem of material that remains online while the person who created or managed it is no longer there.
That changes his language.
He begins to see ProCuration as a custos role — a custodial role.
The shift matters.
Curation can sound like taste: finding good things and putting them where others can see them.
Custodianship introduces responsibility across time.
Who keeps the route open?
Who preserves the context?
Who makes ownership visible?
Who notices when a platform disappears or a link decays?
Who knows enough about the object to say what it is without silently becoming its new author?
Clyde Street does not answer all of those questions in ProCurate.
It makes the obligation visible.
A knowledge object can outlive the moment that produced it. Somebody may then have to care for the conditions under which it remains findable and intelligible.
Persistence solves one problem
By 2015, another piece of infrastructure enters Clyde Street: the digital object identifier.
Keith wrote that he had been slow to use DOIs, partly because much of his work lived in public spaces outside formal DOI registration systems. A post by Laurence Horton helped him think about their value.
The attraction is straightforward.
A DOI can provide a stable, persistent and resolvable reference even when the web address or location of an object changes.
That sounds like exactly the solution the IACSS09 migration lacked.
But the same source introduces a vital boundary.
A DOI is not a mark of quality.
An object can be persistently identifiable and still be poor, misleading, incomplete or wrong.
This is one of the most useful distinctions in the whole Wave 4 knowledge trail.
Infrastructure can solve a specific problem without solving knowledge itself.
A persistent identifier helps answer:
Can I still get back to the object?
It does not answer:
Should I trust it?
That second question still requires evidence, method, context, judgement and often another return to source.
Keith also notices work on linked identification and metadata networks. The direction is becoming clearer. An object needs not only a durable address but information that helps other systems and readers understand how it relates to other things.
Persistence keeps the door from moving.
It does not tell you what is in the room.
Reproducibility needs a route, not only a result
By 2017, the knowledge problem becomes more technical.
Keith is following leads shared by Mara Averick when he encounters Joris Muller’s writing about reproducible computational research and, through it, work by Geir Sandve and colleagues.
The principles are practical.
Keep track of how results were produced. Avoid undocumented manual manipulation. Archive software versions. Version-control scripts. Retain intermediate results. Preserve raw data behind plots. Connect claims to the results that support them. Make scripts, runs and results public where possible.
Keith also follows Jenny Bryan’s work on naming files so that they are machine readable, human readable and sortable in useful ways.
Again, most of these prescriptions are not Keith’s inventions.
Clyde Street makes the route of ownership visible: Mara points; Joris interprets; Sandve and colleagues supply the rules; Jenny supplies the naming principles; Keith encounters them and asks what they mean for his own sport-analytics practice.
What makes the post important here is the type of problem it exposes.
A final graph is not enough for somebody trying to reproduce or question an analysis.
They may need the raw data.
They may need the script.
They may need to know which version of the software was used.
They may need intermediate states and naming conventions that let them understand how one file became another.
In other words, the result needs a route back through its own construction.
That route is another thing knowledge may have to carry.
There is also a limit inside the post.
The reproducibility example Keith follows acknowledges that not everything can be shared publicly.
That matters because “carry more context” can easily turn into “publish everything”.
Clyde Street does not justify that leap.
The source itself gives us only the narrower limit: not everything can be shared publicly. The wider questions of privacy, confidentiality, permission and ownership belong to the separate openness-boundary route.
For this one, the point is narrower.
When material can be shared, reusability often depends on preserving more than the polished endpoint.
The journey through the data may be part of the knowledge object.
Metadata is not decoration
In June 2018, Keith described a week of “grazing on the periphery”.
The post moves through a network of openly shared resources in the R community. Mara Averick points Keith towards work by Alison Hill, Matt Dancho, Ulrike Grömping and Guangchuang Yu. Presentations, packages and detailed vignettes make it possible for him to wander into new techniques and ideas.
This is knowledge travelling socially.
But another source in the same post shifts attention to the infrastructure around those resources.
Alice Meadows’s writing leads Keith to Metadata 2020, described as an effort to support richer, connected and reusable open metadata for research outputs.
Metadata can sound peripheral because it is information about the thing rather than the thing itself.
In a long-lived knowledge environment, that distinction breaks down.
A title, creator, date, identifier, version, relationship or descriptive field can help another person discover the object, distinguish it from a neighbour and connect it to the right project.
Metadata cannot make bad data good.
It cannot certify that an analysis is valid.
It can make the object less anonymous to its future.
That is a different kind of value.
Clyde Street’s own structure makes the point visible. Keith’s posts are full of names, links, dates, tags, credits and routes to earlier work. Some of those links later disappear. Some resources remain. The surrounding descriptive information can become part of what lets a later reader reconstruct where an idea came from even when the original path is damaged.
The knowledge is not only the content being carried.
Sometimes it is also the label on the luggage.
The source itself can move
By 2019, Keith encountered a harder version of the same problem.
During the Rugby World Cup, he noticed that figures on the official World Rugby website could change after games had been played, particularly the number of passes. For his own dataset he declared a capture rule: use the values as they stood twenty-four hours after each game, collected in a Google Sheet.
The important point here is not that twenty-four hours was objectively correct. The source gives no basis for that claim.
It is that “World Rugby data” was no longer a complete provenance statement. If the official source had states, the time of capture became part of the route back to what Keith actually analysed.
That different question — what a moving evidence source should require the inquiry itself to change — is followed in What do you do when the evidence is not enough — or changes underneath you? Here the custodial point is narrower: a later reader needs enough temporal provenance to know which state of a mutable source entered the analysis.
They can then question the rule, compare it with another capture point, or reproduce it if those source states remain available.
What survives is not the same as what can be trusted
Across these posts, Clyde Street repeatedly refuses an easy infrastructure story.
Put the conference files in the cloud, and some context still disappears.
Curate openly, and questions of ownership and custody remain.
Give an object a DOI, and its quality is still unresolved.
Make a workflow reproducible, and there may still be legitimate limits on what can be public.
Add metadata, and the underlying data can still be wrong.
Use an official source, and the official numbers can still change.
Each piece of infrastructure solves a different problem.
That is exactly why they should not be collapsed into a single word like “openness” or “sharing”.
The interesting Clyde Street question is not whether Keith liked to share. That is well established elsewhere in the project and is too broad to carry this route.
Nor is the question whether public objects can acquire new lives. Clyde Street contains other work about revising, adapting and keeping public artefacts active.
The question here is custodial.
What information has to remain attached to an object if somebody else is to identify it, locate it, understand its state, question its provenance and reuse it without pretending to know more than the surviving record allows?
Sometimes the answer is a persistent identifier.
Sometimes it is a creator’s name and proper credit.
Sometimes it is context and annotation.
Sometimes it is a version-controlled script and the raw data behind a plot.
Sometimes it is metadata that tells other systems what the object is.
Sometimes it is the date and time at which an official source was captured.
And sometimes the most important information is that part of the original environment has already been lost.
Knowledge needs a carrier
There is a temptation to imagine knowledge as the part that can be detached from all this machinery.
The paper contains the finding. The dataset contains the numbers. The video contains the performance. The proceedings contain the conference. Everything else is packaging.
Clyde Street makes that separation difficult to sustain.
Without a route back, a file can become an orphan.
Without ownership, another person’s work can become detached from its author.
Without version information, two apparently identical analyses may come from different states.
Without context, a preserved object may survive while its meaning thins out.
Without a capture time, a mutable official source can no longer tell us exactly what an analyst saw.
None of these additions guarantees truth.
That is the safeguard running through the whole route.
Provenance is not validity.
Persistence is not quality.
Metadata is not validity.
Reproducibility is not correctness.
An archive is not complete merely because it exists.
But without these things, later inquiry can become unnecessarily blind.
The IACSS09 post begins with a disappearing platform and ends with unavailable material. Nine years later, the Rugby World Cup post responds to a changing official dataset by fixing a declared temporal state.
Between them, Clyde Street passes through curation, custody, persistent identification, reproducible workflow and metadata.
The technologies change.
The underlying responsibility becomes easier to see.
If knowledge is going to travel, somebody has to care about the journey.
Not so that the object arrives untouched.
So that the next person can still tell what arrived.
Question to carry
If this knowledge travels beyond me, what needs to travel with it so the next person can still tell what arrived, where it came from and what may already have changed?
Publication boundary note
This article is about the informational and custodial infrastructure that lets knowledge remain identifiable and interpretable across time. It does not treat availability as use, persistence as quality, metadata as validity, reproducibility as universal publicness or curation as transfer of ownership. External frameworks and recommendations remain owned by their authors. Questions of privacy, consent and disclosure belong to the separate openness-boundary route. Posthumous family or archive stewardship must remain a later custodial layer rather than being represented as continuing Keith-owned activity.
Sources followed
Keith Lyons, “Migratning IACSS09” — 10/08/17
Keith Lyons, “ProCurate” — 12/06/20
Keith Lyons, “Digital Object Identifiers” — 15/05/07
Keith Lyons, “Connecting and Sharing” — 17/08/31
Keith Lyons, “Grazing on the periphery” — 18/06/06
Keith Lyons, “#RWC2019: patterns after 29 games” — 19/10/09
