Intellectual Practices of Keith Lyons

A chart can show a great deal and still leave the inquiry that produced it almost invisible.

That became a problem for the whole-corpus study of Clyde Street.

Keith Lyons often published data, tables, links, code and technical detail. But visibility alone was not the finding.

What had to be exposed before another reader could inspect how the conclusion or object had been made?

A visible result can still hide the route

On 2 January 2014, Keith published a short post about batting partnerships across fourteen Ashes Tests played since 2010.

The final Test of the 2013–2014 series was about to begin in Sydney. Keith presented the partnership profiles for Australia and England.

The post makes the result visible.

But it says very little about how that result was constructed.

There is no fresh account of the source data, the variables collected, the aggregation process or the logic that produced the profiles. The post sits inside an already established analytical series and mainly presents another output from it.

The whole-corpus study therefore retained this kind of record outside the finding.

Sixteen days earlier, Keith had published another post on the same Ashes partnership inquiry.

The surface resemblance is strong. Both posts contain cricket data. Both belong to the same analytical programme. Both make results public.

But the earlier post exposes much more of the construction.

Keith named ESPNcricinfo’s Partnership Graphs as the source. He specified three variables he compiled: the number of batting partners for each player, the total runs scored in each partnership and the total balls faced. He explained that he collected the information from the partnership graph for each innings of the first three Tests, showed how the tables were assembled and stated why he aggregated the material: he wanted to examine the role of partnerships and the shifting partnership curves across three Ashes series.

The reader can see more than the answer.

The reader can inspect the route by which the answer was produced.

That was the distinction the study needed.

Data visibility is not the same as inquiry inspectability

A post does not qualify simply because it contains a chart.

Nor because it links to a spreadsheet.

Nor because a method is mentioned somewhere in the surrounding material.

The stronger cases expose enough of Keith’s own route, construction, source distance, variables, definitions or evidential handling for another person to examine how the public object or conclusion was made.

That does not require every procedural detail to be recoverable. It does require more than the visibility of an output.

This is why the contrast between the two Ashes posts matters. One gives the reader a profile. The other gives the reader enough of the source and construction to inspect where that profile came from.

The route can be inspectable before the inquiry becomes technical

The finding is easy to mistake for a story about open data or reproducible code. The early Clyde Street evidence prevents that narrowing.

In December 2009, Keith wrote a post built around a twenty-four-hour snapshot of his personal learning environment on Twitter.

He did not merely say that Twitter was useful to him. He reconstructed the trail.

A link from Tom Davenport was followed by Typeboard, then Malinka, Kate Caruthers, Alec Couros and his students, the European Graduate School, Mark Drapeau, Graham Attwell, Howard Rheingold and others. Keith identified where material arrived from, what he followed and the sequence in which the material entered his twenty-four-hour record. He placed that trail inside his continuing questions about vicarious learning, network tuning and reciprocal altruism.

The substantive ideas remain with the people who produced them.

Keith’s contribution was that his own learning route became visible enough to inspect.

The object here is not a dataset. It is a path through a distributed learning environment.

That is important because it shows what the behaviour is really about. Inspectability can concern how somebody encountered and assembled an inquiry, not only how they computed a result.

A method can become inspectable through definitions and notation

By January 2013, another form is visible in Keith’s account of the 1991 Rugby World Cup Final.

He made the transcript of his hand notation available for download and pointed readers to the method. He reported game totals, ball-in-play time, elapsed time and performance ratios. More importantly, he explained how the record was compiled.

An activity cycle is defined. Kicks, passes, lineouts, scrums, penalties and injury stoppages are described in relation to the notation. Keith explained what he attempted to record, how sequences were represented and what some of the marks in the original record meant.

The historical game remains the game. The notation is Keith’s.

What makes the source important for this finding is not the quantity of data. It is that the reader can move from the public conclusions back towards the record, definitions and observational choices from which they were made.

The analysis is therefore contestable in a more specific way.

A reader can ask whether the definitions were adequate, whether the notation captured what Keith said it captured, whether a ratio was calculated appropriately or whether an interpretation follows from the underlying record.

Inspectability does not guarantee correctness.

It makes examination possible.

Someone else’s method does not make Keith’s inquiry inspectable

That ownership boundary matters because Clyde Street often points to sophisticated work produced elsewhere.

In June 2014, for example, Keith wrote about World Cup Elo ratings and FiveThirtyEight’s predictions.

He noted that Ritchie King, Allison McCann and Matthew Conlen were using ESPN’s Soccer Power Index, that their model began from 10,000 simulations and that Nate Silver had published a detailed methodological account. Keith compared some of the predictions with Elo rankings.

There is plenty of methodological information in the post.

But most of that methodology belongs to FiveThirtyEight.

Describing another group’s model does not make Keith’s own inquiry construction inspectable. The whole-corpus study therefore retained this record outside the finding.

The distinction is straightforward but consequential.

An external method can be transparent without becoming evidence that Keith exposed his own method.

Code can expose construction rather than merely decorate a result

A 2018 World Cup post shows what a stronger technical case looks like.

Keith was examining player-tracking data shared on the official World Cup website for France and Croatia. He aggregated the tracking data on GitHub and produced R visualisations of the finalists’ distances.

The post includes the source context, aggregate values and thresholds used in the displays. In a postscript Keith stated that he was using R 3.5.1 in RStudio, named the ggplot2 and ggrepel packages, identified the guides he followed and supplied an example of the plotting code.

The code matters because it is connected to the analytical construction.

This is not simply a block of syntax placed beside a result. The reader can see where the data came from, how Keith aggregated them, what environment he used, which packages shaped the representation and an example of how one of the plots was produced.

FIFA retains ownership of the official tracking data. The package authors retain their software. Keith owned the way he compiled and worked with the material.

The source therefore exposes source distance and technical construction without absorbing the work of others.

Tool use sits close to the lower boundary

Not every technical post crossed the threshold so cleanly.

In September 2019, Keith used Rugby World Cup referee data to explore facet_wrap and patchwork in ggplot. He identified the external guides he worked through and presented examples produced with both approaches.

The study retained this at lower confidence rather than placing it in the strongest core.

Why the caution?

Because ordinary tool use can look very much like inspectability.

A reader can see that Keith used the tools and can see where he learned about them, but the source exposes less of a developed inquiry construction than the strongest examples. It sits near the line between showing a technical experiment and making the construction of an analytical object genuinely inspectable.

Keeping that case at lower confidence mattered. Without such controls, any post containing code, software names or a procedural link could have inflated the finding.

Historical verification can be inspectable too

The late corpus broadens the operation again.

In January 2020, Keith published a long account of Neil Lanham’s role in the early use of computers in sport.

Much of the substantive history belongs to Neil. Keith quoted their correspondence, recounted Neil’s work with hand notation and computers, and used documents and published material to place that work in the history of performance analysis.

The important inspectability move appears in how the account was constructed and checked.

Keith made the correspondence visible as evidence. He identified a historical correction concerning Charles Reep’s relationship to managers with whom Neil worked. He said he compiled information about Neil in a Google Doc. In the postscript he explained that he corresponded with Neil throughout preparation of the post to ensure that what he intended to publish had veracity and Neil’s approval.

Neil retains ownership of his testimony, memories and life history.

Keith owned the selection, documentary handling, correction and verification process through which the public account was assembled.

Here inspectability is neither a chart nor code.

It is provenance and verification.

What does not count

The negative cases are part of the finding.

A chart can remain a chart. A table can remain an output. A spreadsheet can remain a file. A repeated tournament update can remain routine reporting. A familiar method can be applied again without becoming newly inspectable. Another person’s transparent model remains theirs. A guest technical analysis remains the guest author’s work. A list of data sources can remain a list.

Across the study, 68 explicitly routed rivals and near-misses were retained specifically to constrain this finding. They are not a denominator of every record that failed to qualify. They are a boundary portfolio showing how often visibility could be mistaken for inspectability.

The recurrent failure mode is simple: the reader can see something, but cannot see enough of Keith’s own construction or evidential route to inspect how it was made.

What the whole corpus established

Across the frozen 2,051-record corpus, 137 records were retained as positive evidence for this operation.

The earliest approved case is from August 2008 and the latest from January 2020. Approved positives occur across forty-three chronological batches.

The evidential core is strong but not uniform. One hundred and fifteen cases form the strongest directly supported core. Eleven primary cases were retained at lower confidence. One historical primary has no separately recoverable strength detail, and ten later secondary co-occurrences have no separate strength assessment for this finding.

That supports a high-confidence claim of recurrence, temporal spread and boundary discrimination.

It does not establish that every transparent-looking Clyde Street post qualifies.

It does not establish that inspectability made the work correct, persuasive or useful.

It does not establish that readers actually reconstructed the analyses.

And it does not establish that Keith was uniquely transparent as a personal characteristic.

The form of the operation changes across the corpus. Early cases often expose routes through people, tools and participation. Later cases include more explicit variables, shared data, code, software environments, documentary provenance and verification. The frozen register permits that within-behaviour description, but not a story of increasing sophistication, importance or inevitable development.

There are also no approved positives after the January 2020 Neil Lanham case in the remaining chronological sequence. The study treats that as descriptive silence, not evidence of decline.

What the distinction leaves with the reader

The easiest version of the lesson would be to say: show your working.

The Clyde Street evidence makes the judgement more precise.

Showing the result is not the same as exposing the route that produced it. Showing code is not enough if the code is disconnected from the claim. Linking to a method is not enough if the method belongs to somebody else. And a completed analysis can still be inspectable if enough of its construction remains visible.

Across the stronger cases, the relation is comparatively clear: the public object or conclusion is identifiable, Keith’s route or construction is identifiable, the originating sources remain properly owned, and another reader can see enough to examine how the account was made.

That leaves a question for the reader — one enabled by the finding rather than taught by Keith as a method:

Question to carry

If someone wanted to inspect how I reached this conclusion or built this object, what could they actually see?

Sources followed

Keith Lyons, “Batting Partnerships After 14 Ashes Tests 2010 2013” — 14/01/02

Keith Lyons, “Partnerships In 2013 2014 Ashes Cricket” — 13/12/17

Keith Lyons, “Vicarious Learning and Reciprocal Altruism” — 09/12/21

Keith Lyons, “1991 Rugby World Cup Final” — 13/01/04

Keith Lyons, “2014 World Cup Elo Ratings 2 June” — 14/06/11

Keith Lyons, “Distances Traversed By The 2018 Worldcup Finalists” — 18/07/13

Keith Lyons, “Referee Appointments At Rwc 2019 Facet And Patchwork Visualisations” — 19/09/20

Keith Lyons, “Neil Lanham Using Computers In Sport” — 20/01/09