Independent researcher on AI governance and the economics of the AI build-out.
Papers and interactive tools on how the AI build-out is financed, powered and regulated. Each study is built from the primary record and archived with a DOI.
CV / one-pager (updating) ORCID 0009-0003-4213-7769 GitHub Zenodo Hugging Face Email
Methods and tools SEC EDGAR / XBRL · PJM and FERC queue data · model evaluation · motive-tiering · Python stdlib, reproducible build.py pipelines
Do open weights let you audit a model, or only say what you ran? Tested on numeric extraction from real SEC filings against each filer's own XBRL tags.
Open the tool →Three operational details of a published data-quality filter, measured across 84 million documents and reproducible from a seeded sample.
Open the tool →Granted and realised pay rank separately, beside the board's own payout against target.
Open the tool →Self-contained, dependency-free front-ends over the frozen, citable datasets. Each is reproducible; the ones built on a frozen dataset are archived with their own DOI.
What twelve consumer AI products publicly evidenced against the EU AI Act's Article 50 transparency clauses, four days after the duty applied. Each cell carries the artefact, access date and capture state behind it, and the one interpretive rule can be reapplied to see which cells move. Not a compliance assessment.
Open the tool →How much announced generation actually gets built. A deflation calculator runs any announced capacity through three measured PJM-queue attrition stages. About one announced MW in five is delivered.
Open the tool →Which water figures are measured, which are planning assumptions, and what kind of water is involved. Reporting years and metrics are selectable, every site carries its evidence status, and unknown water sources stay visible rather than inferred.
Open the tool →Whether the AI buildout's unit economics are holding. Per-edition signal and developments, or one indicator tracked across issues. The signal it watches, a hyperscaler capex-guidance down-revision, has not fired in six issues.
Open the tool →How the field forecasts data-centre electricity demand. Dispersion is measured only within comparable slices, since units and scopes are not interchangeable, and each forecast is traced through a transparency funnel from verified to confirmable to reproducible.
Open the tool →What the S&P 500's highest-paid CEOs took versus what they delivered, across 495 issuers in a dated 503-security universe. Granted and realised pay rank separately, and the board's own payout against target sits beside peer-relative return.
Open the tool →The uncertainty signal an open-weight deployment gets for free, tested as an error detector on numeric extraction from SEC filings. The false-alarm budget is adjustable, and the tool reports what it catches, what it misses, and how many correct answers must be re-read to find one real error. It stays at zero on every model accurate enough to deploy.
Open the tool →The filter behind a widely used training corpus keeps documents scoring 3 or more, and that 3 is a rounded integer. The raw cut is adjustable across the 233 scores the classifier can represent, the rounding rule can be switched to move 458,461 documents across the boundary, and the rounding identity is shown in the page.
Open the tool →Also live, in beta and outside the set above because it runs on a moving feed rather than a frozen dataset: the AI News Board, which assesses AI news coverage by the publisher's incentive in the claim and by whether quoted figures state the base they are measured against (code & data).
Working papers on the finances, energy and risk of the AI build-out, and on how digital and retail systems shape spending and choice. Built from primary public data, with a reproduction script wherever the question is quantitative, and published open access on Zenodo under CC BY 4.0. The underlying datasets are also on Hugging Face.
An open, reproducible referee for data-centre water burden. It rebuilds the Water Consumption Impact index for ten sites from public data and self-checks each value against the source, then adds the two channels that framework leaves open: the hydropower coupling, and the off-site relocation of water that closed-loop cooling produces, which moves the great majority of the footprint to the grid rather than removing it (roughly 92 to 95 per cent on the modelled inputs; the direction holds across every plausible variation of them, the precise band does not). Three of the ten seed sites, across three different operators, turn out to be non-operational planning figures presented as measured. Re-basing the index on the operators' FY2025 reports, holding capacity and peaking constant, moves five of the six primary-verified sites up by 15.6 to 33.3 per cent in a single reporting year. Figures are bands and decompositions, and every row is graded by the standing of its source.
interactive tool · 10.5281/zenodo.21318960 · code & dataA reproducible audit of how the field forecasts data-centre electricity demand: dispersion, revisions and transparency.
10.5281/zenodo.20572928 · code & dataHow much announced generation gets built, a realized-completion anchor from the PJM interconnection queue.
10.5281/zenodo.20559430 · code & dataThe UK counterpart to the work above. Britain's electricity demand connection queue grew from 41 GW to 125 GW in seven months, and roughly 73 GW of it is data centres. How much of that carries a final investment decision is smaller, self-reported, and known only as a band, because the financed count and the current denominator come from different populations five months apart. Meanwhile the announced figure is being used to allocate Green Belt land, and it underpinned a planning approval the government has since conceded was made in error.
Read the article → article, no DOI · written into the Ofgem consultation window, which closes 16 September 2026A transparent order-of-magnitude bound on the electricity used by automated web traffic, and the proof-of-method for the deflation discipline. A naive multiply of the two most-circulated inputs gives about 257 TWh a year; two largely separate methods bracket the real figure at order 10 TWh, and then disagree with each other by a factor of five on the central value. That disagreement is the result, because bridging traffic measurements to energy measurements takes three choices and each one moves the answer by a multiple.
10.5281/zenodo.20512703 · code & dataA supply-and-demand risk assessment of the AI build-out: a likelihood-by-severity register of where the system is fragile.
10.5281/zenodo.20586863 · code & dataA pre-registered, reproducible record of how private AI valuations survive contact with a real public price.
10.5281/zenodo.20672581A reproducible, descriptive audit setting the S&P 500's highest-paid CEOs' realized pay beside what they delivered.
10.5281/zenodo.20680108 · code & dataTracking the unit economics and cost re-pricing of AI: the gap between cheaper tokens and cheaper AI.
10.5281/zenodo.20541643 · code & dataA structured scan of the political, economic, social, technological, legal and environmental pressures on the AI industry.
10.5281/zenodo.20680575A reproducible ledger of how the AI data-centre build-out reaches a single household (device prices, AI-laden subscriptions, electricity, public subsidy and retirement exposure), with each channel's AI-attributable share bounded, the contested energy row given both sides, and water deliberately held outside this household ledger and given its own dedicated tracker instead. Two archetype households; lead with the low bound.
10.5281/zenodo.20941003How individual AI dependency is built, priced, and made hard to leave: a reproducible profile of the commercial machinery (pricing tiers, caps, metering, portability) across US, EU, Chinese and companion providers, with a revealed-elasticity test of which squeezes actually retain users. The consumer leg of the reliance trilogy.
10.5281/zenodo.20807278A reproducible measure of systematic AI dependency across four layers (reliability, downstream attribution, penetration, reversibility), now carrying a dated off-switch register (the June export-control suspension of two Anthropic models, OpenAI's government-gated GPT-5.6, and the July Hugging Face breach) and a demand-side cohort layer. Built from provider status-page JSON and primary government data.
10.5281/zenodo.20763480 · code & dataAn analysis of agent-mediated commerce: where the hype stands against what is delivered, why the “neutral agent” is a false assumption, who controls the resulting chokepoint, and which defences are actually enforceable.
10.5281/zenodo.20790896An evidence review of how physical and online retail environments are engineered to increase spending, why individual vigilance is an inadequate defence, and what works in its place.
10.5281/zenodo.20788844A reproducible map of where AI-fluency becomes part of the job: occupation-level AI-usage penetration crossed with the augment/automate mix, from the open Anthropic Economic Index.
10.5281/zenodo.20778968Three operational details in a published data-quality filter, measured across 84 million documents from 94 Common Crawl snapshots: the documented threshold of 3 is applied to a rounded integer, so the real admission boundary is a raw score above 2.5 and most of the retained corpus was let in by rounding; the classifier reads at most 510 tokens, so it never saw most of what it admitted, a blind spot that widens as web pages lengthen; and it ran in bfloat16, whose grid lands exactly on the tie point, cutting 458,461 documents by rounding parity. None of the three is an error in the filter, and all three are visible only because the scores were published.
interactive tool · 10.5281/zenodo.21740081 · code & data · datasetOpen model weights are advocated on two grounds that get conflated: that they let you say what you ran, and that they let you check whether what you ran was right. This paper separates the two and tests the second on a task where the errors are silent, numeric extraction from real SEC compensation tables scored against each filer’s own XBRL tags. Token-level entropy sits at chance exactly where the errors concentrate and has no threshold that carries between models, and a deterministic source-membership check catches none of the 106 misattributions, because a misattributed figure has its digits in the source. The two interventions that do raise accuracy, model capacity and majority voting over presentations, both run against a closed API. Provenance survives, in a weaker form than usually assumed.
10.5281/zenodo.21543579 · code & dataA motive-neutral referee for the “AI writes most of our code” claims. It grades eight vendor and adopter statements on four evidentiary tests: whether a denominator is stated, whether the figure is audited, whether it is reproducible, and whether a second source confirms it. It keeps a separate dated register of the labs' own research-loop red-lines, reported rather than scored.
10.5281/zenodo.21325372A dated, reproducible snapshot of what twelve consumer AI products publicly evidenced against the EU AI Act's Article 50 transparency clauses on 5 and 6 August 2026, four days after the duty applied. Eleven items tied to clause numbers, every cell carrying a link to an opened artefact and an access date, the interaction disclosure found to vary with account state, and a baseline built for a re-run after the 2 December marking deadline. Not a compliance assessment; the movement between the two runs is the intended result.
interactive tool · 10.5281/zenodo.21819102 · code & dataA published negative result in cosmology, kept on the record because reporting what did not work is part of the discipline.
10.5281/zenodo.20374960