← NM AI Research · research portfolio

Open-Weight Error-Detection Audit

Explore the recorded top-20 token entropy of numeric answers extracted from a small water table. Select a threshold convention and decoding regime to inspect detected misattributions and false alarms. Thresholds are fitted and evaluated on these same stored answers.

Correction notice. The original paper and archived bundle contain claims corrected in the correction record. This explorer uses the first, water-table substrate. The SEC compensation experiment is a separate substrate. The corrected paper is supplied as version 1.1; the original version remains preserved.

Interactive front-end over the frozen v1.0 dataset (Section 5) · canonical record: DOI 10.5281/zenodo.21543579 · code & data · NM AI Research, ORCID 0009-0003-4213-7769 · CC BY 4.0

5.0%

Per model, at your operating point

The threshold is fitted per model on that model’s own correct answers, which is the most favourable case for the detector: it is allowed to calibrate on the answer key. Every rate is shown as a count over its denominator.

modelaccuracymisattributionsthreshold H caughtcatch ratefalse alarmsAUC

Where the errors actually sit

One row per model and one mark per included call, positioned by entropy on a logarithmic axis. Values below the plotted floor share its left edge. Red marks to the right of the threshold are detected misattributions; blue marks there are false alarms.

correct answer misattribution your threshold: anything to the right is flagged

Does a threshold transfer between models?

Fit the threshold on the model in the row, then apply it to the model in the column. Each cell reports flagged correct answers. The outlined diagonal shows the observed rate on the calibration model. Ties, finite counts and percentile convention can make this differ from the nominal target.

What this shows. Top-20 truncated token entropy is compared with the stored labels for correct and misattributed water-table answers. AUC and accuracy are recomputed for the selected decoding regime. AUC 0.5 is the chance-ranking baseline; small error counts give imprecise estimates, and no-error groups have no defined AUC.

Why the convention toggle is here. Nearest-rank and interpolated percentiles can yield different thresholds and detection counts on finite samples. Each aggregate on this page sums per-model counts at the selected operating point. Changing the convention can change individual model results as well as the aggregate.

What it does not show. This is an in-sample exploration of one measured signal. It does not test all thresholds, establish deployment accuracy or compare open and closed models. Semantic entropy and response-level probability estimators are different methods and were not evaluated here. The separate SEC experiment uses filing-derived targets and a documented source-sign exception; it is not the data displayed on this page.

Denominators. 594 model calls, 9 models over 2 decoding regimes. Counts, not percentages, are the primary display throughout; where an N is small the count makes that visible.

Reproducing this. The page is generated by build.py from rescored.json in the repository. The build checks the frozen input digest and selected reference statistics before writing. The original Zenodo deposit does not contain this explorer build. A successful build does not validate source truth, every paper claim or complete historical replay.

Verification. The displayed rates derive from the frozen stored-answer records and the selected regime. The build checks input identity and selected reference values. This does not certify the historical prompts or labels, external accuracy, or the whole paper.

Conflict of interest. The original implementation was Anthropic-assisted. OpenAI GPT-6 assisted the correction and explorer repair. OpenAI competes with providers discussed. The Google reviewer has a direct Gemma interest. No closed model was tested in the benchmark; separate-model advice does not establish every historical claim.

The paper. This page accompanies the first-substrate entropy results from The Model Is a Dependency: Provenance and Error Detection in Financial Data Extraction. Read the paper and bundle series or the original version 1.0. Read the correction record before reusing its claims.

NM AI Research · ORCID 0009-0003-4213-7769 · CC BY 4.0 · independent research, not investment advice.