Operations case study

Sovereign Watch

Operating an AI-assisted monitoring pipeline.

An independent workflow that turns source documents into labelled alerts for human inspection. The work covers operating requirements, AI-assisted implementation, failure handling and release checks.

Role: project owner and operator, with AI-assisted implementation.

Setting: an independent regulatory-monitoring service supporting a public research newsboard.

01The operating problem

A scheduled job could appear successful while sources were unavailable, assessments had failed or document evidence was missing. Failed items could lose their retry opportunity. Later checks also found slow model requests and an empty article capture being treated as though it contained no figures.

The operational requirement was to show which work had actually been completed, preserve unfinished work and give a reviewer enough evidence to inspect an alert.

02Your route through the workflow

  1. 01Source intake
  2. 02Capture evidence
  3. 03Model assessment
  4. 04Save labelled alerts

The automated route publishes labelled candidates for human inspection. A manual-review state is outstanding work, not a record of completed review. This overview omits conditional branches; the detailed diagram and source trace below retain their scope.

Explore the Sovereign Watch workflow: follow source intake, automated assessment and evidence capture through to labelled candidates for human inspection. The recovery view shows how failed and over-limit inputs are handled. On a phone, scroll the diagram horizontally.

Diagram source trace · Editable source and reproduction instructions

03Ownership and delivery

The project owner commissioned the diagnostic and repair, set the operating constraints and approved the releases. AI coding assistants performed implementation and engineering checks. Separate-model review challenged the later release package. The work did not establish sole manual coding authorship or deployment within an employer's organisation.

The service now retains source status, pending work and manual-review states; captures document text and supporting quotations; saves alerts before advancing checkpoints; and coordinates writers through a shared lock. Empty captures remain unassessed. Model requests have bounded input, output and time allowances. These controls make failures visible and support recovery; they do not certify a legal interpretation.

04Evidence of delivery

The later release passed 138 recorded tests. Four deliberately reintroduced empty-capture defects were detected, N = 4 selected regression cases. The published page and the 24-file change set matched the approved release bytes after deployment.

A bounded diagnostic used one cached public document in three requests. The original request timed out; two revised requests returned schema-valid, source-quoted responses. Their priorities differed. This demonstrates a repaired execution path on that document, not stable classification or a measured improvement across a representative workload.

The released repair retains incomplete-coverage reporting. The public board is the output; later scheduled updates may change it.

05Limits and next decision

The latest local checkpoint inspected for this case, recorded on 22 September 2026, remains incomplete. It records pending and manual-review work, deferred documents and an FTC access failure. This is a stored checkpoint, not a fresh endpoint or classification audit.

The next operating decision is how to handle blocked sources and the manual-review queue within an approved resource budget. A completed process is not automatically completed surveillance. No comparable end-to-end reliability, classification-quality or time-saving baseline was run, and no workplace adoption claim is made.

06Verification and AI assistance

The figures and release claims were checked against retained test, diagnostic and publication receipts. Public source and tests are linked above. OpenAI Codex assisted with implementation and this case study; Gemini supplied separate static review of the later repair package. The reviewer did not rerun those tests. OpenAI and Google are also subjects of the wider board. Review and a supporting quotation do not eliminate model errors or establish legal correctness.