The credibility of this tool rests on the depth and provenance of each record, so we measure and publish it rather than asking you to take it on trust. Every figure below is computed live from the 171 policy ideas currently served.
Snapshot generated 2026-09-25. Scores recompute on every request.
Mean record-completeness score
A weighted average across 8 dimensions, weighting substantive descriptions and parliamentary provenance most heavily. It measures completeness of the record, not whether a reform is good policy.
Policy ideas are original synthesis. A usable analysis needs real substance, so we track how many records clear the 500-character bar.
Last checked 2026-09-08 · 760 unique PMG meeting links · 760 OK · 0 broken · 0 unverified.
No dead links found in the last check.
The lowest-completeness records, so gaps are visible rather than buried. This is a worklist, not a disclaimer.
no dated parliamentary meeting linked · no fiscal/economic-impact estimate · no attributed key quote
no dated parliamentary meeting linked · no fiscal/economic-impact estimate · no attributed key quote
no dated parliamentary meeting linked · no fiscal/economic-impact estimate · no attributed key quote
no dated parliamentary meeting linked · no fiscal/economic-impact estimate · no attributed key quote
18 golden questions are put to Ask every week, each answer is checked deterministically (do its citations resolve, do its figures appear in the retrieved data) and then graded by a second, stronger model against the same evidence for faithfulness and completeness (1–5). Scores, and the claims the judge could not find support for, are published here; the answers themselves are never stored. The method is in methodology.
Automated evaluation dated 2026-09-21: 18/18 answers completed; 18 judged. Citation IDs resolving does not establish support for a claim. Judge flags require human adjudication.
Faithfulness
4.5 / 5
18 of 18 answers judged
Completeness
4.1 / 5
11 unsupported claims listed
Citation IDs resolved
100%
18 of 18 cite a tracked source
Figures grounded
89%
of figures in the prose found in the retrieved data
Last run 2026-09-21 on data version 2026-09-20.a0ec9d82; answers by claude-sonnet-4-6 (prompt 2026-09-09.v4); judge claude-opus-5 (rubric 2026-09-10.v3).
Faithfulness over runs
4.5 / 5
How these records are sourced and assessed is documented in the methodology. Policy ideas are original synthesis grounded in parliamentary proceedings — meeting records link back to PMG as the authoritative source.
no dated parliamentary meeting linked · no fiscal/economic-impact estimate · no attributed key quote
no dated parliamentary meeting linked · no fiscal/economic-impact estimate · no attributed key quote
no dated parliamentary meeting linked · no fiscal/economic-impact estimate · no attributed key quote
no dated parliamentary meeting linked · no fiscal/economic-impact estimate · no attributed key quote
no dated parliamentary meeting linked · no fiscal/economic-impact estimate · no attributed key quote
no dated parliamentary meeting linked · no fiscal/economic-impact estimate · no attributed key quote
no dated parliamentary meeting linked · no fiscal/economic-impact estimate · no attributed key quote
no dated parliamentary meeting linked · no fiscal/economic-impact estimate · no attributed key quote
no dated parliamentary meeting linked · no fiscal/economic-impact estimate · no attributed key quote
no dated parliamentary meeting linked · no fiscal/economic-impact estimate · no attributed key quote
no dated parliamentary meeting linked · no fiscal/economic-impact estimate · no attributed key quote
8 judged runs on file
Unsupported claims over runs
11
across 18 golden questions; target ≤ 5
| Run | Prompts | Faithfulness | Completeness | Unsupported claims | Spread | Answer prompt | Answer model |
|---|---|---|---|---|---|---|---|
| 2026-09-21 | 18/18 | 4.5 / 5 | 4.1 / 5 | 11 | — | 2026-09-09.v4 | claude-sonnet-4-6 |
| 2026-09-14 | 18/18 | 4.5 / 5 | 4.0 / 5 | 11 | — | 2026-09-09.v4 | claude-sonnet-4-6 |
| 2026-09-10 | 18/18 | 4.5 / 5 | 4.1 / 5 | 13 | — | 2026-09-09.v4 | claude-sonnet-4-6 |
| 2026-09-09 | 18/18 | 4.4 / 5 | 3.9 / 5 | 19 | ≤1 over 3 | 2026-09-09.v4 | claude-sonnet-4-6 |
| 2026-09-09 | 18/18 | 4.5 / 5 | 4.2 / 5 | 20 | — | 2026-09-09.v4 | claude-sonnet-4-6 |
| 2026-09-09 | 18/18 | 4.5 / 5 | 3.9 / 5 | 20 | — | 2026-09-04.v3 | claude-sonnet-4-6 |
| 2026-09-07 | 16/16 | 4.6 / 5 | 4.3 / 5 | 13 | — | 2026-09-04.v3 | claude-sonnet-4-6 |
| 2026-09-04 | 16/16 | 4.6 / 5 | 4.2 / 5 | 12 | — | 2026-09-04.v3 | claude-sonnet-4-6 |
These are the judging model’s findings against the evidence each answer was shown — a phrase the retrieved records did not support, quoted from the generated answer. They are not corrections to the underlying record, and they are what the next answer-prompt version is written against. The answers themselves are never stored.
eskom-latest — faithfulness 4 / 5
education-reform — faithfulness 4 / 5
eastern-cape-sanitation-audit — faithfulness 4 / 5
municipal-finance-distress — faithfulness 4 / 5
nhi-health-reform — faithfulness 4 / 5
crime-policing-reform — faithfulness 4 / 5
post-office-saa-bailouts — faithfulness 4 / 5
road-accident-fund — faithfulness 3 / 5
Human adjudication (2026-09-25): 11 of 11 flagged claims carried as regression cases have a human verdict — 8 judge findings upheld, 3 overturned because the claim was supported after all. Recorded in data/ask_regression_cases.json.