Method · measured state
What this tracker measures, and where it is wrong
The goal is not to build a system that reads the news well. It is to make every factual claim this tracker publishes carry a measured error rate — so a reader can tell the difference between “the government said this” and “we verified this”, and so can we.
That is a narrower and less comfortable goal than “add AI”. It means every stage needs a labelled sample, an agreement figure, and a known failure mode. Where a measurement has contradicted an earlier claim — including claims made on this page — the correction is stated rather than quietly dropped.
01
What we are actually trying to do
Three claims the site makes, and what backs each one today.
1. “This reform is at stage X.”
Currently produced by asking a model to infer delivery from news tone. It agrees with the human editor on 14 of 38 items (37%), and its confidence does not track its accuracy (45% agreement at ≥0.8 confidence versus 28% below, against a 25% chance rate). This claim is not yet backed. Meanwhile 35 of 39commitments with a road‑map deadline are past it with no completion recorded — arithmetic that needs no model at all.
2. “This article relates to item N.”
Measured on 283 hand-labelled pairs across 45 articles: 70.7% precision, 76.3% recall. Roughly one match in three is wrong. The dominant failure is a dozen “umbrella” items whose titles admit almost any policy story.
3. “Coverage of this item is positive/negative.”
Measured on 200 labelled rows: the stored corpus reads +0.235 more positive than the labels, with κ 0.369. A revised rubric and model cut the bias to +0.128 and lifted agreement to κ 0.41— better, still only “fair” to “moderate”.
None of these three is currently good enough to be presented as settled fact. That is the honest status, and it is why this page exists.
02
The scorecard
Every number below comes from a labelled sample committed to the repository, and is reproducible by re-running the harness.
| Measure | Stored corpus | Current pipeline |
|---|---|---|
| News-to-item precision | 42.2% | 70.7% |
| News-to-item recall | —% | 76.3% |
| Sentiment κ vs labels | 0.369 | 0.41 |
| Sentiment accuracy, re-weighted | 63.3% | 76.5% |
| Sentiment score MAE | 0.368 | 0.197 |
| Sentiment bias (signed) | +0.235 | +0.128 |
| Delivery status vs editor | 37.0% | 37.0% |
“Stored corpus” is the 11,917matches accumulated since April 2026 under earlier prompt revisions. “Current” is the present pipeline measured on the same labelled rows. Samples: 200 sentiment rows, 283 match pairs.
03
Corrections this work produced
Measurements that overturned earlier statements, including ones made here. Stated because a method page that only reports its successes is not a method page.
- “Over half of all item matches are wrong” — withdrawn. The first instrument validated only one item per article, so a pair matched to a different item counted as an error even when it was a genuine second match. Pair-level labelling showed true precision is 70.7%, not 42.2%. The defect was overstated roughly twofold.
- “The sentiment rubric halved match recall” — withdrawn. That comparison used the stored corpus as a baseline, and that corpus spans several earlier and looser prompt revisions. Re-measured against the prompt’s own output, the difference was noise.
- “Retrieval-first is the fix” — withdrawn. Embedding retrieval was proposed as stage 1. Measured, cosine top-5 found 11 of 38 valid pairs (29%) against the matcher’s 32 (84%). Short abstract item titles embed badly against concrete news. Retrieval must supplement the matcher, never replace it.
- “The tracker shows 62 items not started” — fixed. Those 62 had no assessment recorded at all. Conflating “not looked at” with “not done” was the most consequential error found on the site; it is corrected and “not started” is now a recordable state.
- Thinking mode silently disabled our temperature setting. DeepSeek enables reasoning by default and its docs confirm
temperaturethen has no effect at all. Thinking is now explicitly disabled. This entry previously ended “the pipeline is now deterministic at temperature 0” — that claim was wrong and is withdrawn. See the next item. - “Temperature 0 makes the pipeline deterministic” — withdrawn, and this is the most important correction on this page. Five identical runs of the shipping configuration — same code, same input, temperature 0, thinking explicitly disabled — produced F1 between 72.5 and 76.5, a spread of 4 points, with true positives ranging 29–31. The instrument is not reproducible, and no earlier round knew that.The consequence is that every comparison on this page that rests on a single run is weaker than it was presented as being. Differences smaller than about 4 F1 points are not distinguishable from run-to-run variation. The large swings survive — a 20-point precision shift is far outside this floor — but the moderate ones, including the headline 73.4F1 and the “none beat the baseline” conclusion drawn from it, need repeats before they can carry a verdict. Reproduce with
npx tsx --env-file=.env scripts/scope-eval.mts --repeat=5.
04
What is actually blocking accuracy
Four separate attempts to raise matching accuracy all failed the same way. Understanding why changed what we are doing.
| Attempt | Precision | Recall | F1 |
|---|---|---|---|
| Plain prompt (shipped) | 70.7% | 76.3% | 73.4 |
| Prompt hardened with anti-patterns | 84.6% | 57.9% | 68.8 |
| Curated per-item scope rules in the prompt | 90.5% | 50.0% | 64.5 |
| Scope rules as a separate verifier stage | 95.0% | 50.0% | 65.5 |
Every precision intervention cost about as much recall as it bought, and none beat the plain prompt. The precision/recall frontier is real, and prompt engineering does not move it.
The scope experiment isolated the mechanism. With scope rules in the prompt, recall split like this: items with rules kept 54% of their correct matches, but items without rules fell to 40% — and those items had zero false positives to begin with. The loss was not strictness on the targets; it was collateral damage from a longer prompt making the model globally more conservative.
The real blocker is that 16 commitments have never been operationally defined.#49 is not “any infrastructure” — its text names the NPC project pipeline, the review of stalled projects, and the Fast-Track Mechanism. #85 is not “any health story” — it names free hospital beds, the Free Health Portal and pharmacy affordability. Until somebody states which projects and instruments count, nobody — human or model — can apply these items consistently. We have been building machinery to work around an undefined input.
05
The decision — made, and what it changed
This was the blocker. It is now settled, and settling it changed what the site can honestly measure.
A decision sheet covered all 16umbrella items. Each showed the commitment’s own approved text, a reading derived strictly from that text, the concrete false positives the matcher actually produced, and any pair where the proposal and the existing labels disagreed. On 2026-09-26 every item was decided: 15 confirmed as proposed, 1 amended.
The one amendment, #90, mattered. The proposal read “storage”. The operator tightened it to cold storage and added an instrument the proposal had missed entirely — the soil health card for commercial farms within three months, which is named in the commitment text. It also kept the explicit exclusion of daily commodity price bulletins, so “fair prices” resolves to the minimum support price mechanism and not to commodity price reporting.
That is the shape of every one of these: the title admits a subject area, the approved text names instruments, and only the operator can say which governs. No model could have chosen the soil health card. It is not a harder inference; it is not an inference at all.
The decisions are now project state, not a browser artefact. They live in data/item-decisions.json, are carried with provenance in src/lib/item-definitions.ts — so downstream code can tell a definition the operator confirmed from one a machine proposed — and the sheet opens showing them.Review or revise the item decisions →
What it unblocked. The per-pair scope verifier had been written, measured and deliberately left switched off, because it failed on a definitional mismatch rather than a bug: it demanded the item’s named instrument while the tracker implicitly wanted any bearing on the commitment. The operator has now resolved that mismatch in favour of the named instrument, so the rubric is finally the product’s rubric and can be measured fairly.
Measured on the 70 scoped pairs (13 valid, 57 invalid), it does exactly what it was built to do and the trade is brutal: 11 false positives suppressed, 11 true positives lost. Precision on scoped items goes to 100%, F1 collapses to 14.3. Meanwhile no unscoped item was touched — gating the rules per item, instead of putting them in the matching prompt, is what removed the collateral damage measured earlier.
Read the lost true positives literally and the verifier looks broken. Read them against the operator’s own wording and most are correct rejections of stale labels: the pair labelled valid for #49 is the Beni–Jomsom–Korala road, which the confirmed definition excludes; the two for #5 are a caste-housing story and a sexual-violence story, both of which its confirmed exclusion names outright. The gold set was labelled under the looser reading and has not been re-adjudicated since.
Round two: 10 disputed pairs were adjudicated. 7 came back definition too narrow, 3 stale label, and 0 verifier errors. So the verifier was not misapplying the rules — the rules were under-specified. 5 items were amended and 3 labels corrected, and the same instrument was re-run.
Audit result — worse than the question that was asked
| judged | precision | 95% interval | |
|---|---|---|---|
| the 16 defined items | 19/30 | 63.3% | 45.5–78.1% |
| the other 84 items | 9/30 | 30% | 16.7–47.9% |
| everything sampled | 28/60 | 46.7% | 34.6–59.1% |
The 95% the gold pool claimed for the undefined items is refuted outright — the interval does not come near it. But two things must be said plainly. The sixteen treated items are only 63.3% precise, so verifying their definitions improved them without finishing the job. And the two intervals overlap, so this sample cannot show the defined items are better than the undefined ones: a two-proportion power calculation puts the requirement near 35 judged matches per group and this is 30. An earlier eight-match read of the kept matches, which I reported as clean, was luck — at the measured rate, eight straight hits is under 3%.
Extrapolated, roughly 4,353 of 6,456 matches are wrong: the news evidence behind this tracker is about a third accurate. Confidence is weakly informative at the bottom and uninformative above it — matches scored 0.6 are 7% precise, 0.7 55%, 0.8 62% — so a 0.7 floor would drop 1,381 matches and lift precision only to about 59%. A threshold is a mitigation, not a fix.
What the sample does support is the cause. The wrong matches are overwhelmingly undefined commitments attracting any story that touches their subject — flood relief matching a Gen-Z business package, a tourist trail matching “strategic new areas”, a border tax exemption matching import/export transport. That is the same defect the sixteen were fixed for, and the honest next step is a top-up of roughly 35 per group to settle the comparison before acting on it.
Scoped items, before and after the definitions were fixed
| precision | recall | F1 | |
|---|---|---|---|
| matching prompt alone | 33.3% | 80.0% | 68.4 |
| verifier, before | 100% | 7.7% | 14.3 |
| verifier, after | 100% | 80% | 88.9 |
Across the whole gold set the verifier now reaches F1 87.5 at 96.6% precision against the plain prompt’s 68.4 — the first configuration in this project to beat the baseline on F1 rather than trade precision for recall. The trade is 16 false positives suppressed per true positive lost, against 1-for-1 before.
The residual single loss is not a definitional failure: that pair was proposed by the matcher only in that one run. Measured over three identical runs the verifier’s own spread is 1.8 F1 points — narrower than the baseline’s 4 — so the gap is far outside run-to-run variation and further tuning of the definitions would be fitting noise.
Then applied to the live corpus — and the live data was worse than any sample
The verifier now runs on the sixteen scoped items only. Re-judging the 5,988 matches those items carried, it rejected 5,490 and kept 498 — 92% of the news attached to these commitments was not about them. Match rows fell from 11,946 to 6,456, of which the 5,958 on the other 84 items were untouched.5,615 auto-generated comments describing the rejected matches were removed with them, and 4 commitments now carry no coverage at all.
This is a bigger correction than the gold set predicted, and the reason is worth stating: the sampled pool under-represented how bad production was. On the pairs the verifier was actually shown it rejected 16 of 16 invalid and kept 8 of 8 valid — a clean separator — and the audit of its live decisions agreed, so that sample was measuring the instrument correctly and still understated the defect in the corpus.
Read the live corpus audit → A deterministic stratified sample of the matches in the database, thirty from each population, every one judged by hand.
The adjudication record → The set is narrowed to the commitments that actually carry a definition — an earlier version included a pair the verifier never sees, which had entered only because the matching stage is not reproducible run to run. Each is shown with the article itself, a link to the original, the approved commitment text, the definition it failed, and the boundary cases either side, so it can be traced to a stale label, a verifier error, or a definition that is too narrow. Also written to docs/scope-recheck.md.
Settling a definition therefore does not unblock evidence-typed status grading on its own — it unblocks measurement, and the first thing measurement revealed is that the instrument itself was never trustworthy. That is section 03.
06
The pipeline as it now stands
Revised by measurement. Each stage states what it does and how it is currently judged.
Ingest and identify
Fetch 33 of the 60 registered outlets, resolve each through one shared register, detect language, deduplicate. Stores the publisher teaser — averaging 330 characters, which is the real ceiling on every later stage.
Judged by · Feed liveness; share of articles with usable text.
Match articles to items
One model call per batch of 8, plain prompt, temperature 0, thinking disabled. Deliberately unembellished: every attempt to add rules cost recall elsewhere.
Judged by · Measured 70.7% precision / 76.3% recall on 283 labelled pairs.
Score sentiment toward the matched item
A separate call — because coupling it to matching was measured to move match counts. The model reports valence, intensity and ambivalence; the label is derived from them by a pure function, so incoherent label/score pairs cannot occur.
Judged by · κ 0.41, bias +0.128, on 200 labelled rows.
Deadline arithmetic
Whether a roadmap deadline has passed is not a judgement call. 35 of 39 scheduled commitments are past deadline with no completion recorded.
Judged by · Unit-tested boundaries; free and deterministic.
Evidence typing (not yet built)
Grade delivery from typed evidence — legal instrument, official report, budget release, operational artefact, news claim only — instead of from news mood. Blocked on §05, because "which instrument counts" is the same undefined question.
Judged by · Target: beat the current 37% agreement with the editor.
Calibration and abstention (not yet built)
Publish "insufficient evidence" instead of guessing. Needs calibration first, and a human-labelled gold set large enough to fit it on.
Judged by · Target: expected calibration error, risk–coverage curve.
07
How accuracy is judged
Two labelled samples, both committed and both reproducible. Any method — old prompt, new rubric, different provider — is graded on identical rows.
- 200 sentiment rows stratified across classes and languages, with 50 unmatched articles to test the matching decision. Rare classes over-sampled so “mixed” is measurable; scores are re-weighted to corpus priors and say so.
- 283 match pairs across 45 articles, every candidate labelled individually. The candidate pool unions the matcher, the stored corpus and embedding retrieval, so recall is measured against something wider than the matcher’s own output — with the pool’s coverage limitation stated in the output.
Two limits, stated before any result. First, these labels were produced by the same class of tool the pipeline uses, so they are a weak reference: enough to compare methods and catch gross errors, not enough to claim human-level agreement — every κ here is an upper bound. Second, the samples are small. Repeat runs of identical code vary by 5–8 points, and at these sizes the harness cannot resolve differences below roughly ±0.15 in κ. Results smaller than that are labelled indicative, not real.
08
Cost
Measured token counts. Nepali runs 3.81 characters per token against 4.56 for English, so a corpus that is 79% Nepali costs more than an English one of the same size.
| Job | Cost, off-peak cached |
|---|---|
| Re-score sentiment on existing matches | ≈ $0.73 |
| Full re-classification of the corpus | ≈ $3.30 |
| Backfill the 35,177 missing embeddings | ≈ $0.14 |
| Status suggestions for all items | < $0.01 |
The conclusion these numbers force: money is not the constraint on accuracy here. The entire corpus can be re-scored for the price of a meal. What is scarce is validated ground truth and a settled definition of the commitments.
Two levers matter more than model choice. Tariff timing: DeepSeek charges double during 01:00–04:00 and 06:00–10:00 UTC, and the daily job was moved out of both windows. Caching: the item list and the rubrics are byte-identical on every call and placed first so they cache, at roughly a thirty-first of a fresh token.
09
What this method will not claim
- It will not present a stage assessment as verified while it agrees with a human only 37% of the time.
- It will not claim state-of-the-art sentiment for Nepali: published Nepali sentiment work reports roughly 74–80% accuracy on social-media data, which is not news, and no off-the-shelf model is better than a calibrated pipeline here.
- It will not treat a model’s confidence as meaningful until calibration is measured — the current system’s confidence already fails that test.
- It will not read a commitment’s title as its meaning. Where the text names instruments, those instruments are the test.
- It will not report “no evidence” as “not delivered”. 61 items carry no roadmap deadline and are reported as unscheduled, never as on track.
- It will not hide a correction. Where a measurement overturned an earlier claim, §03 says so.
10
Configuration in force
Read from the running environment, so this page cannot describe a pipeline that is not deployed.
- Chat provider
- openai
- Model
- gpt-4o-mini
- Embeddings
- text-embedding-3-small@openai
- Feeds / register
- 33 / 60
- Items with scope rules
- 16 (drafted, not wired in)
- Decision mode
- temperature 0, thinking disabled
Milestone days wired: 7, 15, 30, 45, 60, 90, 100. Labels and harnesses are committed, so every figure on this page can be reproduced by re-running them.