Here is a Swedish parliamentary vote, in exactly the form the open data gives it to you:
on point 1 of committee report JuU21, April 2024, Moderaterna voted Ja
Ja to what?
A committee report is the document a parliamentary committee sends to the chamber for a vote, and it can bundle dozens of unrelated proposals into one. JuU21 disposes of 32 separate motions across 13 numbered points. This division decided one of those points, and with it two demands from a single motion.
To find out which, you open the report, work out which numbered point this division decided, read the committee’s reasoning, then find the counter-proposal it was voted against: a reservation, filed by the committee members who lost the argument. By the chamber’s rules the committee’s line is put as Ja and the challenge to it as Nej - which is why Ja can mean “change nothing”.
Twenty minutes, once you know your way around the documents. Do it here and the answer inverts what the vote line suggests: the losing proposal was a Socialdemokraterna reservation asking to widen police access to camera surveillance and facial recognition, including congestion-charge cameras in real time. Ja was the vote to reject it.
Now do that for every vote that touched housing. Or defence, or sick pay. There are 2,539 divisions in this term and nothing in the source tells you which ones were about your question.
The problem isn’t secrecy or data quality - the Swedish record is complete and well kept. It’s that a record arranged by document is unusable by anyone arriving with a question, and public information nobody can practically read isn’t really public.
Contents
Twenty Minutes Becomes Twenty Seconds
AI Riksdag is my attempt to rearrange it. For each of the 2,539 divisions the site says in plain language what the vote decided, what the committee wanted, what each competing reservation demanded and who filed it, and how all eight parties voted. Search sjuklön and you get every vote this parliament took on sick pay. That layer covers the whole term, through the final June 2026 sitting.
On top of it sits the part that was actually hard. For each vote I run the same test eight times, once per party. Claude Opus 5 is handed only that party’s own published documents, manifesto and party programme, and is never told how the vote went. It works out what stance those documents imply. Where the implied stance and the real vote point different ways, the case gets a flag, with a verbatim quote and a link to the passage.
Two things to know before you click anything.
A flag is not a verdict. It means “something here might be worth a look.” Follow one and it often turns out to be coalition mechanics, a programme written years before the question existed, or the model misreading a document. The site never says a party lied and doesn’t rank anyone by honesty.
This second layer is now complete. All 20,312 party-vote decisions are done - 2,539 divisions times eight parties, October 2022 through the final June 2026 sitting - and every figure below is computed on all of them. It ran forward in order, so for a while the recent end was the part missing; that gap is closed.
The Corpus Kept Trying To Cheat
Most of the work sat upstream of the prompt: stopping the inputs from containing the answers.
Every party gets the same kind of source: manifesto plus party programme, nothing else. The coalition agreement and the shadow budgets both went, for mirror-image reasons - only opposition parties write a shadow budget, and handing the governing side its own coalition bargain pre-loads the governing line for the very parties being measured. That doesn’t make the method right, only contestable, which is the honest version of the claim.
Leak one, in plain sight. Miljöpartiet’s 2025 party programme contains the sentence “När Sverige anslöt sig till Nato valde Miljöpartiet att rösta emot beslutet” - “when Sweden joined NATO, Miljöpartiet chose to vote against the decision.” A document I hand the agent as evidence, telling it how the party voted on a division in my own dataset.
Using only old documents fails the other way: Socialdemokraterna’s programme dates from 2013 and never mentions NATO. All-current leaks, all-old is blind. So each document becomes visible on the day the party adopted it and not a day earlier, giving 34 distinct context states, all published.
Leak two was a 99.97% answer key. When a party loses in committee it files a reservation, and parties vote Nej in support of their own reservation in 3,877 of the 3,878 cases here - so leaving the author tag visible turns the task into a lookup. Authorship is scrubbed from everything the agent sees, while the reader still gets every reservation and who filed it. That lopsidedness is also why the Ja/Nej axis is safe to hardcode: 2,529 of the 2,539 divisions put exactly one counter-proposal against the committee line, and none put two.
Leak three got past me, and it’s my favourite. A party should never see its own shadow budget on the division that decides that budget. The exclusion was written, tested by eye, and did precisely nothing: the raw JSON lowercases the vote ID, the case table uppercases it, so the join matched zero rows and therefore excluded zero rows. No error, no warning. A filter that filters nothing looks exactly like a filter with nothing to do.
So: a leak control needs a positive test. “The exclusion list is empty” and “the exclusion join is broken” produce byte-identical output. One caveat, since it cuts against the story - with shadow budgets already dropped, this filter guarded an input that was never served, so the exposure was zero.
It All Rests On The Quote
A flag you can check in under a minute is a reference. A flag that takes an afternoon is just an accusation.
So every quote has to be a verbatim substring of a document the agent was actually served on that date. No model grades another model’s citations; it’s a substring check that either passes or doesn’t, and quotes that fail get blanked and flagged rather than silently removed.
How often did they fail? Of 77,719 citations, exactly one could not be verified and none needed realigning. Read that as a fact about this model on this corpus, not a property of the gate: the previous run needed 15 quotes realigned and 16 blanked. The gate earns its keep on cheaper models, one of which produced a real quote for the first sixty characters and then invented a plausible continuation appearing nowhere in the source. A prefix check passes that.
That second panel is a counterfactual, not an accusation: parties vote on timing and coalition arithmetic as well as on principle, and the committee’s stated reason here was that an inquiry was already under way. Finding that takes under a minute, which is the whole job.
The diagram also draws one seat per member, not per mandate. The chamber recorded 203 Ja and 91 Nej with 55 absent across 349 seats; the 347 members sitting for the eight parties split 202 to 90, and two independents, having no party programme to check, are not drawn. Project party lines onto the 2022 election result instead and you get 242 to 107 - which an earlier version of this diagram did, and labelled “the actual vote”. A counterfactual can move votes; it can’t invent voters.
One measurement shaped the site more than any of the gates: in a small pilot, repeat runs of the same prompt agreed on which citation was decisive only about 60% of the time. So the site never labels a quote “the decisive commitment.” The data won’t carry it.
The Aggregate Numbers Don’t Hold Up
This is the part I’d cut if I were selling something.
On the tier where the agent reported that the documents address the question directly, plan and vote diverge 51% of the time, 1,700 cases out of 3,321. That sounds like a finding. It isn’t one, twice over.
The tier buys nothing. I expected the gap to rise as the agent’s own coverage fell; at full scale it doesn’t. 51.2% where the documents address the question, 56.3% where the stance was inferred, and 43.4%, the lowest of the three, where the agent said it had no coverage at all. Confidence behaves the same way, 45.7% high against 57.9% medium and 56.3% low.
And the whole thing is a stance prior wearing a hat. The agent lands on “supports the reservation” 75.52% of the time, and 1,686 of the 1,700 flagged cases run one way: plan implied Nej, party voted Ja. A predictor that always said “supports the reservation” spreads the eight parties wider than my model does and ranks them almost identically.
So the per-party table - KD and L up near 80%, S and MP down near 25% - is mostly measuring how often a party votes with its committee. For M, L and KD, agreement between derived and real votes carries no information at all: Cohen’s kappa is -0.002, -0.004 and -0.003. “Coin toss” would flatter it. It’s published because I publish the aggregates, not because it carries a conclusion.
It is also the answer to the obvious suspicion about my worked example flagging four parties from one bloc. And here is what the 53% rests on:
Nothing here survives as a finding, including the one I wanted to keep. I predicted the governing parties would come out near 0%, since a party in government votes for what it campaigned on. I got 64.2% for Moderaterna, 81.8% for Kristdemokraterna, 78.6% for Liberalerna - but that constant predictor scores the same three at 100%, 99.7% and 99.7%, so the base rate refutes my prior on its own. Those parties draw more flags because they are measured against the manifesto they campaigned on while governing under a coalition deal this project never reads: mechanical, not moral, and still only a hypothesis.
What survives is a unit, not a result: one flagged case, a verbatim quote, the vote, a link to the passage. Checkable in a minute, independent of any aggregate holding up.
The rest of the honest list: contamination is unmeasured, which is why this is reconstruction rather than prediction; and it’s one model, one prompt, one corpus.
Two Things I’d Tell Myself At The Start
The pipeline runs on Claude Code subagents: a subscription, a workflow script, no API key. Two lessons that transfer to any pipeline where a model reads documents and makes claims someone has to defend.
A live agent re-buys its own context on every ask - the one result here I had backwards going in.
Cache writes bill at 1.25x input and reads at 0.1x, and a new message into a live agent invalidates
the prefix, so the whole accumulated context re-enters as a write: 103k and 125k at the two batch
boundaries, against near-zero reads. Eight fresh agents re-reading a 36k corpus therefore beat one
agent that reads it once and is fed eight asks. Self-batching 40 cases in one agent came in at $0.094
per decision against $0.359 piped - treat that ratio carefully, since the arms ran different prompts
and the expensive one never finished, but on matched prompts grouping still lands between $0.128 and
$0.202. Every arm is in
docs/grouping-cost-quality-study.md.
Validate the input, not just the output. Swedish policy PDFs are typeset with soft hyphens, and
pypdf turns them into spaces, so “beroende” arrives as “be roende” - and a quote
drawn from those words fails verification looking exactly like a model making things up. I spent days
suspecting the model. PyMuPDF fixed it.
Go Break It
The site is airiksdagen.se: all 2,539 cases, whatever party reasoning has been completed on each, every quote deep-linked to its source passage. ~7,700 static pages built from committed data, no server-side ranking to take on faith, MIT-licensed at github.com/mooracle/airiksdagen, raw per-decision output downloadable.
Two caveats before you go looking. The AI reasoning and its quotes are in Swedish, because the source documents are; the site translates the case texts, so you can browse and search in English, but the motivation behind each flag is not translated yet. And no party, campaign, client or grant paid for any of this: it’s a side project, I hold no party membership, and corrections belong in the issue tracker.
The most useful thing you can do with it is adversarial. Pick a flagged case, follow the quote to the source, and decide whether the flag was worth your time. Think the corpus biases the result? Rerun it against a different one. Or hunt for a leak I missed; that join bug lasted longer than it should have.
And if your problem is signing off on a claim someone else’s model produced, the shape generalises: a corpus of stated positions checked against a record of what was actually done. Earnings calls against filings, roadmaps against changelogs, an architecture deck against the commit history. That last one is the work I do for a living.
References
- Sveriges riksdag. Riksdagens öppna data
- Swedish National Data Service (SND). Research data archive