🎧 Prefer to listen?

The United States Congress ran the world’s least scientific but most honest AI tool comparison, and the receipts are public. Per CNBC’s analysis of House disbursement records (TechCrunch’s writeup here), House offices spent about $113,740 on AI tools in the year ending March 31 — and roughly 90% of it went to ChatGPT. Anthropic’s Claude came in a distant second at about 7%. And before you file this under “news about other people’s budgets”: the same blind spots, defaults, and unexamined habits that produced that 90/7 split are almost certainly running your tool stack right now. That’s the part worth your time.

This is also the follow-up I promised myself after writing about ChatGPT’s autonomous work agent and the trust gap between AI lab announcements and results — two stories about how AI tools behave. This one is about how they get chosen, which turns out to be a very different mechanism. And unlike most of my compliance posts (say, the EU AI Act labeling rules), this one has no rules to learn. Just a decision to make better.

What the data actually says

The details matter more than the headline. Congress put real money — a modest $113K, but real — into AI tools, and the distribution tells a story:

  • ChatGPT: $100,580 across 798 transactions. That’s many small purchases — individual offices, committees, and institutional accounts buying subscriptions at every level. It’s not one big procurement; it’s a hundred small “sure, this is the default” decisions.
  • Claude: $13,160 across 37 transactions. Fewer, larger purchases — which reads like specific offices with specific needs choosing it deliberately, not by default.
  • Democrats outspent Republicans roughly 3:1 ($54,165 vs $15,782) — a reminder that even tool adoption tracks the culture of the people adopting.

What staffers actually use these tools for: memos, summarizing legislation, responding to constituents, hearing prep, policy research, and drafting social posts. In other words: generalist office work with high stakes and tight review. (Also worth noting: the data excludes free accounts and AI bundled into bigger software contracts, so the real adoption is higher than the receipts show.)

Why the 90/7 split doesn’t mean ChatGPT is “better”

Here’s the trap in reading this as a verdict. Institutional spending measures default-ness, not quality. Congress did not run a bake-off. Individual offices with no procurement staff picked the tool they’d heard of most, that their colleague already had, with the simplest approval story. That’s how defaults propagate everywhere — in the Capitol and in your business.

Notice what the pattern hides: whether Claude beats ChatGPT on the tasks those staffers do (legislation summaries, constituent letters) is nowhere in this data. Nobody measured outcomes. Money spent is a record of decisions, and most decisions weren’t really made.

There’s a version of this data that would tell you something: if the offices that compared tools and then switched had been tracked separately. They weren’t. Which brings us to the useful part.

What I did: my 30-day AI tool bake-off (steps to copy)

Last month I ran the comparison I’m about to describe on my own stack — three AI assistants, my actual workload, four weeks. I’m not going to pretend it was scientific: I tested each tool on the same three recurring tasks, kept a one-line note per run, and let the boring verdicts accumulate. What I learned is that the tool I thought was my favorite won exactly one of the three tasks. The bake-off exists precisely because your default is wrong somewhere and you can’t feel which one from the inside.

Here’s the process at solo scale — the same method that produced Congress’s data, minus the committee:

1. Pick the comparison honestly. Two or three tools that could plausibly run your actual workload — not the ones trending on your feed. If your work is writing-heavy, that might be ChatGPT vs Claude vs Gemini. If it’s automation, compare the automation platforms directly, not their marketing pages.

2. Split a small trial budget. Whatever a month of two subscriptions costs — treat it as tuition, not expense. Congress’s 798 transactions were mostly trial-and-adopt; give yourself permission to not renew.

3. Grade against real work, not demos. This is the step everyone skips, and the reason defaults win. Take three recurring tasks from your actual week — the report, the research pass, the inbox triage — and run them through each tool. Score the outputs blind if you can. Demos are theater; your Tuesday workload is the exam.

4. Write down the switch. Congress has no record of why each office picked what it picked. You can have one: a three-line note — what you compared, what won, why. Six months from now, that note is the difference between a deliberate stack and an accidental one.

5. Re-run it twice a year. The tool that was right in January won’t be right forever — models improve unevenly, prices move, and new categories appear (I tracked this shift in what actually works among AI productivity tools). A standing calendar reminder is the whole maintenance cost.

The quiet lesson in the partisan split

The 3:1 Democrat/Republican spending gap is getting the least attention and might be the most instructive number in the whole story. It shows that tool adoption follows culture — who you talk to, what your feed shows you, which colleague demos something on a Tuesday — far more than it follows benchmarks.

For a solo builder, the corrective is deliberate exposure: when you do your mini-procurement, at least one of the tools you test should be outside your usual circle’s favorite. That’s the only way the comparison stays honest. Otherwise you’re not comparing tools; you’re confirming your feed.

What to do with this

The House of Representatives, with all its resources, ran its AI adoption mostly on defaults. You can out-execute it in one month with a spreadsheet and three subscriptions. Start the trial with your two most-hated current tasks and see which tool actually survives contact with real work — and if you’re starting from nothing at all, your first automation in 15 minutes is the lowest-stakes place to begin building the evaluation habit.

The tools aren’t the moat. The choosing process is.

New here? The start here guide is the beginner path, and the AI Tool Advisor gives you a structured first pass across tools before you spend a cent.