Essay · For engineers

About Is Not Evidence

Why AI systems confuse topical relevance with warrant: a distinction human languages independently built grammar to enforce.

≈ 7 min read  ·  Brent A. Seeley

A conspiracy theory article about vaccines is intensely about vaccines.

It may be the most about-vaccines document you will read all year. Every paragraph concerns vaccines. Its keywords are vaccine keywords. Its embedding, in any modern vector database, sits squarely in the vaccine neighborhood.

None of that makes it evidence about vaccines.

A paper refuting a claim is about that claim. A history of a false belief is about the belief’s subject matter. A fabricated diary can be about Napoleon. Aboutness is cheap. Warrant, the property of actually giving you a reason to believe something, is expensive. And most of the AI infrastructure built in the last few years quietly buys the cheap one and spends it as if it were the expensive one.

That gap has a structure, and once you see it, you see it everywhere.

The valve and the manual

Start with something small.

The valve is closed.

Now wrap it:

The manual says the valve is closed.

The second sentence has not stopped being about the valve. If someone asks whether anything in the room bears on the valve’s position, this sentence is obviously relevant. The valve is right there, visible inside it.

But the second sentence is true or false because of a fact about the manual. The valve could be wide open and the sentence would still be true, provided the manual says otherwise. The subject matter stayed put; the thing that makes the sentence true moved.

Wrap it differently and the same thing happens:

Alice believes the valve is closed.
The sensor reports the valve is closed.
The model predicts the valve is closed.

Every version keeps the valve as its topic. Every version relocates its truth conditions somewhere else: to a mind, an instrument, a statistical artifact. “About the valve” is the one property shared by all five sentences (the bare claim and all four wrappers alike), and it is precisely the property that tells you nothing about whether the valve is closed.

This is the about/assert gap: a claim can remain fully, visibly about a subject while asserting nothing about it directly. Humans navigate this gap constantly and mostly well. We know the difference between “the report says X” and “X.”

Mostly.

Retrieval buys aboutness

Now look at how a retrieval-augmented AI system works, because the architecture is disarmingly simple.

A user asks a question. The system converts the question into a vector and searches a database for passages whose vectors sit nearby. Nearness in that space means, roughly, this passage concerns the same subject matter as the question. The passages are handed to a language model. The model writes an answer, drawing on them as its source material.

Notice what happened at the handoff.

The retrieval step found documents that are about the question. The generation step used them as evidence for an answer. Nobody decided to treat topical relevance as warrant. No line of code says “assume relevant means reliable.” The conversion happens by default, in the gap between two components, because nothing in the pipeline represents the difference.

Embedding similarity is an aboutness detector. It is a rather good one. But the thing it measures, topical similarity, simply doesn’t track warrant. The conspiracy article and the systematic review live in the same region of vector space, because they are about the same thing. That is what the space measures. Asking it to distinguish them is asking a topic model to do epistemology.

This is not a bug in any particular product. It is a category error poured into infrastructure, and it explains a family of failures that otherwise look unrelated: the chatbot that cites a paper criticizing a method as support for the method. The summary that reports a debunked claim in the system’s own voice because the debunking article discussed the claim at length. The assistant that treats “many sources mention X” as “X is well-supported,” when the sources mention X in order to trace, question, or ridicule it.

The field is not blind to this. Serious pipelines now bolt on rerankers, stance classifiers, entailment filters, and citation-faithfulness checks; these help, genuinely. But notice what they are: patches applied downstream of the point where the distinction was lost, each one nudging a relevance score by a signal correlated with warrant. Nowhere in the stack is there a place to write down “this passage discusses X in order to refute it” and have that fact govern what the passage may be used for. The patches compensate for a representation the architecture still doesn’t hold.

In every case the failure has the same shape: the system had material about the subject and behaved as if it had warranted evidence for a conclusion. About is not evidence. The pipeline never learned the difference because the pipeline has nowhere to keep it.

Languages built grammar for this

Here is the part I find genuinely startling, and it comes out of cross-linguistic research I have been doing on how languages encode truth-related meaning.

Human languages did not leave the about/assert gap to good judgment. They built grammar for it.

In Quechua, you cannot state a proposition without marking how you stand behind it. The language has obligatory evidential suffixes: one marker for I directly witnessed this, another for I was told this, a third for I infer or conjecture this. A Quechua speaker relaying secondhand news is grammatically prevented from asserting it as if firsthand. The sentence keeps its subject matter (it is still about the harvest, the neighbor, the valve), but the warrant is stamped on it, in the morphology, every single time. Omitting the marker is not casual; it is ungrammatical.

That is source-marking: where does this claim come from, and how far does that get you? Languages also built a second, distinct piece of equipment, policing a different substitution: resemblance-marking. Across tradition after tradition we find the same coinage: a word for truth, compounded with a marker for likeness, producing a term that means looks true without thereby being true. Latin: vērisimilis. Greek: alēthophanēs, truth-appearing. Sanskrit: -ābhāsa, the mere semblance of a thing. Mandarin: 逼真, pressing-on-real. Tagalog: parang totoo, like true. Quechua, again: cheqa hina. (Some of these languages are distant cousins: Latin, Greek, and Sanskrit share an ancestor. But the compounds themselves are independent formations, not shared inheritance. Six separate traditions each built the word from local parts.)

These are two different disciplines, and it matters that they’re different. Was reported and resembles the real thing are distinct ways a claim can fall short of the thing itself; hearsay is not forgery. What they share is the underlying rule both kinds of marking enforce: neither reportedness nor resemblance may be silently cashed in as the thing itself. Languages found that rule worth violating often enough to build separate equipment against each violation.

Languages are ledgers of what went wrong often enough to be worth preventing. Speakers in tradition after tradition needed to talk about a subject while suspending commitment to claims about it, and needed it badly enough, for long enough, that the distinction fossilized into grammar and word-formation alike.

Now the uncomfortable observation. Large language models are trained on the output of these languages. The marking is in the data: every “allegedly,” every “according to,” every reported-speech construction. But nothing in the architecture enforces what the marking means. A retrieval pipeline is, in effect, a communication system with the evidentials stripped out: everything arrives marked only as relevant, which is to say, marked only as about. We have built information infrastructure that is grammatically poorer than Quechua.

Topical continuity is the camouflage

There is a deeper reason to care, and it comes from looking at the classical fallacies, the catalog of ways arguments go wrong that logicians have been compiling since Aristotle.

Run down the list and a large family stands out. Call them the substitution fallacies. In each of them, the subject holds still while something else moves. The ad hominem stays on topic (it is still about the claim) while quietly swapping the speaker’s character in for the claim’s truth. The appeal to authority stays on topic while substituting someone said it for it is so. The appeal to popularity: still about the claim; the warrant swapped for a headcount. Post hoc reasoning: still about the two events; mere sequence swapped in for causation.

Equivocation and its cousins run the same trick in reverse: they appear to stay on topic while the actual subject drifts under a stable word.

Not every fallacy fits this mold: false dilemmas and level-confusions are different animals. But across the substitution family, the largest and most rhetorically effective one, topical continuity is doing the same job. It is the camouflage. An argument that visibly changed its subject would fool no one; what makes these fallacies persuasive is that the argument keeps feeling like a conversation about the same thing while the relation to that thing is silently replaced. For this family, aboutness is not incidental to the bad reasoning. It is the delivery mechanism.

Which reframes the retrieval problem rather sharply. A system whose only organizing principle is topical continuity, whose entire notion of relevance is “concerns the same subject,” has built the camouflage layer into its foundations and forgotten to build anything underneath it. It is structurally exposed to precisely the substitution family, at industrial scale, not because it reasons badly, but because the distinction those fallacies exploit is the distinction its architecture cannot represent.

The fix, in one sentence

None of this means embedding search is bad. Aboutness detection is exactly what retrieval should do: finding material that concerns the question is a real and necessary function.

The fix is a demotion:

Relevance should nominate candidates. It should never confer weight.

A passage’s similarity to the question earns it a place in the room. Nothing more. Whether it then counts as evidence, and for what, and how strongly, has to be decided by something that actually tracks the questions similarity ignores: Who asserts this? Firsthand, reported, or conjectured? Does it support the claim, merely discuss it, or exist to refute it? These are the distinctions Quechua marks in three suffixes, and they are checkable, but only by a component that represents them, at the boundary where retrieved material becomes believed material.

Some of us are working on architectures that enforce exactly this: treating “about,” “asserts,” and “supports” as different relations that a system must never silently interconvert. But the principle matters more than any implementation.

Language after language learned to keep two ledgers: what the talk is about, and what the talk establishes. The manual can say the valve is closed all day long.

Someone still has to check the valve.


The research program behind this piece