The journal

Verify, Don't Generate: Why AI-Written Product Content Is a Liability You're Paying For

AI-written product content asserts facts nobody verified. Why generation is a growing liability for merchants — and why verification is the answer.

A product record with claims and a provenance column — two sourced from admin data, one sourced from nowhere, flagged in red

Lena runs a pet-supply store, and last winter she did what the ads suggested: pointed an AI content tool at her catalog and let it "enrich" four hundred listings. Fill the gaps, complete the attributes, make everything machine-readable. It finished in an hour, the previews read beautifully, and she approved the batch.

Months later a customer email stopped her cold. A chew toy — one her supplier rates for adult dogs — carried a line she'd never written: suitable for puppies from 8 weeks. She searched the batch. The tool had been thorough. Materials nobody had specified ("BPA-free"). Weights nobody had measured. Origins nobody had confirmed. Hundreds of confident, specific, machine-readable claims — with no source behind them, published under her brand, quoted by whoever asked a machine about her products.

The vendor's answer, when she asked how a fact nobody supplied got into her catalog: you approved the batch.

That sentence is the whole industry's business model in five words, and this file is the argument against it. Not against AI — against a specific, identifiable malpractice: letting a generative system originate facts and publish them into records of truth. The alternative has a name, it's the discipline this entire series has been building, and the line between the two is sharp enough to write down.

The seduction, and why it works

The pitch lands because the problem is real. FILE 04 established it: catalogs are full of gaps — missing attributes, empty fields, thin records — and gaps get you skipped by every machine that matters. Meanwhile generation is fast, cheap, and fluent. "AI fills your missing attributes" sounds like the remediation you've been told you need.

But walk it through the fixable framework. A missing attribute is missing because the fact lives outside the data — in the supplier's spec sheet, in the merchant's head, on a scale in the warehouse. That makes it the "only the merchant" class: unresolvable by any engine, because resolving it means knowing it. A tool that fills it anyway hasn't closed the gap. It has replaced an honest absence with a confident fabrication — and an absence, as the legibility file showed, makes machines skip you, while a fabrication makes them trust you wrongly and then find out.

One of those failure modes is recoverable.

The line: language work versus fact origination

Here's the distinction that keeps this from being an anti-AI sermon, because it isn't one — we use language models, this series exists in an ecosystem of them, and they are genuinely extraordinary at what they're actually for.

Language models are masters of language work: taking information that exists and expressing it — rewriting for clarity, restructuring prose into attributes, mapping a supplier's messy spec sheet into schema fields, summarizing, translating, suggesting what a human should confirm. In all of that, the model is transforming facts that were present in its input. The fact's provenance survives the transformation. You can point at where it came from.

Fact origination is the other thing: producing a specific, checkable claim — a material, a dimension, a compatibility, an age rating, an identifier — that appeared nowhere in the input. The model isn't lying, exactly; it's doing what generative systems do, completing patterns plausibly. "BPA-free" is an extremely plausible thing for a pet toy listing to say. Plausible is the product. True was never on the menu.

So the standard, stated the way this series states standards: a tool may express, map, suggest, and summarize — it may never assert an unverified fact into a published record. Generate the sentence; never the fact. And the operational test is one question asked of every claim before it ships: where did this come from? An admin field, a merchant's answer, a supplier's document — fine, generate beautiful language around it. Nowhere — then the tool's only honest output is a question addressed to the human who knows, not an answer addressed to the public.

Two columns — a tool may transform what is present, versus a tool may never originate what is absent, with examples under each FIG.02 — Language work versus fact origination. Generate the sentence; never the fact — provenance has to survive the transformation.

Why this liability is compounding right now

Four forces, all tightening at once, turn yesterday's shrug into tomorrow's incident.

Machines check. The legibility file covered this from the exclusion side: identifiers and attributes get cross-referenced against manufacturer data, feeds, and other retailers. An invented fact isn't a quiet gap-filler — it's a checkable inconsistency, and a record caught wrong on one claim gets its whole testimony discounted. Fabrication doesn't just risk being caught; in a cross-referencing ecosystem, being caught is the expected outcome.

Courts attribute. In Moffatt v. Air Canada (2024), a Canadian tribunal held the airline liable for a policy its website chatbot invented — rejecting, in memorable terms, the argument that the AI was somehow a separate entity responsible for its own words. The published output of your systems is your statement. Lena's "suitable for puppies" isn't the tool's claim in any way that matters; it's her store's, with everything that implies when the claim is safety-adjacent and wrong.

Platforms police scale. Google's spam policies now explicitly target scaled content abuse — mass-produced content generated primarily to manipulate rankings, however it's produced — and its structured-data policies have always required that markup reflect accurate, verifiable information about the actual product. Generated-at-scale plus unverifiable is the exact intersection those policies exist to punish.

Answers amplify. The cruelest force: everything this series taught you about being quotable works on false facts too. AEO doesn't verify your claims; it attributes them. An invented specification, once legible, gets extracted and repeated with your name attached — your catalog's fabrications, distributed by the most trusted answer surfaces your customers use, at exactly the moment the invisible-channel file showed those surfaces deciding purchases.

Four stacked forces — machines check, courts attribute, platforms police scale, answers amplify — each with the consequence that follows FIG.03 — Why the liability compounds now. Four forces tightening at once, and in a cross-referencing ecosystem being caught is the expected outcome.

Verification: the alternative, operationally

"Verify, don't generate" isn't a vibe; it's a pipeline shape. Three properties define it:

Every published claim traces. Each attribute in the record carries an answerable provenance: this admin field, this merchant confirmation, this supplier document, this specification. Not stored ceremony — just the discipline that nothing enters a published record without a source that could be pointed to if a machine, a customer, or a tribunal asked.

Gaps surface; they don't get paved. When the source of truth is silent, the correct output is the "only the merchant" routing from FILE 07: a precise question to the human who knows — what is this toy actually made of? — with the field held empty until answered. Slower than generation. Also the only version that's true.

Language generation runs with facts locked. The legitimate, powerful use of models in this pipeline: hand them the verified facts as fixed inputs and let them do language work — fluent descriptions, restructured attributes, mapped schemas — under the constraint that the facts are variables the model expresses but cannot alter or extend. All of the fluency, none of the origination.

That's the whole architecture of the "verifier, not generator" position, and you can now see it isn't a marketing tagline — it's a provenance policy. It's also, not coincidentally, the pipeline that survives all four compounding forces at once.

The fifteen-minute provenance audit

If AI tooling has ever touched your catalog, run this today.

Pick five AI-touched listings — prioritize anything with materials, dimensions, compatibility claims, age or safety phrasing, or identifiers, in that order of urgency. For every specific claim, ask the one question: where did this come from? Match it against the three acceptable answers — an admin field that predates the tool, something you or your team supplied, a supplier document you hold. Anything that traces to none of them is a fabrication wearing your brand, and the response is quarantine: strip or correct the claim now, ask the human who knows, republish when sourced. Safety-adjacent phrases don't wait for the audit to finish — Lena fixed the puppy line before she finished reading the batch, and that was the right order of operations.

A provenance audit card listing the claim types to check first and the three acceptable answers to where did this come from, with quarantine instructions boxed in red FIG.04 — The provenance audit. One question per claim, three acceptable answers — anything else is a fabrication wearing your brand.

Then adopt the standard forward, in the same spirit as the scope test and the rollback questions: before any tool writes words into your records, ask the vendor where facts come from in their pipeline — and treat "our AI intelligently completes your data" as a confession, not a feature.

That completes the trust standard this block has been assembling: measured honestly, fixable provably, reversible always, scoped minimally, and now — sourced or silent. One piece of technical ground remains under all of it: the markup specifications themselves, where half the advice online is years out of date. What Google actually validates today versus what your theme vendor shipped in 2021 — that's next.

Sources & further reading


Rank Sniper — Field Notes. Catalog legibility and AI-visibility verification for Shopify. We verify what AI agents actually receive from your store; we don't generate content and hope. Verifier, not generator.

Run a free scan Get the next dispatch
Keep reading
◎ Scan Free