z-of-a Zone of Avoidance

proverbs textual criticism

Lachmann Had Three

The New Testament survives in about 25,000 manuscripts. Somewhere between three witnesses and several thousand, counting documents stops being the same activity as counting sources.


The diagnostic model
You are seeing
  • A claim appears to have many sources and the sources are rewordings of each other
  • Coverage counts rise while the number of distinct positions does not
  • Deduplication is done by string similarity because nothing else is available
  • Consensus tightens without any new information arriving
The mechanism
Surface variation conceals identity and surface identity conceals derivation, so a count of documents diverges from a count of independent sources — and the gap widens faster than the corpus grows.
The older apparatus
Paremiology, which faced uncountable surface variants and responded by building a type index that assigns one identifier to a pattern rather than to any wording.
The false friend
Genuine independent agreement. Several analysts can reach the same view separately, and the tell is whether the wordings share arbitrary specifics — a rounded figure, an identical framing — that convergence would not produce.
The discriminating test
Collapse the corpus to distinct claims rather than distinct documents, then recount. If the consensus survives the collapse it is real; if it thins to three sources it was always three sources.
On your own data
Take a sentiment or news-count feature and recompute it over type-collapsed claims rather than documents. Compare the two series; the divergence is the feature's overcount.

Lachmann edited Lucretius from three manuscripts.

Three is a workable number. Every shared error among them can be checked by hand, the relationships drawn as a single tree, and the result defended line by line.

The Greek New Testament survives in more than 5,800 manuscripts. Add roughly 10,000 in Latin and about 9,300 in Syriac, Slavic, Ethiopic and Armenian, and the tradition runs to something on the order of 25,000 witnesses.

The method does not scale, and the reason is arithmetic rather than principle. Detecting every shared error requires comparing witnesses against each other, and the number of pairwise comparisons grows far faster than the number of witnesses. The logic that works perfectly at three is still correct at twenty-five thousand and cannot be executed.

That is why large traditions moved to eclectic and coherence-based methods, and why the Dead Sea Scrolls corpus — around 800 manuscripts, of which roughly 220 are Hebrew Bible — sits in a useful middle: enough witnesses to compare, few enough to collate closely.

Somewhere between three and several thousand, counting documents stopped being the same activity as counting evidence.

The problem in its pure form #

Folklorists hit a harder version, because their material has no fixed text at all.

Is an English proverb and its German near-equivalent the same proverb? There is no original to compare them against. There may be no written form older than either. The wordings share no vocabulary. And yet the answer is often obviously yes, in a way that no string comparison recovers.

Matti Kuusi’s response was to stop naming sentences. His international type system assigns a numeric identifier to the underlying pattern rather than to any language’s particular wording, so a single type number covers the English, German and Finnish versions simultaneously — provided a paremiologist judges them to express the same thing. He built it out of his own Proverbia Septentrionalia, 900 Balto-Finnic types with Russian, Baltic, German and Scandinavian parallels, before it generalised.

What that system produces is a name that identifies a cross-language cluster rather than a specific sentence. Ask how many proverbs a corpus contains and you get two numbers: how many wordings, and how many types. They are not close to each other, and only one of them is a count of ideas.

The cost is visible in the definition. Somebody has to judge sameness, and the judgment is recorded as an identifier rather than derived from the text. There is no mechanical route to it, which is why the system is maintained by specialists and not by a similarity threshold.

Both errors, in both directions #

The two disciplines are guarding against opposite failures and need each other’s instrument.

Textual criticism worries that identical wording conceals derivation. Two manuscripts reading the same is not two witnesses agreeing; it may be one witness copied twice. The count of documents overstates the evidence.

Paremiology worries that different wording conceals identity. Two proverbs sharing no words may be one proverb. The count of documents understates the pattern and overstates the diversity.

A corpus of financial claims has both problems simultaneously, and the standard treatment addresses neither. A syndicated wire story republished across forty outlets is one witness appearing forty times. The same view expressed by a strategist in a note, an analyst in a model, and a journalist in a paragraph is one claim in three registers with almost no lexical overlap.

Count documents and you get a number that moves with republication volume and is insensitive to how many distinct positions exist. Feed it into a sentiment measure and the measure tracks the intensity of restatement rather than the distribution of opinion — and those two quantities diverge most sharply exactly when a view is becoming crowded, which is when the measure is being relied on.

What it would take #

The fix is not clever and it is expensive, which is presumably why nobody has done it.

Collapse the corpus to types before counting anything. Not by string similarity, which handles the easy cases and misses both hard ones, but by a judgment about whether two statements express the same claim — recorded, once, as an identifier that subsequent counting uses.

Then recount. If a consensus survives type-collapse, it is a consensus. If it resolves to three sources and their echoes, it was always three sources, and it was Lachmann’s problem all along, wearing a larger number.

The number of pairwise comparisons is the reason nobody does this by hand, and it is also the reason it is now worth doing. Judging whether two statements make the same claim is exactly the operation that has recently become cheap.

Kuusi did nine hundred by himself.

Diagram: Lachmann Had Three

Questions

How many independent sources does a large corpus actually contain?

Fewer than it contains documents, and the ratio is not stable. Textual criticism treats witness count as a parameter rather than a total, because a witness copied from another witness adds a document without adding evidence. The same holds for any corpus where items can be derived from each other.

What is a proverb type number?

Matti Kuusi's international system assigns a numeric identifier to an underlying proverb pattern rather than to any single language's wording, so one number can cover an English, German and Finnish version at once. The name identifies a cross-language cluster rather than a sentence. It is the same logic Thompson's Motif-Index applies to narrative elements.

Why does the stemmatic method stop working at scale?

Not because the logic changes. The number of pairwise comparisons needed to detect every shared error grows far faster than the witness count, so a tradition of three manuscripts can be resolved by hand and one of twenty-five thousand cannot. Large traditions move to eclectic and coherence-based methods for that reason alone.

Can you deduplicate claims by text similarity?

Only the easy cases. Two statements of the same claim in different languages, or in different registers, have no useful string overlap, while two genuinely independent claims can be worded almost identically. Paremiology solved this by having a specialist judge whether two wordings express the same pattern, then recording the judgment as an identifier.