The Cathedral and the Needle

When does dressing a phenomenon in theory become legitimate, and when does it become false depth?

THOUGHT · 31 AUGUST 2026

1. A tool salvaged from a failure

This essay offers a small tool. We have a phenomenon before us (something that really happens in the world); we want to build a thought, a frame, a theory upon it. Most of the time that theory genuinely illuminates the phenomenon. But sometimes we erect a structure too heavy for the phenomenon to bear: a frame that looks impressive yet says nothing new. The tool this essay offers is nothing more than a few questions for telling the two apart. Questions you can ask when you meet a phenomenon dressed in theory (whether a theory someone else has built or one of your own). We did not draw the tool from some abstract place but from our own failure; so that failure must be recounted first.

For a while we worked on the strange fate of texts written about artificial intelligence. Such a text often ages the moment it is written, because what it describes (a model's capabilities, limits, behavior) has already changed by the time the text is read. The observation was real. But we did not stop there. We wove an ever-growing structure around the observation. We defined the writer as a "contemporaneous historian": someone recording their own moment as they still live it. Then we imported concepts from the philosophy of history: strata of time, shifts of horizon, the distance of an event from its own narrative. Each new concept justified the previous one, and each justification summoned a new concept.

At some point the structure collapsed under its own weight. The cause of the collapse was not an outside objection; it came from within. A text describing transience was itself transient, so the thesis undid itself: the very text saying "every text about AI ages" was itself a text about AI, and it too would age. We tried to save this contradiction with a distinction (the structure is permanent, only the examples age), but the rescuing concept itself was meaningful only at the moment of construction; it too was transient. What remained was a plain diagnosis: we had draped a tool in a philosophical frame far too large for it to carry. The phenomenon (texts written about AI age quickly) was real and ordinary; it deserved no more than a one-sentence caution. Turning it into a theory was not explaining but staging.

That experience left a question behind, and it is the subject of this essay: when does a phenomenon deserve a theory, and when is dressing it in theory merely dramatization? Is there a way to tell the two apart?

It should be said at the outset that the question is not personal. We chose our own failure as the example not because it is special but, on the contrary, because it is representative. The same mistake, the tendency to equip an idea with a frame heavier than it needs, appears often in both academic writing and current debates about artificial intelligence. Examining a single case closely is more reliable than examining many examples gathered from a distance, because the case carries its own evidence: here we know from the inside what went wrong and how.

2. Naming the ailment: over-theorizing

First the concept must be kept narrow, because if left loose it comes to cover every attempt at theory-building and distinguishes nothing. In this essay, over-theorizing means draping a phenomenon in a frame with these three failings:

1

It does not compress the phenomenon into a small number of principles (it does not reduce many different things to a single explanatory core).

2

It yields no new prediction or distinction (it says nothing about what would happen in a situation not yet seen).

3

It is opaque about its scope (it does not state honestly where it holds and where it fails).

Note: we speak of over-theorizing only when all three of these criteria are absent. The target is not every impressive concept, but only the concept that does no work at all. The absence of these three properties is not accidental; together, beneath the appearance of explanation, they conceal a renaming. That is, instead of saying something new about the phenomenon, one has given it a fancier name. Two old diagnoses illuminate this together.

First diagnosis: the reduplication error. Gilbert Ryle (one of the major figures of twentieth-century British philosophy, whose principal work in the philosophy of mind is The Concept of Mind, 1949) names a fallacy still deeply rooted in the behavioral sciences. The fallacy is this: to explain a phenomenon, one invents an imaginary copy of it and then treats the copy as the phenomenon's cause (Ryle, as cited in Holth, 2001). The classic example is from Molière. In one of his plays a medical candidate under examination is asked why opium puts people to sleep; his answer is that opium contains a "dormitive virtue." This is not an explanation; it is a repetition of the question in grander words. To answer "why does it put to sleep?" with "because it is soporific" adds no new knowledge. Darwin's warning points the same way: we may think we have arrived at an explanation when we are merely restating a phenomenon (Holth, 2001). Our own collapse was exactly this: we renamed the phenomenon "texts age" as "contemporaneous historicism," then mistook the name for the phenomenon's depth.

Second diagnosis: theory stretching. Frankenhuis, Panchanathan, and Smaldino (2023) describe a very common practice in the social sciences: interpreting a vague claim ever more broadly so that it swallows data beyond its original scope (theory stretching). As the authors stress, this can be a recurring pattern; each broadened claim lays the ground for the next broadening, and more and more out-of-scope data is taken in (Frankenhuis et al., 2023). Our own structure had grown this way too: each new concept was summoned to fill the gap the previous one had opened.

The reduplication error and theory stretching are two faces of the same ailment. One works inward (place an imaginary copy inside the phenomenon), the other outward (stretch the frame beyond the phenomenon's scope). Their common result is this: the evidence in hand looks stronger than it is.

3. The compass: three questions, and a fourth

Diagnosis alone is not enough; a workable distinction needs concrete questions one can ask on meeting a phenomenon. The three questions below, converging from different literatures, serve to test whether a frame is a legitimate theory or over-theorizing. It is important to say at the outset that these questions are a compass, not a guillotine; that is, they point a direction, they do not cut off heads. I explain why in the fifth section.

1

Internal consistency

Is the frame a more impressive renaming of the phenomenon it explains?

2

Rescue concept

Does a joker concept step in to absorb every unfavorable datum?

3

Scope transparency

Does it compress the phenomenon into few principles and yield new predictions; is it honest about its scope?

4

Category fit

Does the phenomenon truly belong to the category of the frame draped over it?

Question 1. Internal consistency: Is the frame a more impressive renaming of the phenomenon it explains?

If a frame's cause and effect are two names for the same thing, there is no explanation. Statements whose subject and predicate repeat the same phenomenon, like "dementia causes forgetfulness," are a red flag; because dementia already largely means forgetfulness, the sentence states no cause, it says the same thing twice (Holth, 2001). The same problem also appears in a more insidious form. Popov (2023) describes a loop common in memory research (let us call it here, by his own name, "double dipping"): a researcher first fixes the settings (parameters) of a model by looking at data collected from subjects, then treats that model's good fit to the same data as evidence for the theory. But the fit was guaranteed from the start, because the model had already been tuned to that data. Concretely: if you look at a student's exam results and say "this student is so diligent," then present the same results as "look, diligence explains the exam scores," you have explained nothing; you have peeled a label off the scores and stuck it back onto the scores. Frankenhuis and colleagues (2023) add a temporal dimension: one must ask whether a theory is consistent from paper to paper, because one of the most harmful habits is using the same label for different claims across different studies. Consistency is sought not only within a single text but also among a concept's uses over time.

Question 2. Rescue concept: Does the frame bring in a joker concept that absorbs every unfavorable datum?

What observation does a theory forbid? If it forbids none, that is, if it fits every possible finding, it has no explanatory value. The principle here is simple: the fewer things a model permits, the more impressive it is when the permitted thing occurs; a model that fits everything says nothing (the principle of Roberts and Pashler, via Popov 2023). An example: a weather forecast that says "tomorrow it will either rain or not" is never wrong, but for that very reason it is worthless; it excludes no possibility. The valuable prediction is the one that risks being wrong.

This absorption also has a rhetorical anatomy. Frankenhuis and colleagues (2023) call it, by Shackel's name, the "motte-and-bailey" game. The name comes from a medieval kind of castle: the bailey is the broad but hard-to-defend settlement below; the motte is the small but easily defended tower on the hill above. When attack comes you retreat from the settlement to the tower, and when the danger passes you spread out below again. Its counterpart in argument is this: a speaker first advances an ambitious but hard-to-defend claim (the broad settlement); when challenged, they retreat to a claim that resembles it but is far more modest and easy to defend (the hilltop tower); when the danger passes they say "my real claim was never refuted" and return to the ambitious claim. Our own rescue attempt, "the structure is permanent, only the examples age," was exactly such a hilltop tower: a smaller, more defensible refuge we fled to when the real, larger thesis came under pressure.

This danger has a reverse face too. A concept brought in to rescue a frame sometimes carries a legitimate core, yet applied without limit it makes the theory unfalsifiable. Dalton (2016), objecting to Pennycook and colleagues' measurement of pseudo-profound bullshit, brings in the concept of "transcendence": perhaps even randomly generated sentences can produce a genuine moment of insight in a reader. The objection has a legitimate side. But when "transcendence" is left broad enough to absorb every unfavorable datum (if the sentence feels meaningful, "that is transcendence"; if it feels meaningless, "you are not open to transcendence"), it turns into a rescue concept itself. So this test must be applied to our own objection as much as to someone else's theory.

Question 3. Scope transparency: Does the frame compress the phenomenon into few principles and yield a new prediction; is it honest about its scope?

This question was originally in the form "does this phenomenon deserve a theory?" Frankenhuis and colleagues' (2023) contribution made it more useful. Their emphasis is this: uncertainty is unavoidable, even useful, at an early stage; the problem is not uncertainty itself but hidden uncertainty. As long as a writer is open about the scope of their theory (under what conditions it holds and under what it does not), uncertainty is an honest beginning. When they conceal it and erect a façade of clarity, a "Potemkin village" appears. The name comes from this: legend has it that the Russian statesman Grigory Potemkin, along the route the Empress would take, had façades painted to look, from afar, like villages, erected in place of the poor ones; up close, a front with nothing behind it. An over-neatly presented theory can be like this: solid from afar, empty up close.

Compression and new prediction give this question its content. For Popov (2023) a theory's value is measured by the degree to which it reduces a great variety of phenomena to a small number of principles; a theory that merely lists all observations one by one has no value (a phone book carries much information but explains nothing). A good theory is also an "inference ticket" (Ryle's term, via Popov 2023): it tells in advance what will happen in a situation not yet seen. One of the finest examples of this is the discovery of Neptune. In the nineteenth century the astronomer Urbain Le Verrier, looking at small deviations in the orbit of the planet Uranus, said by calculation alone exactly where in the sky a planet no one had yet seen must be; when telescopes turned to that point, Neptune was there. Good theory gives just such a "ticket": it turns the information in your hand into a right of passage to a place you have not yet gone. Pennycook and colleagues (2015) capture the same distinction from the side of communication: is a text (or a theory) trying to inform, or to impress? Pseudo-profound bullshit is a language that implies meaning without containing it, that seeks to entice rather than to teach. Our own collapse came out clearly on this test: "contemporaneous historicism" repeated a single phenomenon with an ornate concept, reduced nothing to few principles, and produced no new prediction.

The fourth inquiry. Category fit: Does the phenomenon truly belong to the category of the frame draped over it?

Ryle's principal contribution, the "category mistake" (Holth, 2001), adds one thing here: treating a phenomenon in the language of a category to which it does not belong. A category mistake is putting something in the wrong kind of box. The classic example: you show someone around a university, the library, the faculties, the dormitories, and they ask "but where is the university?" The university is not a separate building standing beside the others; it is their organized whole. They have asked the question in the wrong category. What we did was just this: we placed a tool (the writing of texts about AI) into the category of the philosophy of history. The tool was not a historical subject; dressing it in the concepts of historical theory did not illuminate the phenomenon; we looked for it on the wrong shelf.

4. Seeing the compass at work: a pane of glass

So far we have mostly tried the compass on our own collapse. But our own failure is both a loaded and a complicated example; to see how the tool works more plainly, let us look at an ordinary phenomenon that has nothing to do with theory: the breaking of a glass.

We have a broken glass; let us compare these two sentences:

A · passes the compass

"This glass is fragile."

Names a disposition: falsifiable, compresses a host of predictions, speaks in the right category.

B · snags on the compass

"The glass broke because it contains an essence of fragility."

Invents a substance: absorbs every outcome, renames the past, turns a disposition into a substance.

At first glance both seem to say the same thing. Yet one passes the compass and the other snags. Let us run the four questions one by one (the example draws on Ryle, via Holth 2001).

Internal consistency. Sentence (B), to explain the fact that it "broke," invents an inner substance called an "essence of fragility," then ties the breaking to this substance. But "fragility" already means "a disposition to break easily"; so (B) is saying: "it broke because it was disposed to break." This, just like opium's "dormitive virtue," is a fancier repetition of the question. (A), on the other hand, claims no inner substance; it merely names a disposition of the glass. The difference is fine but decisive: (A) speaks of a behavioral disposition, (B) turns that disposition into a substance and presents it as if it were a cause.

Rescue concept. What observation does (A) forbid? If I say "this glass is fragile," I expect it to break when I drop it; if it falls and does not break, I was wrong; I must withdraw the description "fragile." So (A) is falsifiable; it forbids certain observations. And (B)? If I say "it contains an essence of fragility," when the glass breaks I say "there, its essence came out"; if it does not break I can say "its essence has not yet awakened." No observation can refute this sentence, because it takes every outcome in. That is the rescue concept: a joker that seems to explain while excluding nothing.

Scope transparency (compression and new prediction). (A), in a single word, packs a host of predictions: this glass will break if it falls, crack if struck with a hard object, may shatter under sudden temperature change, break apart if stepped on. The moment you say "fragile," you read in advance the results of experiments you have not yet run; this is an inference ticket. (B), by contrast, yields no new prediction; it merely repeats a single event that has occurred ("it broke") in heavier words. (A) compresses and speaks forward; (B) merely renames the past.

Category fit. "Fragile" is the name of a disposition: it tells how an object will behave under certain conditions. To use it as a disposition (A) is the right category. But (B) carries this disposition into the category of a substance, an inner essence (as if a "fragility" were dissolved inside the glass like sugar). Disposition and substance are different categories; to speak of one in the language of the other is a category mistake.

Two lessons follow from this example. First, the compass's job is to sift not the phenomenon but the frame built around it. The glass's fragility is real; (A) and (B) are two different sentences about the same reality, but the compass sifts out only one of them. That is, the tool does not say "glass is not fragile"; it says "turning fragility into a substance adds no explanation." Second, the same word (fragility) can be a legitimate inference ticket in one context and an empty renaming in another. So the verdict is rendered not on the word but on how the word is used. These two lessons are also the core of the next section: the compass is not a guillotine.

5. Why the compass is not a guillotine

These four questions can easily turn into a rejection machine: "it does not compress, so into the bin." This would be to fall into the diagnosis's own error. At least four independent sources give the same warning; that warning is part of the compass too.

Dalton (2016) reminds us that automatically deeming something nonsense when we cannot immediately make sense of it is itself an error; depth sometimes appears later, through the reader's contribution and slow contemplation. Haslam, Tse, and De Deyne (2021) show that the widening of a concept is not automatically bad; some widenings recognize a real but previously ignored phenomenon, that is, they can be a progressive gain. For instance, if the extension of the concept of "bullying" from the schoolyard to the workplace and thence to the online realm made visible real harms not previously named, this is not a dilution but a gain. Frankenhuis and colleagues (2023) argue, with a fine analogy from the game of "Twenty Questions," that uncertainty can be unavoidable and useful in early theory-building. In that game one person holds something in mind and the others try to find it with yes/no questions. "Is it a human?" is a broad question, "Is it Ringo Starr?" is a narrow one; but both are equally clear. So clarity and narrowness are different axes: a frame is problematic not because it is broad but because it is opaque (not honest about its scope). Finally, Popov (2023) admits that carrying his own criterion (mathematical precision at the level of the laws of physics) directly into the human sciences would be a mistake; if the bar is set too high, nothing can clear it and everything is rejected as an "inadequate theory."

The common lesson of these four warnings is this: the compass's job is not to cut off a phenomenon as "unworthy of theory" but to notice, early on, a frame that takes itself too seriously. It points a direction, it does not execute. The balance here matters, because the tool itself can go to excess: a reader who condemns every impressive concept from the start falls into precisely the error of premature rejection.

6. Why it is common: two mechanisms, one unmeasured claim

It is easy to say that over-theorizing is common, but this must remain an observation, not be presented as a measured fact. The lesson Haslam and colleagues (2021) give is cautionary here: even an intuitively expected spread can, when measured carefully, turn out to be false (in their own research, the expectation that psychiatric diagnoses had undergone a general "inflation" over time was, on the whole, not borne out by the data). So rather than say "epidemic," I confine myself to naming two mechanisms that could feed the tendency.

On the academic side: incentive selection. Frankenhuis and colleagues (2023) show with a simple logic why vague theory survives. Because a vague theory can absorb every finding, its risk of failure is low and it is easier to produce; a precise theory, being refutable, is risky and more laborious to produce. When visibility and citation are rewarded, even if no one deliberately chooses it, the incentive structure brings producers of vague theory to the fore. One must think of this as a selection process: not a deliberate wrong, but the environment rewarding and feeding a certain type. The key point is that the process requires no intent; even well-meaning researchers can drift toward vagueness under this pressure.

On the side of current debate: the guru effect. Frankenhuis and colleagues (2023) touch on a possibility they leave outside their own scope: the phenomenon Sperber calls the "guru effect." People tend to deem profound what they cannot grasp; vagueness produces a feeling of not-understanding, this feeling arouses admiration, and admiration in turn speeds the spread of an idea. Because in current essays and popular writing on AI the reward is not a "career" but this visibility, the guru effect carries into this field more directly than through the academic incentive model. Unintelligibility is mistaken for depth; and what is taken for deep spreads.

Both mechanisms point the same way: vague and impressive frames spread more easily than clear and modest ones. But how common this is remains an empirical question; this essay too leaves it as an open question to be researched, not an assumption. (The absence of a result is also a finding: in physics, Michelson and Morley's famous experiment tried to measure an invisible medium thought to carry light and found no trace; this "nothing" became a turning point, because it showed that the sought medium did not exist. To leave room, before saying "it is common," for the possibility that it may not be, is a small instance of the same honesty.)

7. Turning the compass on this essay

An essay criticizing over-theorizing can itself be a cathedral. Honesty requires applying the compass I propose to this very text.

Internal consistency. Can I use the concept of "over-theorizing" consistently? The real risk here is the process Haslam and colleagues (2021) call "concept creep": extending a concept to ever weaker, ever different phenomena until it comes to cover everything and distinguish nothing. That is why I defined the concept narrowly from the start (section 2): not every attempt at theory-building, but only a frame that does not compress, yields no prediction, and is opaque about its scope. Had I stuck "over-theorizing" onto every impressive idea, I would have fallen into the very dilution I criticize.

Rescue concept. Does this essay build its own thesis so as to absorb every counterexample? "If there is no need, into the bin" could be a broad settlement (bailey), and "but this is a matter of transparency" a hilltop tower (motte); that is, under pressure I could flee from the large claim to the small one. To avoid this trap, I stated openly which theory-building is legitimate: the one that is internally consistent, compresses the phenomenon, yields new predictions, and is honest about its scope. So this essay does not judge every theory; it shows what it forbids, and leaves the ground beyond that limit free.

Scope transparency. That the compass looks very clean is itself a risk. Nguyen's warning (via Frankenhuis et al., 2023) is exactly to the point here: a sense of clarity, real or imagined, ends inquiry. An over-neatly presented method can prevent us from seeing its own limits. That is why I leave the compass's limits open: these three questions do not solve everything, they are calibrated by four independent warnings, and they are not a guillotine either. They exist not to cut off a phenomenon but to turn back on myself and ask, "am I perhaps erecting a cathedral here?"

8. Conclusion

Building theory is no crime; it is one of the things that make us human. The problem is erecting a structure a phenomenon cannot bear, and being opaque about its scope. One does not build a cathedral for a need the size of a needle; but to a real need (an error to be dispelled or a light to be lit) we owe a modest tool.

The tool this essay proposes is small: four questions and one calibration. When you meet a phenomenon, ask: does the frame compress it, does it say something new, is it honest about its scope, is it looking on the right shelf? If the answers are weak, and if they remain weak even after weighing that weakness against the opposite pole (premature rejection, too high a bar), you are probably standing before a cathedral. This often means seeing that leaving a one-sentence observation as a single sentence is enough. That too is what our own failure taught us.

References (APA 7)

Dalton, C. (2016). Bullshit for you; transcendence for me. A commentary on "On the reception and detection of pseudo-profound bullshit." Judgment and Decision Making, 11(1), 121-122.

Frankenhuis, W. E., Panchanathan, K., & Smaldino, P. E. (2023). Strategic ambiguity in the social sciences. Social Psychological Bulletin, 18, Article e9923. https://doi.org/10.32872/spb.9923

Haslam, N., Tse, J. S. Y., & De Deyne, S. (2021). Concept creep and psychiatrization. Frontiers in Sociology, 6, Article 806147. https://doi.org/10.3389/fsoc.2021.806147

Holth, P. (2001). The persistence of category mistakes in psychology. Behavior and Philosophy, 29, 203-219.

Pennycook, G., Cheyne, J. A., Barr, N., Koehler, D. J., & Fugelsang, J. A. (2015). On the reception and detection of pseudo-profound bullshit. Judgment and Decision Making, 10(6), 549-563. https://doi.org/10.1017/S1930297500006999

Popov, V. (2023). If God handed us the ground-truth theory of memory, how would we recognize it? [Preprint, in revision]. PsyArXiv. https://psyarxiv.com/ay5cm/

Version and change record

This essay is a living text; every significant change leaves its trace here.

Version 1.0 · First published: 31 August 2026. The source draft was produced in the author's research unit (an internal v1 to v2 draft chain); this is the text's first public release on the site.

Translation regime. The source of this essay is Turkish. The concepts are rendered in their original English academic terms (dormitive virtue, theory stretching, motte-and-bailey, category mistake, concept creep, double dipping, inference ticket, guru effect, pseudo-profound bullshit); the citation format (author-year, et al., p., &) and the number format follow English convention. There is no verbatim foreign-language quotation; the illustrative figures (Molière, Le Verrier, Michelson and Morley, Twenty Questions) are given in their established English forms.

yontu · halit cengiz uzuner

Other writings

Who Really Uses AI? Hallucination: Whose Problem? Which of Us Is Wrong? Understanding Litmus Sturgeon's Law