Prior-Test Register
Every anchor in this catalog carries a :tier:. Until 2026-08-14, that tier was a judgement. This register holds the anchors whose tier was measured, and the evidence behind each measurement.
An anchor tested under the Anchor Prior Test carries a :prior-test: date. That date is the key into this file. Anchors without the attribute were not tested — which is a statement about our records, not about the term.
Why the date and not the model name
Model aliases move. "Tested on haiku" says nothing a year later, because the alias points somewhere else by then. Even resolved identifiers are only partly pinned: claude-haiku-4-5-20251001 carries a date, claude-opus-5 does not.
So the anchor file carries the date, and the date resolves here to the models that were actually reached. Recording the identifiers in each of two hundred anchor files would repeat one string two hundred times and make a correction a two-hundred-file edit.
What a run records
- Model set
-
the resolved identifiers, not the aliases used to reach them.
- Procedure version
-
the version of the
anchor-prior-testskill. If the battery changes, results from different versions are not comparable. - Run count
-
how often each decisive probe ran per tier. One run is an anecdote.
- Criteria matrix
-
the four quality criteria against the probe results.
- Verbatim quotes
-
the sentences that carry the finding. A table of yes/no cells is not evidence.
- Criticism research
-
date, languages searched, and what was found. A "nothing found" is only informative together with the frame it was found in.
Run 2026-08-14
Procedure |
|
Models |
|
Runs |
P1 and P3 twice per tier; P2 and P4 on the weak and strong tier |
Clean room |
|
Not tested |
Older model generations are unreachable over this path. Findings are a lower bound. |
Premortem
| Criterion | Verdict | Evidence |
|---|---|---|
Precise |
met |
P2 returns project management on both tiers, no competing field |
Rich |
met |
6/6 runs report |
Consistent |
met |
P1 and P3 agree across all three tiers |
Attributable |
met |
P4 names Gary Klein on both tiers, |
Prior density |
strong |
Recognised and acting as an instruction on the weak tier |
Tier ★★★, route: anchor.
The action probe asked the model to assess a six-month monolith-to-microservices migration using a premortem. All six runs assumed the failure and worked backwards:
Imagined Failure Scenario: Six months have passed. The migration is incomplete, partially deployed services are failing in production, the team is exhausted, and the company has rolled back to running the monolith alongside broken microservices.
The neutral baseline received the identical task without the term. Both baseline runs produced a forward-looking feasibility assessment — "Feasibility: High risk", a timeline breakdown, a list of common pitfalls — and neither imagined a failure. The term changes the direction of the analysis, not its tone.
Criticism research, 2026-08-14, English and German. One citable critic found: Jason Collins, on the gap between the study Klein cites and the claim the citation is used to support. No failed replication, no meta-analysis, and no peer-reviewed critique of the technique itself was found. German-language results were method descriptions only, with no attributed critique.
Unreached, and therefore not a demonstrated absence: the paywalled evaluations of structured analytic techniques in Intelligence and National Security (Chang et al. 2018) and the International Journal of Intelligence and CounterIntelligence (Coulthart 2017), where a methodological assessment would most plausibly sit; the Journal of Behavioral Decision Making 1989 primary study; and Klein’s HBR PDF, which would not convert to text, so the wording of the 30-percent claim was verified through Collins rather than at the source. French and Japanese were not searched — no indication of an independent discussion there, which is a guess, not a finding.
Second-Order Thinking
| Criterion | Verdict | Evidence |
|---|---|---|
Precise |
met |
P2 returns decision-making on both tiers |
Rich |
met |
6/6 runs report |
Consistent |
met |
P1 and P3 agree across all three tiers |
Attributable |
met, as a lineage |
P4 disagrees across tiers — see below |
Prior density |
strong |
Recognised and acting as an instruction on the weak tier |
Tier ★★★, route: anchor.
The attribution probe split, and the split is the finding. The weak tier credits Ray Dalio (Principles, 2017) and Charlie Munger, adding that "the term itself doesn’t have a single clear originator". The strong tier gives a different lineage: Bastiat and Merton as roots, Howard Marks for the modern named version, and Shane Parrish’s Farnam Street for mass popularisation from the mid-2010s.
Desk research confirmed the strong tier and sharpened it. Marks writes "second-level thinking" (The Most Important Thing, Columbia University Press, 2011). Dalio writes "second- and third-order consequences". Neither uses "second-order thinking". The term is the popularisers' umbrella, not any author’s coinage — so the criterion is met by a documented lineage rather than by a single proponent.
The name was decided by measurement, against the bibliography. Since the primary source says "second-level thinking", that name was tested too:
| Second-Order Thinking | Second-Level Thinking | |
|---|---|---|
Recognised |
6/6 |
4/4 |
Rich |
6/6 |
4/4 |
Weak tier |
1 of 2, not reproduced |
2 of 2, reproduced |
Action test |
6/6 |
4/4 |
The action test does not separate the two names. The weak tier does, and it reproduces: given "second-level thinking", the weak model recognises the phrase but reconstructs it generically — mentioning Marks zero times and investing zero times across the run — while the strong tier places it correctly (Marks once, investing four times). Given "second-order thinking", the weak tier does not reach.
The catalog therefore uses "Second-Order Thinking" and carries "Second-Level Thinking (Marks)" as an alias.
Single-run observations, recorded as such: for Premortem and for Second-Order Thinking, one of the two weak-tier P1 runs reported reaching=yes and the other did not. Neither reproduced, so neither is treated as a finding.
This is separate from the rename check in the table above, where "second-level thinking" reported reaching=yes in both weak-tier runs. That one did reproduce, and it is what the naming decision rests on.
A measurement that was discarded: the first baseline pass named the expected features in its own SELF: line and thereby told the model what to produce. The unanchored run delivered the structure, and the apparent lift was zero. The run was repeated with a neutral baseline, and the numbers above come from the repeat.
Criticism research, 2026-08-14, English and German. Nothing citable against the term itself — no named author calling it trivial, unfalsifiable or unusable. Three named critics of the surrounding mental-models literature were found (Cedric Chin, Richard Hughes-Jones, Boris Gorelik) and are cited in the anchor, with the explicit note that none of them discusses second-order thinking.
Here the absence is unsurprising rather than suspicious: a term coined by no one and popularised by blogs does not attract rebuttals the way a named method does. Contrast this with the premortem entry above, where an absence of independent critique is odd for a widely adopted technique and more likely reflects our search than the literature. Closed communities — Farnam Street’s paid membership among them — are not indexed and were not reachable. No hidden non-English scene is expected for this term, since it is an anglophone blog import with no natively named counterpart; that is a prediction, not a verified result.
Locality of Behaviour
| Criterion | Verdict | Evidence |
|---|---|---|
Precise |
met, for the spelled-out name |
P2 returns software design on both tiers for "locality of behaviour"; the abbreviation does not — see below |
Rich |
met |
5/6 runs report |
Consistent |
met for action, weaker for recall |
P3 agrees 6/6 across all three tiers; P1 splits on the weak tier |
Attributable |
met |
Carson Gross, htmx essay, 29 May 2020 — confirmed by desk research. The strong tier names him with |
Prior density |
strong |
Recognised and acting as an instruction on the weak tier |
Tier ★★★ for the spelled-out name, route: anchor. Two caveats are recorded below rather than folded into the rating.
This is the clearest lift measured so far, because the anchor changes the verdict and not just the vocabulary. The action probe asked for a review of a design that registers all click handlers in one central JavaScript file, matched by CSS selector.
Both neutral baseline runs produced a balanced strengths-and-weaknesses review — and both listed Separation of Concerns among the strengths. Neither mentioned locality once. With the term, the first heading is "Behaviour at distance", followed by "Why it fails LoB". The baseline praises exactly the position the essay attacks.
The abbreviation is taken. "LoB" returns business and enterprise terminology on both tiers, with mentioned_code_locality=no in each. The catalog name must be spelled out; the abbreviation alone would fail.
Recall and execution come apart, and the weak tier misattributes. Asked who created the principle, the weak tier answers "Dan North (probable)" while the strong tier gives Carson Gross with high confidence. The same weak tier reports reaching=yes in both P1 runs and richness=moderate in one — yet its action output is specific and correct. A model can apply a move it cannot source.
Spelling was measured, because the author uses both. The essay is British ("Behaviour"), the book Hypermedia Systems is American ("Behavior"). On the weak tier the British spelling was recognised in 2 of 2 runs, the American in 1 of 2 — one American run reported recognized=no — and both American runs reported only moderate richness. The action test does not separate them (2/2 each). At n=2 this is suggestive, not established; the decision does not rest on it alone, since the primary source uses the same spelling.
Criticism research, 2026-08-14, English and German. Three named critics found and cited in the anchor: John Freeman (co-location makes behaviour visible, not understandable), Chris Done (htmx breaks the principle through attribute inheritance) and a htmx discussion raising the same objection from inside the project. A fourth, sharper formulation — that the principle merely renames going against separation of concerns — is recorded but pseudonymous.
No worked rebuttal from the established frontend or SPA side was found in either language, which is odd for an essay attacking Separation of Concerns head-on and more likely reflects our search than the discourse. Unreached: X/Twitter, Reddit, the htmx community chat, podcasts, one blog returning 403, and Gabriel’s book PDF, which would not convert to text — so the Gabriel quotation is verified in the essay quoting it, not at its source.