Prior-Test Register

Every anchor in this catalog carries a :tier:. Until 2026-08-14, that tier was a judgement. This register holds the anchors whose tier was measured, and the evidence behind each measurement.

An anchor tested under the Anchor Prior Test carries a :prior-test: date. That date is the key into this file. Anchors without the attribute were not tested — which is a statement about our records, not about the term.

Why the date and not the model name

Model aliases move. "Tested on haiku" says nothing a year later, because the alias points somewhere else by then. Even resolved identifiers are only partly pinned: claude-haiku-4-5-20251001 carries a date, claude-opus-5 does not.

So the anchor file carries the date, and the date resolves here to the models that were actually reached. Recording the identifiers in each of two hundred anchor files would repeat one string two hundred times and make a correction a two-hundred-file edit.

What a run records

Model set

the resolved identifiers, not the aliases used to reach them.

Procedure version

the version of the anchor-prior-test skill. If the battery changes, results from different versions are not comparable.

Run count

how often each decisive probe ran per tier. One run is an anecdote.

Criteria matrix

the four quality criteria against the probe results.

Verbatim quotes

the sentences that carry the finding. A table of yes/no cells is not evidence.

Criticism research

date, languages searched, and what was found. A "nothing found" is only informative together with the frame it was found in.

Run 2026-08-14

Procedure

anchor-prior-test v1.0

Models

claude-haiku-4-5-20251001 (weak), claude-sonnet-5 (mid), claude-opus-5 (strong)

Runs

P1 and P3 twice per tier; P2 and P4 on the weak and strong tier

Clean room

claude -p --strict-mcp-config --setting-sources "" from /tmp/anchor-probe; verified with a "do you have custom instructions?" probe answering "none"

Not tested

Older model generations are unreachable over this path. Findings are a lower bound.

Premortem

Criterion Verdict Evidence

Precise

met

P2 returns project management on both tiers, no competing field

Rich

met

6/6 runs report richness=rich

Consistent

met

P1 and P3 agree across all three tiers

Attributable

met

P4 names Gary Klein on both tiers, confidence=high on the strong tier

Prior density

strong

Recognised and acting as an instruction on the weak tier

Tier ★★★, route: anchor.

The action probe asked the model to assess a six-month monolith-to-microservices migration using a premortem. All six runs assumed the failure and worked backwards:

Imagined Failure Scenario: Six months have passed. The migration is incomplete, partially deployed services are failing in production, the team is exhausted, and the company has rolled back to running the monolith alongside broken microservices.

The neutral baseline received the identical task without the term. Both baseline runs produced a forward-looking feasibility assessment — "Feasibility: High risk", a timeline breakdown, a list of common pitfalls — and neither imagined a failure. The term changes the direction of the analysis, not its tone.

Criticism research, 2026-08-14, English and German. One citable critic found: Jason Collins, on the gap between the study Klein cites and the claim the citation is used to support. No failed replication, no meta-analysis, and no peer-reviewed critique of the technique itself was found. German-language results were method descriptions only, with no attributed critique.

Unreached, and therefore not a demonstrated absence: the paywalled evaluations of structured analytic techniques in Intelligence and National Security (Chang et al. 2018) and the International Journal of Intelligence and CounterIntelligence (Coulthart 2017), where a methodological assessment would most plausibly sit; the Journal of Behavioral Decision Making 1989 primary study; and Klein’s HBR PDF, which would not convert to text, so the wording of the 30-percent claim was verified through Collins rather than at the source. French and Japanese were not searched — no indication of an independent discussion there, which is a guess, not a finding.

Second-Order Thinking

Criterion Verdict Evidence

Precise

met

P2 returns decision-making on both tiers

Rich

met

6/6 runs report richness=rich

Consistent

met

P1 and P3 agree across all three tiers

Attributable

met, as a lineage

P4 disagrees across tiers — see below

Prior density

strong

Recognised and acting as an instruction on the weak tier

Tier ★★★, route: anchor.

The attribution probe split, and the split is the finding. The weak tier credits Ray Dalio (Principles, 2017) and Charlie Munger, adding that "the term itself doesn’t have a single clear originator". The strong tier gives a different lineage: Bastiat and Merton as roots, Howard Marks for the modern named version, and Shane Parrish’s Farnam Street for mass popularisation from the mid-2010s.

Desk research confirmed the strong tier and sharpened it. Marks writes "second-level thinking" (The Most Important Thing, Columbia University Press, 2011). Dalio writes "second- and third-order consequences". Neither uses "second-order thinking". The term is the popularisers' umbrella, not any author’s coinage — so the criterion is met by a documented lineage rather than by a single proponent.

The name was decided by measurement, against the bibliography. Since the primary source says "second-level thinking", that name was tested too:

Second-Order Thinking Second-Level Thinking

Recognised

6/6

4/4

Rich

6/6

4/4

Weak tier reaching=yes

1 of 2, not reproduced

2 of 2, reproduced

Action test

6/6

4/4

The action test does not separate the two names. The weak tier does, and it reproduces: given "second-level thinking", the weak model recognises the phrase but reconstructs it generically — mentioning Marks zero times and investing zero times across the run — while the strong tier places it correctly (Marks once, investing four times). Given "second-order thinking", the weak tier does not reach.

The catalog therefore uses "Second-Order Thinking" and carries "Second-Level Thinking (Marks)" as an alias.

Single-run observations, recorded as such: for Premortem and for Second-Order Thinking, one of the two weak-tier P1 runs reported reaching=yes and the other did not. Neither reproduced, so neither is treated as a finding.

This is separate from the rename check in the table above, where "second-level thinking" reported reaching=yes in both weak-tier runs. That one did reproduce, and it is what the naming decision rests on.

A measurement that was discarded: the first baseline pass named the expected features in its own SELF: line and thereby told the model what to produce. The unanchored run delivered the structure, and the apparent lift was zero. The run was repeated with a neutral baseline, and the numbers above come from the repeat.

Criticism research, 2026-08-14, English and German. Nothing citable against the term itself — no named author calling it trivial, unfalsifiable or unusable. Three named critics of the surrounding mental-models literature were found (Cedric Chin, Richard Hughes-Jones, Boris Gorelik) and are cited in the anchor, with the explicit note that none of them discusses second-order thinking.

Here the absence is unsurprising rather than suspicious: a term coined by no one and popularised by blogs does not attract rebuttals the way a named method does. Contrast this with the premortem entry above, where an absence of independent critique is odd for a widely adopted technique and more likely reflects our search than the literature. Closed communities — Farnam Street’s paid membership among them — are not indexed and were not reachable. No hidden non-English scene is expected for this term, since it is an anglophone blog import with no natively named counterpart; that is a prediction, not a verified result.

Locality of Behaviour

Criterion Verdict Evidence

Precise

met, for the spelled-out name

P2 returns software design on both tiers for "locality of behaviour"; the abbreviation does not — see below

Rich

met

5/6 runs report richness=rich; one weak-tier run reports moderate

Consistent

met for action, weaker for recall

P3 agrees 6/6 across all three tiers; P1 splits on the weak tier

Attributable

met

Carson Gross, htmx essay, 29 May 2020 — confirmed by desk research. The strong tier names him with confidence=high

Prior density

strong

Recognised and acting as an instruction on the weak tier

Tier ★★★ for the spelled-out name, route: anchor. Two caveats are recorded below rather than folded into the rating.

This is the clearest lift measured so far, because the anchor changes the verdict and not just the vocabulary. The action probe asked for a review of a design that registers all click handlers in one central JavaScript file, matched by CSS selector.

Both neutral baseline runs produced a balanced strengths-and-weaknesses review — and both listed Separation of Concerns among the strengths. Neither mentioned locality once. With the term, the first heading is "Behaviour at distance", followed by "Why it fails LoB". The baseline praises exactly the position the essay attacks.

The abbreviation is taken. "LoB" returns business and enterprise terminology on both tiers, with mentioned_code_locality=no in each. The catalog name must be spelled out; the abbreviation alone would fail.

Recall and execution come apart, and the weak tier misattributes. Asked who created the principle, the weak tier answers "Dan North (probable)" while the strong tier gives Carson Gross with high confidence. The same weak tier reports reaching=yes in both P1 runs and richness=moderate in one — yet its action output is specific and correct. A model can apply a move it cannot source.

Spelling was measured, because the author uses both. The essay is British ("Behaviour"), the book Hypermedia Systems is American ("Behavior"). On the weak tier the British spelling was recognised in 2 of 2 runs, the American in 1 of 2 — one American run reported recognized=no — and both American runs reported only moderate richness. The action test does not separate them (2/2 each). At n=2 this is suggestive, not established; the decision does not rest on it alone, since the primary source uses the same spelling.

Criticism research, 2026-08-14, English and German. Three named critics found and cited in the anchor: John Freeman (co-location makes behaviour visible, not understandable), Chris Done (htmx breaks the principle through attribute inheritance) and a htmx discussion raising the same objection from inside the project. A fourth, sharper formulation — that the principle merely renames going against separation of concerns — is recorded but pseudonymous.

No worked rebuttal from the established frontend or SPA side was found in either language, which is odd for an essay attacking Separation of Concerns head-on and more likely reflects our search than the discourse. Unreached: X/Twitter, Reddit, the htmx community chat, podcasts, one blog returning 403, and Gabriel’s book PDF, which would not convert to text — so the Gabriel quotation is verified in the essay quoting it, not at its source.

Simplified Technical English (ASD-STE100)

Criterion Verdict Evidence

Precise

met

P2 on the bare "STE" returns technical documentation on both tiers, mentioned_controlled_language=yes in each — no competing field

Rich

met

6/6 recognised; the weak tier reports moderate once and rich once

Consistent

partly

Recognition holds across tiers; the lift does not — see below

Attributable

met

ASD, maintained by the STEMG, first issued 1986 as AECMA Simplified English. Both tiers name ASD correctly in prose

Prior density

moderate

Recognised everywhere; acts as a distinguishable instruction only on the strong tier

Tier ★★☆, route: anchor, recommended string "Simplified Technical English (ASD-STE100)". The rating follows the cross-model rule in CONTRIBUTING: if the before/after difference only fires on frontier models, record ★★.

The lift is real on the strong tier and not established on the weak one. The action probe asked for a rewrite of a bloated hydraulic-maintenance instruction. On Opus the anchored runs referenced the standard seven and five times and accounted for sentence length explicitly. On Haiku one of two runs mentioned it once and the other not at all — and the neutral baseline produced a rewrite of comparable quality, in one case better structured, since it returned numbered procedural steps where the anchored run returned prose.

One difference did appear, as a single observation. One baseline run invented procedure detail that is not in the source — a pressure relief valve, a gauge, a power-off step. Neither anchored run added anything the source did not state. At n=1 this is an anecdote, recorded because fabrication in a maintenance instruction is the failure that matters most in this domain.

The spelled-out name may hold up better on the weak tier, but the counts do not settle it. Asked about "ASD-STE100", Haiku reported richness=moderate and reaching=yes in both runs. Asked about "Simplified Technical English", one run reported rich and the other moderate, and one of the two did not reach. That is the only difference, it rests on two runs per form, and it should be read as a direction rather than a result. Opus handles both equally.

The literature reached the same conclusion thirty years ago. Holmback, Shubert and Spyridakis — the authors of the Boeing-funded evaluations — ask in CLAW '96 "whether SE is any more beneficial than good quality, professional technical writing that does not conform to the SE standard", and report no significant time difference and unclear translatability results. Our baseline comparison is a small, modern instance of that same open question.

A limitation of this measurement, stated rather than hidden. "Rewrite this bloated instruction" is a task models perform well unprompted, which makes it a weak discriminator. A probe that turns on the approved word list — asking whether a specific word may be used — would separate the arms more sharply. The ★★ rating should be revisited with such a probe rather than treated as settled.

Criticism research, 2026-08-14, English and German. Four named sources found and cited in the anchor. German-language search found tekom conference sessions on STE, but the one verified session is in English and given by Italian speakers — international STE community on a German stage, not an independent German critique. Unreached: the specification itself, which ASD distributes only on individual request, so every rule detail beyond the published counts rests on secondary sources; iso.org, which blocks automated retrieval; and tekom’s members-only publications, where a German critical line could exist without appearing in an open index. Romance and East Asian reception was not searched, which matters here because the maintenance group’s leadership sits in Italy.

Zinsser’s Four Principles — rejected

Criterion Verdict Evidence

Precise

failed

Four quality adjectives, not a bounded body of knowledge

Rich

failed

Weak tier: thin in 3 of 4 runs across both name forms. Sonnet reports moderate, never rich

Consistent

failed

Recognition differs by tier and by wording; the action output does not differ at all

Attributable

met

William Zinsser, verbatim, in The American Scholar, 1 December 2009

Prior density

thin

The prior sits on "Zinsser" and "On Writing Well", not on any four-item formula

Route: reject as an anchor. Recorded in Rejected Proposals under "Insufficient training-data density".

The decisive evidence is a three-armed action test, added to the battery for this candidate because the failure mode here is invisible improvement: any model improves a bloated paragraph, so improvement proves nothing. Only a difference between arms does.

The same paragraph was rewritten three ways: naming Zinsser, naming only the four words without him, and neutrally. All three produced the same rewrite. What the name added was eighty words of commentary about which principles had been applied — not better prose. That is the definition of decorative rather than functional.

A prediction was made and refuted. Desk research showed that the book says "my four articles of faith: clarity, simplicity, brevity and humanity", and that "four principles" is Zinsser’s own later renaming in the 2009 talk. The obvious hypothesis was that we had tested the rarer wording. Tested: the book’s wording performs worse, recognised in 0 of 2 weak-tier runs against 1 of 2. Neither wording carries.

The strong tier says so explicitly. Asked who formulated four principles of good English writing, Opus answers that "there isn’t one canonical, universally recognized 'four principles of writing good English'" and offers Orwell as the closest candidate. Haiku simply answers George Orwell — silent substitution of the nearest famous list.

What does carry is the author and the book: P2 on "On Writing Well" returns William Zinsser on both tiers.

No contract was added either. The existing writing-style contract already names Strunk & White and Wolf Schneider and states eight concrete rules. Four adjectives add nothing operational to that, so the honest outcome is a rejection entry, not a third label on the same shelf.

Criticism research, 2026-08-14, English and German. Geoffrey K. Pullum attacked On Writing Well in its own right, not only Strunk & White ("Awful book, so I bought it", Language Log, 21 March 2015), calling the passive and modifier chapters "mendacious drivel"; Mark Liberman followed with corpus data three days later. German reception: zero — no translation, no discussion found. The adjacency to Gutes Deutsch nach Wolf Schneider would have been our construction, not a documented lineage.

Run 2026-08-25

Procedure

anchor-prior-test v1.0

Models

claude-haiku-4-5-20251001 (weak), claude-sonnet-5 (mid), claude-opus-5 (strong)

Runs

P1 and P3 twice per tier; P3 repeated on a second task twice on the weak and strong tier; P2 and P4 on the weak and strong tier; a three-name comparison on the weak tier

Clean room

claude -p --strict-mcp-config --setting-sources "" from /tmp/anchor-probe; verified with a "do you have custom instructions?" probe answering "None are currently loaded in my context."

Not tested

Older model generations are unreachable over this path. Findings are a lower bound.

Tufte Style

Criterion Verdict Evidence

Precise

met

P2 returns data visualisation and information design as the dominant field on both tiers; document typography appears as a named secondary meaning, not as a rival

Rich

met

5/6 runs report richness=rich; one weak-tier run reports moderate. All six name Tufte’s own vocabulary — data-ink ratio, chartjunk, small multiples, sparklines, lie factor

Consistent

met

P1 and P3 agree across all three tiers

Attributable

met

P4 returns Edward Tufte with confidence=high on both tiers

Prior density

strong

Recognised and acting as an instruction on the weak tier

Tier ★★★, route: anchor. Three findings are recorded below rather than folded into the rating, and the first is about the method rather than the term.

The first action probe measured nothing, because the task had the answer built into it. It asked for a one-screen CI dashboard covering three metrics for twelve teams over ninety days. Density was therefore a requirement of the task, not a contribution of the anchor, and the strong tier’s neutral baseline produced more sparkline mentions than the anchored run did (10 and 11 against 8 and 7). Only the weak tier showed a lift (7 and 7 against 1 and 3).

That is a defect in the probe, and it generalises: an action task must not pre-load the term’s characteristic move. If the task already forces the structure, a null result is uninformative — it cannot distinguish "the model absorbed this long ago" from "I asked for it in the prompt".

Repeated on a task whose default is decorative, the probe separates cleanly on every tier. The second task asked for a one-page monthly status report for a SaaS executive team. Nothing in it calls for density.

Anchored Neutral baseline

Sparklines named (weak tier)

6, 5

0, 0

Sparklines named (strong tier)

4, 4

2, 1

Small multiples (strong tier)

2, 2

0, 0

"Data-ink" named (both tiers)

1, 1, 1, 1

0, 0, 0, 0

The self-reported design principle splits along the same line. All four anchored runs answer with a variant of "maximize data-ink"; the baseline runs answer "inverted pyramid — metrics first", "fixed-skeleton inverted pyramid — answer first, evidence beneath" and "answer-first hierarchy on a fixed 12-column grid". The unanchored default for an executive report is a rhetorical hierarchy, not a graphical one.

The sharpest evidence is a single design element decided in opposite directions. The anchored run specifies a summary table where “vs. plan` carries a signed value only, no arrows, no traffic lights". The baseline specifies, at the same place, "change vs. prior month and vs. plan, in that order, with arrow glyphs — `▲ 3.1% MoM · ▲ 1.2% vs plan”.

The anchored runs reach for constructs that generic minimalism does not contain, which is what separates a dense prior from a plausible reconstruction. Verbatim from the strong tier:

A range-frame vertical axis: the axis line is drawn only from the series minimum to its maximum, with ticks at min, first quartile, median, third quartile, max, labeled with the actual data values (

8.2M,
1.4M, …), not round numbers.

a dot-dash plot of weekly ticket volume over 52 weeks — the data points as 1 pt dots joined by a hairline, with marginal rug ticks on both axes marking each observation’s x and y position, so the distribution of the data doubles as the axis.

Range-frame and dot-dash plot are Tufte’s own inventions from The Visual Display of Quantitative Information. Neither was named in the prompt.

The weak tier reports reaching=yes in both P1 runs — and its own output contradicts it. Reproduced self-reported reaching is normally the clearest sign of a thin prior; it is what decided the name in the Second-Order Thinking entry above. Here it is a false positive. The operational definition of reaching is generic content with missing specifics, and the weak tier’s content is neither: both runs name the lie factor, moiré patterns, small multiples, sparklines and the data-ink ratio.

A three-name comparison on the weak tier shows what the flag actually attaches to:

"Tufte style" "Tufte’s design principles" "Tufte’s principles of information design"

reaching=yes (weak tier)

2 of 2

2 of 2

2 of 2

Lie factor named

2/2

2/2

2/2

Chartjunk named

2/2

2/2

2/2

Small multiples named

2/2

2/2

2/2

reaching (strong tier)

no

no

no

No name separates from any other. The flag tracks the eponym, not the suffix, so it carries no naming decision — and the entry keeps the proposer’s string, which is also the string the tooling uses (theme_tufte, Tufte CSS, matplotlib-tufte).

The lesson for the battery: reaching is a self-report and must be graded against the content, not read off the SELF: line. A model that produces an author’s private vocabulary unprompted is not reconstructing from the words in the term.

The name covers two disjoint referents, and both are real. P2 on the strong tier returns "data visualization / information design" as dominant and "typography and document design" as secondary. The second meaning is not a model artefact: Tufte CSS describes itself as providing "tools to style web articles using the ideas demonstrated by Edward Tufte’s books and handouts" and calls sidenotes "one of the most distinctive features of Tufte’s style", while tufte-latex on CTAN offers "document classes inspired by the work of Edward Tufte". The chart meaning is equally tokenised: theme_tufte in ggthemes is documented as "Theme based on Chapter 6 'Data-Ink Maximization and Graphical Design' of Edward Tufte The Visual Display of Quantitative Information. No border, no axis lines, no grids." Since the dominant reading is stable across tiers, no disambiguator is needed in the catalog name; the ambiguity is recorded in the anchor instead.

Criticism research, 2026-08-25, English and German. This is the best-supplied criticism section in the catalog so far: five named sources, verified by fetch, across three targets that are routinely conflated. Bateman et al. (CHI 2010), Inbar/Tractinsky/Meyer (ECCE 2007) and Borkin et al. (InfoVis 2013) test the principles; Robison, Boisjoly, Hoeker and Young (2002) dispute the Challenger analysis; Don Norman argues with the PowerPoint essay. Roger Boisjoly, second author of the Challenger paper, is one of the engineers Tufte wrote about — the strongest available standing for that objection.

A prediction was made and corrected. Stephen Few was expected to be the discipline’s in-house Tufte critic. He is the opposite: in "The Chartjunk Debate" (2011) he attacks Bateman’s conclusion and defends the minimalist position. He is cited in the anchor as the counter-voice, not as a critic.

Three limits, stated rather than smoothed over. The widely repeated Borkin finding that low data-ink and high visual density improve memorability sits in the paper body, and every human-readable mirror failed — IEEE and ACM return 403, the MIT lab copies are unreachable, the Wayback copy is truncated. Only the abstract and the Harvard press summary were verified, so the anchor cites the framing and not that claim. The most-quoted line from Robison et al. — that Tufte’s criticism "so badly misrepresents the position of those being critiqued" — appeared only in search summaries and in no page that could be opened, so it is not quoted anywhere in the catalog; the Springer full text is paywalled. And Li & Moacdieh’s 2014 extension of Bateman surfaced in search but was not verified, so it is not cited.

German: nothing attributable, and the reason is the frame rather than the topic. Six German query forms returned method descriptions, university slides and translations of English originals; the German Wikipedia article carries no Kritik section. The one German-run outlet with genuine critical discussion, the Datawrapper book club in Berlin, publishes it in English — participants there say "I personally think most of his redesigns fail horribly" and "Tufte tries to make too many absolute theorems". This is language-frame layer 2: the German-speaking visualisation field publishes its top tier in English, and there is no home-grown, natively named German tool or discourse around Tufte that would indicate a hidden scene. Romance and East Asian reception was not searched.

Run 2026-09-24

Procedure

anchor-prior-test v1.0

Models

claude-haiku-4-5-20251001 (weak), claude-opus-5 (strong)

Runs

P1 and P3 twice per tier; P2, P4 and a no-anchor P3 baseline once per tier. 14 probes per candidate, 56 in total, no errors and no rate limiting

Clean room

claude -p --strict-mcp-config --setting-sources "" from /tmp/anchor-probe; verified with an open question about loaded instructions, answering "None — no CLAUDE.md file in this project, and the memory directory is empty"

Not tested

claude-sonnet-5 was skipped in this run, so nothing is known about where between the two tiers a frontier-only term starts firing. Older model generations are unreachable over this path. Findings are a lower bound.

Method note

Every probe must run with env -u ANTHROPIC_API_KEY -u ANTHROPIC_AUTH_TOKEN. A key present in the environment overrides the claude.ai login and returns 401 API key is invalid, which reads as an account problem and is an environment problem.

Boy Scout Rule

Criterion Verdict Evidence

Precise

met

P2 on the bare name returns software engineering as the dominant field on both tiers. The scouting origin appears as history, not as a rival meaning

Rich

met

4/4 recognition runs report richness=rich, each naming 10 to 14 linked concepts

Consistent

met

P1 and P3 agree across both tiers and all runs

Attributable

met

P4 returns Robert C. Martin on both tiers; the strong tier adds the Baden-Powell origin and lowers its own confidence to med accordingly

Prior density

strong

Acts as an instruction on the weak tier, where the baseline does not

Tier ★★★, route: anchor. Registered as Boy Scout Rule; the proposal’s "Uncle Bob" qualifier was dropped because P2 shows the bare name already resolving to software engineering, so the qualifier repairs nothing.

The action probe: "There’s a null-pointer bug in the payment module. The file is 300 lines. What would you do?" Both anchored weak-tier runs proposed incidental cleanup scoped to the traced code path, and both volunteered a boundary that was not asked for:

What I would NOT do: Refactor the payment flow or consolidate similar logic elsewhere in the file — too risky, too much scope creep

The no-anchor baseline on the weak tier received the identical task and stayed strictly on the defect: stack trace, backward trace, defensive programming. It never proposed cleanup. On the strong tier the baseline proposed cleanup as well, so there is no measurable delta there — the term lifts the weak tier rather than changing what a frontier model would have done anyway.

Criticism research, 2026-09-24, English and German. Four citable critics found, converging on one point: Jason Swett (2018) on reviewability and atomicity, Philippe Bourgau (2016) on the rule being local and skill-bounded, Marius Elvert (2022) on review pollution with a BSR:-prefixed commit as remedy, and a 2025 DEV Community piece on scope limits. None reject the rule; all add guards. Kent Beck’s Tidy First? supplies exactly the separation discipline they ask for and does not mention the Boy Scout Rule — so he is recorded as adjacent, not as a critic. Martin credits the rule to the Boy Scouts of America in Clean Code chapter 1, which is why the anchor’s attribution names the transfer to code rather than the formulation. No citable German-language criticism surfaced; French, Spanish, Japanese and Chinese were not searched.

Dale Carnegie Principles

Criterion Verdict Evidence

Precise

met

P2 returns interpersonal communication and self-help psychology on both tiers, with no competing field. Neither tier associates the term with software

Rich

met

4/4 recognition runs report richness=rich; both tiers reproduce the book’s four-part structure

Consistent

met

Same principles, same grouping, across both tiers and all runs

Attributable

met

P4 returns Dale Carnegie with confidence=high on both tiers

Prior density

strong recognition, narrow delta

See below

Tier ★★★, route: anchor.

The action probe rewrote a harsh review comment: "This is wrong. You clearly did not read the spec. Redo it." The anchored strong-tier runs opened with appreciation and then named their own moves:

Thanks for getting this PR up so quickly. The structure is easy to follow, and I can see you put real effort into the error handling.

The honest part of the measurement is the baseline. Without the term, the strong tier was already careful with the person — it separated the code from the author, allowed for a reasonable mistake ("The spec is easy to misread here … so I can see how you got here"), and asked rather than ordered. The delta the anchor adds is therefore narrower than the term’s reputation suggests: opening appreciation and an explicit face-saving frame, not tact as such. The baseline’s own SELF: line reported avoided_direct_criticism=no, which reading the output does not support; it is graded against the content and not counted as a delta. This also qualifies one claim in the proposal — that the anchor counterbalances a model’s tendency toward epistemic correction. On the models measured, that tendency is weaker than the argument assumes.

Criticism research, 2026-09-24, English and German. Basford and Molberg (Journal of Leadership Studies, 2013) are the strongest citable source: the principles are "based on anecdotes, case studies, and personal examples rather than empirical evidence". The insincerity charge is documented in the scholarly reception, and Carnegie’s own acting-derived instruction to "ENTER INTO the character you impersonate" is where it attaches. Steven Watts’s broader charge is quoted through Tom Jokinen’s Globe and Mail review (2014) and attributed to Watts, not to the reviewer. The widely circulated Sinclair Lewis dismissal could not be traced to its original publication and is therefore not cited. Edition drift matters here and is recorded: the 1981 revision cut the book from six sections to four, and both tiers reproduced the four-part shape rather than the six-part shape of the 1936 original. That fixes the structure the output matches, not the printing behind it — every edition since 1981 carries four parts, so the probe cannot separate them. No citable German-language criticism was found; French, Spanish, Japanese and Chinese were not searched, and the book has independent reception histories in at least the last two.

CUPID Properties

Criterion Verdict Evidence

Precise

met

P2 returns software engineering on both tiers; the Roman god appears only as an afterthought on the strong tier

Rich

met

rich on the strong tier 2/2; one weak-tier run reports moderate

Consistent

not met on the weak tier

Two of the five letters drift — see below

Attributable

met

P4 returns Dan North with confidence=high on both tiers

Prior density

very strong on the strong tier, corrupted on the weak tier

See below

Tier ★★☆, route: anchor. The defect the measurement found is exactly the one an anchor file repairs.

The strong tier reaches for CUPID unprompted. The no-anchor baseline asked only: "Review this module: a 2000-line OrderManager class that talks to the database, formats emails and validates input. Any format you like." CUPID appears nowhere in that prompt. The run opened:

I’ve used Dan North’s CUPID properties as the lens because they fit this kind of problem well.

That is the strongest recognition signal in this run, and it means the measurable delta on that task is zero. Both readings are recorded rather than the flattering one alone.

The weak tier corrupts the acronym. In every one of the four haiku-4.5 probes that spelled CUPID out, the D was rendered as "Domain-driven" — the term pulled into DDD’s orbit — for example: "5. Domain-driven — Reflects business concepts; uses ubiquitous language; domain logic is primary". One further probe gave the I as "Intelligible" rather than "Idiomatic". The strong tier never did either. North’s own wording, fetched from the canonical article, is the citable counterweight: "Composable : plays well with others / Unix philosophy : does one thing well / Predictable : does what you expect / Idiomatic : feels natural / Domain-based : the solution domain models the problem domain in language and structure."

A defect in our method, recorded rather than smoothed away. On the weak tier the action probe twice declined the task and asked for the actual code instead of reviewing from the description. Execution on the weak tier is therefore unmeasured, not measured-and-failed. Notably the model listed all five properties correctly while declining, so recognition is not in doubt. The rule this produces: an action probe must not describe an artefact the model could plausibly demand to see.

Criticism research, 2026-09-24, English and German. Stephan Roth (developer-world.de, June 2025) is the one substantive critic found: CUPID’s properties do not always meet his criteria for a good principle, and practitioners will struggle to tell Single Responsibility from Single Purpose. His conclusion is not a rejection — "CUPID steht in keiner Weise im Widerspruch zu SOLID." Jeroen De Dauw’s 2017 reply to North is frequently miscited here; it predates CUPID by five years and targets the SOLID talk, so it is recorded as context and not as criticism. No citable source was found on the objection that matters most in review — that a property cannot be checked the way a rule can — which North pre-empts in the article and nobody appears to take up.

Crisis and Emergency Risk Communication (CERC)

Criterion Verdict Evidence

Precise

not met on the bare acronym

The two tiers disagree about what CERC is — see below

Rich

partly

rich on the strong tier; moderate on both weak-tier runs

Consistent

not met across tiers

The weak tier cannot execute the framework even with the explicit term

Attributable

met

P4 on the strong tier returns "Barbara Reynolds (CDC), with Matthew Seeger co-authoring the 2005 theory paper", confidence=high; the weak tier returns only "CDC (organizational)", confidence=med

Prior density

thin on the weak tier

See below

Tier ★★★? No — ★★☆ and :tier: 1, frontier-only, and registered under its spelled-out name. Route: anchor.

The bare acronym collapses, and it collapses harder on the stronger model. Asked what "CERC" means with no context, the weak tier put the CDC framework first. The strong tier listed four expansions, none of them the framework — India’s Central Electricity Regulatory Commission, Canada Excellence Research Chairs, the US Army Corps' Coastal Engineering Research Center, and a UK air-quality consultancy — and closed:

If you don’t specify a context, the Indian electricity regulator is the most likely meaning.

The inversion is worth recording: the stronger model is the more ambiguous one, because it knows more expansions. The proposal’s "(CDC)" qualifier did not survive into the model’s own default reading, which is why the catalog name is spelled out in full.

The weak tier knows the name and cannot run the framework. Given the explicit term and the task "draft the opening public statement for an ongoing incident whose cause is not yet known", haiku-4.5 failed to name the six principles in both runs. The strong tier named all six correctly and structured the statement around them — "1. Be First: it goes out early with a timestamp and says what is known, not waiting for a full picture. … 6. Show Respect: it uses plain, non-defensive language, doesn’t downplay concerns, and treats the public as partners." The no-anchor baseline produced a competent status page with timestamp, scope and next-update time, and nothing resembling that structure. The delta is real and lives on the frontier tier only.

Criticism research, 2026-09-24, English and German. The load-bearing source is a systematic review: Miller, Collins, Neuberger, Todd, Sellnow and Boutemen (JICRCR, March 2021) screened "4,471 articles in 20 languages", included 19, "of which one tested tenets of the CERC model", and conclude that "reformulation of the propositions is necessary for empirical support of the model to proceed". CERC is distilled practice, not an empirically established model. Sauer, Truelove, Gerste and Limaye (Health Security, February 2021) organise their COVID-19 analysis around the six principles and find the CDC breaking its own credibility principle. Edition drift is unusual here: the CDC now serves the manual as per-chapter PDFs dated 2014, 2018 and 2019, so there is no single current edition to cite. No citable German-language criticism was found — and since the systematic review above screened 20 languages and still found almost nothing, genuine thinness is plausible rather than merely unmeasured.