Correction & corpus rebuild — SOoL v2.0.0, 6 August 2026
The v1 pipeline derived Nodes 7–8 from the court’s disposition, so the predictive result reported here (AUC ≈ 0.98) was true by construction and the contradiction‑type distribution was distorted. All predictive claims are withdrawn, and the full civil corpus has been re‑annotated: 5,186 cases in a single v2 regime, 5,103 of them with determinate outcomes. Any figure still labelled v1 is superseded; the v2 figures are stable and citable, while the hypotheses built on them remain provisional.
What changed, in full
Correction — August 2026
Earlier versions of this page reported that Contradiction Debt predicts appellate outcomes at AUC ≈ 0.98. That figure was an artifact of the annotation pipeline: Nodes 7 and 8 were derived from the court’s disposition, and the outcome variable was in turn derived from those nodes. The association was true by construction. All predictive claims have been withdrawn pending a validation against independently coded dispositions. The structural findings — node failure distributions, contradiction-type patterns, and cross-domain comparisons — were described here as unaffected. That was wrong, and is corrected below.
Corpus rebuild COMPLETE — SOoL v2.0.0 released 6 August 2026
The same annotation rule that manufactured the predictive result also distorted the contradiction-type distribution, so the structural findings are not unaffected. Under v1, Recognition Failure (RF) was forced wherever Node 7 failed; it and RPF therefore appeared in roughly 64% of cases, while Correlativity Contradiction (CC) — the qualified-immunity pattern in which a court assumes a violation and bars the remedy anyway — was absorbed into RF and recorded just 3 times in 5,598 cases. The rule has been removed and the full civil corpus has been re‑annotated (v2).
All eight domains are rebuilt (First Amendment, Employment Discrimination, Administrative Law, Criminal Procedure, Immigration, Civil Rights §1983, Contract, Family Law), and the v1 figures do not survive: RF falls from ~64% to 2.9%, RPF to 5.5%, CC rises to 18.1% overall, and Norm Indeterminacy is now the most common type at 53.8%. The Node 7/Node 8 coupling that v1 recorded at φ = 0.953 was rule‑imposed; measured freely it is 0.75, with the two nodes diverging in 13.6% of cases. Most consequentially, the CD gap reverses direction: under v1 denied cases carried ~13× the debt of granted ones, whereas on v2 data granted cases carry slightly more (0.145 vs 0.126, a ratio of 0.87). Any contradiction-type frequency or CD figure still labelled v1 on this page is superseded.
The rebuild is finished. All eight domains are annotated under v2 — 5,186 cases in a single regime, with no v1 or masked‑pass rows remaining anywhere in the corpus. Of those, 5,103 resolve to a determinate outcome and carry every statistic on this site; the remaining 83 are extraction failures, excluded and disclosed rather than silently averaged in. Every tab now reports the same complete dataset; the transitional mixture described here in August is gone. Each record still carries an ann provenance code, and every published statistic names its annotation regime and its contradiction‑debt weight profile (core-v1).
The figures below are therefore stable and citable as figures. The hypotheses built on them remain provisional for reasons that completing the corpus does not touch — several rest on rule‑forced contradiction types, Node 7 failure is closely related to the case outcome without being independent of it, and the contradiction‑debt ordering depends on weights nobody has calibrated. Each card under Hypotheses states what would falsify it.
What is SOoL?
The Structural Ontology of the Law
treats every legal proceeding as a chain of eight structural conditions —
the Minimum Legal Chain (MLC).
All eight nodes must close for a claim to succeed. When nodes fail,
Contradiction Debt accumulates,
giving a structural account of where a claim breaks down
(masked annotation, n=4,934).
Key Terms
CD — Contradiction Debt:
structural stress score (0 = clean chain, ~0.55 = maximum observed). MLC — Minimum Legal Chain:
8 nodes from authority source to remedy. CT — Contradiction Type:
one of 13 named structural failure modes (CF, AI, JC, FM, RCL…). Protection denied / granted —
whether the claimant's right was recognized by the court.
What the Numbers Mean
CD ratio — how much more
structural stress cases with a derived outcome of "denied" carry than "granted" ones.
Both quantities are computed from the same node closures, so this is a measure of
internal consistency, not a predictive result. The 14.0× figure previously
quoted here was a v1 number produced largely by an annotation rule that has since
been removed; the corpus has since been fully re-annotated and the ratio has inverted — granted cases now carry marginally more debt than denied ones.
The live value is computed from the released v2 corpus. Mean CD by domain — which
areas of law carry the most structural stress on average. N7 ★ pivotal — the Legal
Effect node is the primary failure point in 67% of reversals. ⚖ Contested — cases flagged by adversarial audit
(Pass 3) as structurally ambiguous. 17–18% of high-CD cases across both corpora. The 67% and 17–18% figures above were measured under the v1
annotation regime. “High-CD” is defined in terms of a CD score whose
composition has changed, so both are provisional until the re-annotation completes.
MLC 8 Nodes
N1
Source of Authority
N2
Norm
N3
Actor in Role
N4
Triggering Facts
N5
Legal Act
N6
Target
N7 ★
Legal Effect pivotal
N8
Remedy
Contradiction Debt by Domain
closed
failed
partial
indeterminate
Click domain to filter Browse
Corpus Growth
cases added per month · 1990–present
Navy = monthly additions · ● Cumulative total
Derived Outcome Distribution ⓘ
Contradiction Type × Domain Heatmap
Click cell to filter Browse
SOoL · Structural Ontology of the Law · Koepsell (2026) · SEAL Lab, Texas A&M University ·
A Structural Ontology of the Law (Palgrave Macmillan, forthcoming) ·
Internal consistency check (SOoL v2.0.0, corpus re-annotation complete): CD/derived-outcome ratio 14.0:1
(protection_denied avg CD 0.252 vs. protection_granted 0.017)
Filters◀
Domain
Derived Outcome ⓘ
Contradiction Debt (CD)
0.00
0.60
Year Range
to
Contradiction Types
Node 3 — Role
— cases
Build structural queries across the corpus. Results appear in the Browse tab.
All conditions are combined with AND logic.
Structural Signature Query
Find cases with a specific MLC failure pattern — the core cross-domain structural query.
Cross-Domain Family Resemblance
The core validation query: find cases sharing a structural signature across
doctrinally unrelated domains. Select a contradiction type to see which
domains it appears in and how many cases share the structural signature.
CD Comparative Query
Compare mean Contradiction Debt across any two groups.
MLC Reference
Mean Contradiction Debt by Domain, 1990–2024 3-year rolling avg · click domain to toggle
CT co-occurrence (13×13 phi) click cell for interpretation
Normalized CT rates by domain
% of cases in domain with each CT active — click bar to filter
Click a cell for details.
CD distribution by outcome
Are denied/granted distributions separable?
Node failure rates by chain outcome
Which nodes most predict adverse outcomes?
Top structural signatures most frequent compound CT patterns · colored by avg CD
Structural impact of major SCOTUS decisions mean CD 2 years before vs 2 years after · affected domain only
Click a decision for details.
Corpus coverage — every collected domain
What has been collected, what has usable opinion text, and what is actually analysable. Domains absent from the statistics elsewhere on this site appear here with the reason.
Annotation quality & Node 7 / Node 8 coupling
Node 7 / Node 8 divergence by domain
Under v1 a rule forced N8 to fail whenever N7 failed, so divergence counted as an annotation error. That rule is gone: v2 assigns Node 8 on its own evidence, and divergence is now a finding, not a fault. Remedy failure almost always implies effect failure (98.9%), but effect failure leaves a remedy intact about a fifth of the time.
Annotation confidence by domain
High / medium / low / needs review
CD/outcome ratio by domain does 13:1 hold within each domain independently?
Correction — August 2026
Earlier versions of this page reported that Contradiction Debt predicts appellate outcomes at AUC ≈ 0.98. That figure was an artifact of the annotation pipeline: Nodes 7 and 8 were derived from the court’s disposition, and the outcome variable was in turn derived from those nodes. The association was true by construction. All predictive claims have been withdrawn pending a validation against independently coded dispositions. The structural findings — node failure distributions, contradiction-type patterns, and cross-domain comparisons — were described here as unaffected. That was wrong, and is corrected below.
Corpus rebuild COMPLETE — SOoL v2.0.0 released 6 August 2026
The same annotation rule that manufactured the predictive result also distorted the contradiction-type distribution, so the structural findings are not unaffected. Under v1, Recognition Failure (RF) was forced wherever Node 7 failed; it and RPF therefore appeared in roughly 64% of cases, while Correlativity Contradiction (CC) — the qualified-immunity pattern in which a court assumes a violation and bars the remedy anyway — was absorbed into RF and recorded just 3 times in 5,598 cases. The rule has been removed and the full civil corpus has been re‑annotated (v2).
All eight domains are rebuilt (First Amendment, Employment Discrimination, Administrative Law, Criminal Procedure, Immigration, Civil Rights §1983, Contract, Family Law), and the v1 figures do not survive: RF falls from ~64% to 2.9%, RPF to 5.5%, CC rises to 18.1% overall, and Norm Indeterminacy is now the most common type at 53.8%. The Node 7/Node 8 coupling that v1 recorded at φ = 0.953 was rule‑imposed; measured freely it is 0.75, with the two nodes diverging in 13.6% of cases. Most consequentially, the CD gap reverses direction: under v1 denied cases carried ~13× the debt of granted ones, whereas on v2 data granted cases carry slightly more (0.145 vs 0.126, a ratio of 0.87). Any contradiction-type frequency or CD figure still labelled v1 on this page is superseded.
The rebuild is finished. All eight domains are annotated under v2 — 5,186 cases in a single regime, with no v1 or masked‑pass rows remaining anywhere in the corpus. Of those, 5,103 resolve to a determinate outcome and carry every statistic on this site; the remaining 83 are extraction failures, excluded and disclosed rather than silently averaged in. Every tab now reports the same complete dataset; the transitional mixture described here in August is gone. Each record still carries an ann provenance code, and every published statistic names its annotation regime and its contradiction‑debt weight profile (core-v1).
The figures below are therefore stable and citable as figures. The hypotheses built on them remain provisional for reasons that completing the corpus does not touch — several rest on rule‑forced contradiction types, Node 7 failure is closely related to the case outcome without being independent of it, and the contradiction‑debt ordering depends on weights nobody has calibrated. Each card under Hypotheses states what would falsify it.
Plain Language Guide
When the Architecture of Law Predicts Its Outcomes
A layperson's guide to the SOoL structural jurimetrics study — no legal or technical background required
13
Contradiction Types
Weights were fixed before any data was collected, so the coding scheme could not be tuned to fit the corpus.
14.0×
CD Ratio (Denied / Granted) · v2 only
Computed live from the v2 rows only, so it no longer blends annotation regimes. A value below 1.0 means granted cases carry more contradiction debt than denied ones — the reverse of the v1 figure, which was produced by a rule that forced debt onto losses. The v2 re-annotation is complete, so this value is stable. It is not a predictive claim.
5,103
Federal Appellate Cases
8
Doctrinal Domains
"An architect asks whether the building is beautiful. A structural engineer asks whether it will stand. SOoL is the structural engineering approach applied to law."
The Core Idea: Structural Soundness vs. Doctrinal Beauty
Standard legal analysis asks: does the rule apply to these facts? SOoL asks something more fundamental: is the architecture of this legal claim structurally sound? A claim can have excellent facts and strong doctrine and still fail — if the structure holding it together is broken.
The Eight-Step Structural Checklist (The MLC)
MLC stands for Minimum Legal Chain — the minimal sequence of eight conditions that must all be satisfied for a legal claim to be structurally complete. Think of it as eight load-bearing walls. If any wall is missing or cracked, the structure is at risk.
N1 — Source of Authority
Is there a legitimate legal source — a statute, constitutional provision, or binding precedent — that authorizes this rule?
N2 — Norm
Is there a determinate, intelligible rule that specifies what conduct is required, permitted, or prohibited?
N3 — Actor in Role
Is there an identifiable person bearing a specific legal role — employee, officer, regulated entity — to whom the rule applies?
N4 — Triggering Facts
Are the facts that activate the rule actually present and accurately established? Fabricated or suppressed facts fail here.
N5 — Legal Act or Omission
Did the relevant conduct actually occur? Courts often disagree about whether particular behavior counts as the legally relevant act.
N6 — Target
Is there an identifiable party against whom the obligation runs, and do they have legal standing to be in the case?
N7 — Legal Effect ★ PIVOTAL
Does the right, duty, or legal consequence actually attach? This is where most claims ultimately succeed or fail — the recognition gateway.
N8 — Remedy
Is there an available remedy if the obligation is breached? A right without a remedy is structurally incomplete. In the current corpus, N8 failure following N7 failure is imposed by annotation rule rather than observed.
Contradiction Debt (cd_case): The Structural Stress Score
CD stands for Contradiction Debt — a single number that measures how structurally stressed a legal claim is. It is calculated by identifying which structural failure modes are active in the case and adding up their assigned weights. A CD of zero means the legal architecture is sound. The higher the CD, the more structural stress.
Which quantity is being reported
Three different measures have circulated under the single name “CD”. They are not interchangeable, and mixing them has produced published figures that cannot be compared. Each is now named explicitly wherever it appears.
cd_case
A weighted flow. The sum of the sool:cdWeight values of the contradiction types active in one case. Ceiling 1.52 under weight profile core-v1; observed mean 0.1375, observed maximum 0.7800. Every contradiction-debt figure on this page is cd_case.
ct_count
An unweighted count. How many of the 13 contradiction types are active in a case (0–13), ignoring weights. Corpus mean 1.41; 25.7% of cases carry none. The banded scale on the printed Quick Reference card (0 / 1–2 / 3–4 / 5–6 / 7+) is a ct_count scale — it was never a cd_case scale, and reading it as one puts all 5,103 cases in band “0”.
cd_stock
An accumulated stock. Debt carried by an institution over time, clamped to [0, 1], used by the SimLex simulator to trigger regime shifts. It is an accumulation of cd_case increments, not a cd_case value. No figure on this page is a cd_stock.
Because cd_case is a per-adjudication flow and cd_stock is an institutional accumulation, the corpus mean (0.1375) cannot be read against SimLex's incoherence thresholds (0.32–0.56) — those are thresholds on cd_stock. Establishing that bridge is open work, not a published result.
An important limitation. CD is not an independent predictor of case outcomes, and earlier versions of this page wrongly claimed that it was. In the current pipeline the recognition and repair nodes (N7, N8) are derived from the court's disposition by an annotation rule, and the reported outcome variable is derived from those same nodes. Any strong association between CD and outcome is therefore true by construction.
The corpus makes this visible. Restricting CD to the upstream nodes (N1–N6) — the part of the chain not touched by the derivation rule — separates outcomes at AUC = 0.5518, near chance. The same collapse shows up in the raw means:
Measure
Denied
Granted
Ratio
Full CD (includes N7/N8)
0.254
0.020
13.0×
CD restricted to N1–N6
0.019
0.014
1.4×
About 88% of all Contradiction Debt in this corpus comes from the two nodes derived from the disposition. Remove them and the headline separation falls from 13× to 1.4×. This is reported as a limitation, not as evidence against circularity — the cd_upstream field is now published alongside contradiction_debt so the check is reproducible.
Superseded — recomputation complete (SOoL v2.0.0, 6 August 2026). Every figure in this box was computed under the v1 annotation regime, in which an annotation rule forced Recognition Failure (RF) wherever Node 7 failed. That rule was the direct cause of the concentration described above: RF and RPF fired on roughly 64% of cases, and both attach to N7/N8. The rule has been removed and the entire corpus has been re‑annotated under v2. On the completed corpus RF falls from 64% to 3.0%, and the share of contradiction debt contributed by N7/N8 falls from about 77% to 27.1% — so the 13.0×, 1.4× and 88% figures above do not survive. The CD gap does not merely shrink, it reverses direction: granted cases carry marginally more debt than denied ones (0.145 vs 0.126, a ratio of 0.87). The superseded figures are left visible rather than deleted so the correction stays auditable. The replacement figures are published: see the Hypotheses tab, where P3 reports the reversed CD ordering and P8 reports that the debt weights add no discriminative power over a raw contradiction count.
This section previously reported four "independent validation tests" of CD's predictive power. They have been withdrawn. None of the four was independent of the annotation rule that derives the outcome variable from the node closures, so all four measured the same circularity rather than testing it.
SUPPORTED
That an eight-node chain can be applied consistently across eight doctrinal domains by independent annotation passes; that failures concentrate at identifiable nodes; and that those concentration patterns differ by domain in ways the coding scheme did not anticipate.
NOT SUPPORTED
That CD predicts how a court will rule. Establishing that requires comparing CD against the court's recorded disposition, coded independently of the node closures. That work is in progress; until it lands, no predictive claim should be drawn from this corpus.
Where Law Fails: One Shared Profile, Different Stress Points
Superseded — fully refuted 6 August 2026
This section previously reported three failure clusters, assigning each domain a distinct primary failure node: N2 for Contract and Family Law, N5 for First Amendment, Employment and §1983, N6 for Administrative Law, Immigration and Criminal Procedure. With the corpus complete, all three have now been tested and none survives.
N5 in its three assigned domains runs 8.8% against 8.5% elsewhere (p = 0.20); N6 runs 2.7% against 3.0% (p = 0.40). The N2 cluster could not be tested in August because Contract and Family Law had not been rebuilt; they now have been, and it fails too — Norm Indeterminacy fires in 59.7% of those two domains against 54.2% elsewhere, a ratio of 1.10. Statistically detectable on 5,103 analysed cases, but nowhere near a cluster, and Node 2 is not the most-failed node in a single domain. Node 7 is, in all eight. The accompanying bar chart displayed hard-coded v1 odds ratios and now computes live from v2 rows.
What the rebuilt data show instead is a shared failure profile — every domain breaks at the same place, in the same order — with domains distinguished not by where the chain fails but by which kind of contradiction shows up there. The chart below plots failure rates at the three formerly-claimed cluster nodes alongside Node 7; the flat N2/N5/N6 series against a uniformly high N7 series is the point.
D1 FIRST AMENDMENT — ROLE CONTRADICTION
RC3 in 25.3% of cases · 3.6× the rate elsewhere
The speaker holds two roles at once — public employee and citizen — and the same act of speech produces opposite legal effects through each. Garcetti v. Ceballos is the paradigm. The chain does not fail because the norm is unclear; it fails because the actor is two people at the same moment.
D3 ADMINISTRATIVE LAW — AUTHORITY INFLATION
AI in 14.1% of cases · 5.4× the rate elsewhere
The sharpest domain signal in the corpus. Authority Inflation fires at N1→N2: the decision-maker's will displaces the governing norm — the agency substitutes its own preference for the rule it was empowered to apply. Administrative law is where that substitution is most visible, even though it is not where the chain finally breaks.
Correction, 6 August 2026. This panel previously labelled AI “Authority Incompleteness” and glossed it as incomplete delegation. That was wrong. Both the ontology of record and the annotation prompt define AI as Authority Inflation — close to the inverse claim. The 14.7% measurement is unaffected, because the annotator worked from the correct definition; it was the reading published here that did not match what was measured.
D6 CIVIL RIGHTS §1983 — CORRELATIVITY
CC in 30.7% of cases · 1.8× the rate elsewhere
The qualified-immunity shape: the court assumes or acknowledges the violation and still bars liability. The right is recognised; the correlative duty is refused. Under v1 this pattern was absorbed into Recognition Failure and recorded just 3 times corpus-wide — it is visible only now that the forcing rule is gone.
D8 FAMILY LAW — JURISDICTIONAL CONTRADICTION
JC in 38.4% of cases · 1.7× the rate elsewhere
New with the completed corpus. Two authority sources issue competing norms over the same conduct — state against state under the UCCJEA, state against federal, tribal against both. Family law is where a claimant is most often governed by two systems at once. It also carries the corpus's highest mean contradiction debt (0.176), consistent with a domain in which the governing authority is itself contested.
D7 CONTRACT — NO SIGNATURE
Contract is the one domain with no contradiction type elevated at all (nothing above 1.2× at p < 10⁻⁴; Norm Indeterminacy runs 55.0% against 55.5% elsewhere, a ratio of 0.99). Employment Discrimination, Criminal Procedure and Immigration are likewise unmarked. That is a finding rather than a null: if every domain carried a signature, "domain-invariant architecture" would be doing no work. Half the corpus sits on the shared baseline with nothing distinctive, which is what the invariance claim predicts.
Domains 2, 4 and 5 (Employment, Criminal Procedure, Immigration) show no contradiction type elevated at a comparable margin; their profiles sit close to the corpus baseline. Note that RC3, CC and NI are rule-forced types — the annotation prompt compels them under stated conditions — so these rates measure how often those conditions obtain, not an unprompted discovery. Domains 1–8 all receive an identical prompt, which is what makes the comparison across them legitimate.
What This Means for Lawyers, Theorists, and Students
⚖ For Practitioners
The MLC gives a vocabulary for describing where a claim breaks down — which link in the chain carries the defect, rather than a global impression that a case is weak. It is a descriptive and diagnostic tool. It is not a predictive one, and nothing here should be used to price settlement risk or forecast how a court will rule.
📚 For Theorists
This corpus offers a structural vocabulary for comparing failure patterns across doctrinal areas that are usually studied in isolation. It takes no position on legal realism. An earlier version of this page claimed the data challenged realism by showing structure outpredicts judicial ideology; that claim rested on a circular outcome variable and has been withdrawn. Adjudicating it would require a predictive test against recorded dispositions, which has not yet been run.
🎓 For Students
IRAC tells you what the law says. The MLC tells you whether the structure holding the argument together is sound. You will spend your legal career applying rules — understanding the structural architecture beneath those rules makes you a better analyst of why some arguments win and others lose.
Why the Weights Are What They Are
Every structural failure mode carries a weight that determines how much it contributes to the Contradiction Debt score. These weights were assigned from first principles before any data was collected — they were not calibrated to predict outcomes. They reflect three theoretical criteria: how fundamentally the failure violates the BFO-typed constraint it targets; how far down the chain the failure occurs and how irreversible it is; and how difficult the structural problem is to repair without external intervention.
The weights range from 0.07 to 0.18 — a ratio of 2.57:1 between the lowest and highest. This spread is deliberate: narrow enough that all thirteen types remain informative, wide enough to reflect genuine theoretical differences in severity.
CODEWEIGHTNAME & PLAIN-LANGUAGE RATIONALENODE
RCL0.18Recognition Collapse — the most severe violationN6→N8
The highest weight because RCL destroys the symmetry of the legal relationship itself. The agent loses the capacity to hold rights while retaining the capacity to bear obligations — the same person is simultaneously unable to assert protections and can still be held to duties. This violates the fundamental Hohfeldian principle that rights and correlative duties must exist together. No other failure type destroys this ontological symmetry so completely. It also carries the heaviest repair burden: restoring someone's legal subjecthood requires institutional reconstitution, not just doctrinal clarification. Paradigm: an immigrant in removal proceedings who is regulated as an object of enforcement but cannot assert rights against removal.
CF0.15Conferral Failure — the chain has no foundationN1→N3
The second highest because it fails at the very first link of the chain and does so irreversibly within the proceeding. If the authority source failed to properly vest the actor in the required role, there is no legitimate actor — the entire subsequent structure has no bearer. You cannot fix this by clarifying the norm or reassessing the facts. The chain simply has no foundation. In patent law, this fires when the patent is found invalid: the USPTO failed to properly confer the right, so the authority source at Node 1 never properly existed. High repair burden, earliest possible chain position, complete structural void.
RC30.14Role Contradiction — incompatible obligations for the same actN3
The same agent bears two roles whose obligations are genuinely incompatible for the same act — not merely difficult to reconcile, but structurally irreconcilable. The structural problem is not vagueness but genuine incompatibility: no amount of normative clarification resolves it. The only structural resolution is role recharacterization, which typically requires a new legal framework rather than better application of the existing one. Garcetti v. Ceballos is the paradigm: Richard Ceballos simultaneously holds the citizen-speaker role (whose speech is protected) and the official-duty public employee role (whose work product is not) for the exact same memo. No interpretation of the Pickering/Garcetti framework can honor both roles for the same act. This is why RC3 is common in First Amendment corpus cases — the structural problem is endemic to the doctrinal architecture. (Frequency revised: 19.5% was measured under the v1 regime; with First Amendment now fully rebuilt under v2 it stands at 25.2% — higher than either earlier figure, and 4.1× the rate across the other rebuilt domains, making it the second-strongest domain signal in the corpus. Note RC3 is rule-forced: annotation RULE 2 compels it wherever two roles produce incompatible effects from one act, so this measures how often that configuration arises rather than an unprompted discovery. The rebuild is complete — see the notice at the top of this page.)
CC0.14Correlativity Contradiction — right acknowledged, duty refusedN7
Matches RC3 because both sever a Hohfeldian correlation, but CC does it at Node 7: the court acknowledges the right exists and was violated, then refuses to enforce the correlative duty through immunity or structural bar. The court is effectively saying: your right was violated, and the person who violated it faces no obligation to answer for it. This is ontologically impossible in a coherent legal system — a right without a correlative duty is not a right in any meaningful sense. The qualified immunity doctrine in civil rights cases is the paradigm: "Assuming a constitutional violation occurred, the right was not clearly established, so the officer is immune." Right acknowledged; duty refused. The 0.14 weight reflects that this is a recognized structural contradiction built directly into doctrine — not an accident but a deliberate structural choice with serious ontological consequences.
High weight but below CF and RCL because, under the v1 annotation rules, RPF almost always cascaded from Recognition Failure rather than firing independently. When Node 7 failed (right not recognized), Node 8 was stipulated to fail too — the claimant has no recognized right to repair. This cascade was imposed by annotation RULE 1, not discovered: the measured phi of 0.953 records rule compliance, not an empirical regularity. The rule has since been removed. Across the completed corpus, RPF no longer cascades from Node 7 and fires in 5.6% of cases rather than 64%; measured freely the Node 7/Node 8 association falls from 0.953 to 0.75, with the two nodes diverging in 13.9% of cases. That residual coupling is not symmetric, and the asymmetry is the interesting part: remedy failure almost always implies effect failure (98.9%), while effect failure leaves a remedy intact in roughly a fifth of cases. The 0.953 figure is a v1 artifact retained here only to document what the rule did. The passage below, on RPF firing independently, describes what v2 is now measuring directly. The weight is set high enough to reflect that blocked repair is genuinely serious, while acknowledging the cascade nature of most RPF findings. The important exceptions — where RPF fires independently — are the most structurally interesting cases: habeas corpus (the entire proceeding challenges prior repair adequacy), FRAND patent licensing (the remedy pathway is explicitly blocked by licensing obligations), and some qualified immunity cases where relief is theoretically available but procedurally inaccessible.
AI0.12Authority Inflation — personal will displaces the normN1→N2
Higher than NI and JC because AI is not a problem of vagueness or competing norms — it is a displacement of the institutional norm by personal will. When a decision-maker's preferences substitute for the governing rule, the chain loses its institutional grounding at the N1→N2 link. The authority source continues to exist formally, but the norm it issues is no longer traceable to that source — it is traceable to the individual's own disposition. This severs the relationship that makes norms institutional rather than personal. Post-Loper Bright, AI is expected to become more common in administrative and securities cases as agencies attempt to extend their authority through enforcement action rather than rulemaking — using prosecutorial discretion to create new norms without going through the statutory authority process.
RF0.11Recognition Failure — court denies the right (v1 frequency withdrawn)N6/N7
Lower than RCL and RPF despite appearing in roughly 65% of corpus cases. This reflects a core theoretical principle: frequency does not imply severity. The 65% figure is withdrawn. It was produced by v1 annotation RULE 5, which forced RF wherever Node 7 failed; RF was not the most common contradiction type, it was an automatic one. With the rule removed, RF stands at 3.0% across the completed corpus, and much of what it had absorbed is properly Correlativity Contradiction — CC, which v1 recorded just 3 times corpus-wide, now appears in 18.9% of cases and reaches 30.7% in §1983. The weight rationale below is unaffected — it never depended on the frequency — but the “most common CT” label above does not hold: the most common type on v2 data is Norm Indeterminacy at 53.8%. RF fires when the court denies the claimant's asserted right — the legal effect fails to be recognized. This is structurally serious and consequential, but no Hohfeldian correlation is violated: there is no acknowledged right without a correlative duty, because the right is simply denied. The court is operating coherently within the system — it has evaluated the claim and found it insufficient. The relatively modest weight reflects that RF is the normal operation of a legal system that evaluates claims and sometimes rejects them, rather than a fundamental ontological rupture.
FM0.11Fact Manipulation — the factual foundation is corruptedN4
Matches RF because both represent serious but structurally bounded failures. FM fires when the triggering facts are fabricated, suppressed, or made indeterminate — the factual foundation of the entire chain is corrupted. The paradigm is a Brady violation: the prosecution withholds evidence that would help the defendant, making the triggering facts structurally indeterminate for the defense. FM is weighted at 0.11 rather than higher because it is in principle remediable — the structural problem is an epistemic failure about facts that, if correctly established, might support a coherent chain. This is qualitatively different from RCL (which requires reconstituting legal subjecthood) or CF (which voids the authority chain). A retrial with correct facts can repair FM. Nothing comparable repairs RCL.
JC0.10Jurisdictional Contradiction — two valid norms, incompatible claimsN2→N6
Moderate weight because both competing norms are individually legitimate — the problem is their coexistence, not their individual validity. Two authority sources each issue a norm that claims to govern the same conduct, and the norms pull in incompatible directions. This is structurally serious but in principle resolvable through conflict-of-laws analysis, preemption doctrine, or legislative clarification. Neither norm is invalid; the system simply has not yet resolved which one governs. Federal/state conflicts and civil/criminal regime overlaps are the paradigm cases. JC is more remediable than RC3 (where a single actor's roles are irreconcilable) because the resolution mechanism — deciding which norm prevails — is well-established in legal doctrine even if applying it is difficult.
SE0.10Self-Undermining Effect — the right destroys the system that recognizes itN7
Matches JC because both involve systemic rather than case-level structural problems. SE fires when the legal effect at Node 7, if fully realized, would destroy the authority source at Node 1 that generates it — a self-referential collapse. The realized right undermines the institutional conditions for its own recognition. This is ontologically precise: the effect's realization destroys the continuants that ground its own authority. SE is weighted moderately because it is rare, highly context-dependent, and philosophically complex — it requires a case where recognizing the right would destabilize the system's capacity to recognize any rights. Cases involving claims against the structural prerequisites of judicial authority itself are the paradigm.
PC0.09Procedural Contradiction — the norm blocks its own fulfillmentN2→N5
The act required by the norm cannot be performed through the norm's own procedure — a self-blocking realization pathway. Lower weight than NI (which questions what the norm is) and JC (which involves competing authority sources) because PC is a procedural rather than substantive failure. The substantive norm may be perfectly sound; the procedural mechanism that implements it is defective. This is remediable through procedural reform without touching the substantive norm itself. An example: a statute that requires claimants to exhaust administrative remedies before suit, where the administrative remedy process is itself unavailable or takes longer than the statute of limitations allows. The norm is valid; the procedure defeats it.
TC0.08Temporal Contradiction — norm and facts misaligned in timeN2↔N4
Second-lowest weight because temporal misalignment is the most readily remediable structural failure. TC fires when the norm and triggering facts are misaligned in time — either the norm was not yet in force when the facts arose (retroactivity) or the facts predated the norm's existence (ex post facto). The structural fix is obvious and requires no fundamental ontological reconstitution: the norm simply should not apply to these facts. TC requires temporal boundary clarification, not institutional reconstitution. This is qualitatively different from RCL or CF. A statute that purports to apply retroactively has a structural TC problem, but the remedy — holding it inapplicable — is clean and leaves the rest of the legal system undisturbed.
NI0.07Norm Indeterminacy — the governing rule cannot be specified (lowest)N2
The lowest weight for two reasons. First, courts routinely resolve NI in the very opinion that identifies it — they specify the indeterminate norm and apply it, turning an NI into a closed node within the same proceeding. The Garcetti majority reservation (declining to decide academic speech) is a rare case of genuine persistent NI; most NI findings are resolved by the time the opinion issues. Second, NI is often not a failure of the legal system but a recognition that the system has reached the frontier of its current normative development. A norm that is indeterminate today may become determinate through legislative or judicial specification tomorrow. This is qualitatively different from RCL, which strips rights that the system cannot simply declare back into existence, or CF, which voids an authority chain that must be rebuilt from the institutional ground up. NI is structurally significant — an indeterminate norm cannot activate the chain — but it is the most remediable of the thirteen failures.
CD WEIGHT SPECTRUM — CLICK ANY ROW ABOVE TO EXPAND RATIONALE
NI0.07
TC0.08
PC0.09
JC0.10
SE0.10
RF0.11
FM0.11
AI0.12
RPF0.13
RC30.14
CC0.14
CF0.15
RCL0.18
Weights range from 0.07 (NI — most remediable) to 0.18 (RCL — most severe). The 2.57:1 spread keeps all thirteen types informative while reflecting genuine theoretical differences. These weights were fixed before any data was collected and have not been tuned to the corpus — a genuine methodological safeguard against overfitting. Note that the 14.0:1 CD/derived-outcome ratio does not validate them: both sides of that ratio are computed from the same node closures, so it reflects internal consistency rather than empirical confirmation.
Quick Glossary
AUCArea Under the Curve — predictive accuracy measure. 0.5 = coin flip; 1.0 = perfect.
BFOBasic Formal Ontology — a philosophical framework for classifying what kinds of things exist.
CDContradiction Debt — a numerical score measuring structural stress in a legal chain.
CTContradiction Type — one of thirteen specific structural failure modes indexed to MLC links.
IRACIssue, Rule, Application, Conclusion — standard legal analysis framework taught in law school.
MLCMinimum Legal Chain — the eight-node structural checklist for legal claim completeness.
RFRecognition Failure — court denies the claimed right exists or applies. (Previously labelled the most common contradiction type; that frequency was an artifact of a removed annotation rule and is withdrawn pending the v2 rebuild.)
RPFRepair Failure — remedy is structurally blocked. Fires on its own textual evidence (5.5% of v2 cases); the v1 cascade from Recognition Failure was a rule artifact and has been removed.
RC3Role Contradiction — same person bears two roles with incompatible obligations. Garcetti paradigm.
CCCorrelativity Contradiction — right acknowledged, correlative duty refused (qualified immunity).
NINorm Indeterminacy — court cannot specify what rule governs (scope uncertainty, not application).
SOoLStructural Ontology of the Law — the framework (Koepsell, Palgrave Macmillan, forthcoming).
RCLRecognition Collapse — subject's capacity to bear rights stripped while obligations continue.
FMFact Manipulation — triggering facts fabricated or suppressed. Brady violation is the paradigm.
A Structural Ontology of the Law · David R. Koepsell · SEAL Lab, Texas A&M University · Palgrave Macmillan, forthcoming [email protected] · seal.tamu.edu/legal-kernel · n=5,103 v2 cases · 1990–2024
SOoL Criminal Norms Corpus
Structural Ontology of Criminal Law
Loading...
—
Total Cases
6
Offence Categories
—
CD Ratio (Rev/Conv)
—
% of Target (3,000)
Inverted perspective:
In this corpus the state is the claimant asserting criminal liability.
protection_granted = conviction upheld ·
protection_denied = reversal or acquittal.
High CD predicts reversal, not conviction.
CD by Offense Category
Defense Type → CT Profile
Mens Rea Gradient — CD
Corpus Growth
Browse Cases
Structural Ontology of Criminal Norms ·
David R. Koepsell · SEAL Lab, Texas A&M University ·
loading...
OWL Inference Engine
Structural Reasoner
Loading...
—
Cases Analyzed
—
Inconsistencies
—
Quality Flags
—
Novel Findings
How this works:
Each annotated case is evaluated against the SOoL structural axioms derived from the OWL ontology.
Inconsistencies are structurally impossible annotations (e.g., conviction but N7 failed).
Quality flags are suspicious patterns that may indicate annotation errors.
Novel findings are structurally interesting cases that warrant closer examination.
Paste any legal text — opinion, oral argument transcript, brief — and receive a full SOoL structural annotation: MLC node closures, active Contradiction Types, CD score, and outcome prediction.
Anthropic API KeyKey stored locally in your browser — never sent to our server
MLC Node Closures
Active Contradiction Types
Structural Analysis
Adversarial Assessment (Pass 3)
Corpus Comparison
SOoL Corpus 2.0.0 — released · hypotheses remain provisional
This is corpus version 2.0.0, the first release in which every in-scope domain has been annotated under the v2 regime. The v1 forcing rules (RULE 1 / RULE 5, which manufactured Recognition Failure from a Node 7 failure) are gone, and no v1 or masked-pass rows remain in the analysis scope: 5,103 cases across 8 domains, weight profile core-v1, ontology sool_bfo_mlc_core.ttl sha256 b8a0160a64afa0d5…. Figures on this page are stable and citable as figures.
What the reconstruction bought. The v1 corpus could not support these questions at all: two annotation rules manufactured Recognition Failure and Repair Failure from any Node 7 failure, so the headline results measured rule compliance rather than law. Removing them and re-annotating every case has produced a corpus that is internally consistent, single-regime, and 98.4% blind-annotated — the disposition was withheld from the annotator on almost every case. Findings drawn from it now stand or fall on their own evidence. Several already have: the cross-domain profile correlation rose when the final two domains were added, and the one claim that failed (the three-cluster finding) was refuted on the record rather than quietly dropped.
The hypotheses below are still hypotheses, and each card states what would falsify it — that is what makes them usable, not a hedge. The remaining limits are specific and tractable: several contradiction types are rule-forced by the prompt, Node 7 is non-independent of the case outcome, and the debt weights are uncalibrated — P8 below now quantifies what that last one costs. None of these needs more data; they need targeted tests, which a stable corpus finally makes possible.
last v2 annotation written: 2026-08-06T07:12:51Z · regenerate with gen_provisional.py --splice
Hypotheses suggested by the partial v2 data
P1Failure concentrates at the terminal nodes, in the same order, in every domain.
Across all 8 analysed domains the pooled failure rate is 61.3% at Node 7 (Legal Effect) and 48.7% at Node 8 (Remedy), against at most 2.2% at the three constitutive nodes (N1 Authority 0.3%, N3 Actor-in-Role 0.5%, N6 Target 2.2%). The eight-node failure profile correlates across domain pairs at mean Spearman ρ = +0.960 (min +0.898, 28 pairs), and the three most-failed nodes are the same three, in the same order (N7 > N8 > N5), in all 8 domains. Within this corpus, chains do not break where they are built; they break where they are supposed to bind.
Scope: this is a fact about the appellate tier, not about law. The near-floor upstream rates (N1 0.3%, N3 0.5%) cannot be read as evidence that authority and role rarely fail, because an appellate corpus cannot show upstream failure. A case in which authority or actor-in-role is genuinely contested is resolved below, settles, or is never filed; what reaches a court of appeals has already had its constitutive nodes conceded by both sides. The corpus pre-filters for terminal stress — the same selection logic that the SCOTUS-tier backtest exhibits one level further up. The claim P1 can support is that conditional on reaching appellate review, failure concentrates terminally and in a domain-invariant order. Whether that ordering is a property of legal structure or of the filter is not decidable from this corpus, and the test that would decide it is a district-court or agency-adjudication sample where upstream contest actually appears. Until that sample exists, P1 is a tier-scoped finding.
domain
n
N1
N2
N3
N4
N5
N6
N7
N8
D1 First Amendment
641
0
2
0
3
5
1
58
47
D2 Employment Discriminat
655
0
1
0
7
10
1
57
47
D3 Administrative Law
722
0
3
1
3
11
4
59
43
D4 Criminal Procedure
619
0
0
0
1
6
1
70
58
D5 Immigration
656
0
1
0
4
12
2
59
47
D6 Civil Rights §1983
657
1
3
1
3
10
3
66
53
D7 Contract
640
0
2
0
3
10
2
60
45
D8 Family Law
513
1
4
1
4
9
4
61
51
POOLED
5103
0
2
0
3
9
2
61
49
What would sink it. Node 7 failure and protection_denied agree on 90.2% of cases. That is not definitional, and the reason matters. 98.4% of the corpus was annotated blind: the opinion is truncated at the first disposition signal and outcome words are redacted, so the annotator could not read the result off the page. Among those 5,022 cases concordance is 90.0%; among the 81 where the disposition was visible it is 97.5%. Seeing the outcome does push agreement up, but blind annotation still reaches 90.0%. The right objection is therefore not tautology but non-independence: Node 7 and chain_outcome are two readings of one text by one annotator, so their agreement is internal consistency rather than corroboration. The non-tautological content is the rest of the ordering (N8, then N5) and the near-floor upstream rates, neither of which is fixed by the outcome. Restricting to the 2,940 cases not flagged for review raises N7 failure to 69.1%, so the pattern is not an artifact of low-confidence annotations.
P2Domains share one contradiction ranking but differ by a few sharply elevated types.
The 13-type contradiction profile is near-identical across domains — mean pairwise ρ = +0.914 (min +0.769); restricted to the 9 types not compelled by a mandatory inference rule (CF, AI, PC, TC, FM, RF, RCL, SE, RPF), still +0.879. Against that shared baseline, a small number of types spike in exactly one domain:
type
domain
rate
95% CI
elsewhere
ratio
p
AIAuthority Inflation
D3 Administrative Law
14.1%
11.8–16.9
2.6%
5.4×
4e-45
RC3Role Contradiction⚠RULE 2
D1 First Amendment
25.3%
22.1–28.8
6.9%
3.6×
2e-50
RFRecognition Failure
D6 Civil Rights §1983
5.6%
4.1–7.7
2.6%
2.2×
4e-05
Reading: doctrinal domains are not structurally distinct architectures; they are the same architecture under different characteristic stress. Domain identity appears to be carried by a handful of elevated types rather than by a different ranking — which is the SOoL domain-invariance claim in its strongest testable form.
One of these elevations is a prediction the framework made in advance, and it landed. Authority Inflation is the contradiction the ontology treats as master: a legal act that claims more authority than was conferred on it. If that reading is right, AI should concentrate in the one domain whose contested object is the scope of conferred authority. It does. AI fires in 14.1% of Administrative Law cases against 2.6% everywhere else — a 5.4× elevation (χ² = 199, p = 4e-45, n = 5,103). AI is not a rule-forced type: unlike RC3, NI, JC, CC, nothing in the v2 system prompt tells the annotator when to assign it, and domains 1–8 all receive the same prompt with no domain-specific hints. So this is not a domain-difference observed after the fact — it is the placement the master-contradiction claim commits to, tested against a corpus that had no way to know about it. It is the strongest confirmatory result on this page. Its limit is that a single hit on a single domain is one observation: the framework should be made to name the expected primary type for each remaining domain before the next corpus cut, so the next test is scored against a pre-registered prediction rather than a retrofitted one.
What would sink it. Types marked ⚠ are rule-forced: the v2 system prompt compels their assignment under stated conditions (RC3 RULE 2, NI RULE 3, JC RULE 4, CC RULE 6), so their rates measure how often those conditions obtain, not an unprompted discovery. The elevations are nonetheless comparable across domains because domains 1–8 receive an identical system prompt with no domain-specific hints (per-domain supplements exist only for domains 9–11, which are outside this cut). A high ρ among 13 points is also easy to obtain when one type dominates; the unforced-only ρ is the figure to trust.
P3Contradiction Debt tracks how contested a case was, not which way it came out.
Ordered by mean CD, outcomes line up by depth of merits engagement, not by winner:
chain outcome
n
mean CD
95% CI
mean tag count
95% CI
moot
81
0.0140
0.0023–0.0256
0.148
0.028–0.268
dismissed
186
0.1009
0.0861–0.1157
1.070
0.920–1.219
protection_denied
2822
0.1256
0.1213–0.1299
1.237
1.198–1.277
protection_granted
570
0.1447
0.1358–0.1536
1.519
1.431–1.607
remanded
656
0.1512
0.1433–0.1590
1.691
1.611–1.770
partial
788
0.1850
0.1768–0.1931
1.918
1.843–1.992
Unweighted replication. The ordering above is produced by the core-v1 weight profile, and P8 shows those weights are less discriminative than a raw count of active contradictions. So the same table is recomputed with the weights removed — identical rows, identical denominators, per-case statistic changed from Σw to a plain count. The rank order is unchanged (Spearman ρ = +1.00 between the two orderings), so the contestedness reading does not depend on the weight profile. The count also explains more of the between-outcome variance than CD does (η² = 0.0791 vs 0.0555), and separates granted from denied more sharply (d = +0.261 vs +0.166, t = +5.72, p = 2e-08). Inert weights would give parity on both measures; worse-than-parity means some weights are pulling against the signal. See P8.
Cases that never reach the merits carry almost no debt; cases that split carry the most. Merits-reached (n = 4,180, mean 0.1394) vs not-reached (n = 267, mean 0.0745): Welch t = +10.2, p = 2e-21. Critically, protection_granted (0.1447) is not belowprotection_denied (0.1256) — the difference runs the other way (t = +3.79, p = 2e-04, Cohen’s d = +0.17). This reverses the direction reported under v1, where high CD was read as predicting denial.
What would sink it. The v1 CD gap was manufactured by deleted RULE 1/RULE 5, which forced Recognition Failure and Repair Failure wherever Node 7 failed — mechanically loading debt onto losses. The v2 direction is the one to test, but note the granted/denied difference is small (d = +0.17) and the whole CD range is compressed: corpus mean CD is 0.1375 against a ceiling of 1.52 — 9% of the scale — with 1.41 contradictions per case on average and 25.7% of cases carrying none at all. Every CD figure on this page uses weight profile core-v1 (the 13 sool:cdWeight values in sool_bfo_mlc_core.ttl, summing to a ceiling of 1.52). CD is a weighted sum, not a count of contradictions, so the ordering of outcomes above is a function of those weights and could move under a different profile. No empirical calibration of the weights exists; they have not been varied to test whether this ordering is robust to them. Re-interpreting CD as a contestedness index rather than a merits predictor is a reframing of the construct and needs the completed corpus before it is asserted.
P4Remedy failure presupposes effect failure — but not the reverse.
The N7/N8 relation is strongly asymmetric. When Node 8 fails, Node 7 has almost always failed too: P(N7 fail | N8 fail) = 98.8% (2,452/2,483). But the converse is much weaker: P(N8 fail | N7 fail) = 78.4% (2,452/3,128) — in 676 cases the claimant’s asserted effect did not attach yet a remedy survived, against only 31 cases with the opposite divergence. That is a near one-way conditional: remedy is structurally downstream of effect, and the chain has a direction that the ontology asserts but had not previously been measured.
This is the strongest result on the page, and the reason is the shape of what was instructed. The ontology does not merely say N7 and N8 are related; it says remedy is downstream of effect, which is a claim about direction and therefore a claim that can come out backwards. The v2 prompt instructs the coupling and then explicitly invites divergence in both directions, with a worked example of each. So the annotator was free to produce a symmetric split, and a symmetric split is what an instructed correlation with no underlying ordering would produce. What came back is 676 against 31. Under a null of no direction — divergences equally likely to fall either way — that imbalance has probability 4e-159. Most findings on this page are measurements of a corpus; this one is a measured confirmation of a structural commitment the framework made before the corpus existed, on a dimension the prompt left free.
Prompt-variant control. "The prompt left it free" is an argument, not a measurement, so it was measured. One sample of cases was re-annotated under three system prompts identical in every respect except the directional sentence: the production wording, the sentence removed, and the sentence inverted so the annotator is told Node 7 usually follows Node 8.
prompt arm
n
effect failed, remedy survived
remedy failed, effect survived
p vs 1:1
production, verbatim
145
19
0
4e-06
directional sentence removed
141
39
0
4e-12
directional sentence inverted
141
31
1
2e-08
The inverted arm is the load-bearing one: a prompt actively arguing for the opposite direction still returns 31 against 1, p = 2e-08. A finding that survives an instruction to find the reverse is not a finding the instruction produced. The removed-sentence arm is the more interesting one: it yields 39 divergences where the production prompt yields 19 on the same cases, all in the same direction. The coupling sentence was suppressing divergence, not creating it, so the published 676:31 understates the asymmetry rather than manufacturing it. What the control cannot rule out is a disposition the model carries independently of this prompt — that would need a second model or a human-coded subsample, and neither exists yet. Reproduce with p4_prompt_control.py --report.
What would sink it. The v2 prompt states that "Node 8 usually follows Node 7," so the coupling is partly instructed. The prompt does not instruct the asymmetry — it explicitly invites divergence in both directions and gives examples of each (independently available relief; independently blocked relief). The 22:1 imbalance is therefore the finding, not the correlation. Distinguishing a real ordering from an anchoring effect needs a prompt-variant control that the current pass does not include.
P5Forum may matter more than doctrine.
Node 7 failure varies more across courts than across doctrinal domains. Between the 12 courts with n ≥ 60 the rate spans 46.5%–74.7% (28.2 points; Cramér’s V = 0.166, p = 3e-24). Between the 8 domains it spans only 56.8%–70.3% (13.5 points; V = 0.088). The court effect persists within domains in 5 of 8 domains tested separately, so it is not merely a docket-composition artifact. If this holds, the domain-invariance thesis gains an unexpected companion: structural failure is roughly domain-independent but markedly forum-dependent.
What would sink it. Circuits differ in what reaches them and in what plaintiffs file there; this is a selection effect at least as much as a judicial one, and nothing here separates the two. Both V values are small in absolute terms. The comparison also inherits P1’s problem — if N7 largely encodes "claimant lost," this may restate known circuit-level affirmance differences rather than reveal anything structural. Needs the full corpus and a case-mix control.
The version that is not about outcome. The objection above is fatal as stated: N7 failure is close to "the claimant lost," and a forum effect on N7 is circuit affirmance variation in SOoL vocabulary. So the same test is run on a variable that cannot be outcome, because outcome is held fixed. Among the 3,128 cases whose effect node has already failed — every one of them a loss at N7 — the remedy node still closes in 9.8%. Whether it does varies by court from 6.1% (ca7, n = 380) to 19.1% (cadc, n = 236) across 12 courts with n ≥ 60 (χ²(11) = 41.9, p = 2e-05, V = 0.116, n = 3,108). The same variable also varies by domain (V = 0.100, p = 6e-05), and the two are confounded — the D.C. Circuit and Administrative Law are largely the same cases — so a case-mix control is still required before the forum reading is preferred to the subject-matter one. But the quantity itself is structural rather than dispositional: it asks whether a forum treats a broken legal effect as repairable, on a subset where the claimant has already lost. That is the P5 claim worth defending, and it is not restatable as an affirmance rate.
withdrawnP6 — Norm Indeterminacy is rising over time.
Withdrawn as a finding; the numbers are kept here for the record. Three reasons, any one of which is sufficient. First, it contradicts H6 on the same page, which reports the temporal profile as stable — a reader reaches both and has no way to reconcile them, and the disagreement is not a substantive dispute but an artifact of two different aggregations of the same variable. Second, NI is rule-forced and its trigger clause, "the norm’s application to these facts is genuinely contested," is close to a definition of appellate litigation, so a rate that rises tracks drift in the annotator’s reading of "contested" at least as readily as anything about courts. Third, and decisively, NI fires at 42.4% when Node 2 is closed, 97.2% when Node 2 is partial, 23.7% when Node 2 is indeterminate. A tag named for indeterminacy of the norm that fires at radically different rates across the norm node’s own closure states — and near-ceiling in one of them — is not measuring what it names. That is a construct-validity failure, not a caveat, and it disqualifies the trend regardless of its slope. Reinstating P6 requires re-specifying the NI trigger and re-annotating; it is not fixable by a control.
Norm Indeterminacy is the modal contradiction overall (55.4% of cases) and its rate climbs monotonically by decade of decision, as does mean CD:
decade
n
NI rate
N7 fail
mean CD
1990s
832
48.7%
52.4%
0.1187
2000s
1156
50.1%
65.7%
0.1301
2010s
1770
57.2%
62.0%
0.1403
2020s
1310
62.0%
62.3%
0.1518
Year of decision vs NI: Spearman ρ = +0.109, p = 9e-15 (n = 5,068). The trend survives a within-domain control — NI rises from first to last measurable decade in 6 of 8 domains taken separately, so it is not driven by the corpus’s domain mix shifting over time. Candidate reading: appellate courts increasingly decline to specify the governing rule, deciding on narrower or more contested grounds.
What would sink it. NI is rule-forced (RULE 3) and its trigger includes the broad clause "the norm’s application to these facts is genuinely contested" — which is close to a description of appellate litigation as such. Its construct validity is also loose: NI fires at 42.4% when Node 2 is closed, 97.2% when Node 2 is partial, 23.7% when Node 2 is indeterminate, i.e. largely independently of whether the norm node itself closed. A rising rate may track drift in what counts as "contested" rather than anything about the courts. This is the weakest card here and the most likely to be withdrawn.
P7The same failure shapes recur across every domain, not just the same ranking.
P1 and P2 compare domains on aggregate profiles. This is the stronger claim the ontology actually makes: a structural signature — the exact set of node/type failures a case exhibits — should be a recognisable object that recurs independently of doctrine. It does. The corpus contains 210 distinct signatures, of which 32 appear in at least 6 of the 8 domains with n ≥ 20, and 23 appear in all 8. Those broadly recurring shapes account for 63.3% of all cases — so a clear majority of the corpus fails in a way that some other area of law also fails in.
What would sink it. Signatures are built from the same node closures and contradiction types as everything else, so this inherits their dependencies — including the rule-forced types. A signature is also a set, not a sequence: it records which failures co-occurred, not how one produced another, so recurrence is not evidence of a shared causal path. The count of distinct signatures is bounded by the taxonomy, so some recurrence is guaranteed by construction; the reportable part is the share of cases carrying a widely-shared shape, not the existence of shared shapes.
P8The contradiction-debt weights do not earn their keep.
Contradiction Debt is a weighted sum, so it should separate outcomes better than simply counting how many contradictions a case carries. It does not. Separating denied from granted cases, a raw count gives Cohen’s d = -0.261 (p = 2e-08), while CD under profile core-v1 gives -0.166 (p = 2e-04) — the weighted measure is no better than the unweighted one. The 13 weights are theoretically motivated but have never been empirically calibrated, and on this corpus they add nothing discriminative.
What this does and does not mean. It does not show the weights are wrong — CD was never designed as an outcome predictor, and a measure of structural burden need not track who won. It does mean that any published figure resting on the weighted scale is resting on an unvalidated choice. Both effects are small in absolute terms.
Not merely inert — actively worse. Weights that carried no information would give parity with the count. 0.166 against 0.261 is worse than parity, which means some weights are pulling against whatever signal exists rather than simply failing to add to it. P3 now reports the same table on unweighted counts: the outcome ordering is unchanged (ρ = +1.00), so the ordering is not a weight artifact, but the count explains more of the between-outcome variance than CD does (η² = 0.0791 vs 0.0555).
Which weights, and against what. The right calibration target is not outcome — calibrating CD against who won would turn it into the outcome predictor this rebuild has spent its effort establishing it is not. It is repair adequacy: among the 3,128 cases whose effect node has already failed, does the remedy node still close? Every case in that subset lost at N7, so the question is purely structural. Per-type point-biserial correlations with remedy survival, against the core-v1 weight each type carries:
type
core-v1 weight
n in subset
r with remedy survival
p
AIAuthority Inflation
0.12
103
+0.313
4e-72
PCProcedural Contradiction
0.09
191
+0.285
2e-59
FMFact Manipulation
0.11
183
+0.133
7e-14
NINorm Indeterminacy
0.07
1585
+0.097
6e-08
TCTemporal Contradiction
0.08
119
+0.013
0.458
JCJurisdictional Contradiction
0.10
605
+0.005
0.782
RC3Role Contradiction
0.14
219
-0.040
0.026
RFRecognition Failure
0.11
132
-0.058
0.001
RPFRepair Procedure Failure
0.13
244
-0.076
2e-05
CCCorrelativity Contradiction
0.14
795
-0.130
2e-13
The types split by sign, and the split is not the one the weights encode: AI, PC, FM, NI are associated with the remedy surviving a failed effect, while RC3, RF, RPF, CC are associated with it failing too (TC, JC show no association either way). Rank-correlating the core-v1 weights against damage-to-repair gives ρ = +0.518 (p = 0.125) — the weights are close to uncorrelated with the one structural quantity they might have been calibrated against. Notice also that this sign split lines up with the families P9 recovers from co-occurrence, which is two independent measurements agreeing on the same partition.
noteFurther structure worth testing.
Contradictions are not independent of each other. Some co-occur far more than chance allows: AI+PC at 3.0× expected, RC3+CC at 2.0× expected, PC+FM at 1.8× expected. Others actively exclude one another: FM+CC at 0.46×, PC+CC at 0.37×. If the thirteen types were independent tags this would not happen. This observation has been promoted to P9 above, where the full matrix is tested pair by pair, clustered into families, and checked for stability — it is the corpus test of the claim that the thirteen types are derived from a smaller kernel rather than enumerated.
Remedy survival varies by domain and by forum. Across the corpus a remedy survives Node 7 failure in 21.6% of cases (P4), but the rate runs from 17.9% in D4 Criminal Procedure to 29.1% in D3 Administrative Law. P5 now runs the same variable across courts, because remedy survival conditional on effect failure is the one forum comparison that cannot be restated as an affirmance rate. Whether either split reflects genuinely different remedial architecture or different pleading practice is untested.
The temporal trend is broader than Norm Indeterminacy. The withdrawn P6 reported NI rising by decade; the total number of contradictions per case rises too (Spearman ρ = +0.111, p = 2e-15), while the rate at which claimants lose does not move materially (ρ = +0.029). Cases are being recorded as structurally more contested over time without the outcomes changing to match — which is either a real trend in adjudication or drift in what the annotator treats as contested, and those two are not yet separable.
P9The thirteen contradiction types are not independent tags: they attract and repel each other in families.
If the types were thirteen separate labels applied to thirteen separate phenomena, every pair would sit at observed/expected = 1.0 — co-occurring exactly as often as their marginals imply. If instead they are derived from a smaller set of kernel primitives, two things follow that independence does not predict: types sharing a primitive should co-occur above chance, and types whose primitives are incompatible should co-occur below it. Mutual exclusion is the load-bearing half — independent tags have no mechanism to repel. Across 45 testable pairs among the 10 types with n ≥ 50, 16 depart from independence at p < 0.001: 12 attracting and 4 repelling.
pair
obs
exp
obs/exp
χ²(1)
p
AI + PCAuthority Inflation / Procedural Contradiction
NI + RPFNorm Indeterminacy / Repair Procedure Failure
127
158.0
0.80×
14.0
2e-04
FM + CCFact Manipulation / Correlativity Contradiction
31
66.9
0.46×
24.8
6e-07
PC + CCProcedural Contradiction / Correlativity Contradiction
31
83.9
0.37×
44.2
3e-11
AI + CCAuthority Inflation / Correlativity Contradiction
15
40.9
0.37×
20.3
7e-06
Recovered families. Average-linkage clustering on the lift matrix, cut at 4: {AI, FM, PC} | {CC, RC3} | {JC, NI, RF} | {RPF, TC}. This partition is recovered in 6 of 10 resamples (two split-halves and 8 leave-one-domain-out runs, 60.0%). Pairs that never separate in any run: AI+FM, AI+PC, CC+RC3, FM+PC. Forcing in the 3 near-vacant types (CF, RCL, SE) gives a different partition ({AI,JC,RCL} | {CC,RC3} | {CF,FM,NI,PC,TC} | {RF,RPF,SE}), so the family structure below is a claim about the 10 types with real support, not about all 13.
What this tests, and what it does not. The framework holds that the thirteen types are derived theorems over a smaller kernel, not a flat list. That is a falsifiable claim and this is the corpus test of it: the derivation predicts which types should cluster, and the recovered families either match the primitive-sharing structure or they do not. The prediction cannot be scored here — the kernel-primitive-to-type mapping and the A/B/C/D strata are not in this repository, so the families above are reported as an unmatched empirical partition. Supplying that mapping turns this card from a description into a confirmation or a refutation, and it is the single cheapest confirmatory test available on the current corpus. Two caveats hold either way. Lift is sensitive to marginal prevalence, and NI alone appears in 55.4% of cases, so pairs involving it have little room to move upward. And co-occurrence is measured within a single annotator pass: two types that one prompt tends to emit together will look coupled whether or not they share a primitive. The exclusions are the safer evidence, because a prompt that over-emits a pair inflates lift but has no comparable mechanism for suppressing one below chance.
governanceTwo artifacts in circulation state contradiction debt on scales this corpus cannot produce. This is a measurement incommensurability, not a labelling nuisance.
Three distinct quantities have circulated under the one name “CD”: cd_case (the weighted per-case flow reported throughout this page, ceiling 1.52 under core-v1), ct_count (the unweighted count of active types, 0–13), and cd_stock (an institution's accumulated debt over time, clamped to [0, 1], used by the SimLex simulator). The two problems below are both consequences of not distinguishing them.
1. The Quick Reference bands are a count scale; the pipeline computes a weighted sum. The card in circulation bands debt as 0 / 1–2 / 3–4 / 5–6 / 7+ — a ct_count scale. Under profile core-v1, cd_case is Σw over active contradictions and its arithmetic ceiling is the sum of the 13 sool:cdWeight values, 1.52. Observed mean is 0.1375 and observed maximum is 0.7800. So under the published bands every one of the 5,103 cases is band "0", and bands "3–4", "5–6" and "7+" are not merely empty but unreachable — no assignment of contradictions to a case can ever enter them. The bands are coherent under one reading only: ct_count, a plain count of active contradictions, which does populate them.
band
share as ct_count
share as cd_case (Σw, core-v1)
0
25.7%
25.7%
1-2
57.1%
0.0% (unreachable)
3-4
17.0%
0.0% (unreachable)
5-6
0.2%
0.0% (unreachable)
7+
0.0%
0.0% (unreachable)
Every figure on this page uses the Σw reading, because that is what the ontology defines and what compute_cd() writes to the database. The card therefore cannot be applied to any number here. One of the two has to move: either the card is reissued on the Σw scale, or CD is redefined as a count and every published CD figure is restated. P3 and P8 bear on which — the unweighted count is the better discriminator of both outcome and adversarial flagging, so the count reading is not obviously the one to give up.
2. SimLex carries a different thirteen-type list under the same name. Parsed from SimLex.html at build time: 13 types summing to 1.63 against core-v1’s 1.52. 6 codes are shared (AI, CC, CF, JC, RF, RPF) and carry identical weights, so the two look interchangeable on inspection. The other 7 are different types entirely — AE Act–Effect Mismatch (0.09), CS Chain Severance (0.25), EN Effect Nullification (0.10), RC Retroactive Contradiction (0.11), RGR Recht/Gesetz Rupture (0.18), RV Role Vacancy (0.08), TA Triggering Ambiguity (0.07) — with no counterpart here, while FM, NI, PC, RC3, RCL, SE, TC exist only in the ontology. A case scored 0.40 under one list and 0.40 under the other is not the same measurement, and nothing in either number says which list produced it. This is the four-version divergence surfacing as an incommensurable metric rather than as a naming inconsistency, and it is the item on this page most likely to be caught by a reviewer. Minimum fix: every CD figure, everywhere, carries its profile identifier — the way this page names core-v1 — and SimLex either adopts core-v1 or declares its own profile id and stops calling the result CD without qualification.
noteTwo of the 13 contradiction types are empirically near-vacant.
In 5,103 annotated cases: RCL (Recognition Collapse) 2 cases, SE (Self-Undermining Effect) 11 cases. Mean contradictions per case is 1.41 and 25.7% of cases carry none. The operative taxonomy may be closer to 11 live types than 13. Whether these types are theoretically real but rare, or simply undetectable from published opinions, is not resolvable from this data.
Standing methodological disclosures
Coverage. 5,103 cases, all eight domains complete. Fully rebuilt: D1 (641), D2 (655), D3 (722), D4 (619), D5 (656), D6 (657), D7 (640), D8 (513). Un-annotated domains are missing wholesale, not sampled out, so this is not a random subset of the corpus. Because the pass is live, these counts move: the section is regenerated from the database, not hand-written.
Excluded rows. 83 of the 5,186 v2 rows written so far carry a chain_outcome outside the documented enum (indeterminate 55, (empty) 28). All are validation-failed and review-flagged, and the indeterminate group carries CD = 0.0 uniformly — they are failed extractions, not findings. They are dropped from every figure above, leaving n = 5,103.
Rule-forced types. RC3 (RULE 2), NI (RULE 3), JC (RULE 4), CC (RULE 6) are assigned by mandatory inference rule when stated conditions hold. Their prevalence measures rule trigger frequency; it is not an independent discovery.
Claimant-centric framing. RULE 0 defines Nodes 7–8 from the claimant’s perspective, and Node 7 failure correlates 90.2% with protection_denied. That correlation is not definitional: 98.4% of the corpus was annotated with the disposition redacted, and among those cases the figure is 90.0% (against 97.5% where the disposition was visible). The two are non-independent — one annotator, one text, two readings — rather than identical by construction.
What v2 removed. v1 RULE 1/RULE 5 forced RF and RPF from Node 7 failure. Under v2 these are assigned on textual evidence only — RF is now 3.0% and RPF 5.6% of cases. Published v1 figures (13.0× CD ratio, RF ~64%, RPF φ = 0.953) are artifacts of those rules and do not survive.
Annotation quality. 42.4% of v2 rows carry a review flag. Excluding them does not weaken P1 (N7 failure 61.3% → 69.1%).
Single annotator. All v2 annotations come from one model under one prompt. No inter-annotator agreement statistic exists for this pass; nothing here is validated against human coding.
Empirical Testing
Hypothesis Laboratory
Run statistical tests on pre-defined structural hypotheses, or construct and test your own using the builder below. All tests run on the loaded corpus — no server required.
Pre-defined Structural Hypotheses
H1
Definitional check — Pass-2 rule consistency
Not an empirical hypothesis. The Pass-2 rule derives the outcome variable from Node 7, so this necessarily returns ~98.5%. It is retained only as a consistency check that the derivation rule was applied uniformly across the corpus.
H2
Cascade Propagation
Upstream MLC node failures (Nodes 1–4) predict downstream failures (Nodes 5–8) at rates exceeding base rates — the chain fails as a chain, not independently per node.
H3
Adversarial Divergence by Domain
Structural ambiguity (adversarial flagging) clusters in specific doctrinal domains rather than distributing uniformly — revealing which areas of law are genuinely indeterminate at the structural level.
H4
Circuit Jurisdiction as Structural Variable
Structural failure patterns vary systematically by circuit — law is not structurally uniform across jurisdictions, even controlling for domain composition.
H5
Structural Signature of Remand
Remanded cases have a structurally distinct profile from denied or granted cases — specifically involving mid-chain failures (Nodes 4–6) rather than the terminal Node 7–8 failures that characterize denial.
H6
Temporal Drift in Contradiction Composition
The distribution of active contradiction types shifts over time (1990–2024) even as the structural architecture stays constant — which nodes fail changes with doctrine, but the fact that they fail does not.
H7
Rare Contradiction Type Clustering
Rare contradiction types (CF, TC, PC, CC, FM) are not uniformly distributed but cluster in specific doctrinal domains where the legal structure makes them theoretically expected.
H8
Adversarial Strength as Structural Ambiguity Indicator
Adversarial strength (how defensible the opposite annotation is) identifies genuine structural indeterminacy, not annotation noise — adversarially flagged cases carry structurally distinct CD profiles from non-flagged cases.
H9
Protection Rate by Domain — CD Mediation
Domain-level variation in protection grant rates is mediated by CD score — domains grant protection less often because they carry more structural stress, not due to exogenous doctrinal favoritism.
User-Defined
Custom Hypothesis Builder
Define a grouping variable and metric, apply optional filters, and run a statistical comparison across the corpus.
AI-Assisted
Generate New Hypotheses
Describe a research question or area of interest. Claude will propose testable hypotheses grounded in the corpus structure and pre-configure the builder for each one.
Anthropic API KeyShared with Case Analyzer · never sent to our server