YUCLAW v5.0.0 Signal Validation Lab
Validation LabOpen Index EvidenceCanada Resources EvidenceForward TrackingGitHubPyPILedgerMethodologyHome
Per-label hit-rate ledger: Forward Tracking →
Data through 2026-07-23 (last completed U.S. trading day) · regenerated daily after market close · last build 2026-07-23 23:00 UTC
Forward OOS
EARLY · n=42 periods
no significant alpha yet — underpowered, accruing daily
In-sample replay
OPTIMISTIC
parametric look-ahead — educational only, collapsed below
Ledger
LIVE · 47 blocks
git-anchored daily · root 46f20ef5…
Reproducibility
ONE-COMMAND VERIFY
stdlib script reproduces every statistic + ledger roots
Disclaimer — Research cohort analysis — hypothetical illustration, not investment advice, not performance advertising. Classifications, not recommendations.
Honest reading

No forward alpha has been statistically proven yet. What is running: the point-in-time infrastructure is live, every daily signal set is anchored to a public, git-committed ledger, and every statistic on this page reproduces from published derived data with one command. The evidence-risk channel (C6) is the next out-of-sample proof target — its in-sample sign is known and its forward test is defined below. Everything here is research and education, not investment advice; research classifications, not recommendations.

Data integrity note — Jun 26 – Jul 3 outage disclosed. Price-derived inputs were stale. No snapshots were retroactively edited. Full log ↓
What is proven · what is not proven
Proven (verifiable today)
  • Point-in-time snapshot discipline — daily as-of writes, zero retroactive edits (outage window included)
  • Git-anchored replayable ledger — daily sha-256 roots committed publicly before pages update
  • One-command reproducibility — every statistic + 2,847 leaf hashes re-derive from published data
  • Deterministic evidence grounding measurement — corpus grounding 0.75, citation fidelity 0.85 (definitions footnoted below)
Not proven (open, tracked)
  • Forward alpha — n=28 periods, underpowered; not significant at 5%
  • IC significance — forward 5d IC +0.09 loses significance after overlap (HAC) correction; 20d descriptive only
  • Evidence→price lead — event-study CAR is adverse at the current backfill-era sample; live-era n too small
  • C6 risk gate out-of-sample — rareness confirmed OOS 2026-07-06; sign confirmation pending (elevated arm n=2; accrual live from 2026-07-16)

A Fama–French-style decile-cohort event study: does YUCLAW's composite signal score carry forward information about subsequent realized returns? Cohorts are grouped by score decile or signal label and tracked as equal-weighted research cohorts against two references: the equal-weight universe (all scored tickers, same rebalance dates) and SPY. Derived statistics only — no raw prices. This is an event study, not portfolio management.

Latest rolling record
continuous display across the regime boundary · statistics stay strictly per-regime
Updated through 2026-07-23 · last build 2026-07-23 23:00 UTC · daily after U.S. market close
High-minus-low cohort spread — latest rolling record · updated daily after U.S. market close
0%-8.6%+7.0%Jun 23Jun 30Jul 8Jul 16Jul 23High−Low spread -4.2%
0%-5.8%+10.3%May 26Jun 10Jun 25Jul 9Jul 23High−Low spread -1.2%
forward out-of-sampleMay 180%-14.4%+0.2%Apr 29Jun 3Jun 18Jul 8Jul 23High−Low spread -10.2%
in-sample replayforward out-of-sampleMay 180%-5.2%+16.4%Feb 18May 21Jun 12Jul 2Jul 23High−Low spread -0.6%

Continuous for display only: the dashed May-18 line marks the in-sample replay → forward out-of-sample regime boundary. All statistics are computed per-regime and never blended across the boundary. Windows are rebased to their start date.

Label cohorts vs SPY — latest rolling record · updated daily after U.S. market close
0%-5.1%+5.7%Jun 23Jun 30Jul 8Jul 16Jul 23Positive-label -2.4%Risk-flag -1.8%SPY +0.6%
0%-5.0%+11.0%May 26Jun 10Jun 25Jul 9Jul 23Positive-label +2.4%Risk-flag -1.7%SPY -1.9%
forward out-of-sampleMay 180%-3.9%+16.7%Apr 29Jun 3Jun 18Jul 8Jul 23Positive-label +7.7%Risk-flag +9.8%SPY +4.0%
in-sample replayforward out-of-sampleMay 180%-4.5%+111.5%Feb 18May 21Jun 12Jul 2Jul 23Positive-label +95.2%Risk-flag +16.7%SPY +7.8%

Continuous for display only: the dashed May-18 line marks the in-sample replay → forward out-of-sample regime boundary. All statistics are computed per-regime and never blended across the boundary. Windows are rebased to their start date. Label cohorts have small, variable membership — see the archived replay for per-date n.

Universe & Coverage
what is scored, what enters the decile study, and the inclusion rule
79
tickers scored daily
49
U.S. large-cap equities
30
ETFs / ETNs
79/79
priced coverage (median day)
SegmentCountMembers
Equities (tech, financials, health care, energy, staples, industrials, …)49AAPL, MSFT, NVDA, … (full list in v3/universe.json)
Sector ETFs15XLK, XLF, XLE, XLV, XLU, XLI, XLY, XLP, XLB, XLRE, XLC, SMH, KRE, IBB, XBI
Broad-market ETFs5SPY, QQQ, IWM, DIA, MDY
Macro ETFs/ETNs (rates, metals, dollar, China/EM, volatility)10TLT, IEF, GLD, SLV, UUP, FXI, EEM, VIXY, VXX, TAIL
Inclusion rule. A signal date enters the decile study only if it scored at least 40 universe tickers (a "decile" of a handful of names is meaningless). 1 partial-universe date was excluded under this rule (2026-05-31, a non-trading Sunday on which an ad-hoc run scored 3 tickers). Within an included date, a ticker contributes to its cohort's period return only if closing prices exist at both the entry and exit dates — currently 79 of 79 tickers on the median rebalance (min 79, max 79). Top/bottom decile = the highest/lowest ~10% of tickers by composite score, i.e. 8 of 79.
Evidence-tier names (49, Canada Resources). Evidence-tier names are covered for filings evidence and research dashboards only. They are not scored and are not part of the Lab decile study or the 79-ticker forward-track universe. See the Canada Resources Evidence page and the methodology's evidence-tier boundary section.
Panel 1 · Forward (Out-of-Sample)LOOK-AHEAD-FREE
is_backfill = false · Day 0 = 2026-05-18 · the honest panel
⚠ Early forward period — 44 trading days, 42 rebalances. Not yet statistically meaningful.

Forward Day 0 = 2026-05-18; return window 2026-05-20 → 2026-07-23 (signals 2026-05-20 → 2026-07-23). A 44-trading-day out-of-sample window is far too short for any statistical inference — this is a directional illustration shown for transparency as the forward record accrues, not evidence of skill. In this early window the top-decile cohort trails the equal-weight universe; that is what the data shows and it is shown unblended.

0%-9.5%+5.3%May 20Jun 5Jun 23Jul 9Jul 23High-score -7.9%Low-score -2.9%Universe EW +3.6%SPY +0.2%
CohortCumulative returnMax drawdownVolatility (periodic)Hit-rate vs SPYn (min/med/max)
High-score cohort (top decile by composite score)-7.93%-13.02%1.97%40%8/8/8
Low-score cohort (bottom decile by composite score)-2.88%-9.51%1.86%43%8/8/8
Equal-weight universe (all scored tickers)+3.59%-3.22%0.83%60%79/79/79
SPY benchmark (broad-market reference)+0.18%-4.49%0.90%
Current high-score (top-decile) cohort membership · as of 2026-07-23 · 8 of 79 tickers
TickerComposite scoreSignal labelEvidence gradeFilings cited
TMO+0.6138POSITIVE_RESEARCH+Grade C1
MRK+0.5381POSITIVE_RESEARCHGrade C0
ABBV+0.5183POSITIVE_RESEARCHGrade C1
XLV+0.5070POSITIVE_RESEARCHGrade C0
BMY+0.4983POSITIVE_RESEARCHGrade C0
ABT+0.4981POSITIVE_RESEARCHGrade C2
DHR+0.4753POSITIVE_RESEARCHGrade C2
UNH+0.4102POSITIVE_RESEARCHGrade C0

Research classifications, not recommendations. Membership is recomputed at every signal date; the cohort above is today's snapshot and changes daily. Evidence grades follow the public grading rubric (confidence × evidence depth); "Insufficient" appears when composite confidence < 0.30.

Top-minus-bottom cohort spread (research spread statistic — not a position, not tradeable)
cumulative -6.18% · max drawdown -14.57%
0%-10.5%+4.8%May 20Jun 5Jun 23Jul 9Jul 23Top−Bottom spread -6.2%
Panel 4 · Evidence-Qualified Protocol CandidateFORWARD-ONLY
same decile methodology, restricted to names meeting minimum evidence criteria as of each date · window 2026-06-01 → 2026-07-23 · 36 rebalances

Qualification (point-in-time): grade A/B OR >=1 cited filing OR SourceLock-accepted event within 30d; grade 'Insufficient' excluded. Evidence-quality fields exist only from the v4.0 launch (2026-06-01), so this panel is structurally forward-only — no in-sample variant is possible without fabricating grades.

37/79
qualified pool (median/day)
47%
evidence coverage
7d
median evidence age
47%
C6 coverage of pool
Cohort (n)Cumulative returnMax drawdownVolatility (periodic)
Qualified pool, equal-weight (n=25–43/day)-3.59%-7.72%1.71%
Qualified high-score decile (n=2–4/day) ⚠ too small for inference-7.17%-18.74%3.90%
Equal-weight universe (n=79, reference)-0.30%-3.22%0.85%

Honest result: over its first 36 periods the qualified pool trails the universe (-3.59% vs -0.30%). The cohort is too young and (at the decile level) too small for inference; it accrues as evidence coverage grows. Criteria will NOT be loosened to fatten n.

Current qualified membership · as of 2026-07-23 · 44 names (top 4 = decile cohort)
TickerScoreDisplay labelGradeCited filingsEvidence ageC6
TMO+0.614POSITIVE_RESEARCH+Grade C10d+0.00
ABBV+0.518POSITIVE_RESEARCHGrade C117d+0.00
ABT+0.498POSITIVE_RESEARCHGrade C27d+0.34
DHR+0.475POSITIVE_RESEARCHGrade C23d+0.00
LLY+0.396NEUTRALGrade C142d-0.26
PYPL+0.344NEUTRALGrade B437d-0.02
XOM+0.336NEUTRALGrade B122d+0.00
COP+0.324NEUTRALGrade B130d-0.02
WFC+0.320NEUTRALGrade B19d+0.00
JNJ+0.304NEUTRALGrade B40d-0.46
PSX+0.278NEUTRALGrade B32d-0.46
PFE+0.253NEUTRALGrade C135d-0.00
AXP+0.219NEUTRALGrade B38d-0.22
CVX+0.178WATCHGrade B555d-0.46
JPM+0.161WATCHGrade B50d-0.31
SLB+0.133WATCHGrade B143d-0.13
PG+0.116WATCHGrade C19d+0.00
COST+0.090WATCHGrade B215d+0.08
HPE+0.085WATCHGrade B31d-0.37
V+0.076WATCHGrade B58d-0.46
AAPL+0.059WATCHGrade B155d-0.44
PEP+0.036WATCHGrade B215d+0.00
MA+0.033WATCHGrade B58d-0.46
INTC+0.033WATCHGrade B20d+0.44
MS+0.011WATCHGrade B56d-0.46
GS-0.004RISK_FLAG (weakening)Grade B52d-0.46
TSLA-0.006RISK_FLAG (weakening)Grade B41d+0.74
DELL-0.012RISK_FLAG (weakening)Grade B510d-0.46
MSFT-0.041RISK_FLAG (weakening)Grade B441d-0.46
KO-0.045RISK_FLAG (weakening)Grade B57d-0.46
AMAT-0.046RISK_FLAG (weakening)Grade B522d-0.46
META-0.059RISK_FLAG (weakening)Grade B51d-0.46
AMD-0.065RISK_FLAG (weakening)Grade B56d-0.46
WMT-0.093RISK_FLAG (weakening)Grade B56d-0.46
MRVL-0.112RISK_FLAG (weakening)Grade B57d-0.38
AMZN-0.116RISK_FLAG (weakening)Grade B514d-0.46
NVDA-0.119RISK_FLAG (weakening)Grade B521d-0.53
RKLB-0.128RISK_FLAG (weakening)Grade C515d-0.41
MU-0.137RISK_FLAG (weakening)Grade B517d-0.44
CRCL-0.140RISK_FLAG (weakening)Grade C513d-0.39
LUNR-0.151RISK_FLAG (weakening)Grade C58d-0.46
ARM-0.169RISK_FLAG (weakening)Grade B549d-0.46
LRCX-0.178RISK_FLAG (weakening)Grade B59d-0.46
GOOGL-0.271RISK_FLAG (event)Grade B50d-0.46

Evidence age = days since the last SourceLock-accepted filing event for the ticker (stale flag at >90d). Limitations: 36 periods; decile cohort of 2–4 names; insider-evidence stream live since 2026-07-16 (batch coverage ended 2026-05-15; gap backfilled with ingestion-time as-of) — evidence ages for some names reflect the coverage gap until live events accrue; qualification uses the public grade rubric and is recomputed point-in-time daily.

Panel 2 · In-Sample Replay — Educational replay only (n=13 rebalances) · collapsed

⚠ Educational replay only. The evidence-extraction model's training cutoff overlaps this window — in-sample results carry an unavoidable parametric look-ahead bias and are systematically optimistic. Label cohorts can be as small as a single name on some dates — treat all label-cohort figures below as illustrative, not evidence. The replay's final holding period is capped at forward Day 0 (2026-05-18) so this window never overlaps Panel 1's.

Archived replay (methodology view) · return window 2026-02-18 → 2026-05-18 · 63 trading days
Decile cohorts vs equal-weight universe and SPY (n=8-name cohorts)
0%-4.5%+20.1%Feb 18Mar 11Apr 1Apr 29May 18High-score +20.1%Low-score +11.5%Universe EW +11.5%SPY +7.6%
CohortCumulative returnMax drawdownVolatility (periodic)Hit-rate vs SPYn (min/med/max)
High-score cohort (top decile by composite score)+20.06%-3.34%2.40%62%8/8/8
Low-score cohort (bottom decile by composite score)+11.52%-6.09%3.21%54%8/8/8
Equal-weight universe (all scored tickers)+11.54%-2.35%1.62%69%79/79/79
SPY benchmark (broad-market reference)+7.63%-5.47%1.86%
High-minus-low cohort spread (research spread statistic — not a position, not tradeable)
cumulative +5.96% · max drawdown -12.84%
0%-2.0%+16.4%Feb 18Mar 11Apr 1Apr 29May 18High−Low spread +6.0%
Label cohorts vs SPY (n= per-date membership in the table below — as small as 1; look-ahead-optimistic figures)
0%-4.5%+85.3%Feb 18Mar 11Apr 1Apr 29May 18Positive-label +85.3%Risk-flag +14.0%SPY +7.6%
CohortCumulative returnMax drawdownVolatility (periodic)Hit-rate vs SPYn (min/med/max)
Positive-label cohort ⚠ thin+85.29%-3.87%8.01%54%1/3/10
Risk-flag cohort+14.03%-5.36%3.30%69%4/12/26
SPY benchmark (broad-market reference)+7.63%-5.47%1.86%
Panel 3 · Statistical Rigor
every statistic carries its n and window · * = p < 0.05 · bootstrap: percentile, 10,000 i.i.d. period resamples, fixed seed · IC t-stats are Newey–West (Bartlett) HAC-corrected for overlapping return windows
Forward OOS · 2026-05-20 → 2026-07-23 · look-ahead-free
Decile spread tests (per-rebalance arithmetic spreads)
SpreadMean / periodBootstrap 95% CItpn periods
Top − Bottom decile-0.125%(-0.85%, +0.56%)-0.340.73242
Top decile − EW universe-0.265%(-0.78%, +0.25%)-1.000.32342
Information coefficient (per-date cross-sectional Spearman, Fama–MacBeth mean)
HorizonMean ICNewey–West tpshare IC>0T datesmedian cross-section
1-day+0.0106+0.30 (lag 0)0.76348%4679
5-day+0.0017+0.03 (lag 4)0.97750%4279
20-day ⚠ too few independent blocks+0.0140(+0.28 — not interpretable)N/A — descriptive only65%2679
Market model (top-decile cohort per-period returns, OLS)
Regressionα / periodα annualizedβt(α)pn
vs equal-weight universe-0.279%-49.0%+1.17-1.040.3050.24442
vs SPY-0.186%-36.1%+1.06-0.690.4930.23442
Statistical power meter
n=42 periods
≈532% detectable |α| (annualized, 80% power)
UNDERPOWERED

This panel is not failing. It is too young. Alpha estimates are shown for completeness and are not significant; detectability improves as n accrues daily.

In-Sample Replay · 2026-02-18 → 2026-05-18 · parametric look-ahead — optimistic
Decile spread tests (per-rebalance arithmetic spreads)
SpreadMean / periodBootstrap 95% CItpn periods
Top − Bottom decile+0.553%(-1.94%, +3.08%)+0.410.68613
Top decile − EW universe+0.586%(-0.98%, +2.17%)+0.700.49613
Information coefficient (per-date cross-sectional Spearman, Fama–MacBeth mean)
HorizonMean ICNewey–West tpshare IC>0T datesmedian cross-section
1-day+0.0289+0.80 (lag 0)0.43854%1379
5-day+0.0143+0.26 (lag 0)0.80262%1379
20-day-0.0026-0.04 (lag 3)0.97338%1379
Market model (top-decile cohort per-period returns, OLS)
Regressionα / periodα annualizedβt(α)pn
vs equal-weight universe+1.555%+123.1%-0.13+1.960.0750.00813
vs SPY+1.559%+123.6%-0.20+2.160.0540.02413
Statistical power meter
n=13 periods
≈245% detectable |α| (annualized, 80% power)
UNDERPOWERED

This panel is not failing. It is too young. Alpha estimates are shown for completeness and are not significant; detectability improves as n accrues daily.

Honest reading: at the current sample sizes, no forward spread, IC, or alpha is statistically significant at the 5% level once overlap is corrected. The forward 20-day IC is positive on all observed dates but has too few independent blocks to test. "Not yet significant" is the finding — the statistics accrue daily and this panel recomputes with them.

Protocol readiness · maturity gates
infrastructure-ready, statistically young — 3 of 6 gates passed
GateStatusEvidence / requirement
Gate 1 · Point-in-time infrastructure + public ledgerPASSEDlive since 2026-05-18 · 47 anchored daily blocks
Gate 2 · Honest measurement disciplinePASSEDper-regime statistics, HAC-corrected inference, power reporting, adverse results published
Gate 3 · Independent reproducibilityPASSED (young)one-command replay live since 2026-07-05; awaiting first external replication
Gate 4 · Forward statistical significanceNOT YETno spread, IC, or alpha significant at 5% with adequate power — requires more forward data
Gate 5 · Evidence→price lead, out-of-sampleNOT YETlive-era event sample n=4; needs ≥30 live directional events for a first read
Gate 6 · C6 risk-gate OOS confirmationNOT YETrareness confirmed OOS 2026-07-06 (22% fire rate, n=9 held-out); sign confirmation pending (elevated arm n=2; accrual live from 2026-07-16)
Reproduce this page
one command, fresh environment · derived data only (no vendor market data bundled or required)
# packaged (v5.0+)
pip install yuclaw
yuclaw replay-lab

# or fully standalone (stdlib only, nothing to install)
curl -sO https://yuclawlab.github.io/yuclaw-brain/replay/lab_replay_bundle.json
curl -sO https://raw.githubusercontent.com/YuClawLab/yuclaw-brain/main/tools/replay_lab.py
python3 replay_lab.py lab_replay_bundle.json

The script (Python ≥3.10, standard library only; pip install yuclaw optionally adds the full SDK) rebuilds the decile cohorts from the bundled composite scores, re-derives every cohort period return, recomputes all Panel-3 statistics (same bootstrap seed), and — the tamper-evidence step — recomputes every forward snapshot's sha-256 leaf hash from disclosed derived inputs and rolls them into daily roots that must match the public yuclaw-trust ledger. It exits non-zero on any mismatch.

this build derives from: source commit 4e9c15cde2f1 · ledger block 2026-07-23 · daily root 46f20ef576a0d6ba… · 47 public ledger blocks

Compliant data path: the bundle contains YUCLAW-derived data only — scores, locked labels, component scores, content hashes, and derived period returns. No raw vendor OHLCV rows are exported (data-provider terms). Analyses requiring raw prices need the user's own licensed price feed.

Download evidence packet

Everything this page derives, as files — YUCLAW-derived data only (derived statistics, counts, classifications; never raw vendor price/options data). Regenerated in the daily chain.

Download packet (.zip)

Cite this page
YUCLAW Validation Lab, v5.0.0, data through 2026-07-23, built 2026-07-23, commit 4e9c15cde2f1; evidence-tier names excluded from scoring universe.
Use YUCLAW in your research
1 · Verify the record

pip install yuclaw then yuclaw replay-lab.
No install: tools/replay_lab.py (stdlib only) against the published bundle.
Exit 0 = every statistic and ledger root reproduced. How to report a replication →

2 · Inspect one evidence trace

One real Suncor 6-K, end to end:
filing → exhibit → extracted prose → event type → grade → C6 posture.
Open the trace → · example evidence memo (Suncor) →

3 · Cite a research lens

Every evidence packet ships a ready citation snippet
(version, data-through, build date, source commit).
Get the citation →

Status — proven · not proven · accruing

Rendered from one shared source (v3/web/useful_blocks.py) on every page that shows it, so the copies cannot drift. Statuses are measured, not aspirational.

Proven (verifiable today)
  • Replay works — one command reproduces every Lab statistic and ledger root from published data
  • Ledger anchored daily — sha-256 daily roots committed to a public git repository before pages update
  • Evidence traces to filings — every accepted event carries a source URL, accession number, and verified excerpt
  • Coverage measured — SEC-filer weight per lens is stated as measured, never rounded up
  • Snapshots are point-in-time — daily as-of writes, zero retroactive edits (outage window disclosed, not repaired)
  • Evidence-tier names are never scored — enforced by positive gating and a standing negative check
Not proven
  • Forward alpha — no spread, IC, or alpha significant at 5% with adequate power
  • C6 risk-gate sign — rareness confirmed OOS 2026-07-06 (22% fire rate, n=9 held-out); sign confirmation pending (elevated arm n=2; accrual live from 2026-07-16)
  • Peer-model CAR lead — event-study lead over peer models is not established; live-era sample remains small
Accruing
  • · Forward out-of-sample record — one period per trading day, accruing daily
  • · Matured CAR events — each accepted event matures into the event study after its forward window completes
  • · C6 elevated arm — live Form-4 ingestion since 2026-07-16 restores the insider stream to production inputs
  • · External replications — the replication log accrues as independent runs are reported
Reproduction challenge

Independent replication is invited. Run the three commands in "Reproduce this page"; the script exits non-zero on ANY mismatch between recomputed statistics and this page, or between recomputed hash roots and the public ledger. If you find a mismatch, the ledger is broken and we want to know: open an issue with the script output. Replications that confirm are equally welcome — independent verification is the point of publishing the bundle.

Roadmap · Risk Gate Lab (next proof target)
no new claims — the tests are defined before the data can answer them
Data Integrity Log

Infrastructure note (Jun 26 – Jul 3, 2026) — a network outage on the research host interrupted external data feeds. Daily signal snapshots continued to be written on-box, point-in-time, throughout the window — but from Jun 26 to Jul 2 their price-derived inputs were frozen at Jun 25 closes (the price feed was unreachable), and this page was not republished during the outage. Price history and SEC filing ingestion were restored and backfilled on Jul 3, and the filing window was re-checked against EDGAR on Jul 5 (no missing filings). No snapshot or ledger row was retroactively edited: the outage-window snapshots stand exactly as written, stale inputs and all.

Live Form-4 ingestion enabled 2026-07-16. Insider-event stream restored to production inputs (batch coverage previously ended 2026-05-15; the gap is backfilled with ingestion-time available_as_of and cannot affect past replays). C6 elevated-arm accrual for the out-of-sample sign study begins from this date.

Policy: disclosures are never deleted; presentation may be compressed, substance may not. Fabricating retroactive point-in-time data would invalidate the replayable ledger; the disclosed staleness is the honest record. See also methodology.

System properties · numbers and status, no adjectives
each row: the measured number and its honest maturity label
PropertyMeasuredStatus
Git-anchored replayable ledger47 daily blocks · latest root 46f20ef576a0… (2026-07-23)LIVE — anchored daily before pages publish
Evidence grounding (v5 Layer-1 corpus)corpus grounding 0.52 → 0.75 · citation fidelity 0.66 → 0.85 after the prose-first extraction fix (commit f130983e)MEASURED on the v5 Layer-1 filing corpus
C6 evidence/risk channelfires on 35% of in-sample and 26% of forward snapshots (rare by construction) · in-sample within-class IC +0.36 on material non-insider events (n=38)rareness confirmed OOS 2026-07-06 (22% fire, n=9 held-out); sign confirmation pending (elevated arm n=2; accrual live from 2026-07-16)
Event-type extraction specialists10 dedicated extractors (v5 Layer 1) — earnings, guidance, M&A, insider, governance, …LIVE for 8-K and Form-4 streams · Form-4 live since 2026-07-16 (batch 2026-02-18 → 05-15; gap backfilled, ingestion-time as-of)
Point-in-time disciplinedaily as-of snapshots; outage of Jun 26 – Jul 3 disclosed (snapshots continued point-in-time on frozen price inputs; zero retroactive edits)LIVE — the disclosed gap is the proof it isn't backfilled

Definitions (exact internal rubric, deterministic verifier — no LLM in the loop):
Corpus grounding = points_grounded / points_total across the filing corpus, where an agent key-point is grounded iff it carries ≥1 citation that verifies as a verbatim (whitespace/case-normalized) span of the source filing AND every numeric token in the point appears within those verified quotes; ungrounded points are discarded with the reason recorded.
Citation fidelity = citations_verified / citations_total — the share of quoted spans an agent cites that locate as verbatim spans of the source filing after whitespace/case normalization. Verifier source: v5/swarm/grounding.py.

Methodology summary

Equal-weighted cohorts, rebalanced at each signal date, ranked by composite total_score; top/bottom decile (~10%, currently 8 of 79). References: the equal-weight universe cohort (all scored tickers, identical rebalance schedule) and SPY. Returns are close-to-close from internal price_history (derived statistics only — raw prices never shown). Inclusion rule: a signal date enters the study only if it scored ≥40 tickers; a ticker contributes only when entry and exit closes both exist. The in-sample replay's final holding period is capped at forward Day 0 so the two panels' return windows never overlap. The two panels are never blended. Annualized figures are intentionally omitted — annualizing a weeks-long window is misleading; cumulative return over N trading days is shown instead. Full methodology, including the in-sample look-ahead disclosure, is in methodology/validation_lab.md.

Disclaimer — Hypothetical research illustration. Not investment advice, not performance advertising, not an offer of any product. Research classifications, not recommendations. Past results — in-sample or forward-tracked — do not predict future performance.