No forward alpha has been statistically proven yet. What is running: the point-in-time infrastructure is live, every daily signal set is anchored to a public, git-committed ledger, and every statistic on this page reproduces from published derived data with one command. The evidence-risk channel (C6) is the next out-of-sample proof target — its in-sample sign is known and its forward test is defined below. Everything here is research and education, not investment advice; research classifications, not recommendations.
A Fama–French-style decile-cohort event study: does YUCLAW's composite signal score carry forward information about subsequent realized returns? Cohorts are grouped by score decile or signal label and tracked as equal-weighted research cohorts against two references: the equal-weight universe (all scored tickers, same rebalance dates) and SPY. Derived statistics only — no raw prices. This is an event study, not portfolio management.
Continuous for display only: the dashed May-18 line marks the in-sample replay → forward out-of-sample regime boundary. All statistics are computed per-regime and never blended across the boundary. Windows are rebased to their start date.
Continuous for display only: the dashed May-18 line marks the in-sample replay → forward out-of-sample regime boundary. All statistics are computed per-regime and never blended across the boundary. Windows are rebased to their start date. Label cohorts have small, variable membership — see the archived replay for per-date n.
| Segment | Count | Members |
|---|---|---|
| Equities (tech, financials, health care, energy, staples, industrials, …) | 49 | AAPL, MSFT, NVDA, … (full list in v3/universe.json) |
| Sector ETFs | 15 | XLK, XLF, XLE, XLV, XLU, XLI, XLY, XLP, XLB, XLRE, XLC, SMH, KRE, IBB, XBI |
| Broad-market ETFs | 5 | SPY, QQQ, IWM, DIA, MDY |
| Macro ETFs/ETNs (rates, metals, dollar, China/EM, volatility) | 10 | TLT, IEF, GLD, SLV, UUP, FXI, EEM, VIXY, VXX, TAIL |
Forward Day 0 = 2026-05-18; return window 2026-05-20 → 2026-07-23 (signals 2026-05-20 → 2026-07-23). A 44-trading-day out-of-sample window is far too short for any statistical inference — this is a directional illustration shown for transparency as the forward record accrues, not evidence of skill. In this early window the top-decile cohort trails the equal-weight universe; that is what the data shows and it is shown unblended.
| Cohort | Cumulative return | Max drawdown | Volatility (periodic) | Hit-rate vs SPY | n (min/med/max) |
|---|---|---|---|---|---|
| High-score cohort (top decile by composite score) | -7.93% | -13.02% | 1.97% | 40% | 8/8/8 |
| Low-score cohort (bottom decile by composite score) | -2.88% | -9.51% | 1.86% | 43% | 8/8/8 |
| Equal-weight universe (all scored tickers) | +3.59% | -3.22% | 0.83% | 60% | 79/79/79 |
| SPY benchmark (broad-market reference) | +0.18% | -4.49% | 0.90% | — | — |
| Ticker | Composite score | Signal label | Evidence grade | Filings cited |
|---|---|---|---|---|
| TMO | +0.6138 | POSITIVE_RESEARCH+ | Grade C | 1 |
| MRK | +0.5381 | POSITIVE_RESEARCH | Grade C | 0 |
| ABBV | +0.5183 | POSITIVE_RESEARCH | Grade C | 1 |
| XLV | +0.5070 | POSITIVE_RESEARCH | Grade C | 0 |
| BMY | +0.4983 | POSITIVE_RESEARCH | Grade C | 0 |
| ABT | +0.4981 | POSITIVE_RESEARCH | Grade C | 2 |
| DHR | +0.4753 | POSITIVE_RESEARCH | Grade C | 2 |
| UNH | +0.4102 | POSITIVE_RESEARCH | Grade C | 0 |
Research classifications, not recommendations. Membership is recomputed at every signal date; the cohort above is today's snapshot and changes daily. Evidence grades follow the public grading rubric (confidence × evidence depth); "Insufficient" appears when composite confidence < 0.30.
Qualification (point-in-time): grade A/B OR >=1 cited filing OR SourceLock-accepted event within 30d; grade 'Insufficient' excluded. Evidence-quality fields exist only from the v4.0 launch (2026-06-01), so this panel is structurally forward-only — no in-sample variant is possible without fabricating grades.
| Cohort (n) | Cumulative return | Max drawdown | Volatility (periodic) |
|---|---|---|---|
| Qualified pool, equal-weight (n=25–43/day) | -3.59% | -7.72% | 1.71% |
| Qualified high-score decile (n=2–4/day) ⚠ too small for inference | -7.17% | -18.74% | 3.90% |
| Equal-weight universe (n=79, reference) | -0.30% | -3.22% | 0.85% |
Honest result: over its first 36 periods the qualified pool trails the universe (-3.59% vs -0.30%). The cohort is too young and (at the decile level) too small for inference; it accrues as evidence coverage grows. Criteria will NOT be loosened to fatten n.
| Ticker | Score | Display label | Grade | Cited filings | Evidence age | C6 |
|---|---|---|---|---|---|---|
| TMO | +0.614 | POSITIVE_RESEARCH+ | Grade C | 1 | 0d | +0.00 |
| ABBV | +0.518 | POSITIVE_RESEARCH | Grade C | 1 | 17d | +0.00 |
| ABT | +0.498 | POSITIVE_RESEARCH | Grade C | 2 | 7d | +0.34 |
| DHR | +0.475 | POSITIVE_RESEARCH | Grade C | 2 | 3d | +0.00 |
| LLY | +0.396 | NEUTRAL | Grade C | 1 | 42d | -0.26 |
| PYPL | +0.344 | NEUTRAL | Grade B | 4 | 37d | -0.02 |
| XOM | +0.336 | NEUTRAL | Grade B | 1 | 22d | +0.00 |
| COP | +0.324 | NEUTRAL | Grade B | 1 | 30d | -0.02 |
| WFC | +0.320 | NEUTRAL | Grade B | 1 | 9d | +0.00 |
| JNJ | +0.304 | NEUTRAL | Grade B | 4 | 0d | -0.46 |
| PSX | +0.278 | NEUTRAL | Grade B | 3 | 2d | -0.46 |
| PFE | +0.253 | NEUTRAL | Grade C | 1 | 35d | -0.00 |
| AXP | +0.219 | NEUTRAL | Grade B | 3 | 8d | -0.22 |
| CVX | +0.178 | WATCH | Grade B | 5 | 55d | -0.46 |
| JPM | +0.161 | WATCH | Grade B | 5 | 0d | -0.31 |
| SLB | +0.133 | WATCH | Grade B | 1 | 43d | -0.13 |
| PG | +0.116 | WATCH | Grade C | 1 | 9d | +0.00 |
| COST | +0.090 | WATCH | Grade B | 2 | 15d | +0.08 |
| HPE | +0.085 | WATCH | Grade B | 3 | 1d | -0.37 |
| V | +0.076 | WATCH | Grade B | 5 | 8d | -0.46 |
| AAPL | +0.059 | WATCH | Grade B | 1 | 55d | -0.44 |
| PEP | +0.036 | WATCH | Grade B | 2 | 15d | +0.00 |
| MA | +0.033 | WATCH | Grade B | 5 | 8d | -0.46 |
| INTC | +0.033 | WATCH | Grade B | 2 | 0d | +0.44 |
| MS | +0.011 | WATCH | Grade B | 5 | 6d | -0.46 |
| GS | -0.004 | RISK_FLAG (weakening) | Grade B | 5 | 2d | -0.46 |
| TSLA | -0.006 | RISK_FLAG (weakening) | Grade B | 4 | 1d | +0.74 |
| DELL | -0.012 | RISK_FLAG (weakening) | Grade B | 5 | 10d | -0.46 |
| MSFT | -0.041 | RISK_FLAG (weakening) | Grade B | 4 | 41d | -0.46 |
| KO | -0.045 | RISK_FLAG (weakening) | Grade B | 5 | 7d | -0.46 |
| AMAT | -0.046 | RISK_FLAG (weakening) | Grade B | 5 | 22d | -0.46 |
| META | -0.059 | RISK_FLAG (weakening) | Grade B | 5 | 1d | -0.46 |
| AMD | -0.065 | RISK_FLAG (weakening) | Grade B | 5 | 6d | -0.46 |
| WMT | -0.093 | RISK_FLAG (weakening) | Grade B | 5 | 6d | -0.46 |
| MRVL | -0.112 | RISK_FLAG (weakening) | Grade B | 5 | 7d | -0.38 |
| AMZN | -0.116 | RISK_FLAG (weakening) | Grade B | 5 | 14d | -0.46 |
| NVDA | -0.119 | RISK_FLAG (weakening) | Grade B | 5 | 21d | -0.53 |
| RKLB | -0.128 | RISK_FLAG (weakening) | Grade C | 5 | 15d | -0.41 |
| MU | -0.137 | RISK_FLAG (weakening) | Grade B | 5 | 17d | -0.44 |
| CRCL | -0.140 | RISK_FLAG (weakening) | Grade C | 5 | 13d | -0.39 |
| LUNR | -0.151 | RISK_FLAG (weakening) | Grade C | 5 | 8d | -0.46 |
| ARM | -0.169 | RISK_FLAG (weakening) | Grade B | 5 | 49d | -0.46 |
| LRCX | -0.178 | RISK_FLAG (weakening) | Grade B | 5 | 9d | -0.46 |
| GOOGL | -0.271 | RISK_FLAG (event) | Grade B | 5 | 0d | -0.46 |
Evidence age = days since the last SourceLock-accepted filing event for the ticker (stale flag at >90d). Limitations: 36 periods; decile cohort of 2–4 names; insider-evidence stream live since 2026-07-16 (batch coverage ended 2026-05-15; gap backfilled with ingestion-time as-of) — evidence ages for some names reflect the coverage gap until live events accrue; qualification uses the public grade rubric and is recomputed point-in-time daily.
⚠ Educational replay only. The evidence-extraction model's training cutoff overlaps this window — in-sample results carry an unavoidable parametric look-ahead bias and are systematically optimistic. Label cohorts can be as small as a single name on some dates — treat all label-cohort figures below as illustrative, not evidence. The replay's final holding period is capped at forward Day 0 (2026-05-18) so this window never overlaps Panel 1's.
| Cohort | Cumulative return | Max drawdown | Volatility (periodic) | Hit-rate vs SPY | n (min/med/max) |
|---|---|---|---|---|---|
| High-score cohort (top decile by composite score) | +20.06% | -3.34% | 2.40% | 62% | 8/8/8 |
| Low-score cohort (bottom decile by composite score) | +11.52% | -6.09% | 3.21% | 54% | 8/8/8 |
| Equal-weight universe (all scored tickers) | +11.54% | -2.35% | 1.62% | 69% | 79/79/79 |
| SPY benchmark (broad-market reference) | +7.63% | -5.47% | 1.86% | — | — |
| Cohort | Cumulative return | Max drawdown | Volatility (periodic) | Hit-rate vs SPY | n (min/med/max) |
|---|---|---|---|---|---|
| Positive-label cohort ⚠ thin | +85.29% | -3.87% | 8.01% | 54% | 1/3/10 |
| Risk-flag cohort | +14.03% | -5.36% | 3.30% | 69% | 4/12/26 |
| SPY benchmark (broad-market reference) | +7.63% | -5.47% | 1.86% | — | — |
| Spread | Mean / period | Bootstrap 95% CI | t | p | n periods |
|---|---|---|---|---|---|
| Top − Bottom decile | -0.125% | (-0.85%, +0.56%) | -0.34 | 0.732 | 42 |
| Top decile − EW universe | -0.265% | (-0.78%, +0.25%) | -1.00 | 0.323 | 42 |
| Horizon | Mean IC | Newey–West t | p | share IC>0 | T dates | median cross-section |
|---|---|---|---|---|---|---|
| 1-day | +0.0106 | +0.30 (lag 0) | 0.763 | 48% | 46 | 79 |
| 5-day | +0.0017 | +0.03 (lag 4) | 0.977 | 50% | 42 | 79 |
| 20-day ⚠ too few independent blocks | +0.0140 | (+0.28 — not interpretable) | N/A — descriptive only | 65% | 26 | 79 |
| Regression | α / period | α annualized | β | t(α) | p | R² | n |
|---|---|---|---|---|---|---|---|
| vs equal-weight universe | -0.279% | -49.0% | +1.17 | -1.04 | 0.305 | 0.244 | 42 |
| vs SPY | -0.186% | -36.1% | +1.06 | -0.69 | 0.493 | 0.234 | 42 |
This panel is not failing. It is too young. Alpha estimates are shown for completeness and are not significant; detectability improves as n accrues daily.
| Spread | Mean / period | Bootstrap 95% CI | t | p | n periods |
|---|---|---|---|---|---|
| Top − Bottom decile | +0.553% | (-1.94%, +3.08%) | +0.41 | 0.686 | 13 |
| Top decile − EW universe | +0.586% | (-0.98%, +2.17%) | +0.70 | 0.496 | 13 |
| Horizon | Mean IC | Newey–West t | p | share IC>0 | T dates | median cross-section |
|---|---|---|---|---|---|---|
| 1-day | +0.0289 | +0.80 (lag 0) | 0.438 | 54% | 13 | 79 |
| 5-day | +0.0143 | +0.26 (lag 0) | 0.802 | 62% | 13 | 79 |
| 20-day | -0.0026 | -0.04 (lag 3) | 0.973 | 38% | 13 | 79 |
| Regression | α / period | α annualized | β | t(α) | p | R² | n |
|---|---|---|---|---|---|---|---|
| vs equal-weight universe | +1.555% | +123.1% | -0.13 | +1.96 | 0.075 | 0.008 | 13 |
| vs SPY | +1.559% | +123.6% | -0.20 | +2.16 | 0.054 | 0.024 | 13 |
This panel is not failing. It is too young. Alpha estimates are shown for completeness and are not significant; detectability improves as n accrues daily.
Honest reading: at the current sample sizes, no forward spread, IC, or alpha is statistically significant at the 5% level once overlap is corrected. The forward 20-day IC is positive on all observed dates but has too few independent blocks to test. "Not yet significant" is the finding — the statistics accrue daily and this panel recomputes with them.
| Gate | Status | Evidence / requirement |
|---|---|---|
| Gate 1 · Point-in-time infrastructure + public ledger | PASSED | live since 2026-05-18 · 47 anchored daily blocks |
| Gate 2 · Honest measurement discipline | PASSED | per-regime statistics, HAC-corrected inference, power reporting, adverse results published |
| Gate 3 · Independent reproducibility | PASSED (young) | one-command replay live since 2026-07-05; awaiting first external replication |
| Gate 4 · Forward statistical significance | NOT YET | no spread, IC, or alpha significant at 5% with adequate power — requires more forward data |
| Gate 5 · Evidence→price lead, out-of-sample | NOT YET | live-era event sample n=4; needs ≥30 live directional events for a first read |
| Gate 6 · C6 risk-gate OOS confirmation | NOT YET | rareness confirmed OOS 2026-07-06 (22% fire rate, n=9 held-out); sign confirmation pending (elevated arm n=2; accrual live from 2026-07-16) |
# packaged (v5.0+) pip install yuclaw yuclaw replay-lab # or fully standalone (stdlib only, nothing to install) curl -sO https://yuclawlab.github.io/yuclaw-brain/replay/lab_replay_bundle.json curl -sO https://raw.githubusercontent.com/YuClawLab/yuclaw-brain/main/tools/replay_lab.py python3 replay_lab.py lab_replay_bundle.json
The script (Python ≥3.10, standard library only; pip install yuclaw optionally adds the full
SDK) rebuilds the decile cohorts from the bundled composite scores, re-derives every cohort period
return, recomputes all Panel-3 statistics (same bootstrap seed), and — the tamper-evidence step —
recomputes every forward snapshot's sha-256 leaf hash from disclosed derived inputs and rolls them
into daily roots that must match the public
yuclaw-trust ledger.
It exits non-zero on any mismatch.
this build derives from: source commit 4e9c15cde2f1 · ledger block 2026-07-23 · daily root 46f20ef576a0d6ba… · 47 public ledger blocks
Compliant data path: the bundle contains YUCLAW-derived data only — scores, locked labels, component scores, content hashes, and derived period returns. No raw vendor OHLCV rows are exported (data-provider terms). Analyses requiring raw prices need the user's own licensed price feed.
Everything this page derives, as files — YUCLAW-derived data only (derived statistics, counts, classifications; never raw vendor price/options data). Regenerated in the daily chain.
YUCLAW Validation Lab, v5.0.0, data through 2026-07-23, built 2026-07-23, commit 4e9c15cde2f1; evidence-tier names excluded from scoring universe.
pip install yuclaw then yuclaw replay-lab.
No install: tools/replay_lab.py
(stdlib only) against the published bundle.
Exit 0 = every statistic and ledger root reproduced. How to report a replication →
One real Suncor 6-K, end to end:
filing → exhibit → extracted prose → event type → grade → C6 posture.
Open the trace → ·
example evidence memo (Suncor) →
Every evidence packet ships a ready citation snippet
(version, data-through, build date, source commit).
Get the citation →
Rendered from one shared source (v3/web/useful_blocks.py) on every page that shows it,
so the copies cannot drift. Statuses are measured, not aspirational.
Independent replication is invited. Run the three commands in "Reproduce this page"; the script exits non-zero on ANY mismatch between recomputed statistics and this page, or between recomputed hash roots and the public ledger. If you find a mismatch, the ledger is broken and we want to know: open an issue with the script output. Replications that confirm are equally welcome — independent verification is the point of publishing the bundle.
Infrastructure note (Jun 26 – Jul 3, 2026) — a network outage on the research host interrupted external data feeds. Daily signal snapshots continued to be written on-box, point-in-time, throughout the window — but from Jun 26 to Jul 2 their price-derived inputs were frozen at Jun 25 closes (the price feed was unreachable), and this page was not republished during the outage. Price history and SEC filing ingestion were restored and backfilled on Jul 3, and the filing window was re-checked against EDGAR on Jul 5 (no missing filings). No snapshot or ledger row was retroactively edited: the outage-window snapshots stand exactly as written, stale inputs and all.
Live Form-4 ingestion enabled 2026-07-16. Insider-event stream restored to production inputs (batch coverage previously ended 2026-05-15; the gap is backfilled with ingestion-time available_as_of and cannot affect past replays). C6 elevated-arm accrual for the out-of-sample sign study begins from this date.
Policy: disclosures are never deleted; presentation may be compressed, substance may not. Fabricating retroactive point-in-time data would invalidate the replayable ledger; the disclosed staleness is the honest record. See also methodology.
| Property | Measured | Status |
|---|---|---|
| Git-anchored replayable ledger | 47 daily blocks · latest root 46f20ef576a0… (2026-07-23) | LIVE — anchored daily before pages publish |
| Evidence grounding (v5 Layer-1 corpus) | corpus grounding 0.52 → 0.75 · citation fidelity 0.66 → 0.85 after the prose-first extraction fix (commit f130983e) | MEASURED on the v5 Layer-1 filing corpus |
| C6 evidence/risk channel | fires on 35% of in-sample and 26% of forward snapshots (rare by construction) · in-sample within-class IC +0.36 on material non-insider events (n=38) | rareness confirmed OOS 2026-07-06 (22% fire, n=9 held-out); sign confirmation pending (elevated arm n=2; accrual live from 2026-07-16) |
| Event-type extraction specialists | 10 dedicated extractors (v5 Layer 1) — earnings, guidance, M&A, insider, governance, … | LIVE for 8-K and Form-4 streams · Form-4 live since 2026-07-16 (batch 2026-02-18 → 05-15; gap backfilled, ingestion-time as-of) |
| Point-in-time discipline | daily as-of snapshots; outage of Jun 26 – Jul 3 disclosed (snapshots continued point-in-time on frozen price inputs; zero retroactive edits) | LIVE — the disclosed gap is the proof it isn't backfilled |
Definitions (exact internal rubric, deterministic verifier — no LLM in the loop):
Corpus grounding = points_grounded / points_total across the filing corpus,
where an agent key-point is grounded iff it carries ≥1 citation that verifies as a verbatim
(whitespace/case-normalized) span of the source filing AND every numeric token in the point
appears within those verified quotes; ungrounded points are discarded with the reason recorded.
Citation fidelity = citations_verified / citations_total — the share of quoted
spans an agent cites that locate as verbatim spans of the source filing after
whitespace/case normalization. Verifier source: v5/swarm/grounding.py.
Equal-weighted cohorts, rebalanced at each signal date, ranked by composite total_score; top/bottom decile (~10%, currently 8 of 79). References: the equal-weight universe cohort (all scored tickers, identical rebalance schedule) and SPY. Returns are close-to-close from internal price_history (derived statistics only — raw prices never shown). Inclusion rule: a signal date enters the study only if it scored ≥40 tickers; a ticker contributes only when entry and exit closes both exist. The in-sample replay's final holding period is capped at forward Day 0 so the two panels' return windows never overlap. The two panels are never blended. Annualized figures are intentionally omitted — annualizing a weeks-long window is misleading; cumulative return over N trading days is shown instead. Full methodology, including the in-sample look-ahead disclosure, is in methodology/validation_lab.md.