Methodology
How signals are built, how they're evaluated, what was pre-registered, and where the hard limits of this study are.
Pre-registration
Before any evaluation code ran on real data, the full study design was committed
to study/PREREGISTRATION.md in git. This file defines:
hypotheses, signal definitions, evaluation metrics, the universe, return targets,
multiple-testing procedures, and failure criteria.
The commit hash that locked the pre-registration is recorded in the amendments log.
Any change to the study design after that commit must be documented in
study/AMENDMENTS.md with a date and reason.
Silent edits to the pre-registration are not permitted.
Universe
S&P 500 core: Point-in-time membership tracked from Wikipedia's S&P 500 historical changes table. We record each membership change with its effective date. Constituents at date t are those in the index as of close on that date, not today's membership. GICS sector is stored at each date. Limitation: Wikipedia's change table may contain errors and has incomplete history before ~2000.
Retail-attention basket: A separate, fixed cohort analyzed independently.
Inclusion rule: elevated retail attention or short interest as of the study start date.
The cohort is defined in config/universe_retail.yaml and was fixed in the pre-registration.
Signals that work only in this cohort are noted as such and not generalized to the S&P 500.
Ticker changes are handled via a stable internal security_id (SEC CIK where available)
mapped to tickers over time. Each ticker's identity is verified before inclusion.
Adjustments for splits and dividends use the adjusted close price from the price source;
raw close is also retained.
Signal construction
For each signal family and each trading day t, we compute a raw value using only
data with observed_at strictly before the 16:00 ET decision cutoff on date t.
No future data is ever used.
Step 1 — Time-series z-score: For each security, compute z = (value_t − mean(value_{t−60,t−1})) / std(value_{t−60,t−1}). Securities with fewer than 20 valid observations in the rolling window are excluded.
Step 2 — Cross-sectional rank: Within each date and GICS sector, rank securities by their z-score and normalize to [0, 1]. This removes sector-level mean reversion and standardizes across signals with different distributions.
Step 3 — Shrinkage: For signals derived from counts (e.g., insider transactions), apply empirical Bayes shrinkage toward the cross-sectional mean: shrunk = n·raw / (n + κ) + κ·mean_cs / (n + κ), where κ is the shrinkage parameter chosen by cross-validation on the training set.
Price and momentum are never part of any alt-data signal. They appear only as controls in Fama-MacBeth regressions and as baselines.
Evaluation
The primary evaluation metric is the rank information coefficient (IC): the Spearman rank correlation between the predicted cross-sectional rank of securities (from the signal) and the realized cross-sectional rank of subsequent excess returns.
Return targets are excess returns vs SPY (market) and vs sector ETF, at horizons of 1, 5, and 21 trading days. We use adjusted close prices and overlapping return windows purged by the embargo (see Walk-forward below).
Statistical significance is assessed via a two-sided t-test on the time series of daily ICs, with Newey-West standard errors to account for autocorrelation from overlapping windows (lag = horizon + 5 trading days).
Secondary metrics include: quintile long-short spread (Q5 − Q1 mean excess return), hit rate (fraction of predictions in correct direction), and Fama-MacBeth cross-sectional regressions with controls for size (log market cap), momentum (12-1 month return), volatility (21-day realized vol), and sector fixed effects.
Walk-forward validation
We use a walk-forward split with embargo (also called purged k-fold). The training set ends at date t; the evaluation window starts at t + embargo_days, where embargo_days = horizon + 5 trading days. This prevents target leakage from overlapping return windows.
We never use random splits, which would allow future data into the training set for time-series problems with overlapping labels.
Results labeled backfilled use historical data that was collected after the fact (e.g., Wikipedia pageviews from 2020–2024 fetched in 2026). Results labeled live use data actually collected in real time. These are always reported separately. Backtested results are not presented alongside live results without an explicit label.
Multiple-testing correction
This study tests multiple signals × horizons × cohorts simultaneously. All tests are listed in the pre-registration. We apply the Benjamini-Hochberg (BH) procedure at a false discovery rate of q = 0.05 across all tests. A signal is declared "significant" only after passing this correction.
The number of trials (N_trials) is counted honestly in the pre-registration. Every test—including null results—is reported. We do not selectively report only significant findings.
We also report the pre-correction p-values alongside BH-adjusted decisions, so readers can apply their own correction if desired.
Baselines
Every signal is compared against at least these baselines:
- Zero: Predict zero excess return for all securities (naive).
- Momentum: Past 12-1 month return rank as predictor.
- AR(p): Autoregressive model of the signal itself (does the signal predict itself?).
- Sector mean: Predict each security earns its sector's average excess return.
A signal must beat its best baseline out-of-sample, after multiple-testing correction, to be considered "working." Improvements over the zero baseline but not momentum are interpreted as capturing a momentum effect, not an independent alt-data signal.
Power analysis & timeline
The minimum detectable IC at a given sample size can be estimated from: t = IC × sqrt(N) / sqrt(1 − IC²). For a two-sided test at α=0.05 after BH correction (adjusted α ≈ 0.05/N_trials), we need IC ≥ 0.03 and N ≥ 500 observations (date-security pairs per test) for 80% power at a typical IC of 0.05.
With ~500 S&P 500 securities and daily observations, we accumulate ~500 obs/day. The minimum history needed for stable walk-forward evaluation is approximately 60–90 trading days (~3 months). With <1 year of live data, no result should be considered conclusive.
Failure criteria
A signal is declared to have no usable predictive power if, after accumulating the minimum required history (see Power above), the BH-corrected test fails to reject the null at q=0.05, AND the point estimate of IC is below 0.02 for every horizon and cohort.
Failure is reported prominently on the signal's page. We do not remove failed signals from the site or from the evaluation framework. The pre-registration defines these criteria; they cannot be changed post-hoc.
Amendments log
Changes to the study design after the pre-registration commit are documented in
study/AMENDMENTS.md. Each entry includes: the date, what changed,
why it changed, and whether it was a pre-specified decision rule or a new decision.
Data sources
Full source documentation including terms, what we store, and what we publish:
We only ingest sources whose terms permit our use. Sources that prohibit bulk redistribution are used to generate derived statistics only — the raw files are not published. Sources with unclear terms are flagged and skipped by default.
Known limitations
- Short history: The live data archive starts October 2026. Backfilled data is labeled separately and treated with appropriate skepticism.
- Survivorship bias: The point-in-time S&P 500 membership list mitigates this but is imperfect for older dates.
- Entity matching: GDELT and Wikipedia mapping to companies is approximate. Precision is audited quarterly.
- Transaction costs: Quintile spreads are gross of transaction costs. A stated TC assumption of 5 bps per leg is applied in sensitivity checks.
- Market microstructure: At 1-day horizon, execution timing assumptions strongly affect results.
- Signal decay: If this work becomes widely known, any genuine signal will decay as more participants act on it.
- Sector concentration: The retail-attention basket is concentrated in technology and consumer discretionary.