Four-Factor Macro Tail-Risk Model
Companion white paper for the dashboard
Version: v0.9
Date: 2026-07-17
Dashboard path: index.html
Machine-readable summary payload: model_summary.json
Full dashboard payload: dashboard_payload.json
Abstract
The model began as a liquidity-impulse thesis: financial regimes can become self-reinforcing when earnings growth, credit creation, buying power, asset prices, and margin capacity reinforce each other. The empirical work narrowed that idea.
The primary model is now a decile rank model. This is more legible than a named-regime model because each decile contains roughly the same number of historical months. The question becomes simple:
Where are we in the historical distribution of macro tail-risk pressure?
Most of the distribution is noisy. The useful signal is concentrated in the upper tail, especially the ninth and tenth deciles.
Epistemology And Scope
This model is not trying to forecast the next recession catalyst. It is trying to measure whether the equity market is being priced and financed in a way that leaves it fragile if a catalyst arrives.
That distinction matters. The Federal Reserve's financial-stability framework separates shocks from vulnerabilities: shocks are often surprises and hard to predict, while vulnerabilities build over time and determine how badly the system can respond under stress. That is also the working distinction in this dashboard. COVID, an AI capex cycle turn, a geopolitical oil shock, or a bank run can be the visible trigger. The model is aimed at the pre-existing risk-compensation and leverage state, not the narrative form of the trigger.
The model therefore makes a narrower claim:
We usually cannot predict the shock.
We may be able to measure whether the market is priced and levered in a way
that makes disappointment more dangerous.
This is why the dashboard is framed as a tail-risk and beta-budgeting tool, not a general return forecast. In most months, the model should not be expected to have strong signal. The historical evidence is concentrated in upper-tail states where several partially distinct causal pressures are elevated at the same time.
The practical use case is therefore:
How much long beta do I want?
How much long-volatility / put exposure do I want?
Is the current setup historically ordinary, elevated, or extreme?
It is not:
Can this model forecast next month's return?
Can this model identify the exact exogenous shock?
Can this model replace security-level analysis?
Causal Intuition
The original hypothesis was a liquidity-cycle mechanism:
earnings growth and optimism
-> credit creation and financing capacity
-> buying power
-> asset-price appreciation
-> higher collateral and margin capacity
-> more buying power
That loop can be rational for a long time. If growth expectations are genuinely improving, low risk premiums and rapid capital formation may be justified. The danger comes when the same loop pushes risk compensation low and leverage high enough that a modest disappointment can have an outsized effect.
The mechanism is close to the leverage-cycle literature: Adrian and Shin emphasize that credit availability can vary through intermediary leverage, while Baron and Xiong argue that credit expansions can coincide with neglected crash risk. The final model keeps that intuition but stops pretending that "liquidity" is one measurable scalar. The literature and the local backtests both argued against that simplification. Credit impulse, public liquidity, bank balance-sheet capacity, securities leverage, spreads, and valuation all move through related but different channels.
The model therefore uses four causal legs:
| Leg | Dashboard Factor | Causal Role |
|---|---|---|
| Equity risk compensation | Synthetic ERP tightness | Are equity holders being paid enough for risk, after rates and expected growth? |
| Credit risk compensation | Baa spread tightness | Are credit investors accepting unusually little compensation for default and liquidity risk? |
| Balance-sheet saturation | Broad stock saturation | Is broad non-broker debt/leverage stock stretched relative to history? |
| Demand-side securities leverage | Broker-customer leverage saturation | Is securities financing already heavily used by customers and leveraged buyers? |
The dashboard is intentionally about the interaction between these legs. A single elevated factor can be noisy. Several elevated factors at once are more concerning because they describe a market with lower compensation, tighter credit pricing, and less unused balance-sheet elasticity.
Why We Do Not Test Every Plausible Factor
Macro backtests are a small-sample problem. The post-1960 monthly dataset sounds large, but true independent crisis episodes are scarce. Twelve-month forward outcomes overlap. Many macro series revise. If we test dozens of plausible indicators, then reweight the winners, then add interactions, then tune thresholds, we will almost certainly create a model that explains history better than it will help in real time.
The research policy is:
1. Start with a causal mechanism before seeing the result.
2. Prefer public, refreshable data.
3. Prefer simple primitives over fitted composites.
4. Normalize with the same rolling z-score discipline where appropriate.
5. Use live prior thresholds for historical flags.
6. Compare against simple benchmarks.
7. Check overlap so breadth is not just double-counting one latent variable.
8. Reject or quarantine candidates that only work after repeated tuning.
This is not anti-empirical. Data should change priors. The rule is that data can discipline a first-principles hypothesis; it should not let us manufacture a new hypothesis after seeing every table. That is why the final dashboard keeps a small number of pre-specified causal factors and uses a simple 75/25 score-plus-breadth calibration instead of a fitted regression.
The local research notes preserve rejected or diagnostic work instead of silently deleting it. Growth variables, forecast revisions, earnings-expectation proxies, NFCI/yield-curve variants, and pure liquidity impulse screens were researched. Some were directionally interesting. They were not promoted because they were too noisy, too overlapping with ERP, too coincident with already-visible stress, too short-history, or too dependent on ex-post variable selection.
Signal, Explanation, And Calibration Layers
The dashboard separates three concepts that are easy to confuse.
Explanation layer:
the economic story behind the factors and historical episodes
Signal layer:
the four normalized factor scores
Calibration layer:
historical deciles and forward-outcome base rates
The explanation layer is necessary because the model should not be a black box. The signal layer is deliberately simple: four comparable z-scores, equal-weighted. The calibration layer asks what happened historically when the 75/25 score-plus-breadth rank was in a similar zone.
The old named-regime taxonomy was useful for research but too verbose as a user-facing primary output. The decile model is cleaner:
D1-D7: mostly noisy
D8: elevated watch zone
D9: elevated tail-risk zone
D10: historically ugly zone
That language is less satisfying than a precise forecast, but it is more honest.
Current Read
Current model date: 2026-06-30
Primary decile: D10
75/25 risk rank: 93.2%
4-factor score: +0.93
Score rank: 91.5%
Breadth rank: 98.6%
Breadth: 3/4 active factors
Current factor states:
| Factor | Score | Live Flag Threshold | Flagged |
|---|---|---|---|
| Synthetic ERP tightness | +1.29 | +1.17 | Yes |
| Baa spread tightness | +1.28 | +1.13 | Yes |
| Broad stock saturation | -0.06 | +0.60 | No |
| Broker-customer leverage | +1.20 | +0.71 | Yes |
The current state is still D10, but it has moved deeper into that bucket. At 93.2%, it is no longer just barely over the top-decile cutoff: breadth has risen from 2/4 to 3/4 active factors, and Baa tightness has joined ERP tightness and broker-customer leverage as a live contributor.
Source-freshness note: the July 17, 2026 maintenance run completed the standard --refresh rebuild successfully against the live public inputs. The published payload remains dated June 30, 2026, now reconfirmed from a live refresh, and still preserves coverage for all 40 required public series with no missing inputs or failed fetches. Consumer credit (TOTALSL) remains current through 2026-06-30, keeping the prior stale-source issue cleared. Some slower-moving public inputs still naturally lag the model cut, including the S&P CoreLogic Case-Shiller U.S. home price index (CSUSHPINSA, last observation 2026-04-30) and the synthetic ERP backcast (2026-04-30), but they remain inside the payload's freshness rules and are not currently flagged stale.
Current Interpretive Overlay
This section is deliberately separated from the model definition. It records the current Bayesian / qualitative overlay on top of the statistical read. The model output is the model output; this section is the human interpretation of why today's deeper D10 read may or may not deserve a more aggressive risk-management response.
The strict model read is:
Current state: top-decile risk, deeper D10.
Current rank: 93.2%, with 3/4 active factors.
Empirical signal: the model remains inside the part of the calibration where history gets materially worse, and breadth has now risen further.
The working interpretation is that the warning has strengthened since the prior low-end D10 read. The historical tables suggest that the first seven or eight deciles are mostly noisy, D9 is a caution zone, and D10 is where the forward drawdown and volatility tables become materially uglier. The move from 91.4% to 93.2% matters because it is not just a within-bucket drift: breadth has stepped up from 2/4 to 3/4, which is exactly the kind of cross-channel reinforcement the model is designed to detect. In risk-budgeting language:
D9:
stay awake, avoid complacency, consider efficient convexity
Low D10:
stronger case for carrying hedges and trimming marginal beta
Deeper D10:
stronger case for more explicit de-risking
This matters because the current D10 still does not read like a full-stack debt-saturation signal. The active pressure remains mostly market-side, and the July 17 refresh preserved that structure while broker-customer leverage strengthened to +1.20z and broad stock saturation remained slightly below zero:
| Channel | Current State | Interpretation |
|---|---|---|
| Synthetic ERP tightness | Flagged | Equity risk compensation is thin. |
| Baa spread tightness | Flagged | Credit investors are now also accepting unusually little compensation. |
| Broad stock saturation | Not flagged | Broad non-broker debt/leverage stock is still not the binding source of risk. |
| Broker-customer leverage | Flagged | Securities financing / market leverage remains the clearest non-valuation warning. |
The broad-stock leg being benign does not mechanically prevent D10. Because the model is 75% magnitude and 25% breadth, elevated ERP tightness, elevated Baa tightness, and elevated broker-customer leverage can push the average score into the top historical decile even if broad stock saturation remains low.
Example:
ERP tightness: +2.5z
Baa tightness: +2.0z
Broad stock saturation: -0.8z
Broker-customer leverage: +2.5z
Equal score = (+2.5 + 2.0 - 0.8 + 2.5) / 4 = +1.55z
That would likely be a high historical score. With three of four factors flagged, the breadth leg would also contribute materially. So the model can describe a market-side danger regime without requiring the whole real economy to be overlevered. That is effectively what the current reading is saying: the decile escalation is coming from concentrated valuation / credit-pricing / securities-leverage heat, not from a broad debt-saturation confirmation.
The main speculative overlay is that the current AI cycle may allow an elevated regime to persist or intensify before it breaks. The closest historical rhyme is the dot-com period: a transformative technology narrative, major capex build-out, public-market speculation, tight risk compensation, and eventual uncertainty over realized ROIC. The current AI regime differs in at least three important ways:
- Broad non-broker balance-sheet saturation is not currently as stretched as it was around the dot-com/GFC-style episodes.
- The perceived total addressable market and macro impact of AI may be much larger than the internet narrative appeared at comparable stages.
- The market-side speculative channel can be strong even when broad debt/GDP is not the binding constraint.
This overlay does not say that AI will succeed, that the current market is not a bubble, or that low-end D10 is harmless. It says that if AI remains a powerful capex, productivity, earnings-expectation, and animal-spirits driver, then a concentrated market-side regime can stay hot for longer than a simple "top decile equals immediate collapse" rule would imply.
Near-term narrative catalysts can also matter. The public backdrop still looks more consistent with an animal-spirits / issuance-heavy regime than with a clean valuation reset, but the early-July news flow did shift toward a louder exogenous-risk narrative: renewed Iran ceasefire fragility and oil volatility briefly hit global risk assets even while AI leadership continued cushioning the major U.S. indexes. The July 16 live refresh did not reverse the key qualitative point from the prior run: the factor mix is no longer just ERP plus broker-customer leverage. Baa tightness remains in the active set, and the slight cooling in broad stock saturation was not enough to pull the signal out of deeper D10. That narrative observation is still secondary to the model output, but it matters for how seriously the D10 read should be treated. One qualifier from the v0.8 diagnostics layer: orthogonalized against ERP (the latent factor all three market-side legs echo), the 3/4 breadth currently decomposes to ONE independent elevated pressure — the Baa and broker readings sit at the 70th and 63rd percentiles of their ERP-residuals, below flag levels. The raw breadth is real as displayed, but "three independent channels confirming" overstates it; the historical terminal months (2000-02, 2007-06, 2021-01) all read 3/4 on the orthogonalized ladder, while today reads 1/4 — the on-ramp shape, matching the 1997-03 nearest analog.
One more phase consideration (v0.9). The owner's standing thesis is that if the underlying technology is genuinely epochal, this episode should ultimately read like the dot-com era at larger amplitude. The model's data half-agree already: on the dividend-yield metric the market is ALREADY past the 1999 record (0.91-1.09% bounded estimate vs 1.17%), while the composite z (+1.23 vs +2.2-2.6 at the 1999/2021 peaks) and the orthogonalized breadth (1/4 vs 3/4 at every terminal month in the sample) say the speculative-velocity phase has not arrived. If the thesis is right, the melt-up leg ahead could exceed the 1997 analog's +40% twelve-month sequel before the reckoning, and the eventual unwind should be sized above the historical base rates. The disagreement between thesis and model is dated and watchable rather than philosophical: the thesis is confirmed if the escalation dials fire (residuals reaching 3/4, an issuance flood) BEFORE the top; it is weakened if the market tops from the current on-ramp configuration. And the standing discipline: epochal technology lengthens the right tail; it says nothing about entry prices. The 1999 bulls were right about the internet and still wrong about the decade of returns from the 2000 entry.
Owner amendment (2026-07-12, medium confidence, recorded from review). The magnitude leg of the thesis is now explicit: if this episode completes, the owner expects the melt-up to exceed the dot-com arc in amplitude. The causal weight in that claim sits on capital structure, not on the technology itself: transformativeness mainly extends how long disconfirmation can be deferred and puts an earnings floor under the eventual bust, while the height of the run is set by the capital stack. The 2026 stack differs from 1999's in ways the dials already see: structurally negative float in the buyback regime (2026-Q1 was the first positive net-issuance quarter since the 2021 flood, against 1997's immediately-open 474-deal window), a private pipeline funding unprofitable mega-caps at scale before they ever list, a loose global funding channel, retail margin channels engaged in Japan and Korea, and fast money positioned short (melt-up fuel) rather than crowded long.
Two constraints hold the claim at medium rather than high confidence. First, composition: the dividend-yield metric is already past the 1999 record, so a larger-than-dot-com run must be substantially earnings-carried rather than multiple-carried — which makes the hyperscaler capex/OCF dial the load-bearing tell for this specific claim. Operating cash flow catching up to the record 0.75 ratio validates the price; the ratio being sustained above 1.0 on debt issuance while the D&A gap widens is the capex-bust-without-a-1999 branch, the main way the amplitude claim fails. Second, the marginal buyer: asset managers are already near max-long, so incremental melt-up flow must come from retail, leverage, corporate, and foreign channels — partly visible in the hot Japan/Korea margin dials, but the passive-share and options-positioning legs of that handoff are not instrumented here, and that is where the confidence discount lives. The amendment changes no model output and no risk posture; a larger expected melt-up strengthens narrative capture near the top, which raises rather than lowers the authority of the pre-committed escalation dials, and it sizes the eventual unwind above, not below, the historical base rates.
Current qualitative risk posture:
Model-only:
top-decile tail-risk, now driven by valuation, credit pricing, and market leverage
Human overlay:
the regime can still run, but the hedge bar is higher than it was in low-end D10
Risk-management implication:
treat this as a real D10 warning, not just a watch state
prefer explicit convexity / hedge carry over complacent unhedged beta
escalate further if the rank pushes deeper into D10, breadth rises to 4/4,
or broad stock saturation joins the already-flagged market-side factors
sharper escalation dial (v0.8): Baa or broker ERP-residuals crossing
their 80th prior percentiles = independent confirmation (the
2000/2007/2021 terminal shape); the global funding channel arming
(dollar z through +1 with USD credit growth rolling over) = the
transmission risk to non-US sleeves going live
Context And Diagnostics Layer (v0.8)
Beyond the dashboard's frozen primary model, a diagnostics layer now surrounds the current read. Nothing in it is scored; each piece is a labeled, gated, independently-refreshable diagnostic with a standing doc under docs/. The v0.8 state of each dial:
| Dial | Current read | Doc |
|---|---|---|
| Orthogonalized breadth (P7) | 1/4 independent pressures (raw 3/4); terminal months historically read 3/4 | TWO_BLOCK_ORTHOGONAL_BREADTH_v0.1.md |
| Two-block state (P5) | compensation +1.23z, saturation -0.01z; 12m compensation path +0.90 to +1.23 | same doc |
| Nearest analog (P6) | 1996-03 to 1997-06 cluster, closest month 1997-03 (on-ramp, not blow-off) | NEAREST_NEIGHBOR_REFERENCE_v0.1.md |
| JP/KR margin, normalized (P4) | JP 0.51% of TSE cap = 58th pctile (flow-hot, stock mid-range); KR 0.51% = 55th pctile, HALF the 2021 peak — intensity fell through the parabola | MARGIN_NORMALIZATION_v0.1.md |
| Domestic two-leg cells (P4) | Korea = valuation_only (the USA-1999/US-2026 cell); Japan = neither | DOMESTIC_TWO_LEG_STATES_v0.1.md |
| AI-cycle fuel line (P14) | hyperscaler capex/OCF 0.75 TTM = series record; AMZN and ORCL above 1.0; complex D&A growing 10pp faster than revenue | XBRL_CAPEX_GAUGES_v0.1.md |
| Global funding channel (P15) | UNARMED: dollar z +0.33 and falling, global USD credit +8.5% y/y — loose funding is fuel, not squeeze | GLOBAL_FUNDING_CHANNEL_v0.1.md |
| Issuance/supply axis (P12a, v0.9) | net-issuance z +1.28 and rising; 2026-Q1 = the first positive net-issuance quarter since the 2021 SPAC flood (2021 top read +2.82); quality dial (Ritter): 2025 = 90 IPOs at 53% EPS<0 vs 75-81% at every mania year — quantity response started, quality flood NOT arrived | ISSUANCE_SUPPLY_v0.1.md |
| Insider supply (P12b, v0.9) | count sell-share 84.8% (z +1.13) vs +2.02 at the 2021 selling peak — elevated, not manic; the dial's only buy-majority quarters in 20 years are the 2008-09 capitulation | INSIDER_POSITIONING_v0.1.md |
| Futures positioning (P12c, v0.9) | leveraged funds net SHORT -18.3% of OI (z -1.22) — the mirror of the 2018/2021 crowded-long air-pocket setups; asset managers near max-long (+49.6%, z +1.79), a slow-unwind reservoir | same doc |
Composite v0.8 read: a top-decile, valuation-led, single-independent- pressure state — the on-ramp shape — with the AI capex loop still fuel-pressurized but increasingly debt-funded at the margin, domestic Asian melt-ups running on valuation rather than leverage stock, and the international transmission channel not yet armed. The escalation dials that would change this composite are enumerated in the risk-posture block above and in each doc's "changes if" line. Methods, evidence classes, and the finding-by-finding ledger live in docs/improvement_program_notes/.
Primary Model
The dashboard's primary scalar is:
primary risk rank =
0.75 * four-factor score percentile
+ 0.25 * breadth percentile
Then that rank is mapped into equal-sized historical deciles.
This is not a pure flag model. The continuous score carries most of the weight. Breadth is included because the empirical work suggested that multiple partially independent causal pressures firing together is more informative than one large z-score in isolation.
The 75/25 split is a deliberately modest concession to breadth. A pure score model can hide the difference between one very extreme factor and several moderately elevated factors. A pure flag-count model throws away too much magnitude information. The current structure keeps magnitude dominant while still recognizing that breadth is part of the causal story:
75% = how elevated is the average causal pressure?
25% = how many distinct causal pressures are active?
This is not meant to be the mathematically optimal weighting. It is meant to be a disciplined, interpretable weighting that does not require fitting coefficients to a small number of historical crisis episodes.
The four-factor score is:
4-factor score =
average(
synthetic ERP tightness,
Baa spread tightness,
broad stock saturation,
broker-customer leverage saturation
)
How The Calculation Works
The dashboard updates once the input data are refreshed. It cuts the model at the latest completed month, not the current partial month. That avoids treating an incomplete month as if it were a true month-end observation.
The model turns raw macro and market series into comparable z-scores. A z-score means:
z-score = (current value - trailing average) / trailing standard deviation
The trailing window is 120 months where possible, with at least 36 months required. Scores are clipped at +/-3 so that one extreme historical observation cannot dominate the model.
Example interpretation:
+1.0z = about one trailing standard deviation above normal
0.0z = around trailing normal
-1.0z = about one trailing standard deviation below normal
The sign is adjusted so that higher always means more tail-risk pressure:
low ERP -> high ERP tightness score
tight Baa spread -> high Baa tightness score
high broad stock saturation -> high saturation score
high broker leverage -> high broker-customer leverage score
Each month therefore has four comparable factor scores. The dashboard averages them into a four-factor score.
The dashboard also creates four factor flags. A factor is flagged when it is above its own live prior 80th percentile. "Live prior" matters: the threshold at a historical date only uses information that would have existed before that date.
factor flag = current factor >= prior 80th percentile
breadth = count of active factor flags
Then the model converts the continuous score and breadth into percentile ranks:
score rank = where today's four-factor score sits vs prior history
breadth rank = where today's breadth count sits vs prior history
The final rank is:
primary risk rank =
0.75 * score rank
+ 0.25 * breadth rank
That final rank is mapped into deciles:
D1 = 0-10%
D2 = 10-20%
...
D9 = 80-90%
D10 = 90-100%
This is the main dashboard output. The decile tells the user where the current macro risk setup sits in the historical distribution.
There are two related but different ranking modes in the dashboard:
Validation rank:
Uses live prior history and strict filters.
This is the main evidence object.
Historical chart rank:
Uses the frozen current formula to show older episodes on one common scale.
This is a visualization and event-study aid.
That distinction avoids a common dashboard error. The historical chart is meant to make the drawdown map legible; the validation tables are the place to judge predictive signal.
The Four Inputs
1. Synthetic ERP Tightness
synthetic ERP tightness = -rolling_z(synthetic ERP)
Lower ERP means investors are receiving less compensation for bearing equity risk. ERP is preferred to CAPE as the primary valuation variable because it directly incorporates rates and a growth assumption. CAPE remains a benchmark diagnostic.
Raw concept:
equity risk premium = implied expected equity return - risk-free rate
The model wants to know whether equity investors are being paid enough for risk. When ERP is low, future equity returns have less cushion against disappointment. That is why the sign is inverted:
lower ERP -> richer valuation -> higher risk score
The official monthly Damodaran ERP history is short, so the dashboard uses a synthetic Damodaran-style monthly backcast. The backcast uses annual Damodaran ERP anchors, monthly S&P price movement, monthly Treasury rates, and a growth/cash-flow reconstruction. It is not a naive interpolation; monthly price and rate moves affect the signal.
This choice is partly empirical and partly first-principles. CAPE has a deeper literature and a longer clean history, but it is mechanically backward-looking and does not explicitly include the risk-free rate or forward growth. Implied ERP is theoretically closer to the object we care about: the compensation investors demand after incorporating rates and expected cash-flow growth. The synthetic backcast is labeled clearly because it is model-derived, but it is preferred because the underlying economic object is better matched to the question.
2. Baa Spread Tightness
Baa tightness = -rolling_z(Baa spread)
Tight spreads mean credit risk is priced cheaply. The model treats this as low risk compensation, not as immediate stress.
Raw concept:
Baa spread = Moody's Baa corporate yield - 10-year Treasury yield
Wide spreads usually mean credit stress is already visible. This model is using the opposite condition: unusually tight spreads. Tight spreads can be benign, but at extremes they indicate that credit investors are accepting little compensation for default/liquidity risk.
The sign is inverted:
lower Baa spread -> tighter credit pricing -> higher risk score
This factor is not saying "credit is currently breaking." It is saying "credit risk is priced with less cushion."
This is intentionally the opposite of a standard credit-stress screen. Wide spreads often mean stress is already visible. Tight spreads can be a late-cycle risk-compensation signal: credit is calm, capital is available, and investors are accepting low compensation. That can support markets for a while, but it also means less cushion if defaults, funding costs, or growth expectations disappoint.
3. Broad Stock Saturation
broad stock saturation =
0.294 * broad debt / GDP z
+ 0.235 * nonfinancial business debt / GDP z
+ 0.235 * NFCI leverage subindex z
+ 0.235 * NFCI nonfinancial leverage subindex z
broad stock saturation score = rolling_z(broad stock saturation)
This is the broad non-broker balance-sheet saturation leg. It deliberately excludes FINRA margin debt, margin utilization, broker-dealer customer receivables, and broker-customer leverage growth. Those securities-financing inputs belong to the fourth factor.
The causal object is:
Is the broad economy / financial system debt stock stretched,
excluding the broker-customer securities-financing channel?
Its components are:
broad debt / GDP:
all-sector debt securities and loans relative to nominal GDP
nonfinancial business debt / GDP:
corporate and business debt securities and loans relative to nominal GDP
NFCI leverage:
Chicago Fed financial-condition leverage subindex
NFCI nonfinancial leverage:
Chicago Fed nonfinancial leverage subindex
The weights are coarse causal priors, not fitted coefficients. They are close to equal weight. Broad debt/GDP receives a small overweight because it is the broadest total-system stock variable. The other three inputs split the remaining weight across business leverage, financial-condition leverage, and nonfinancial leverage.
The important methodological change in v0.7 is the input firewall:
Factor 3 owns broad non-broker stock saturation.
Factor 4 owns broker-customer securities leverage.
The same primitive should not vote in both places.
Prior easing was tested as both a 50/50 state/path composite and as a separate fifth factor. It remains causally plausible, but it was not additive enough to promote into the primary score. It is better kept as a diagnostic note than smuggled into broad saturation with a compromise weight.
4. Broker-Customer Leverage Saturation, Composite B
Composite B measures broker-customer securities financing saturation:
leverage impulse anchor =
average(
stitched FINRA-like margin debt 12m growth z,
Z.1 broker-dealer customer receivables 12m growth z
)
broker-customer stock =
Z.1 broker-dealer customer receivables / GDP z
Composite B =
average(leverage impulse anchor, broker-customer stock)
Before direct FINRA margin-debt history is available, the model bridges with Z.1 broker-dealer customer receivables using the FINRA/Z.1 overlap ratio. This is a z-score timing bridge, not a precise historical dollar-level margin-debt claim.
Raw concept:
broker-customer financing = securities credit and receivables tied to customer leveraged exposure
This factor is meant to capture the demand-side leverage channel: when customers and leveraged investors have already expanded financing aggressively, the market has less unused balance-sheet elasticity. That matters for your original "rubber band" intuition. A stretched financing stock can support upside while the process continues, but it can also make the system more sensitive when price or funding conditions reverse.
Composite B has two blocks:
flow / impulse:
is broker-customer leverage growing unusually fast?
stock / saturation:
is broker-customer leverage high relative to GDP?
The model averages these two blocks so it does not treat growth and level as four separate hidden votes. The conceptual split is simply:
how fast is leverage expanding?
how stretched is the stock?
Composite B was added only after a separate read-only evaluation showed that it was both causal and reasonably independent from the three-factor stack. The important methodological point is not that margin data is magical. It is that broker-customer leverage measures a different part of the system from ERP or credit spreads: actual securities-financing demand and saturation.
Why Candidate Factors Were Rejected Or Kept Diagnostic
The dashboard is allowed to have supporting diagnostics without promoting them into the primary model. That is how the research process avoids both underfitting and overfitting.
Pure Liquidity Impulse
The original liquidity-impulse thesis was directionally plausible, especially through the Biggs, Mayer, and Pick distinction between the stock of credit and the flow of new credit. The problem was empirical robustness. Pure liquidity impulse screens were suggestive but not strong enough as standalone forward predictors. Stress-confirmed versions worked better, but those often included market stress that was already visible.
So liquidity impulse remains intellectually important and diagnostically useful, but it is not the headline score.
Growth And Growth Expectations
Growth variables were researched because they are the main possible justification for high valuations, tight spreads, and high leverage. In principle, a market can look stretched because the future really is unusually good. This is related to the broader growth-at-risk literature, which treats financial conditions as affecting the distribution of future growth rather than only the mean.
The screen split growth into realized growth, growth impulse, and expected-growth/revision variables. The result was not strong enough to justify a fifth primary factor. Realized GDP moved too slowly. Growth deceleration variables were directionally interesting but noisy. Expected-growth and earnings-revision variables often overlapped with ERP because ERP already embeds growth expectations.
The conclusion is:
Growth matters conceptually.
The available public growth/revision proxies did not add enough independent signal.
Profit Share
Corporate profit share was interesting because it can represent margin-cycle saturation: high profits may be cyclical, politically contested, or vulnerable to mean reversion. It helped some monotonicity tests, but the later four-factor calibration did not clearly need it. It also overlaps with the growth and valuation story that ERP already partially captures.
Profit share remains a research candidate, not a primary factor.
Yield Curve, NFCI, Volatility, And Market Stress
These variables can be useful diagnostics, but they are dangerous as primary predictors for this specific dashboard. If a factor mostly activates after the market is already in freefall, then it may improve a backtest while weakening the model's epistemic claim. The current dashboard wants pre-stress risk-compensation and leverage conditions, not just confirmation that stress is already happening.
This is why the dashboard can show recession bands, drawdowns, and funding context without giving them equal status inside the primary score.
The Named Missing Axis: Supply And Issuance (v0.9)
The four factors measure the PRICE of risk (ERP, Baa tightness) and the STOCK of fuel (broad saturation, margin). None measures the quantity response: how much new paper is being manufactured into the demand, and how aggressively informed holders are selling into it. That axis has the strongest causal pedigree of any candidate not yet in the model — issuer quality predicts subsequent returns (Greenwood and Hanson, RFS 2013), bubbles deflate when supply catches up with demand (Hong and Stein), and the equity share of issuance forecasts returns (Baker and Wurgler) — and it marked both template tops mechanically: the 2000 junk-IPO flood and the 2021 SPAC/unprofitable-IPO record. It also addresses this dashboard's widest error bar: the model times ENTRY into fragile states well and times tops poorly (D10 dwell has run 6-17 months); supply floods are how tops have historically announced themselves.
Status: this is P12a in the improvement program — diagnostic-first (Z.1 net equity issuance plus issuance-quality proxies; aggregate insider Form-4 selling as the informed-supply follow-on; CFTC positioning as a fast dial, never a factor), with promotion gated by the frozen-spec rules like everything else. The other demand-side candidate, household equity allocation, was probed and reads +2.0z at the 96th percentile with no historical precedent for its joint condition; it is reported as a level rather than ranked, because every demand-side variable percentile-ranked on this era reads extreme — it adds level, not discrimination.
First cut executed 2026-07-10 (ISSUANCE_SUPPLY_v0.1.md): the net- issuance dial separates the template tops correctly — 2021-Q1 read z +2.82 (the SPAC flood) while 2007 read -2.03 (LBOs retiring equity: a credit top, not a supply top) — and the current state is the finding: 2026-Q1 printed the FIRST positive net-issuance quarter since the 2021 flood (+$124B; the only other positive quarters of the buyback era are 2020-Q3 and 2021-H1), with the z at +1.28 and rising. Three independent phase dials now agree the episode is on-ramp with the supply response assembling: compensation +1.23, orthogonalized breadth 1/4, issuance +1.28 — each well below its terminal-month readings.
The quality increment (Ritter's data, wired the same day) completes the supply read: every mania year in the 1980-2025 sample printed a negative-EPS IPO share of 75-81% (1999, 2000, 2020, 2021; the non-mania maximum is 2007's 56%), while 2025 read 53% on a cohort of just 90 deals — the IPO window is only now reopening after the post-2021 shutdown. Quantity response started; quality deterioration not yet visible. The mega-cap AI pipeline, dominated by unprofitable issuers, is the cohort that would move this dial when it prices.
P12b and P12c completed the block: aggregate insider selling (SEC structured Form-4 data, count convention) reads z +1.13 against +2.02 at the 2021 selling peak — and the only buy-majority quarters in the twenty-year sample are the 2008-09 capitulation, the dial's face validity at both extremes. Futures positioning (CFTC TFF) shows the MIRROR of the known air-pocket setups: leveraged funds lean net short (z -1.22, vs +1.85/+1.65 into Jan-2018 and late-2021) while asset managers sit near their sample-max long. The completed supply block: quantity started, quality clean, insiders elevated-not-manic, fast money not crowded long, allocators stretched — the fifth and sixth independent dials agreeing on elevated-but-sub-terminal.
Historical Calibration
The model is evaluated with current-vintage public data and strict ex-ante filters:
Nasdaq not already in >10% drawdown
prior 3-month Nasdaq return > -10%
credit stress <= +0.5z
market fragility <= +0.5z
Current D10 outcomes versus base rates:
| Asset | D10 Months | Avg 12M Return | Base Avg 12M Return | 12M DD <= -10% | Base DD <= -10% | 12M Loss <= -10% | Base Loss <= -10% |
|---|---|---|---|---|---|---|---|
| NASDAQ | 19 | -7.5% | +11.0% | 68.4% | 20.5% | 52.6% | 5.4% |
| SP500 | 19 | +1.8% | +9.2% | 52.6% | 15.1% | 31.6% | 3.9% |
| DOW | 19 | +1.1% | +7.9% | 52.6% | 15.1% | 15.8% | 2.9% |
The current bucket is no longer the mixed D9 caution zone. In the historical D10 sample, Nasdaq forward outcomes were outright bad on average and broad-market drawdown odds were materially worse than base rates. This is exactly why the dashboard should be read as a tail-risk model rather than a point-return model: the key information is not that every D10 month immediately crashes, but that the distribution of forward outcomes becomes much uglier once the rank and breadth push into this zone.
The July 12 refresh did not change these calibration tables. What it did confirm is that the live state still belongs in this historical bucket: the current rank remains 93.2% with 3/4 active factors, so the same D10 base-rate framing still applies without any need to re-fit the model.
The SP500 row now uses a merged long/current monthly price series rather than the shortened public FRED response history. This makes the SP500 validation sample comparable to NASDAQ and DOW instead of being dominated by recent history.
Historical Drawdown Event Map
The dashboard now includes a historical event map for interpretation. It uses the S&P 500 as the broad market drawdown tape, overlays NBER recessions, and plots the frozen 75/25 score-plus-breadth risk rank against actual market drawdowns.
The event table is not a separate model. It is a presentation layer asking:
Before the S&P 500 first crossed -10% from a local peak,
had the 75/25 risk rank already moved into an elevated zone?
The table uses a historical calibrated rank so older episodes can be reviewed under the frozen current formula. The rule is:
Yes = max prior 12-month 75/25 rank >= 80%
Partial = max prior 12-month 75/25 rank between 70% and 80%
No = max prior 12-month 75/25 rank below 70%
N/A = model inputs unavailable
| Event | Crossed -10% | Trough | Max Drawdown | Pre-Signal | Max Prior Rank | Ex-Post Catalyst |
|---|---|---|---|---|---|---|
| 1990 recession / Gulf War | 1990-09-30 | 1990-10-31 | -14.8% | No | 62.2% / D7 | Oil shock from Iraq's invasion of Kuwait, late-cycle tightening, S&L stress, and the 1990-91 recession. |
| Russia / LTCM crisis | 1998-09-30 | 1998-09-30 | -11.8% | Yes | 96.2% / D10 | Asian crisis aftershocks, Russia default, LTCM deleveraging, and a brief global funding shock. |
| Dot-com bust | 2000-12-31 | 2001-09-30 | -29.7% | Yes | 98.4% / D10 | Technology valuation unwind, capex reversal, recession, and accounting scandals. |
| Global Financial Crisis | 2008-01-31 | 2009-03-31 | -50.8% | Yes | 88.9% / D9 | Housing bust, structured-credit losses, bank/dealer leverage unwind, and funding-market seizure. |
| Eurozone crisis / US downgrade | 2011-08-31 | 2011-09-30 | -15.5% | No | 25.2% / D3 | Eurozone sovereign crisis, US debt-ceiling shock, S&P downgrade, and global growth fear. |
| Q4 2018 Fed / trade shock | 2018-12-31 | 2018-12-31 | -14.0% | Yes | 86.6% / D9 | Fed tightening, quantitative tightening, US-China trade war, and growth scare. |
| COVID shock | 2020-03-31 | 2020-03-31 | -20.0% | Yes | 86.9% / D9 | Exogenous pandemic shock, lockdowns, forced deleveraging, and emergency policy response. |
| Inflation / Fed-hike bear market | 2022-04-30 | 2022-09-30 | -24.8% | Yes | 98.7% / D10 | Inflation shock, aggressive Fed hiking, rate-duration repricing, and valuation compression. |
The important epistemic point is the COVID row. The model did not and could not forecast the specific pandemic catalyst. But the system was already inside an elevated historical risk zone before the shock. That is consistent with the model's intended use: it is not a shock forecaster; it is a fragility and risk-compensation monitor.
NASDAQ Primary Deciles
| Decile | Months | Avg Rank | Avg Score | Avg Breadth | Avg 12M Return | 12M DD <= -10% | 12M Return <= -10% |
|---|---|---|---|---|---|---|---|
| D1 | 22 | 9.3% | -0.97 | 0.00 | +12.9% | 4.5% | 0.0% |
| D2 | 22 | 17.3% | -0.67 | 0.00 | +15.0% | 27.3% | 0.0% |
| D3 | 21 | 24.9% | -0.50 | 0.05 | +12.1% | 23.8% | 0.0% |
| D4 | 22 | 32.6% | -0.25 | 0.05 | +17.1% | 0.0% | 0.0% |
| D5 | 21 | 46.2% | -0.06 | 0.76 | +10.5% | 9.5% | 0.0% |
| D6 | 22 | 56.0% | +0.17 | 1.00 | +13.1% | 9.1% | 0.0% |
| D7 | 19 | 64.7% | +0.36 | 1.11 | +9.8% | 5.3% | 0.0% |
| D8 | 18 | 74.3% | +0.52 | 1.50 | +10.5% | 33.3% | 5.6% |
| D9 | 18 | 83.5% | +0.68 | 1.83 | +14.1% | 33.3% | 0.0% |
| D10 | 20 | 95.0% | +1.39 | 2.95 | -6.1% | 65.0% | 50.0% |
This is why the decile model is the cleanest dashboard language. The middle of the distribution is noisy. D8 and D9 are elevated risk zones, but not deterministic bearish forecasts. D10 is the clearest risk-off bucket.
Tail Bands As Cross-Check
The named bands remain useful as a diagnostic, but they are no longer the primary model.
| Band | Rule |
|---|---|
| Normal | 0 flags and score below high threshold |
| Watch | 1 flag or high score |
| Elevated | 2 flags, or 1 flag plus high score |
| High risk | 3 flags, or 2 flags plus very-high score |
| Extreme | 4 flags, or 3+ flags plus extreme score |
The decile model is preferred because it is easier to read and uses equal-sized buckets.
Caveats
This is a current-vintage public-data model, not a fully real-time trading simulation. Some input series revise, and quarterly data are carried forward until the next observation arrives. The v0.8 vintage audit prices this caveat: at twelve historical decision dates the revising inputs differed from what was known in real time by 2-9% in level (the Z.1 broker leg by up to 18%) but by under 1pp in the growth transforms the factors consume, and where revisions did matter they ran one way — spike-era first prints overstated and were later revised down, so the revised-data backtest, if anything, understates what a live model would have flagged (see VINTAGE_REVISION_AUDIT_v0.1.md).
Composite B is model-derived before direct FINRA data starts. It is useful for z-score timing, not for precise historical level claims.
The model does not forecast exogenous shocks. It estimates whether the market is in a historically stretched risk-compensation and leverage-saturation state if disappointment arrives.
The historical output should be read as base-rate evidence, not as a guarantee. D9 and D10 have been meaningfully worse than normal historically, but the middle deciles are noisy and even high-risk states can persist while markets keep rising.
References And Source Notes
External research and data anchors:
| Topic | Source | How It Informs The Model |
|---|---|---|
| Shocks vs vulnerabilities | Federal Reserve Financial Stability Report framework | Supports the distinction between forecasting catalysts and measuring vulnerability. |
| Credit / vulnerability early warnings | BIS early-warning indicators of banking crises | Supports the idea that credit, debt-service, and balance-sheet pressure are vulnerability indicators, not exact timing tools. |
| Credit impulse | Biggs, Mayer, and Pick, Credit and Economic Recovery | Supports separating credit flow/impulse from credit stock. |
| Procyclical leverage | Adrian and Shin, Procyclical Leverage and Value-at-Risk | Supports the balance-sheet amplification mechanism behind the leverage-cycle intuition. |
| Credit expansion and crash risk | Baron and Xiong, Credit Expansion and Neglected Crash Risk | Supports the idea that credit expansion can lower perceived risk while increasing future crash risk. |
| Growth-at-risk | Adrian, Boyarchenko, and Giannone, Vulnerable Growth | Supports targeting downside distributions rather than only mean returns or mean GDP. |
| Financial-conditions impulse | Federal Reserve FCI-G | Public analogue for measuring financial conditions as impulse to future growth. |
| Implied ERP data | Aswath Damodaran ERP data page | Anchor source for the official and annual ERP data used in the synthetic ERP backcast. |
| Baa spread | FRED BAA10Y | Public source for Moody's Baa yield relative to the 10-year Treasury yield. |
| Margin statistics | FINRA margin statistics | Public source for customer debit and credit balances in securities margin accounts. |
| Current Iran / market narrative | AP market report, May 26 2026 and Axios Iran/oil note, May 26 2026 | Used only in the interpretive overlay, not the model. |
| AI mega-cap IPO pipeline | Axios on SpaceX, OpenAI, and Anthropic IPOs, May 20 2026 | Used only in the interpretive overlay, not the model. |
| Ukraine ceasefire dynamics | Guardian Ukraine war briefing, May 5 2026 | Used only in the interpretive overlay, not the model. |
Internal project notes used to build this paper:
| Project Note | Role |
|---|---|
tools/macro_model/docs/DASHBOARD_REASONING_CHAIN_CONTEXT_v0.1.md | Captures the research arc from liquidity impulse to four-factor tail-risk model. |
tools/macro_model/docs/ORTHOGONAL_FLAG_RESEARCH_PLAN_v0.1.md | Defines candidate-admission policy and anti-overfit gates. |
tools/macro_model/docs/LIQUIDITY_IMPULSE_LITERATURE_REVIEW_v0.2.md | Summarizes liquidity, credit impulse, global liquidity, and growth-at-risk literature. |
tools/macro_model/docs/SYNTHETIC_ERP_BACKCAST_v0.1.md | Documents the synthetic ERP construction and validation against official monthly ERP. |
tools/macro_model/docs/VALUATION_SUBSTITUTION_COMPOSITE_TEST_v0.1.md | Documents the CAPE vs ERP substitution decision. |
tools/macro_model/docs/BROKER_CUSTOMER_LEVERAGE_COMPOSITE_B_HANDOFF_v0.1.md | Documents the fourth-factor candidate and its independence tests. |
tools/macro_model/docs/GROWTH_INTERACTION_SCREEN_v0.1.md | Documents why growth variables stayed research-only. |
tools/macro_model/docs/EARNINGS_EXPECTATIONS_REVISION_PROXY_SCREEN_v0.1.md | Documents the public earnings-expectations proxy work and why it was not promoted. |
tools/macro_model/docs/FORECAST_REVISION_SCREEN_v0.1.md | Documents SPF forecast-revision work and its limitations. |