Paper 4: The Conditional Probabilities Actually Hold
Signal-Level Calibration and Dashboard Utility of the Conditional Probability Exceedance Framework: Pairwise Validation, Gold Dashboard Evaluation, & Extended Portfolio Tilt Evidence Across 528 T.Ds
*This is the fourth post in a series about a live, publicly-accessible quantitative research framework I’ve been building. Previous posts covered the descriptive atlas (Paper 1), a negative result on SPY (Paper 2), and a mixed-positive result on portfolio tilts (Paper 3). Paper 4 is now a preprint on Zenodo/SSRN.*
The question that was never answered:
For three papers, I’ve been building and testing something called the Conditional Probability of Exceedance (CPE) framework. The idea is simple to state: when gold futures (GC=F) or Bitcoin ETFs are in the top 5% of their historical distribution over the past 6 months, what’s the probability that gold exceeds its historical median return over the next 6 months?
The framework computes these conditional frequencies across 161 instruments and roughly 170,000 predictor-target pairs. It outputs numbers like: “when IBIT’s 126-day return is above its 95th historical percentile, GC=F’s 252-day return exceeds its median with probability 0.94, versus 0.37 unconditionally.”
That’s a 2.54× lift. Whether that lift holds outside the training data is the whole question.
Papers 1 through 3 tested whether *using* these probabilities to make portfolio decisions worked. The answer was: sometimes, partially, with important caveats. But none of those papers ever tested the more fundamental question: **are the stated conditional probabilities actually calibrated?**
Paper 4 does that. And the answer is yes — at least for bullish signals, at least over this window.
What “calibrated” means and why it matters:
A probability is calibrated if things it predicts at 90% happen 90% of the time. A weather forecast that says 90% chance of rain and is right 70% of the time is not calibrated — it’s overconfident.
I test this by doing something simple: for every day in the test period (528 trading days, Jan 2025 to Jun 2026), for every catalog row whose predictor condition fired (e.g., IBIT in its historical upper tail), I record the signal and resolve the outcome the stated number of days later. Did gold actually exceed its historical median return threshold?
Across 103,983 resolved instances: **97.0% realised hit rate against 93.0% stated CPE**. The framework is slightly *conservative* — it understated its own edge by about 4 percentage points. The stated lift of 2.45× over the unconditional baseline was realised as 2.51×.
This matters because it means the probabilities in the catalog are genuine, not artefacts of data mining. When the gold dashboard says “87% probability of exceeding the 63-day median return given current predictor conditions,” that number is trustworthy.
The asymmetry: bullish works, bearish doesn’t
Not everything worked. Bearish signals — conditions that historically predicted gold falling — achieved only 29% hit rate against 83.7% stated CPE. That’s a catastrophic failure.
The specific instruments responsible: UVXY and VIXY, leveraged and unleveraged short-dated VIX ETPs. These were the dominant bearish predictors. In the training data, periods when UVXY and VIXY were in their extreme upper tails (volatility spikes) were followed by gold weakness — investors selling gold to cover equity margin calls.
In 2025, the opposite happened. Volatility spikes (the April tariff shock being the biggest) coincided with gold surges. The safe-haven mechanism dominated, and structural demand factors — central bank buying, de-dollarisation flows — that weren’t present in the training period kept upward pressure on gold even during risk-off episodes.
This isn’t a random failure. It’s a regime change, and it’s informative: the CPE bearish signals for gold encode a specific historical mechanism that broke down when structural demand factors changed. Paper 3 diagnosed the same instruments as having non-stationary quantile structures due to roll decay. Both papers are looking at the same phenomenon from different angles.
The gold dashboard: what it actually did
The gold dashboard at [quantarram.github.io/quant-regime-research/notebooks/gold_dashboard.html](https://quantarram.github.io/quant-regime-research/notebooks/gold_dashboard.html) aggregates these pairwise and joint CPE signals into a net directional view.
Here’s what it did over the test period:
**Throughout all of 2025:** BULLISH. Gold went from $2,629/oz in January to $5,318/oz on 29 January 2026 — a 102% gain. The dashboard correctly identified this as a sustained bullish regime and never wavered.
**From early February 2026:** BEARISH. Gold corrected from $5,318 to ~$4,100-4,400. The signal adapted — not because I changed anything, but because the balance of bullish versus bearish CPE signal weights shifted as the correction unfolded.
This is the dashboard’s central achievement: it maintained conviction through the entire bull run (including the April 2025 shock when gold temporarily dipped with equities) and pivoted at the top. Not because it “knew” the top was coming — it didn’t — but because the same predictor conditions that had been bullish for gold started generating more mixed signals as the price reached historically unprecedented levels.
Directional discrimination at 63 days: BULLISH signal days averaged +7.23% gold return versus +2.46% on BEARISH days. The +4.77 percentage point spread is meaningful for tactical allocation decisions.
The portfolio dashboard: why this paper is stronger than Paper 3
Paper 3’s portfolio tilt strategy found marginal significance — the hold-to-horizon variant passed a randomisation test (pct_exceeding 1.8%) but the main five-sleeve strategy was essentially indistinguishable from its neutral benchmark (Sharpe 0.560 vs 0.563).
Paper 4 finds t-statistics of 4.07 at 63 days (p < 0.001) and 5.51 at 126 days (p < 0.000001). The difference deserves an explanation, not just a number.
**Paper 3’s problem** was at the joint-configuration layer. The greedy search for multi-predictor signal combinations concentrated on UVXY and VIXY for the gold, bonds, and FX sleeves — because those instruments’ decay-inflated thresholds produced artificially high CPE values in training data. Those thresholds (log returns of -307% and -130%) were never re-cleared in live 2025 prices. Result: gold, bonds, crypto, and FX sleeves sat at neutral for all 250 evaluation days.
**This paper** tests pairwise firing rather than joint. IBIT, BITB, and FBTC (Bitcoin spot ETFs) fire regularly and are genuinely calibrated — 97%+ hit rates for gold bullish signals. The tilt scoring now reflects real signal activity across all five asset classes. Combined with per-row catalog filtering, a longer window (528 vs 250 days), and non-overlapping weekly periods that avoid horizon-overlap bias, the result is cleaner and stronger.
**Cumulative:** +30.6% versus +17.4% neutral equal-weight over 1.4 years. Outperforms 60/40, time-series momentum, and risk parity on Sharpe (1.43 vs 1.03, 0.94, 0.71, 1.19 respectively).
What I’m being honest about
The headline numbers are strong. The limitations are real.
**The most important one:** the portfolio t-statistics are computed on overlapping weekly observations. 58 data points for the 126-day horizon, but those 58 periods overlap by about 17 out of 18 weeks. The standard t-test assumes independence. Newey-West HAC-corrected standard errors with ~25 lags would be the technically correct approach and would reduce those t-statistics. I haven’t computed this yet — it’s the thing I’d want fixed before formal journal submission.
**Second:** the 100,397 calibration instances are dominated by IBIT, BITB, and FBTC — three near-identical Bitcoin ETFs. The effective independent count is much smaller.
**Third:** 1.4 years is one regime. Gold was in one of its strongest bull markets in decades. A sustained bear market, credit crisis, or deflationary environment hasn’t been tested.
I’m reporting these not as disclaimers that undo the results but because the series discipline — established in Paper 1 and maintained through Paper 4 — is to report what the data shows and what it doesn’t show, so someone building on this work knows exactly where the floor is.
What comes next
The dashboards update daily. Every prediction is logged with a timestamp. Outcomes resolve as forward horizons elapse. This isn’t a historical backtest I’m claiming will persist — it’s an ongoing, publicly verifiable record.
The most important thing that can happen next is more conditioning events: another volatility spike, another gold correction, another risk-off episode. Each one is a new test of whether the calibration holds. The Paper 3 data-sufficiency analysis was specific: the vol→equity channel needed four independent episodes before producing a selective screen. More episodes make the evidence clearer, in both directions.
Preprint: https://zenodo.org/records/20830462
Live dashboards: [quantarram.github.io](https://quantarram.github.io/quant-regime-research/notebooks/gold_dashboard.html)
https://quantarram.github.io/quant-regime-research/notebooks/portfolio_dashboard.html
Code: [github.com/quantarram/quant-regime-research](https://github.com/quantarram/quant-regime-research/tree/main/notebooks)
*This is independent quantitative research, not investment advice.*
