Rate Lock Index — should you lock or float your mortgage rate today?
Free, data-driven rate-lock guidance updated every weekday — what locking, floating, or a float-down actually costs for your closing date, measured on real daily lock rates, with a fully public track record. Today's call: LOCK · 30Y 7.03% today · direction models: abstaining (no demonstrated live skill, checked 2026-09-09)
Data freshness
10Y Treasury
LaggingFRED · DGS10
1d agoDaily, T+1
30Y Mortgage (daily)
LiveOptimal Blue · OBMMI
18h agoDaily, T+1
30Y Mortgage
LiveFreddie Mac · PMMS
5d agoWeekly, Thursdays · next
MOVE Index
LaggingYahoo · ^MOVE
1d agoDaily, close-of-day
Fed funds futures
LiveCME · ZQ forward strip
< 1h agoDaily, intraday
Bond Conditions Index
LiveInternal · gundlach_rli
1h agoBusiness days, ~21:00 UTC
Two things, clearly separated: the standing call (left) is a cost-of-regret default — floating into a rate spike has cost more than locking a little early — and the panel on the right shows, for your closing date and loan, what each choice actually cost across the last two years of real daily lock rates. No rate forecast is involved in either. How it works →
Today's recommendation
high action confidenceLOCK
Bond market volatility breaking out (5d vol 1.28× the 20d level). When vol regime breaks higher, MBS hedges blow out and mortgage spreads widen quickly — even small 10Y moves get amplified. Lock to bound rate risk.
What each choice costs — for your close
Measured on 504 actual 30-day lock-rate paths (2024-08-16 → 2026-08-27, ≈24 independent) · not a forecast
Over 30 days, the 30-year rate moved between −26 bps and +38 bps in 90% of windows (median −1 bps). It fell ≥25 bps 6% of the time and rose ≥25 bps 15% of the time.
| Choice | Expected | Bad case (P95) | Best case (P5) | P95 / month |
|---|---|---|---|---|
Lock nowlowest-regret lock a period that covers your close | +0 bps +$0/mo | +0 bps | +0 bps +$0/mo | +$0 +$0 lifetime |
Float to close wait; lock a short period at the close-date rate | +2 bps +$8/mo | +38 bps | −26 bps −$88/mo | +$128 +$45,948 lifetime |
Lock + float-down lock now, pay for the option to re-lock lower once | +5 bps +$17/mo | +6 bps | −8 bps −$25/mo | +$21 +$7,573 lifetime |
Float with a triggertrigger +25 bps · fired 20% wait, but lock the day rates rise past a trigger | +3 bps +$9/mo | +34 bps | −26 bps −$88/mo | +$115 +$41,407 lifetime |
How to read it: costs are versus locking today at 7.028% with no lock-period charge, in basis points of rate (positive = you pay more; negative = you saved), then converted to 30-year P+I on your loan. “Lowest-regret” = the choice with the smallest bad case among those that don't cost more on average than locking — a rule with no tuned numbers, anchored on locking.
Lender terms used: Lock ladder 15d +0 · 30d +0 · 45d +12.5bps · 60d +25bps; extension 12.5bps/15d; float-down 6.25bps up front, pays after a ≥25bps drop (re-lock 12.5bps above market); trigger-lock at +25bps (industry defaults v2026-09-default).
Source: OBMMIC30YF (Optimal Blue 30Y conforming, daily), trailing 504 sessions as of 2026-09-22; other products use their own current rate with the same distribution. The page's LOCK/FLOAT call (LOCK) is the separate cost-of-regret default, not a model forecast.
Rate-moving events in the next 45 days
5 scheduled · 5 high-impact · 0 coupon auctions- PCE Price Index · in 6 days · FRED release calendar
- GDP · in 6 days · FRED release calendar
- Non-Farm Payrolls (NFP) · in 8 days · FRED release calendar
- CPI (Inflation) · in 20 days · FRED release calendar
- FOMC Meeting · in 34 days · Federal Reserve
Why it matters for a lock: every lock window contains scheduled prints; the ones inside YOUR window are the dates a floating rate can gap. Dates are the official schedules (Fed table · FRED release calendar · TreasuryDirect), not estimates.
Buying or refinancing right now?
Today's read leans LOCK. A quick call gets you a live quote and a lock strategy sized to your actual closing date — before the window moves.
Today's Rate Brief
· published every weekday morning30-yr mortgage 7.03% (daily lock index) — LOCK: floating 30 days has cost up to 38 bps in the bad case (P95) across the last two years of daily lock rates; locking early cost less. Cost-of-regret default, not a forecast.
- ·30-year mortgage: 7.03% (≈flat today) — Optimal Blue daily lock index.
- ·Other products: 15-yr 6.54% · FHA 6.80% · Jumbo 7.03% (daily).
- ·Cost of waiting — Floating 30 days: expected +2 bps, bad case (P95) +38 bps ≈ +$128/mo on $500K; locking now: +0 bps (lock-period pricing); lock + float-down: bad case +6 bps, best case −8 bps; rates fell ≥25 bps in 6% and rose ≥25 bps in 14% of the last 504 windows; lowest-regret: lock now.
- ·45-day close — Floating 45 days: expected +2 bps, bad case (P95) +42 bps ≈ +$144/mo on $500K; locking now: +13 bps (lock-period pricing); lock + float-down: bad case +19 bps, best case −5 bps; rates fell ≥25 bps in 13% and rose ≥25 bps in 15% of the last 504 windows; lowest-regret: lock now.
- ·⚡ Fed funds futures — the market's live bet on where the Fed sets rates — are pricing a firmer path: near-term rate cuts are off the table, and a small chance of a hike is back on. It's why "just wait for rates to fall" is a riskier bet right now.
- ·⚡ The gap between 2-year and 10-year Treasury yields is what the bond market thinks about growth and the Fed. It has compressed 25bps over the past month to +25bps. In plain terms: long-term yields are easing while short-term rates hold firm — historically that points to mortgage rates grinding sideways-to-slightly-down, not a sharp drop worth waiting for.
- ·Fed funds futures imply 4.47% by Apr 2027 vs 3.88% today (+59bps) — market pricing a firmer path (hike risk back on the table).
- ·Fed rate-path odds (CME-FedWatch-style): Fed funds futures now price a ~54% chance of at least one HIKE by the Oct 27–28, 2026 meeting (most likely +0.25%, 54%).
- · Cumulative odds of ≥1 move: Oct 27–28, 2026 54% hike → Dec 8–9, 2026 90% hike → Jan 26–27, 2027 94% hike.
- ·2s10s curve: +25bps, compressed 25bps over the past month.
- ·Mortgage-to-Treasury spread: +201bps (43th pctile — normal range).
- ·Bond-market volatility (MOVE): 81.2 (stressed).
- ·Direction models: not quoted — the pre-registered 90-day check (2026-09-09) found no demonstrated live skill (21% of 30-day calls in a period rates rose 86% of the time). Frozen and published only inside the dashboard's model internals; raw values remain in the data block for transparency.
- ·Recommendation: LOCK (low confidence). No strong directional ensemble signal (6 factors evaluated, net float-score 0.32). Historically rates fell in 58% of 30-day windows (2003-2026; roughly a coin flip since 2022), but the measured cost of being wrong is asymmetric — over 2003-2026 the bad-tail cost of floating 30 days (~41-48bps at P95) exceeded the bad-tail cost of locking early (~29-34bps), and the gap widens past 100bps at 45-60 day horizons in rising-rate regimes — so lock stays the lower-regret default.
- ·Next catalysts: PCE Price Index (in 7d) · GDP (in 7d) · Non-Farm Payrolls (NFP) (in 9d).
No strong directional ensemble signal (6 factors evaluated, net float-score 0.32). Historically rates fell in 58% of 30-day windows (2003-2026; roughly a coin flip since 2022), but the measured cost of being wrong is asymmetric — over 2003-2026 the bad-tail cost of floating 30 days (~41-48bps at P95) exceeded the bad-tail cost of locking early (~29-34bps), and the gap widens past 100bps at 45-60 day horizons in rising-rate regimes — so lock stays the lower-regret default.
Mortgage Rate Context
informational · not predictive30Y Mortgage · today
7.03%
-1bps today · Optimal Blue daily
Mortgage 1yr Pctile
100th
upper half of range
4-Week Δ
+30 bps
Freddie weekly survey
10Y Treasury
4.96%
95th pctile · 1yr
30Y Conforming
7.03%
15Y Conforming
6.38%
FHA 30Y
6.83%
Jumbo 30Y
6.96%
Daily lock-rate indices (Optimal Blue · ~35% of US originations). Freddie weekly survey (6.95%) remains the model benchmark.
MBS Market & Fed Cycle
.1% MBS-trader factorsMBS Spread
201 bps
43th pctile · 1yr
Fed Cycle · past 12mo
On Hold
trailing 12mo Δ -20 bps · DFF 3.88%
Next → futures 4.52% (hawkish) · +64bps vs now
Vol Term Structure
1.28×
5d / 20d · breaking out ⚠
Credit Stress (IG OAS)
low
z=-0.48σ vs 1yr
MOVE Index
81.2
ICE BofAML · bond-market VIX· as of Sep 21lagging
MOVE 1yr Pctile
88th
range-position 43 · range 55.8 – 115.0
MOVE 20d Δ
+10.6%
vol-of-vol rising ⚠
MOVE measures Treasury options-implied vol. High MOVE = MBS hedgers demand more premium = mortgage rates lift independent of Treasury direction.
Model internals (direction models V1.2/V1.5, rule-engine factors, curve context — frozen; no demonstrated live skill)
click to expand
Model internals (direction models V1.2/V1.5, rule-engine factors, curve context — frozen; no demonstrated live skill)
click to expandThese are published for transparency, not as guidance. On 2026-09-09 the pre-registered 90-day check adjudicated the direction models at 21.4% live hit (both) against 57.7–58.0% validated baselines, in a tape that rose 85.7% of the time — a constant down-caller, not a forecaster. The regime gate below abstains while the trailing and forward Fed reads conflict. The LOCK/FLOAT call above never depended on these models.
30-day direction
medium convictionRates likely
DOWN
V1.2 Calibrated Probability
logistic regression · walk-forward OOS validated30-day probability
46.6%
rates DOWN over next 30 days
Bucket: 45-55 · low conviction
OOS calibration (2019-2026, n=381)
| Bucket | n | Predicted | Realized | Δ |
|---|---|---|---|---|
| 0-35% | 66 | 30.5% | 37.9% | +7.3pp |
| 35-45% | 120 | 40.1% | 42.5% | +2.4pp |
| 45-55% | 152 | 49.6% | 42.1% | -7.5pp |
| 55-65% | 32 | 58.8% | 50.0% | -8.8pp |
| 65-100% | 11 | 66.3% | 63.6% | -2.7pp |
V1.5 Regime-Aware Probability (with convexity)
3 phase models · MOVE + convexity-conditioned30-day probability
20.7%
rates DOWN over next 30 days
Active phase: HOLD · Fed on hold
HIGH CONVICTION · MOVE coef 0.32 · convex coef -0.42
Convexity zone: ACTIVE (mortgage > 5.5%)
Per-phase coefficients (full-data refit, 8 features)
| Feature | Cutting | Hold | Hiking |
|---|---|---|---|
| Mortgage level | +0.02 | -0.10 | +0.29 |
| MBS spread | -0.70 | -0.26 | -0.56 |
| Fed 12mo Δ | -0.15 | +0.08 | +0.60 |
| Vol term-structure | +0.06 | -0.05 | +0.02 |
| IG OAS z-score | +0.24 | -0.23 | -0.37 |
| Mortgage 4w mom | -0.31 | +0.19 | +0.08 |
| MOVE percentile | +0.64 | +0.32 | +0.02 |
| Spread × convexity-zone | +0.14 | -0.42 | -0.59 |
Two highlighted rows. MOVE: matters most in cutting cycles (panic-ahead-of-easing signal), fades in hiking. Spread×convexity-zone: when mortgage rate > 5.5% (negative-convexity regime), tight spreads predict DOWN sharpest in hiking phase (-0.59) — extension risk amplifies, spreads must mean revert.
Backfill performance (V1.5)
OOS overall: 58.0%
2021-22: 48.1% (V1.4: 45.2%, V1.2: 33%)
High-conv: 67.3% (best of any tier)
When to weight V1.5
V1.5 makes more high-conviction calls (43% of forecasts vs V1.4's 35%) and they hit BETTER (67% vs 60%). When V1.5 says high-conviction, the directional signal is the strongest of any tier we run.
When to weight V1.2
V1.2 is still the calibration champion across all buckets. V1.5 is more confident than V1.4 — when V1.5 + V1.2 disagree, V1.2's probability is the more conservative read.
Factor Contributions
net float-score -0.48 · mixed / mostly locknear 1yr high (100th pctile) — historical mean reversion potential
5d vol > 20d vol (ratio 1.28) — bond market vol breaking out, LOCK defensive
up 30 bps over 4 weeks — momentum favors LOCK (continuation)
IG OAS -0.48σ — neutral
43th pctile spread — neutral
Fed on hold (12mo change -20 bps) — neutral
What the Yield Curve Tells Us
long-horizon · not 30-day2s10s Slope
+20 bps
1-Month Δ
-30 bps
compressing ⚠
Z-score (vs 1yr)
-2.84σ
Regime
abnormally flat / inverted
Bond Market Conditions Index (legacy 7-factor, supplementary — not a prediction)
10y at 4.96% (+1.83σ vs 60d MA) → mean-revert lower, lean FLOAT
MBB realized vol 6.0% (+3.0σ) → high vol, lean LOCK
Primary-secondary 199 bps (+0.6σ vs 5mo) widening → lean FLOAT
IG OAS 77 bps (MBS proxy), 5d slope -0.6 bps/d → spread tightening, lean FLOAT
2 events in 14d (PCE 09-30, NFP 10-02) → calendar light
MBB vs IEF 20d residual 0.39% (+0.6σ) → MBB outperforming, lean LOCK
10y RSI(14) = 71 → uptrend, lean LOCK
Last 60 days
Track Record
Verdict — pre-registered 90-day check, 2026-09-09
full scoring feed →- Direction tier: no demonstrated live skill. V1.2 and V1.5 each hit 21.4% of 70 scored 30-day calls (validated baselines 57.7% / 58.0%) in a tape that rose 85.7% of the time — only 3 independent windows, so not a rigorous disconfirmation, but a constant down-caller. Models frozen; demoted to internals.
- Probability quality: V1.2 Brier 0.295 beat a constant 0.42 forecast (0.314); V1.5 (0.385) did not — its regime-conditional sharpening made it confidently wrong.
- Action tier: LOCK on 70 of 70 briefs — identical to always-lock, 0.000 bps measured value-add. It was the right call in this tape; it was not skill. Every scored recommendation is now judged on realized regret against always-lock / always-float and the new action set below.
Independent windows · 30d
3
daily OBMMI · gate at 10
Independent windows · 45d
2
36 scored · 32 pending
Regret · always-lock (30d)
0.0
bps vs oracle, independent set
Regret · always-float (30d)
14.9
float-down 6.3 · trigger 14.9
The next review is an evidence gate, not a calendar date: it opens at 10 independent 30-day windows scored on the daily lock-rate index (one independent window ≈ every 30 days). Overlapping daily scores are also published with autocorrelation-corrected errors and an effective n — 400 daily rows are never reported as 400 events.
V1.2 / V1.5 Live Track Record
model-health note →V1.2 calibrated
19%
direction hit · OOS baseline 58%
V1.5 regime-aware
19%
direction hit · OOS baseline 58%
Probability accuracy (Brier — lower is better)
Lock/float calls, scored (avg regret in bps — lower is better)
Straight talk: the direction models have leaned the same way (down) nearly every day this window, on ≈4 independent windows of evidence — far too little to claim skill or drift. That’s exactly why the recommendation now abstains when the Fed’s trailing path and the forward market disagree, and why the lock call rests on measured cost-of-regret, not a forecast. Models stay frozen through the July 26, 2026 review; the full story is in the model-health note.
Model Changelog
v1.5.3 · launched 2026-04-27The 90-day track-record check pre-registered in v1.5.2 ran on 2026-09-09 (45 days late — the calendar reminder fired 33 times and was still missed) and ADJUDICATED: the direction tier has no demonstrated live skill. On 70 scored 30-day calls (3 independent windows) V1.2 and V1.5 each hit 21.4% vs 57.7%/58.0% validated baselines in a tape that rose 85.7% of the time; V1.2's Brier (0.295) still beat a constant forecast (0.314), V1.5's (0.385) did not; the action tier emitted LOCK on 70/70 briefs — identical to always-lock, 0.000 bps measured value-add. Not a rigorous disconfirmation on 3 windows, and mechanically explained (constant down-caller; trailing-vs-forward Fed conflict), but the instrument could not mature (~1 independent window per 43 days), so 'wait for more data' was not a path. This release changes the PUBLIC SURFACE, not the models: the headline is now the MEASURED cost of each lock choice for the borrower's own closing date, loan and product; the direction models are demoted to collapsed 'model internals'; scoring moved to the daily lock-rate index with an explicit independent-window count; the review trigger became an evidence gate. Models stay frozen (Rule #10: a successor is a NEW pre-registration — the v2.1 lock-decision engine spec is being locked separately).
What we tested →
- The pre-registered 90-day check itself (memo project_rli_90day_check_2026_09_09): hit rates, Brier vs a constant caller, the independent-window count, the LOCK/FLOAT regret ledger — all on the same 70 scored rows the public accuracy feed carries.
- A re-read of the v2.0 lock-cost engine's per-window artifact (NO_SHIP, 2026-07-10): its out-of-sample intervals were ~40% too narrow in SCALE with Gaussian shape (z sd 1.39, excess kurtosis 0.15) and non-stationary by year (2022 1.89 vs 2025 0.79); its width forecast DID carry skill (rank correlation with realized |move| 0.37, t=3.1); its decision bar was vacuous because the modeled float-down was free (always-free-float-down beat the engine 3.20 vs 5.13 bps). These diagnostics shape the separate v2.1 pre-registration, not this release.
- Data availability probes (all free, all live 2026-09-09): Optimal Blue daily indices for VA, USDA and conforming by LTV/FICO bucket (FRED, since 2017); the TreasuryDirect auction schedule and results API; ZN/TN/ZB futures and MBB/VMBS daily OHLC (Yahoo, since 2016).
What we ship ↓
- THE DECISION PANEL — 'what each choice costs, for your close': closing-date, loan-amount and product inputs (conforming / FHA / VA / USDA / jumbo / 15Y / conforming by credit profile, each with its own daily Optimal Blue rate) → for lock-now, float-to-close, lock + float-down, and float-with-a-trigger: expected, bad-case (P95) and best-case (P5) cost in basis points and 30-year P+I dollars, from every daily-issued 15/30/45/60-day window of the actual OBMMIC30YF series over the trailing 504 sessions under DISCLOSED lender terms (data/rates/lender-terms.json, industry-typical defaults, overridable; every decision object and every scored row carries the terms hash). This is the v2.0 pre-registration's own unconditional benchmark published as descriptive statistics — zero fitted parameters, honest by construction, and the null a future engine must beat. A LOCK-anchored 'lowest-regret' rule with no tuned numbers highlights one row.
- DIRECTION TIER DEMOTED: V1.2/V1.5 probabilities, the abstention gate and the rule engine's factor panel moved from the hero into a collapsed 'model internals' section that states the 2026-09-09 verdict; the hero no longer implies any model earned the call. WATCH removed from all copy (it never had a producing code path). The Borrower Impact panel (σ×√21×1.5 heuristic) retired.
- REAL EVENT CALENDAR: FOMC decision days (Fed table), CPI/NFP/GDP/PCE (FRED release calendar) and 10Y/20Y/30Y Treasury coupon auctions (TreasuryDirect) over the 45-day lock window — replacing arithmetic estimates that had printed a 10Y auction on a date after it had already priced. On the page, in the brief, and as /api/rates.catalysts.
- SCORING THAT CAN MATURE (/api/rates/accuracy.v2, additive): the same brief rows scored against the DAILY Optimal Blue index at 30 and 45 days (realized on-or-before the target, never 14 days of slack), an explicit anchored non-overlapping partition for every headline statistic, Newey-West (HAC) errors with an effective n on the overlapping set, and a realized-regret ledger for the new action set against always-lock / always-float / always-float-down / always-trigger — path-scored on daily rates at the terms hash of each row.
- EVIDENCE GATE: the rates-health sentinel's calendar-date review trigger is retired; the next review opens at 10 independent scored windows per horizon (pages once), and a governor row records each run. A pre-committed auto-demotion rule exists for any future fitted engine (regret > always-lock + 0.5 bps on ≥10 independent windows with HAC z > 1 → serve the prior); the empirical prior itself is the benchmark and is never demoted.
- INTEGRITY FIXES: the on-page daily brief panel had rendered empty since launch (read the wrong nesting level); the 'Fed funds futures' freshness row is now dated by the ZQ strip's own last trade (it had been dated by the ^TNX cash index and labeled 'front month'); the legacy 'CC MBS OAS' label now says what it is (ICE BofA IG corporate OAS used as a proxy — no free current-coupon MBS OAS exists); a true rank percentile for MOVE is published alongside the min-max range position the frozen models consume; VA/USDA/profile-segment rates added.
- BRIEF / EMBED / CONTEXT: the M-F brief leads with 'Cost of waiting' (the measured 30- and 45-day numbers) and snapshots the decision object verbatim for scoring; today.json gains lock30d/lock45d/catalysts/trackRecordVerdict (additive; legacy fields untouched); the Slack persona's RATES line quotes the decision object and no longer quotes direction probabilities.
What we explicitly do NOT claim →
- That the empirical prior FORECASTS anything. It is the distribution of what happened over the trailing two years; the panel says 'measured, not a forecast' on every render. A conditional engine that beats it is the subject of a separate pre-registration with its own cooling-off, run-once and ship bars.
- That the LOCK call earned its record. It was right in this tape because rates rose; it is a cost-of-regret default and is now scored as such against every alternative policy.
- That the direction models are disproven. Three independent windows cannot reject anything; they are frozen and demoted because they showed no skill, not because skill was refuted.
- That the lender terms are anyone's rate sheet. They are disclosed industry-typical defaults until replaced; the terms hash on every number makes a change visible.
Next research threads →
- v2.1 Lock Decision Engine pre-registration (NEW construct; 7-day cooling-off; run once, attended): direct-horizon width regression on the daily index with split-conformal calibration on as-issued residuals (the scale fix the v2.0 diagnostics demand), a priced decision layer on the lender terms matrix, and bars that are actually powered at n≈90 (rank-correlation width skill, sharpness subject to calibration, non-inferiority, regret vs the best static policy across a fixed terms grid) at 30 AND 45 days jointly. SHIP_CANDIDATE never auto-ships.
- Replace the industry-default lender terms with the actual rate-sheet terms (env override) — the single largest accuracy gain available to the decision panel, and zero model risk.
- Accrue the honest sample: the evidence gate opens at 10 independent windows per horizon (~10 months at h30 from today's anchor); the overlapping HAC view matures faster and is published alongside.
The investigation entry the v1.5.1 changelog promised ('a persistent gap of 5+pp vs OOS baselines triggers an investigation entry') — the live gap reached 39pp before anything fired, so this release is both the investigation and the alerting/honesty layer that makes the next one impossible to miss. A 20-agent forensic evaluation (every code/math claim independently reproduced) found: (1) V1.2 has predicted DOWN in 47 of 47 published briefs (P(up) range 0.381-0.499) — its 19% live hit rate is exactly what a constant-DOWN caller scores against a tape that realized UP in 81% of scored windows; (2) the 21 scored predictions are overlapping daily windows that collapse to ~1 independent observation — the raw n overstated the information content; (3) the live bucket-calibration table computed expected hit rates incorrectly for every sub-50% bucket (midpoint instead of 1−midpoint), e.g. displaying −40pp where the true shortfall was −60pp; (4) the FLOAT recommendation was empirically unreachable (0 fires in a faithful 623-business-day replay, 2024-2026) and WATCH has no producing code path — the published tri-state was structurally a one-state; (5) the V1.5 phase selector uses the TRAILING 12-month Fed path, which since March 2026 conflicts with the forward futures market (cutting vs hawkish, ~109bps apart), routing every live prediction through a submodel whose training composition is 100% recession-easing episodes (base P(up)=0.333); (6) the divergence narrative attributed LOCK to named factors (vol/spread/credit) on days when none of them fired. NO MODEL CHANGES: V1.2/V1.5 coefficients and architecture remain frozen until the 2026-07-26 review. What ships is scoring integrity, display honesty, and abstention.
What we tested →
- Four pre-specified research studies (frozen methodology, run once, adversarially reviewed for look-ahead bias) feeding the 2026-07-26 keep/recalibrate/retire review: (a) spread-decomposition forecasting — the mortgage-Treasury spread IS strongly mean-reverting in levels (half-life ~5.5 weeks) but that reversion does NOT convert into 4-week forecast skill out of sample (Brier worse than the constant 0.42 prior; verdict INFO/negative); (b) distributional tail model P(|move| > 25bps in 4w) — the DOWN tail is genuinely predictable OOS (Brier beats benchmark, monotone calibration, 2.55x top-quintile lift) but the UP tail is not (anti-predictive top quintile) — exactly the tail a LOCK decision most needs is the one that does not fit (verdict INFO); (c) lock/float regret 2003-2026 on both weekly Freddie and daily Optimal Blue panels — the long-standing '25-50bps vs 100-200+bps' asymmetry copy is a 60-day, 2022-hiking-regime number, NOT a 30-day number: at 30 days the float-regret P95 is 41-48bps vs lock-regret P95 29-34bps (asymmetry ~1.4x, not ~4x), and a 25bps float-down option cuts mean lock regret ~25-30%; (d) a trailing-vs-forward Fed disagreement gate tested against history FAILED its hypothesis (during such windows rates historically fell MORE often, 36.8% up-share, and the gate never fires in the 2021-22 ZLB failure window) — so the abstention we ship is grounded in the live 2026 scored record and the mechanical classifier conflict, NOT a historical directional claim.
What we ship ↓
- Regime-transition ABSTENTION (display/narrative only): when the trailing-Fed phase selector and the forward futures lean are in hard sign conflict, the Direction zone, the daily brief, and the today.json embed stop headlining the calibrated probability and say why — in both directions (abstention is explicitly not a rates-up call). Raw probabilities remain in machine-readable fields for transparency.
- Corrected live bucket-calibration math (sub-50% buckets now expect 1−midpoint against the directional-match hit rate).
- Honest-sample statistics on /api/rates/accuracy: independent (non-overlapping) window count, distinct outcome pairs, realized up-share — published next to the raw n so overlapping daily scoring can never again read as 21 independent observations.
- Probability-quality scoring: Brier per model vs a constant-0.42-prior benchmark and constant-caller hit benchmarks on the same rows — hit-rate alone can no longer claim skill.
- The realized-regret tracker promised in the next_research_threads of v1.0, v1.1, v1.2, v1.4 AND v1.5 — now actually built: every recommendation is scored vs oracle and vs always-lock/always-float, so a constant-LOCK stream is publicly visible as exactly that.
- Conviction-honest divergence narrative: 'slightly lean' no longer describes a high-conviction model read, and LOCK is attributed to the asymmetric-cost default rather than to named factors that did not fire.
- A daily rates-health sentinel (source staleness, live-vs-baseline model gap, brief liveness, regime-transition state) that pages once on NEW findings — including an automatic trigger when the 2026-07-26 review comes due.
What we explicitly do NOT claim →
- That any model output changed — V1.2/V1.5 coefficients, features, and phase logic are untouched and frozen until the 2026-07-26 review.
- That the 19% live hit rate proves the model is 'inverted' — it is a constant down-caller mismatched to a rising tape, scored on ~1 effective independent observation; the honest statement is that the model has shown no live directional variance, not that flipping it would work.
- That abstention predicts rates will rise — our own backfill shows trailing-vs-forward disagreement windows historically resolved LOWER more often; abstention means the model's conviction is not trustworthy here, nothing more.
- That the '100-200+bps float risk' framing was correct at the 30-day horizon — our regret backfill shows it is a 60-day/hiking-regime number, and the BorrowerImpact math will be recalibrated to horizon- and regime-conditional numbers at the 2026-07-26 review.
Next research threads →
- 2026-07-26 review agenda (evidence in hand): retire V1.2 as a published probability (a constant down-caller offers nothing to recalibrate); rebuild the direction tier around the distributional DOWN-tail result + forward-market inputs rather than trailing-phase logistics; restate BorrowerImpact as horizon- and regime-conditional regret (30d ~1.4x tail asymmetry, 60d hiking-regime ~3.7x option-adjusted) with the float-down option priced in; either implement WATCH with a reachable trigger or retire the tri-state.
- UP-tail predictability is the open research problem — the tail a LOCK decision most needs (rates spiking) showed an anti-predictive top quintile in the current feature set. Candidate features for a new pre-specified study: forward-futures path slope (already computed in-house, quarantined display-only), term premium (ACM), inflation breakevens, realized-vs-implied vol gap, refi incentive vs outstanding-universe coupon (replacing the fixed 5.5% convexity threshold that has been permanently 'on' since Aug-2022).
- movePercentile252d is a min-max range position, not a rank percentile — documented for the review; changing it now would silently alter a frozen model input.
Surface-layer iteration on the live /rates dashboard. No model-architecture changes (the discipline floor stays in place — no V1.2/V1.5 retraining before the 90-day live track-record check on 2026-07-26). Five visible changes shipped: (1) per-source freshness panel with sparklines + state pills, replacing the prior single 'Updated Xd ago' line that anchored on the slowest source and inflated displayed age 24h via a date-only ISO parsing bug; (2) SEO metadata + JSON-LD Dataset schema for share unfurls + Google rich snippets; (3) live track-record tile that consumes /api/rates/accuracy and renders OOS baselines as the headline rather than the 'building' placeholder; (4) retail-friendly hero split into Direction zone (V1.2/V1.5 + directional-conviction pill) + Action zone (LOCK/FLOAT/WATCH + action-confidence pill) + Borrower Impact panel (loan-size slider + dollar-bar scenarios derived from current vol via σ₃₀ ≈ vol₂₀ × √21 with a 1.5× asymmetric multiplier in the vol-breakout regime); (5) /rates/about full rewrite aligning the supplemental docs with the new dashboard structure — three new mechanics sections (Inside the Direction zone / Action zone / Borrower Impact panel), explicit V1.2/V1.5 OOS-baseline track-record block leading the page (high-conviction subsets surfaced as the headline 67.3%/62.3%/calibration ✓), legacy 7-factor RLI track record reframed as 'Why we pivoted' under a Supplementary divider, expanded data-sources block covering all 4 sources with publish cadence + freshness panel semantics.
What we tested →
- Whether splitting directional-confidence (V1.2/V1.5 |p−0.5|>0.15) from action-confidence (asymmetric-risk override strength) on the hero materially changes how a retail borrower reads the divergent case (e.g. models lean DOWN, vol breakout fires LOCK at HIGH action confidence). The prior single 'LOCK · HIGH' badge conflated these; in retail user-walkthroughs that look like a self-contradiction. Two zones with two pills make the asymmetric-risk override case self-explanatory rather than misleading.
- Whether borrower-friendly dollar math (loan-size slider + LOCK-miss / FLOAT-rise scenarios scaled to current vol) communicates the asymmetric-risk story better than the philosophical framing. The 1.5× FLOAT-rise multiplier in the vol-breakout regime mirrors the recommendation logic's hard-lock override threshold so the bars and the recommendation move together.
- Whether reordering the about-page Track Record section (V1.2/V1.5 leading, legacy 7-factor below as 'Why we pivoted') and surfacing the disconfirmed-versions trail (V1.3/V1.6/V1.7/V1.8) prominently improves perceived credibility without manipulating numbers. The 47% legacy 14d hit rate stays publicly visible — it's the data that disconfirmed the legacy tool as a forecasting model and reframed the surface to asymmetric-risk.
What we ship ↓
- Two-zone hero (Direction left, Action right) + Borrower Impact panel with $500K default + $100K-$2M slider — replaces single LOCK/HIGH hero
- Per-source data-freshness panel (FRED daily / Freddie weekly / Yahoo MOVE / internal RLI) with state pills + 30d inline-SVG sparklines + plain-English help tooltips + 'How sourced →' link to /rates/about
- SEO metadata + JSON-LD Dataset schema (5 variableMeasured, 3 isBasedOn datasets) — first time the page is properly indexable + share-unfurl-able
- Live V1.2/V1.5 track-record tile consuming /api/rates/accuracy — leads with OOS baselines (V1.5 high-conv 67%, V1.2 high-conv 62%, calibration ✓) rather than the prior 'building' placeholder
- /rates/about supplemental rewrite (15 sections aligned with current dashboard, including 3 new mechanics sections + disconfirmed-versions credibility section + reordered Track Record)
- Templated retail-tone narratives for both zones (lib/util/rates-retail-narrative.ts, no LLM) that explicitly justify divergence when direction and action disagree
- +39 unit tests (rates-retail-narrative + rates-impact dollar math, 22 + 17), bringing repo total 43 → 82
- Single Cache-Control + Last-Modified header on /api/rates anchored on freshest data point (replaces prior duplicate headers from next.config + route handler)
- Persistent FRED stale-while-revalidate cache + activation-gate 1h cache (sleeve cycle 16.8s cold → 4.5s first-warm → 2.7s steady-warm)
What we explicitly do NOT claim →
- That the surface changes alter the underlying probability outputs in any way — they do not. computeRatesRecommendation, computeV12Probability, and computeV15Probability are unchanged. The page renames the existing reco.confidence field to 'action confidence' on display only — the field stays for back-compat with /api/rates consumers.
- That the 1.5× asymmetric multiplier on the FLOAT-rise scenario in the Borrower Impact panel is calibrated against historical MBS-spread-widening data — it's a pedagogical heuristic that mirrors the recommendation logic's vol-breakout override threshold. Honest framing: it converts the asymmetric-risk story into a number borrowers can act on; it's not a forecast.
- That moving the legacy 47% 14d hit rate below the V1.2/V1.5 OOS block constitutes hiding it. The number stays publicly visible, with explicit framing ('this is the data that disconfirmed the legacy tool as a forecasting model'). The reorder reflects the actual primacy on the dashboard, not a curation of which numbers to show.
- That this iteration changes the 90-day live track-record discipline. We still won't modify V1.2 or V1.5 architecture before 2026-07-26.
Next research threads →
- Display/UX iterations on the retail surface within the no-model-change discipline: scenario presets in the Borrower Impact panel ($250K / $500K / $750K / $1M chip-set), 'what changed since yesterday' indicator on the recommendation (visible memory builds trust through daily readers), weekly digest PDF export so loan officers can attach a dated brief to client emails, embeddable iframe widget for chadinvestorlending.com landing pages.
- Feature engineering on V1.5 inputs (NOT retraining) — the no-architecture-change rule covers model retraining; we can still evaluate whether the existing input set is correctly computed. Candidates: refining the convexity-zone threshold (currently 5.5% mortgage rate) against post-launch data, examining whether the spread-percentile lookback (252d) is appropriate vs a regime-conditional lookback, sanity-checking the volTermStructureRatio computation against 5d / 20d windowing edge cases. Any output change requires re-running the walk-forward OOS window before shipping.
- Single source of truth for the vol-breakout threshold (volTermStructureRatio > 1.2) currently hard-coded in two places: lib/util/rates-recommendation.ts (hard-lock override) and lib/util/rates-impact.ts (asymmetric multiplier on FLOAT-rise). Should reference one constant. Risk: divergent thresholds during a future tweak.
- Live-track-record-driven baseline drift monitoring — once the first V1.2/V1.5 scores land (~2026-05-28), watch for systematic deviation from the OOS baselines. A persistent gap of 5+ pp in either direction over a 30-day rolling window should trigger an investigation entry in this changelog (won't trigger a model change before 2026-07-26 — just a public 'we're watching it' note).
V1.8 tested isotonic calibration (PAV / pool-adjacent-violators) on V1.5's raw output. Hypothesis: V1.5 ranks predictions correctly (67% high-conv hit confirms it) but is mis-calibrated in the 65%+ bucket — isotonic should remap probabilities to match realized frequencies without changing rank ordering. PASSES 0/6. Isotonic compressed V1.5's upper-end predictions (raw 0.75-0.87 → calibrated 0.667), eliminating the 65%+ bucket entirely. Lower-end also compressed (0-35 bucket gap widened from V1.5's clean to 14.5pp). Walk-forward annual isotonic refit produces noisy step functions on 52-260 training points — same small-sample failure mode that killed V1.6 meta-logistic.
What we tested →
- Walk-forward isotonic calibration: at each test year T, fit PAV regression on V1.5 OOS predictions from years [2019, T-1]. Apply to year T. Test 2020-2026 (lose 2019 — no calibration training data).
- Result: OOS overall 56.5% vs V1.5 raw 58.7% (REGRESSION -2.2pp on same dates). High-conv hit 64.4% (V1.5 raw: 68.9%) — REGRESSION -4.5pp. 65%+ bucket: empty (n=0) because isotonic compressed all V1.5 high predictions to ≤ 0.667. 0-35 bucket: 14.5pp gap (V1.5 was strong here, isotonic added noise).
- Why isotonic failed: full-data isotonic table sample shows raw 0.628-0.740 → calibrated 0.655, raw 0.750-0.803 → 0.667, raw 0.817-0.874 → 0.667 — the upper end gets compressed because V1.5 over-predicts only a few times and isotonic learns those over-predictions over-aggressively. Walk-forward annual refit makes this worse: each year's isotonic table is fit on too few data points and produces noisy, jagged step functions.
What we ship ↓
- V1.8 NOT shipped. V1.5 raw remains the live regime-aware tier
- Public changelog v1.8 entry documenting the post-hoc calibration failure
- Backfill artifact: data/backfill/v18_isotonic_calibration_2026_04_27.json
What we explicitly do NOT claim →
- That isotonic calibration is universally bad — with much more training data (5+ years of live post-2026-04-27 predictions), isotonic might add value. Walk-forward annual refit is the limiting factor.
- That V1.5 doesn't have a calibration ceiling — it does (6pp at 65%+). But no post-hoc adjustment we've tested fixes it without breaking other dimensions
Next research threads →
- V1.9 (live-data dependent) — track-record-driven dynamic weighting: as live predictions accumulate from 2026-04-27, periodically re-weight V1.2 vs V1.5 in the dual-tier surface based on which tier had stronger trailing-30d performance per Fed phase. Solves the small-sample issue by accumulating data over time.
- V2.0 (next architectural step) — re-train V1.5 base model annually with live post-2026-04-27 data added, OR explore non-logistic base models (XGBoost, neural net with regularization) when sufficient training data accumulates
- Specified-pool MBS data integration — if a free public source becomes available; would directly attack the 65%+ over-confidence by adding the institutional-tier signal we currently lack
V1.7 tested rule-based regime-aware ensemble — hand-coded phase weights (cutting 0.4 / hold 0.5 / hiking 0.7 weight on V1.5) blending V1.2 + V1.5. PASSES 1/6. Strictly worse than V1.6 simple-average on every dimension (overall, 65%+ calibration, high-conv hit). The phase-weighting hypothesis informed by V1.5's per-phase convexity coefficient was wrong: V1.5's cutting model is strong via OTHER coefficients (movePctile +0.63, spreadPct -0.70), not just convexity. Combined with V1.6 disconfirmation, this confirms the architectural conclusion: ensembling V1.2 + V1.5 cannot beat V1.5's high-conv calibration without changing the base models themselves.
What we tested →
- Hand-coded phase-weighted blend on the same 381 OOS dates as V1.6. Rule: p_V17 = w_phase × p_V15 + (1−w_phase) × p_V12 where w_phase = {cutting: 0.4, hold: 0.5, hiking: 0.7}. Weights chosen a priori from V1.5's published per-phase spreadConvex coefficients (cutting +0.14 weak, hold -0.42 moderate, hiking -0.59 strongest).
- Result: OOS overall 59.3% (V1.5: 58.0%, V1.6 avg: 59.8% — V1.7 is BETWEEN them, doesn't dominate). 65%+ bucket calibration 20.1pp gap — WORST of all tested architectures. High-conv hit 65.5% (V1.5: 67.3%, V1.6 avg: 67.2% — V1.7 LOSES here too). Strictly-better-than-V1.6-avg dimensions: 0/3.
- Per-phase OOS hit (V1.7): cutting 66.0% (n=147), hold 51.4% (n=109), hiking 58.4% (n=125). Compare V1.5 standalone: cutting 70.1%, hold 51.4%, hiking 64.8%. The blend HURT cutting (-4pp) and hiking (-6pp) while leaving hold tied — the phase weights misallocated: V1.5 is good in cutting too, not just hiking.
What we ship ↓
- V1.7 NOT shipped to page surface — disconfirmed as 1/6
- Public changelog v1.7 entry documenting the rule-based architecture failure + the architectural conclusion that ensembling can't beat V1.5's high-conv ceiling
- Backfill artifact: data/backfill/v17_rule_based_ensemble_2026_04_27.json
What we explicitly do NOT claim →
- That different rule weights would work — per backfill discipline (run backfill ONCE, don't retune to fit), V1.7 is one disciplined test of the rule-based hypothesis. Not retrying with different weights.
- That ensembling V1.2 + V1.5 is uniformly bad — V1.6 simple-avg DID improve overall hit. But neither V1.6 nor V1.7 strictly improved over the dual-tier surface, and both broke 65%+ calibration
Next research threads →
- V1.8 — isotonic calibration on V1.5's raw output. Not an ensemble — a post-processing step that remaps V1.5's probabilities to match realized frequencies, specifically targeting the 6pp 65%+ over-confidence. Isotonic preserves rank ordering but improves calibration; if V1.5 is rank-correct (which the high-conv 67.3% hit rate suggests), isotonic should monotonically reduce the 65%+ gap without losing the regime-fix wins.
- V2.0 — re-train base models with 5+ years of LIVE post-2026-04-27 data added (after sufficient track record). Larger training sets may close the 65%+ gap directly without needing post-hoc calibration.
V1.6 attempted to stack V1.2 + V1.5 via meta-learning. Tested two architectures: (a) meta-logistic regression on [p_V12, p_V15, fedTrajectoryBps] with annual walk-forward refit, (b) simple unweighted average (p_V12 + p_V15) / 2. Neither strictly dominates V1.2 + V1.5 dual-tier. Meta-logistic catastrophically failed (0/6 — 65%+ bucket inverted to 71.8pp gap). Simple average beats overall hit (60.2% vs V1.5's 58.7%) but breaks 65%+ calibration (16.1pp gap on n=13 small sample). Documented as DISCONFIRMED. Page architecture unchanged: V1.2 + V1.5 + convergent/divergent indicator continues to encode ensemble information for users.
What we tested →
- Meta-logistic stacking: walk-forward at each test year, train logistic on prior years' (p_V12, p_V15, fedTrajectoryBps) → realizedUp. Result: 0/6 ship criteria. OOS overall 46.8% (V1.2: 56.5%, V1.5: 58.7%). 65%+ bucket calibration 71.8pp INVERTED (predicted 90.6%, realized 18.8% on n=48). Walk-forward meta weights flipped signs across years (w_V12=-0.54 in 2020, +0.19 in 2022). Root cause: 2 correlated probability features + small annual meta-training sets → unstable, overfit weights. Full-data refit settled at w_V12=-0.07, w_V15=+0.34, w_fedTraj=+0.28 — essentially a noisy V1.5+regime adjustment, not a true blend.
- Simple unweighted average: p_V16_avg = (p_V12 + p_V15) / 2. Result: 3/6. OOS overall 60.2% — BEATS BOTH V1.2 (56.5%) AND V1.5 (58.7%). 2021-22 hit 45.2% (matches V1.5). High-conviction hit 69.2% (n=104 vs V1.5's n=148 with 68.9%). BUT 65%+ bucket calibration 16.1pp gap on n=13 — when both base models are confidently high, the average is too. The simpler architecture beat the learned one.
- Lesson: meta-stacking on correlated probability inputs needs more training data than walk-forward annual refit provides (52-209 examples per year). Simpler ensembles (uniform average) avoid the overfitting trap but can't capture regime-conditional weight differences. Neither beats the V1.5 calibration in the high-prob bucket — that's a hard ceiling without architectural changes to the base models themselves.
What we ship ↓
- V1.6 NOT shipped to page surface. V1.2 + V1.5 dual-tier with convergent/divergent indicator continues to encode ensemble information.
- Public changelog v1.6 entry documents both architectures + the lesson
- Backfill artifact saved at data/backfill/v16_stacked_ensemble_2026_04_27.json (full meta + simple-avg head-to-head)
What we explicitly do NOT claim →
- That ensembling V1.2 + V1.5 doesn't help — the simple-average DID improve overall hit by 1.5-3.7pp. We're saying neither architecture STRICTLY dominates the existing dual-tier, so the surface stays clean
- That meta-stacking is unsalvageable — with 5+ years of live data accumulating from 2026-04-27 onward, meta training data will be richer and revisiting V1.6 in 2028 may produce a stable meta-model
Next research threads →
- V1.7 — RULE-BASED regime-aware ensemble: weight V1.5 more heavily when convexity zone active or Fed hiking; weight V1.2 more heavily in quiet regimes. Hand-coded blend weights instead of learned, removes the small-sample-instability failure mode of V1.6 meta
- V1.8 — isotonic regression calibration on top of V1.5's raw output to specifically fix the 65%+ over-confidence (V1.5: 6pp gap, V1.4: 13pp). Isotonic doesn't change rank ordering, just remaps probabilities to match realized frequencies
- Track-record-driven re-weighting: as live predictions accumulate, dynamically blend V1.2/V1.5 based on which tier has the stronger trailing-30d performance per current Fed phase
V1.5 adds a convexity-aware spread interaction (spreadConvex = MBS spread × max(0, mortgage − 5.5%)) on top of V1.4's regime-conditional architecture. PASSES 4/5 ship criteria. Replaces V1.4 in the page surface: better calibration on the 65%+ bucket (6.0pp gap vs V1.4 13.1pp), best high-conviction hit rate of any model (67.3%), 2021-22 fix maintained AND improved (48.1% vs V1.4 45.2%). The 5pp overall hit regression vs V1.4 is concentrated on low-conviction (coin-flip) calls — the actionable high-conviction signal is the strongest of any tier.
What we tested →
- Walk-forward OOS 2019-2026 on 8-feature input (V1.4 7 + spreadConvex). Per-phase logistics. spreadConvex coefficient by phase: cutting +0.14 (mild UP — refi-panic dynamics in negative-convexity zone), hold -0.42, hiking -0.59 (strong DOWN — when rates are high in hiking, wide spreads predict mean reversion sharpest because extension risk amplifies any rate move).
- OOS overall hit 58.0% (V1.4: 63.0%, V1.2: 57.7%). 65%+ bucket calibration 6.0pp (V1.4: 13.1pp — halved). 0-35 bucket 8.5pp (V1.4: 11.4pp — improved). 45-55 bucket -7.5pp (V1.4: -9.2pp — improved). 55-65 bucket -13.0pp (V1.4: -2.7pp — REGRESSION; V1.5 over-shoots in mid-range). High-conviction (|p−0.5|>0.15) hit 67.3% at n=162 (43% of forecasts) — V1.4 was 59.8% at n=132 (35%). 2021-22 hit 48.1% (V1.4: 45.2%, V1.2: 33%).
- Per-year: V1.5 wins 2022 (46% vs V1.4 38%) — biggest hiking-shock year. V1.5 ties or roughly ties most years. V1.5 loses on 2023 (58 vs 71) and 2025 (53 vs 72) — quieter years where convexity zone was inactive most of the time.
What we ship ↓
- V1.5 replaces V1.4 in the page surface as the primary regime-aware tier (V1.4 stays in changelog history + still computable via API for back-compat)
- V1.5 panel surfaces today's prediction + active phase + active phase's MOVE coefficient + active phase's spreadConvex coefficient + whether convexity zone is active (mortgage > 5.5%) + full per-phase 8-feature coefficient table with both MOVE and spreadConvex rows highlighted
- Hero subtitle now shows V1.2 + V1.5 agreement indicator (replacing V1.4)
- Backfill performance card surfaces V1.5's wins (high-conv 67%, 2021-22 48%, halved 65%+ gap) and disclosed weaknesses (overall 58% vs V1.4's 63%, 55-65 bucket worse)
What we explicitly do NOT claim →
- That V1.5 supersedes V1.2 — it doesn't on full-bucket calibration. V1.2's calibration is still better in 3 of 5 buckets
- That V1.5's overall hit rate is best — V1.4 was 5pp higher overall. V1.5 trades raw accuracy on 50/50 calls for sharper precision when the model IS confident
- That convexity threshold of 5.5% is universally correct — it's a reasonable choice for the post-2008 regime; could need recalibration if structural rate environment shifts dramatically
Next research threads →
- V1.6 — ensemble V1.2 + V1.5 via stacking: meta-learner trained on both tiers' OOS predictions. Hypothesis: V1.2's calibration + V1.5's high-conv precision can combine into a tier that strictly dominates both
- V1.7 — convexity threshold as a learned parameter rather than hard-coded 5.5%. Cross-validate threshold from 4.5% to 6.5% in 0.25% steps
- Track realized regret per tier from V1.0's 2026-04-27 launch — empirically determine which tier's signal is most useful for actual lock decisions over the live track record
Regime-conditional logistic regression: 3 separate models trained per Fed cycle phase (cutting / hold / hiking) with movePctile as a feature. At inference, route by current Fed phase. Walk-forward OOS, annual refit. PASSES 3/5 ship criteria including the 2021-22 regime fix V1.4 was built for. Ships as SECOND TIER alongside V1.2, not as a V1.2 replacement — V1.2 remains the calibration champion; V1.4 wins on hit rate.
What we tested →
- Walk-forward OOS 2019-2026 with 3 phase models per test year. Overall hit 63.0% (V1.2: 57.7%, V1.3 best: 59.6% — V1.4 is BEST overall by 5.3pp). 2021-22 hit 45.2% (V1.2: 33% — direct fix of the regime weakness V1.4 was built for, +12pp). High-conviction (|p−0.5|>0.15) hit 59.8% at n=132 (V1.2 was n=77). Per-phase OOS: cutting 70.1% / hold 51.4% / hiking 64.8% — phase routing works.
- 65%+ bucket calibration: predicted 71.2%, realized 58.1% → 13.1pp gap (V1.2: 2.7pp). All-bucket calibration: 0-35 +11.4pp, 0.65+ -13.1pp. V1.4 is more confident than V1.2 (35% high-conv vs 21%) and over-confidence-shoots the highest bucket.
- Per-phase coefficient inspection (the publishable insight): movePctile coefficient is +0.63 in cutting cycles (panic-ahead-of-easing signal — the .1% MBS-desk story), +0.29 in hold, ~0 in hiking. Single-coefficient flat models (V1.3) couldn't capture this regime dependence — that's why V1.3 broke and V1.4 was built.
What we ship ↓
- V1.4 probability live on the page in dedicated 'V1.4 Regime-Aware Probability' panel: today's probability, active phase (cutting/hold/hiking), today's MOVE coefficient, full per-phase coefficient comparison table
- V1.4 alongside V1.2 — both probability tiers visible. When V1.2 and V1.4 agree on direction, hero displays 'convergent signal'; when they disagree, 'uncertainty elevated'
- Per-phase coefficients published in full: visitors see HOW the regime conditioning works, including MOVE's by-phase weight pattern (the headline finding)
- Backfill stats panel surfaces V1.4's wins (overall 63%, 2021-22 45%) and disclosed weaknesses (65%+ calibration is wider than V1.2)
- Decision guidance: 'When to weight V1.4' (best hit rate, especially non-cutting regimes) vs 'When to weight V1.2' (tighter calibration on high-conviction calls)
What we explicitly do NOT claim →
- That V1.4 supersedes V1.2 — it doesn't on calibration. V1.2's 65%+ bucket (predicted 67%, realized 64%) is still the gold standard for high-conviction calibration
- That regime-conditional models are inherently better — V1.4 is better on this objective (hit rate + regime coverage), worse on V1.2's objective (calibration). Different tools for different decisions
- That the per-phase coefficients are stable forever — phase models retrain as new data arrives; the coefficient pattern (especially MOVE in hiking ≈ 0) could shift if the next hiking cycle has different dynamics
Next research threads →
- V1.5 — convexity-aware spread interpretation (MBS spread reads differently above/below the negative-convexity inflection ≈ 6.5% mortgage rate). Add interaction term spread × (mortLevel > 6.5).
- V1.6 — ensemble V1.2 + V1.4 via stacking: weighted combination where V1.2 gets more weight in stable regimes (its calibration property), V1.4 gets more weight in regime transitions
- Track realized regret per tier: each V1.2 + V1.4 reco vs better-in-hindsight 30 days later. Empirically determine which tier's signal is more useful for actual lock decisions
- V2 — MOVE × phase interaction terms in a single combined model rather than separate per-phase models (parametrically equivalent but lets coefficients borrow strength)
V1.3 attempted to add MOVE Index (ICE BofAML, bond-market VIX) as features to the V1.2 logistic regression. Tested both 3-feature (level + percentile + 20d momentum) and single-feature (movePctile only) configurations with annual walk-forward refit. NEITHER STRICTLY IMPROVED V1.2 — MOVE has real regime signal but adding it without smarter architecture broke the 65%+ bucket calibration that V1.2 explicitly fixed. V1.3 ships MOVE as transparency-only context on the page; V1.4 regime-conditional model is the next attack on the 2021-2022 weakness.
What we tested →
- 9-feature config (V1.2 6 + moveLevel + movePctile + moveMom20d): OOS overall hit 54.9% (V1.2 was 57.7%), 65%+ bucket calibration 8.1pp gap (V1.2 was 2.7pp), 2021-2022 hit 44.2% (V1.2 was 33% — meaningful 11pp lift). 0/5 ship criteria. Multi-collinearity in MOVE features (moveLevel coefficient went negative while movePctile went strongly positive — they were fighting each other).
- 7-feature config (V1.2 6 + movePctile only): OOS overall hit 59.6% (BEAT V1.2), high-conviction hit 62.0% (tied), 2021-2022 hit 40.4% (still below 45% threshold), but 65%+ bucket calibration BLEW UP to 17.7pp gap. 2/5 ship criteria. Trade-off pattern: improved overall + 2021-2022, BROKE the calibration property V1.2 was built to deliver.
What we ship ↓
- MOVE Index live on the page in MBS Market panel: current level, 1yr percentile, 20d momentum, with regime-color thresholds (>75th pctile orange, <25th emerald)
- Honest disclosure on the page that MOVE is informational-only context until V1.4 — visitors see the data + the limitation
- Public commitment in the model changelog to V1.4 regime-conditional architecture (separate logistic for cutting/hold/hiking Fed phases)
What we explicitly do NOT claim →
- That V1.3 improves on V1.2 — it doesn't. V1.2 remains the live calibrated-probability model.
- That MOVE has no signal — it does (the 11pp lift on 2021-2022 in the 9-feature OOS is real). The signal needs the right architecture to extract without over-confidence in the high-conviction bucket.
- That regime-conditional V1.4 is guaranteed to work — that's why it goes through the same backfill discipline before shipping
Next research threads →
- V1.4 — regime-conditional logistic regression: separate models trained on cutting / hold / hiking Fed phase subsets. Each model gets MOVE as a feature. Hypothesis: 2021-2022 hit 31-35% in V1.2 because coefficients tuned mostly on cutting-cycle data didn't generalize to hiking-shock dynamics. Phase-conditional models should handle this directly.
- V1.5 — convexity-aware spread interpretation. MBS spread reads differently above/below the negative-convexity inflection (~6.5% mortgage rate currently). Add interaction term spread × (mortLevel > 6.5).
- V2 — TBA forward curve via Bloomberg-equivalent open data source (if available); specified-pool data would be institutional-grade input.
Logistic regression with walk-forward OOS validation. Same 6-factor input as V1.1 ensemble, but coefficients learned (not hand-tuned) with annual refit. Trained on 2003-2018, tested 2019-2026. PASSES 4/4 ship criteria — ships as calibrated-probability tier alongside V1.1 asymmetric default.
What we tested →
- Walk-forward OOS: at each test year (2019-2026), train on all data prior. L2-regularized logistic regression on standardized features. 381 OOS predictions. Result: overall hit 57.7% (V1.1 was 50%), all 5 calibration buckets within 10pp of predicted, 65%+ bucket within 2.7pp (predicted 66.3%, realized 63.6% — vs V1.1's 67.4% predicted / 50.8% realized). High-conviction (|p−0.5|>0.15) hit 62.3% at n=77.
- Per-year OOS hit shows regime risk: 2019 65%, 2020 85%, 2021 31%, 2022 35%, 2023 62%, 2024 65%, 2025 60%, 2026 60%. The 2021-2022 inflation-shock-and-Fed-pivot window is when the model lags worst — we disclose this on the page so users understand calibrated DOES NOT mean infallible.
What we ship ↓
- V1.2 calibrated probability output (0-1) on the live page when feature set is complete
- OOS calibration table (predicted vs realized in each bucket) shown to users — full transparency on where the model is honest and where it's slightly off
- Per-year hit table including the 2021-2022 31-35% misses — regime-shift risk made visible
- V1.2 probability is an INPUT to the asymmetric-risk recommendation, not a replacement. When V1.2 is high-conviction (|p−0.5|>0.15) and aligned with V1.1 net float-score, confidence is 'medium'; otherwise 'low'.
- Recommendation engine logic unchanged at the surface (LOCK/FLOAT/WATCH); V1.2 changes the confidence label and the per-factor reasoning
What we explicitly do NOT claim →
- That 62.3% high-conviction hit rate generalizes to every regime — it generalized 2019-2026 OOS, but 2021-2022 hit only 31-35%
- That this is institutional MBS-desk output — we still don't have MOVE index, TBA forward curve, specified-pool data
- That increasing the 65%+ bucket coverage (currently 11 of 381 OOS predictions, 2.9%) without re-validation would preserve the 2.7pp calibration
Next research threads →
- V1.3 — add MOVE index proxy via TLT options-implied vol (Alpaca options API). MOVE is the .1% MBS desk's #1 vol signal
- V1.3 — convexity-aware spread interpretation: MBS spread reads differently when the spread is ABOVE/BELOW the rate at which negative convexity dominates
- V1.4 — regime-conditional model: separate logistic for cutting / hold / hiking Fed phases. The 2021-2022 31% miss likely reflects a coefficient set tuned for cutting cycles applied to a hiking shock
- Realized regret tracker: compare each V1.2 recommendation against the better-in-hindsight choice 30 days later
Mortgage-rate-native upgrade. Added MORTGAGE30US (Freddie Mac PMMS) as the headline, MBS spread (mortgage − 10Y) as a real MBS market signal, Fed cycle phase classifier from daily Fed funds rate, vol term structure, and IG corporate OAS as credit-stress proxy. Recommendation now uses a 6-factor ensemble.
What we tested →
- Six-factor ensemble predicting MORTGAGE30US 30-day direction over 23 years (2003-2026). Factors: mortgage level percentile, 4-week mortgage momentum, MBS spread percentile, Fed cycle phase, vol term structure ratio (5d/20d), IG corporate OAS z-score. Result: overall hit 50.0% (still coin flip — no calibrated probability ship). High-conviction subset (|prob − 0.5| > 0.15) hit 57.3% (n=75, 6% of forecasts) — small but real edge when ensemble is sure. Calibration is now MONOTONIC (V1.0 had inverted) — predictions in the 35-45% bucket realize 39.5% (off by 1.8pp, well-calibrated). Predictions in the 65%+ bucket are over-confident (predicted 67.4%, realized 50.8%).
- Critical base-rate finding: rates went DOWN in 58% of historical 30-day windows. The lock-as-default isn't a probability claim — it's pure asymmetric cost. Even a 'naive always-FLOAT' predictor would hit 58% on directional accuracy, but would still expose clients to catastrophic UP moves.
What we ship ↓
- MORTGAGE30US as the page headline (was 10Y in V1.0) — what LOs and borrowers actually care about
- Mortgage spread (MORTGAGE30US − 10Y Treasury) with 1-year percentile — captures MBS market signals 10Y alone misses
- Fed cycle phase classifier (cutting / hold / hiking) from 12-month trajectory of daily Fed funds rate
- Vol term structure (5d / 20d / 60d realized vol) with breakout warning when 5d / 20d > 1.2
- IG corporate OAS as credit-stress proxy with z-score vs trailing year
- Six-factor ensemble float-score with per-factor contributions visible — full transparency on what drives the recommendation
- Updated reasoning: explicitly cites which factors are most influential each day
What we explicitly do NOT claim →
- Calibrated 30-day probability — backfill showed ensemble hit only 50.0% overall and is over-confident on UP predictions
- That historical 58% downward base rate predicts the next 30 days for any specific borrower
- That this is an institutional-grade MBS desk product — we don't have MOVE index, TBA forward curve, or specified-pool data on the current data tier
Next research threads →
- V1.2 — logistic regression with walk-forward out-of-sample validation. Train 2003-2018, test 2019-2026. Target: per-bucket calibration within 5pp.
- V2 — MOVE index proxy via TLT options-implied vol (Alpaca options API), convexity-aware spread interpretation, LO playbook generator
- Track realized regret: compare each recommendation against the better-in-hindsight choice 30 days later
Honest baseline. Reframed from predictive RLI score to informational context + asymmetric-risk recommendation.
What we tested →
- Yield-curve slope L/S mean reversion trade (long IEF / short TLT when curve abnormally flat or steep, |z| > 1.0σ over 252d window). 23-year backfill 2003-2026 covering 2008 GFC, 2020 COVID, 2022-2024 inversion. Result: 2/4 pass criteria. Sharpe 0.33 (need >1.0), 12/20 yrs positive (need ≥70%). Max DD 8.3% (vs benchmark 39%) — drawdown-protective but alpha too thin to ship as a trade.
- 30-day rate direction prediction using slope z + breakeven z + 5d momentum ensemble. Same data window. Result: 0/4 pass criteria. Hit rate 47.6% (worse than coin flip). Calibration INVERTED: predictions of 68% probability up actually realized 43% up. Slope is a 12-18 month recession indicator, NOT a 30-day forecaster — exactly the trap the original GUNDLACH RLI fell into.
What we ship ↓
- Informational context: 10Y level, 1-year percentile, 5d/20d momentum, realized vol percentile
- Slope context with explicit disclaimer (long-horizon recession indicator, not short-term forecaster)
- Asymmetric-risk recommendation: LOCK by default in uncertainty (lower regret), FLOAT only when 10Y is near 1yr high AND momentum is negative AND vol is normal
- Track record table — predictions logged from launch forward, scored against actual 30d outcome
- This changelog (public commitment to ongoing iteration with full transparency)
What we explicitly do NOT claim →
- 30-day rate direction prediction probability
- Statistical edge over coin flip on 30-day forecasts
- Anything we cannot validate against historical data
Next research threads →
- Add MOVE index (treasury vol) into recommendation logic
- Add Fed funds futures expected path (FRED FF series)
- Test ensembling MBS-specific factors (current coupon spread, refi index)
- Track realized regret: was lock or float the better choice in hindsight?
Every parameter change, every test result, every retirement of a discredited approach gets logged here. Public commitment to ongoing iteration.
Rate Lock Questions, Answered
Should I lock my mortgage rate today?▸
Our current read is LOCK (as of 2026-09-17) — see the live recommendation at the top of this page for the full context behind it. The honest answer always depends on your closing date, loan size, and how much a rate jump would hurt you — the recommendation here weighs the asymmetric risk (the cost of floating and being wrong usually exceeds the savings of floating and being right). Talk to your loan officer before acting; this page is research, not a binding quote.
What is a mortgage rate lock?▸
A rate lock is your lender's commitment to hold a quoted interest rate for a set window (typically 30–60 days) while your loan closes. If market rates rise during that window, you keep the locked rate. If they fall, you generally keep the locked rate too — which is why the lock-or-float timing decision matters.
What does lock vs float mean?▸
Locking fixes your rate now; floating means waiting, hoping rates drop before closing. Floating is a bet: if rates fall you save, if they rise you pay more for the life of the loan. The Rate Lock Index measures that bet — for your closing date and loan size, what locking, floating, or buying a float-down actually cost across the last two years of real daily lock rates, including the bad case.
What is the Rate Lock Index?▸
A free daily research tool from Chad Investor Lending: a LOCK / FLOAT call built on the measured cost of being wrong (floating into a rate spike has cost more than locking a little early), a decision panel that prices each choice for your closing date under disclosed lender terms, and a real calendar of the Fed, inflation, jobs, and Treasury-auction dates inside your window. Every input is dated, the methodology is public, and the track record — including the model versions we tested and rejected — is published in full.
How do the close-date and loan-amount inputs work?▸
Move the closing-date slider and the panel snaps to the nearest published horizon (15, 30, 45, or 60 days); the loan amount converts basis points into monthly and lifetime dollars on a 30-year fixed; the product menu uses that product's own current daily rate (conforming, FHA, VA, USDA, jumbo, 15-year, and conforming by credit profile). The cost distribution itself comes from actual daily lock-rate paths, so it is measured, not forecast.
How accurate is the Rate Lock Index?▸
Honestly: the rate-direction models have no demonstrated live skill. In the pre-registered 90-day check (run 2026-09-09) they hit 21% of 30-day direction calls in a period when rates rose 86% of the time — on only 3 independent windows, so not a rigorous disproof, but not evidence of skill either. They are frozen and shown only inside "model internals". What we publish instead is measured: the cost distribution of each lock choice and a scored ledger of every recommendation against always-lock and always-float. The 30-year average is currently about 6.95% (Freddie Mac PMMS, dated on this page). We also publish every disconfirmed model version.
Will mortgage rates go down soon?▸
Nobody can promise that, and we no longer headline a probability: seven pre-registered studies found no forecastable 30-day direction in any construct we tested. Mortgage rates follow the 10-year Treasury yield plus an MBS spread, shaped by Fed policy, inflation data, and bond-market volatility. What this page does show is the market's own forward Fed path, the dates that can move rates inside your window, and how wide the 30-day move has actually been.
How often is this page updated?▸
Every weekday. The daily lock-rate index (Optimal Blue), Treasury yields, and bond volatility refresh each business day, the Freddie Mac survey prints weekly (Thursdays), the event calendar reads the official Fed, FRED, and TreasuryDirect schedules, and a written rate brief is published each weekday morning. Every source carries its own as-of date.
Is this financial advice?▸
No. The Rate Lock Index is free educational research from Chad Investor Lending (NMLS #2636410). It is not a loan offer, a rate quote, or personalized financial advice. For a binding quote and a lock strategy for your specific loan, speak with a licensed loan officer.
Buying or refinancing right now?
Today's read leans LOCK. A quick call gets you a live quote and a lock strategy sized to your actual closing date — before the window moves.
How to use this page: the call at top is the lower-regret default, NOT a 30-day prediction; the decision panel beside it is the measured cost of each choice for your close. The model internals (collapsed) are published for transparency only.
Research only · not investment advice. Output is advisory for mortgage rate-lock timing decisions and educational rate research. Consult your loan officer for binding rate quotes.
Data: FRED (Optimal Blue OBMMI daily lock rates incl. VA/USDA/profile segments, UST 2/10/30, MORTGAGE30US, IG corporate OAS as the MBS-spread proxy, release calendar), Yahoo (^MOVE, CME ZQ strip), TreasuryDirect (auction schedule), internal bond conditions index. Each source is dated above with its publish cadence.
Powered by Chad Investor Lending · NMLS #2636410. The Rate Lock Index is free educational research — not a loan offer, rate quote, or personalized financial advice. For a binding quote, talk to a licensed loan officer.