Skip to content

Forecasting Architecture

Model

Chronos-Bolt-mini - Amazon's time-series foundation model (smallest variant). Zero-shot forecasting with optional fine-tuning.

Pipeline

raw.sales_daily (historical orders)
        │
        ▼
forecast_all_lifecycle.py
├── Loads per-ASIN daily sales history
├── Applies lifecycle stage detection (launch, growth, mature, seasonal, decline)
├── Runs Chronos-Bolt-mini inference (365-day horizon)
├── Applies event calendar adjustments (Prime Day, holidays, etc.)
└── Writes to forecasts.predictions + forecasts.runs
        │
        ▼
compute_po_recommendations.py
├── Reads forecast + current inventory + lead times
├── Applies MOQ viability gate (won't recommend below MOQ)
└── Writes to raw.po_recommendations

Self-Healing Calibration

Runs on the 1st of each month (calibrate.py): 1. Compares last month's forecast vs actual sales 2. Computes per-ASIN error metrics 3. Adjusts seasonal multipliers if systematic bias detected 4. Tags forecast runs as month_end_baseline_for for tracking

Scoring

score_run.py runs nightly after forecast: - MAPE, WMAPE, bias direction per ASIN - Stored in forecasts.scores - Tracks accuracy trends over time

Key Config

  • Horizon: 365 days (full year forward)
  • Model size: mini (fastest, ~4GB RAM)
  • Lifecycle stages: launch (<90d sales), growth, mature, seasonal, decline
  • Quantiles: p10, p50, p90 for confidence bands (migration 0018)
  • Per-class config: migration 0019 allows different quantile settings per product class

Dashboard Validation

validate_dashboard.py runs nightly (non-blocking): - Samples 20 random ASINs - Cross-checks 7 metrics against SoStocked export + internal Postgres - Emails on failure but doesn't abort pipeline

Seasonality and Demand Tracking (2026-09-06)

The forecast chronically under-called once-a-year spikes. The reference case: Meat Claws (culinary_couture_us / B00UIR2JU4) sold 7,422 units in Dec 2025 on a ~370/mo baseline; the run forecast 1,620 for Dec 2026 and the deployed feed served a flat ~300/mo. Five things were wrong, and each has its own fix.

1. Adaptive seasonality ceiling — scripts/lib_seasonality.py

derive_seasonal_multipliers.py clipped every month of every product at 4.0. Meat Claws' stored December was exactly 4.000 — a cap, not a measurement.

A month now keeps the conservative 4.0 ceiling unless it is proven seasonal: every observed year of that month-of-year cleared 2x the ASIN's own daily rate (PROVEN_MIN_RATIO) on at least 50 units (PROVEN_MIN_MONTH_UNITS), measured only on years with ≥15 observed days of that month. A proven month gets a 30.0 runaway guard instead, and the evidence-weighted shrinkage (w = n_years / (n_years + 0.5)), not a clip, decides its value.

The evidence test is what makes this safe: one freak month cannot unlock the raised ceiling. On the live catalogue it promotes 19 ecomhd ASINs — the Halloween line, on 24-36 months of consistent Octobers — and exactly one Culinary ASIN (Meat Claws, Dec 4.00 → 5.31).

--no-adaptive-cap restores the flat behaviour; --explain <ASIN> prints the per-month raw ratio, per-year ratios, cap and proven flag.

2. Per-product event lifts — raw.product_event_lift

raw.event_calendar carries one lift per event for the whole catalogue: Christmas Peak is 1.8x for a bow tie and 1.8x for meat claws. scripts/derive_product_event_lifts.py learns each ASIN's own lift from its actuals around past occurrences and writes it per event_key (the recurring series, e.g. bfcm, not the per-year bfcm_2025).

Two invariants keep the layers from fighting:

  • Measured against the same-period baseline, not the annual average, so it does not restate what the month multiplier already says. Meat Claws' learned xmas_peak is 1.11 (its whole December is elevated, so the event window is barely above its own neighbours) while its prime_fall is 3.60 against a catalogue 2.0. That is the right decomposition.
  • Applied month-normalized and rescaled to the catalogue vector's horizon mean. A learned lift knows when inside the month volume lands, not how much annual demand the product has — level stays the month multiplier's job and the event layer's run-wide level policy (EVENT_INFLATION_CAP) is unchanged.

--no-product-events restores catalogue-only behaviour.

3. Borrowed seasonality for short-history products

Below --min-months-history (now 12, was 18) a product used to get no seasonality at all — a flat guess. Every Culinary ASIN was in that bucket. Such products now inherit the units-weighted mean shape of their peer group (category → brand → seller-wide, ≥2 peers), normalized to mean 1.0 so it redistributes volume without inventing any, blended with their own partial history at own_months / 12. Live effect: 92 ecomhd and 13 Culinary products that previously had a flat curve.

4. Year-over-year guardrail — forecasts.yoy_flags

scripts/forecast_yoy_guardrail.py compares every fully-forecast product-month against what that calendar month actually did a year earlier and writes a flag below 0.50x (under) or above 2.50x (over). It runs in the nightly before PO recommendations and every sheet push, because the point is to catch the number while it is still a forecast. It flags, it never adjusts.

Comparison hygiene: months where last year was stocked out are recorded as under_uncertain (last year is a censored floor, so the gap is real but its size is not measurable), and partially-observed or partially-forecast months are skipped entirely.

On the Meat Claws case it fires exactly as intended: B00UIR2JU4 2026-12 forecast 1,440 vs 7,422 actual LY = 0.19x under.

The guardrail is the durable answer to the residual gap — and the residual gap turned out to be worth its own investigation. See the next section.

5. Latest complete run selection — forecasts.latest_complete_run()

Consumers selected max(forecasts.runs.id). scripts/backfill_forecast_history.py writes dated historical runs (metrics.backfill = true) with higher ids and cutoffs months in the past — for culinary_couture_us, backfills 217-226 outranked the complete live runs 215/216, which is why the deployed feed served a stale flat December.

forecasts.latest_complete_run(seller, model) (migration 0074) is now the only correct answer: newest training cutoff, backfills excluded, and only runs whose predictions cover every day of their own horizon for every ASIN they claim (forecasts.run_completeness shows why any candidate was rejected). NULL is a real answer — serving nothing beats serving a stale vintage.

Every consumer was moved onto it: push_forecast_feed.py, push_claude_mrp.py, push_forecast_to_sheet.py, compute_po_recommendations.py, build_po_dashboard_xlsx.py, build_halloween_po_sheet.py, ad_driven_products.py, ad_proof.py, runway/order.py, and the three runway/queries/*.sql. The current browser surfaces select planning runs from PO recommendations rather than directly from forecasts.runs; historical backfill runs do not own PO recommendations.

Calibration no longer undoes any of it

forecasts/calibrate.py re-clipped at a flat 4.0, so a proven December derived at 5.31 was pulled straight back to 4.0 on the next monthly run. It now keeps the raised ceiling for any month already above the base one (month_ceiling), and uses a faster smoothing weight (0.60 vs 0.30) for a once-a-year month whose holdout covered the whole month and missed at the residual clip — at alpha 0.30 such a month moves 30% per year and can never catch up.

Tests: tests/test_seasonality_fix.py (22 cases, each named for the failure it locks down).

The residual gap, and why it is NOT a normalization bug (2026-09-06)

After the adaptive cap, Meat Claws December still forecast at 1,440 against 7,422. Chasing it produced the most useful finding of the day, and a null result.

It was never a December problem. The next twelve months forecast at 3,394 units against 11,929 actual in the trailing twelve — 0.28x across the whole year. That is a LEVEL error, not a seasonality error.

The hypothesis was that the multiplier layer normalizes against the wrong thing. Chronos anchors its base on recent context, so its level is "the season we are in now", while --multiplier-normalize horizon divides by the multipliers' own mean, which asserts the base is already an annual average. Two levers were built to test it, both behind flags:

  • derive_seasonal_multipliers.py --multiplier-baseline median — make 1.0 mean "a typical month" instead of "the all-days average that the spike itself inflated" (Meat Claws December: 7.4x on the mean basis, 18.2x on the median).
  • forecast_all_lifecycle.py --multiplier-normalize context — divide by the mean multiplier over the trailing window the model actually anchored on.

On Meat Claws both together looked like a fix: December 1,375 → 3,682 (0.19x → 0.50x of last year) and the year 3,394 → 6,821 (0.28x → 0.57x), with the Culinary portfolio total almost unchanged (56,589 → 57,515) — i.e. redistribution toward the seasonal products, not blanket inflation. Exactly the signature you want.

Then it was backtested, and it lost

forecasts/eval_multiplier_normalize.py holds the model constant (one Chronos-Bolt forward per cutoff) and varies only the normalization policy, with multipliers re-derived per cutoff via --as-of so no arm sees past its own cutoff. Six monthly cutoffs (2025-09-30 → 2026-02-28), six-month horizons, 250 ecomhd ASINs:

baseline / arm overall WMAPE overall portfolio bias
mean / horizon_mean (production) 67.5% −11.5%
median / horizon_mean 69.0% −9.8%
mean / raw (no multipliers) 69.1% −20.7%
mean / annual_mean 70.0% −11.8%
mean / context_90 81.9% −3.6%
mean / median 85.3% +11.6%
median / context_90 103.5% +17.5%
median / median 109.2% +36.3%

Production wins on accuracy at every horizon month. Context anchoring does fix the level bias (−11.5% → −3.6%) but costs 14 points of weighted MAPE, overshooting hard from horizon month 2 onward (+8% to +23% bias by months 4-6). The median basis is a genuine but marginal trade: −1.7 points of bias for +1.5 points of MAPE.

Why context anchoring fails is the useful part. With three years of history Chronos already encodes the season in its own base level, so re-anchoring the multipliers on top double-counts it. The hypothesis was right about the mechanism and wrong about which products it applies to: it only holds where the model cannot see the season in its own context window — a short-history tenant like Culinary, thirteen months old, whose single December is one spike in a 400-day context.

Both flags are kept, defaults unchanged (mean basis, horizon normalize), so the null result stays reproducible — the same treatment the repo already gives per-ASIN model routing.

So how does Meat Claws actually get fixed?

Three layers, in order:

  1. The system fixes above carry it as far as evidence allows: December 4.00 → 5.31 on the cap, and the run/feed selection means the served number is at least the current one.
  2. The guardrail flags what remains. B00UIR2JU4 2026-12 = 0.19x of last December is in forecasts.yoy_flags today, against the run actually being served.
  3. A human resolves it through the override path that already exists — public.forecast_override_insert(seller, asin, 'YYYY-MM', from, to, reason) from migration 0053, which requires a non-blank reason and writes an audited, reversible row in public.forecast_override_audit.

That third step is not a workaround, it is the designed answer. With one observed December there is no estimator that can know whether 7,422 repeats; the defensible data-derived number is materially lower, and the question "does this holiday spike repeat at full size?" belongs to the owner. The system's job — now done — is to make sure that question gets ASKED instead of a flat 300/mo quietly reaching a purchase order.

The durable fix for the class of problem is history: a second December moves the shrinkage weight from 0.67 to 0.80 and the multiplier from 5.31 toward the raw 7.46, with no human in the loop at all.

Pace layer: the served level blends the model with the current pace (2026-09-09)

Ace, 2026-09-08: "rank and deals change the sales pace so much." The model only sees the sales line, so a deal reads as a bump and a rank step as a trend to fade. On the week of Sep 1 to 7 2026, Tie - Black (B0C37Y1K1D) sold 809 units against a plan of 226 at p50 from the run trained through Aug 31.

What was tested (forecasts/eval_trajectory.py, compact results in forecasts/eval_trajectory_summary.json): 14 monthly cutoffs, Jun 2025 to Jul 2026, every ecomhd_us ASIN, scored on calendar-month totals 1, 2 and 3 months ahead, same weighted MAPE as the covariate work.

arm, all cutoffs +1 month +2 +3
production (bolt + multipliers + events) 49.1%, bias -20% 56.8%, -34% 62.0%, -32%
pure pace, cleaned 28-day, stored multipliers as level 54.5%, +5% 68.3%, +6% 77.2%, +7%
pure pace, own last-year shape 62.6%, +7% 66.8%, -1% 74.5%, -1%
blend 0.3 pace (own-year shape) + 0.7 model 45.7%, -12% 53.6%, -24% 59.3%, -23%
blend 0.5 48.7%, -7% 54.9%, -17% 61.0%, -17%

The 0.3 own-year blend also wins on the buy-driving classes (Class A, Relaunch, Bundle: 34.6 / 43.3 / 49.5 vs 37.1 / 45.8 / 51.4), on the recent Feb to Jul 2026 cutoffs (38.7 / 45.9 / 50.7 vs 43.2 / 48.4 / 55.1) and on Tie - Black itself. Pure pace loses at every horizon: the stored month multipliers carried as a cross-month LEVEL blow up on seasonal products (Q1 seasonal at 96%), which is exactly why production normalises them within the horizon. Rank as a reactivity switch (use a 14-day base when the head keyword moved 25%+) was a null result: on the 40 rows where it fired, 68.0% vs 56.9% for the 28-day base. It is not shipped.

What ships (forecasts/lib_pace.py, wired into forecast_all_lifecycle.py, --pace-blend 0.3 default, --no-pace-blend to ablate):

  1. clean_base: mean recorded units/day over the last 28 days before the data watermark, excluding deal days (realised price under 92% of the trailing 90-day median) and stock-out days (no sellable units and zero sales). Under 7 clean days it falls back to the plain 28-day mean.
  2. own_year_ratios: for each forecast month, last year's same-month rate over last year's base-window rate, clipped [0.5, 3.0]. A month whose last-year counterpart is not fully observed, or a product whose last-year base window sold under 20 units, falls back to the stored multiplier ratio clipped [0.6, 1.6].
  3. pace_path: base times ratio times the same event lift the model gets.
  4. blend_quantiles: p50 moves to 0.7 model + 0.3 pace; every other quantile is scaled by the same factor, so the band keeps its relative width and the class-based served quantile still means what it meant.

Run metrics record pace_blend_weight, the counts of months shaped by own-year vs multipliers, deal and stock-out days excluded, and p50_total_model vs p50_total_served. --dry-run --dry-run-asin <ASIN> prints one product's monthly model vs served vs pace.

Validated on ecomhd_us only; it runs for every seller. Culinary should be re-scored with the same script once its second year of history is in the base.