Hippius
maincontainerPublic

cascade7/gen-5523c549930c

sha256:563fa23e1973c18d14e265017befd1d071ad50dc117446a895b582ca2e587f70·Indexed Sep 9, 2026

Layers

20

Total size

248.0 KB

Files

20

Quantization

—

README.md

16.2 KB

weft — spectral-portfolio generator

A cascade challenger submission. Code-only, CPU-only, deterministic.

The bet

Dispersion, not mean, is the binding constraint. On the latest scored round (9018000) the king scores WORSE than the init baseline — -0.0002 / -0.0180 / -0.0058 log-gain at h64/h256/h720, winning only 46% of windows at the two deep rungs — and survives purely on the LCB gate. A challenger at +1.70% mean lost; it needed a 41% cut in cross-feed dispersion, not more mean.

So the corpus is built for COVERAGE, not for edge. A regime the corpus barely trains is a cluster the model is weak in, and on a geometric mean in log space the good feeds cannot compensate for it.

The duel is decided by a lower confidence bound, not a mean:

wins  <=>  LCB >= margin                                cascade/eval/koth.py:379
LCB    =  quantile(rel_bootstrap, alpha / k)            cascade/eval/bootstrap.py:263
rel_b  =  (king_geo_b - chal_geo_b) / king_geo_b        bootstrap.py:216-218

which is well approximated by

LCB ~= mean_rel - z * sigma_boot,    sigma_boot ~= sd_cluster / sqrt(G)

In round 9003600, uid 26 posted the best point estimate in the field (+2.19% against uid 90's +1.66%) and lost the gate outright, because its edge was twice as dispersed across feeds (sd 0.160 against 0.086). The arithmetic closes to the percentage point:

extra dispersion penalty uid 26 pays  = 1.358 pp
mean advantage uid 26 holds           = 0.530 pp
net LCB deficit                       = 0.828 pp   (observed: 0.828 pp)

The exchange rate is z/sqrt(G) ~= 0.13-0.18 — a unit of mean is worth several units of sd, but not infinitely many. So this corpus buys mean and refuses to manufacture dispersion. It does not chase variance reduction at the cost of real mean gain, which would be the opposite error.

Three levers, each measured

1. Constant-prefix conditioning — a known seam, worked harder.

Not a novel idea. The reigning king (vant-energy-lock85) already ships augment.pad_prefix = 0.05. The edge here is 12% + 2.3% paired twins against 5%, and reproducing _prep's exact pad+1 construction — not exploiting an untouched defect. An earlier draft of this file framed it as uncontested; that was wrong.

The inference wrapper left-pads every history to 4096 by repeating the first observation, and does not mask it:

pad = np.full(window_len - h.shape[0], h[0])       toto2_trainer.py:1232-1247
mask[:, ctx_p:] = 1.0                              :1259-1261  (horizon only)
causal_standardize(x)                              :1290       (no mask at all)

so the constant run enters the causal location/scale statistics that anchor un-scaling. The pool keeps only context_length + horizon = 4160 points (pool/builder.py:88) while the cutter takes ctx = min(4096, length - horizon) (windows.py:647):

rung context pad share of metric
h64 4096 0 1/3
h256 3904 192 1/3
h720 3440 656 1/3

Two of three rungs are padded — two thirds of the decision statistic — while training never pads at all (iter_training_batches buckets by patch count and stacks without padding, :284-292). A constant prefix is a distribution the trained model has never seen.

weft.modifiers.apply_prefix reproduces _prep's construction exactly (concatenate([full(k, h[0]), h]), so the identical-value run is k+1 long), weighted toward the two pads the eval actually produces. Measured coverage: 18.4% of emitted series, from slot 0, and prefix_rate is a true master switch — at 0 the mechanism is fully off, twins included.

There is no prologue. An earlier version suppressed prefixes for the first 24,576 slots; that was removed because prefixed rows are in-distribution (two of three rungs are padded), and because at 92.7M points it exceeded corpus_target_points = 67.1M, so the materialised drain — cascade verify, the dedup probe, any cache_reuse round — contained zero prefixed rows while a live stream_cpu round was 99.8% steady state. The feature silently did not exist in one of the two modes.

2. Lengths on the patch grid — a free 0.745%.

iter_training_batches keeps s[:, -w:] with w = 32 * (L // 32) (:338-343), so L mod 32 points are discarded from the head. Under the stock scaffold's uniform [64, 4096] crop that is 15.5 points/series — 0.745% of every point generated, billed against the token budget and thrown away. Every length here is a multiple of 32.

The deeper length effect is real but modest, and widely overstated. Per series, L=4096 yields ~300x the deep-23 masked-run rate of L=768 — but the live feed (corpus_mode = "stream_cpu") streams until the trainer's token budget is met, so longer series means proportionally fewer of them. Measured against the repo's own sample_cpm_masks at a fixed point budget:

policy deep-23 patches / 1e6 pts 4096 point-share
uniform[64,4096] (stock) 82.6 —
six classes incl. 512 (earlier draft) 105.5 64.7%
weft (five classes) 120.7 (1.144x) 85.9%
all-4096 125.1 (1.186x) 100%

The 512 class was dropped outright: at P=16 it cannot host a 23-patch masked suffix at all, so it was pure dead weight against the deep rungs. The remaining 3.6% up to all-4096 is not worth collapsing every optimizer step into a single (P, C) bucket and stranding a partial bucket at the wall.

Head-truncation waste is 0.000% (stock: 0.741%).

3. A domain-matched, capped portfolio — the dispersion discipline.

Twenty families, none above 0.11 of emitted points, split 58% domain-shaped / 42% mathematical. That ratio is the design's central bet.

The eval draws even by domain — DEC-CA-0032's two-tier split is armed at block 8942400 — so each real-world domain carries roughly equal metric mass no matter how many series it holds. And the feeds are specific: aeso_alberta_pool_price, ercot_supply_demand, rki_germany_hospitalizations, ukhsa_covid_cases_daily, noaa_ship_fairweather_underway_meteorology, ixp_rheintal, bitcoin_block_arrivals. A corpus of smooth Gaussian processes is not aligned with that statistic however good those processes are.

The nine domain families reproduce generating mechanisms, not curve shapes:

All seven GIFT-Eval domains now carry exactly 7.0% each. They did not before: healthcare was 26.4% of the domain mass, web_cloudops 24.5%, econ_fin 5.7%, and sales 0.0% — one of the seven had no family at all, which is a cluster the corpus never taught while the design's whole argument was even-by-domain alignment.

family w mechanism it owns
energy_load 0.10 U-shaped temperature response (heating below 15C, cooling above 22C); price variant with log-normal spikes and negative prices
met_station 0.09 diurnal + annual under 2–5 day synoptic fronts, instrument quantisation, Weibull-shaped wind
net_traffic 0.08 multiplicative daily×weekly on a growth trend, heavy-tailed bursts with decay, hard outages
epi_counts 0.07 renewal dynamics with Rt crossing 1, weekend reporting dip + Monday catch-up, integer counts
transport_flow 0.06 bimodal commute peaks; weekends are a different curve, not a scaled one
admin_report 0.06 exact zeros between publication steps, back-revisions
arrival_counts 0.05 inhomogeneous Poisson — the honest lesson that some feeds are near-unpredictable
vitals 0.04 circadian rhythm inside hard bounds that saturate rather than clip
econ_release 0.07 staircase that holds — "the current level persists" is the correct h720 forecast
retail_demand 0.07 intermittent counts, hard business-day zeros, promo lifts with pull-forward, stockout censoring (observed zero, process not)

The eleven mathematical families stay because they are not redundant: they cover the synthetic domain, they supply the long-memory structure (fgn, arfima) that makes the h720 rung learnable at all, and they give shape-robustness on any archetype the domain set failed to anticipate.

Measured under the live contract. gen_seed_mix = 3 is armed and _SeedMixStream yields round-robin one series per child per turn (trainer/stream.py:260-274); slot() is seed-independent, so all three children walk the same slot sequence and each slot is emitted three times consecutively — a 64-row batch draws from only ~21 distinct slots:

families per batch (of 20) max single-family share
gen_seed_mix=1 18–19 10.9%
gen_seed_mix=3 (live) 13–17, median 15 14.1%

An earlier version of this file claimed "every 64-row batch contains all families", measured at mix=1. That is wrong under the live contract. The property that actually carries the dispersion argument is the weaker and true one: no single prior dominates a gradient step, and each step still spans most of the portfolio. Point shares land within 0.1% of target.

Measured properties

property value
cascade verify passes (static + determinism)
corpus digest (seed 0) 773961f2d67ffa64…
throughput, 1 core, through the sandbox stream 4.18M pts/s = 1.13x ref (3.7M)
throughput, 1 core, in-process direct 5.66M pts/s (draw only: 7.28M)
deep-23 signal vs stock 1.144x
head-truncation waste 0.000% (stock: 0.741%)
constant-prefix coverage 18.4% from slot 0, tunable via prefix_rate (0 = fully off)
portfolio split 25% held_rate / 14% TSMixup / 42% domain (7 x 6.0%, even) / 19% mathematical (23 families)
twin-rate spread across families 0.01pp (was 2.42pp)
families per training batch (live, seed_mix=3) 13-17 / 20, max share 14.1%
dependencies numpy only (1 of 16 allowed)

Throughput is measured through open_round_stream with the sandbox — the path the trainer actually uses — pinned to one core with taskset. An earlier draft of this file claimed 2.26x; that number measured only the family draw, excluding the modifier stack, sanitisation, and per-series IPC framing. 1.39x on one core is the honest figure, and it is thin: spec (~30% of generation time) and fgn (~20%) are the two families worth optimising if a lane ever gives you a single contended core. numba and scipy are both on the allowlist.

Layout

generator.py          the only file the AST static guard scans
config.json           every tunable; no bulk numeric payload (DEC-CA-0028)
requirements.txt      hash-locked, allowlisted
weft/
  rng.py              blake2b sub-seed derivation (never hash(), which is salted)
  spectral.py         circulant embedding, Davies-Harte fGn, ARFIMA FIR — O(L log L)
  families.py         the 11 priors, their weights and their cluster archetypes
  schedule.py         O(1) slot -> (family, length, preset, ordinal)
  modifiers.py        prefix / level-scale / outliers / freeze / quantise
  sanitize.py         the backstop for everything validate_series rejects
tools/
  paired_lcb.py       the real verdict, via the validator's own bootstrap
  local_duel.py       one command: train both, score both, print the verdict
tests/                60 tests, ~9s

Running the local duel

cascade score reports a scalar geomean per generator. That is the wrong statistic to iterate on — uid 26 held the best geomean in the field and lost the gate. This runs the paired comparison the validator runs:

cascade fetch king --out ./king

python challenger/tools/local_duel.py \
    --king ./king --challenger ./challenger \
    --pool-dir ./heldout \
    --device cuda --k 10 --tenure 0

It cuts the three-rung ladder (not the flat single-horizon draw), trains both sides under the same contract and seed, scores them on the same window objects — the bootstrap is paired, and pairing is what makes shared window difficulty cancel exactly — then prints mean, LCB, sigma_boot, the implied cross-feed sd, and a per-rung breakdown. Exit code is 0 if the challenger clears the margin.

Set --k to the cohort you expect: the bar moves from ~1.67% at k=1 to ~2.12% at k=17, so field size is a bigger lever than the margin decay. --tenure 0 is a freshly-crowned king (1.0%); --tenure 8 is the floor (0.5%).

Verified against the live ladder — the cutter produces exactly the pads the prefix mechanism targets:

rung context pad to 4096
h64 4096 0
h256 3904 192
h720 3440 656

tests/test_weft.py::test_pad_constants_match_what_the_ladder_actually_cuts re-derives those from ladder_windows so they cannot go stale silently.

Determinism

Every series is a pure function of (seed, family, length, ordinal), and the ordinal is a pure function of the slot index. No wall-clock, no process entropy, no hash() (PYTHONHASHSEED-salted, not reproducible across processes), no global RNG, no torch.

The seeded-prefix property holds: the first N series of a longer draw are byte-identical to a draw of exactly N. This matters because the trainer stops at its token budget wherever the wall lands (DEC-CA-0001: "wall is the law"), so the corpus is always consumed as a prefix.

What has NOT been validated

This has never been trained. This box has no GPU. Everything above is a static or throughput measurement.

Specifically unverified:

  • any actual score against uid 90, or against anything
  • the mean gain, the cross-feed sd, and therefore the LCB
  • whether constant-prefix conditioning helps, hurts, or does nothing
  • whether the portfolio's dispersion discipline survives contact with the real private eval pool

A 150% margin over the king cannot be confirmed without a GPU lane and the private pool, and no amount of local measurement substitutes. Next step is tools/local_duel.py against a fetched king on real hardware.

The empirical base rate is sobering and worth stating: across 85 published scored rounds, 15 of 175 challenger entries cleared the gate (8.6%); in the ladder era (block >= 8992800) it is 2 of 52 (3.8%). One hotkey is one lifetime submission, so this is a single draw.

Remaining gaps

The domain-match gap that dominated the previous version is addressed — weft/domains.py adds nine mechanism-level families at 58% of emitted points, covering energy load and price, epidemiological counts with reporting lag, bounded vitals, network and ops counters, transport flow, administrative reporting and economic releases. Those are the domains where the published receipts show even the king losing to the init baseline.

What is still open:

  • The king is still wider. 82 families against 20, and it models calendar and holiday effects, censoring, and an 8-order-of-magnitude scale mixture in more depth than this does. Breadth is not automatically better — 82 families at even weight means none exceeds ~1.2%, and the cap discipline here is deliberate — but it has covered archetypes this has not.
  • Throughput is thinner than before the domain families: 1.20x ref on one contended core, down from 1.39x. That is the price of calendar arithmetic and it is a conscious trade. numba and scipy are on the allowlist if a lane ever proves core-starved.
  • No pipelining. The king runs prefetch_depth 3 / max_workers 4 / max_inflight_batches 6. This is a single synchronous generator, and it is what caps the TSMixup weight above. scipy and numba are both on the allowlist and unused here.
  • held_rate is under-weighted relative to the king, 0.13 against its 0.271 — the largest single weight in its 62-family config, alongside step_level_exact_frac = 0.85. Whether flat content deserves a quarter of the budget is exactly the kind of question only a scored run answers.
  • TSMixup is under-weighted, deliberately. 0.10 here against the king's 0.18. That gap is a throughput decision: every mixup row costs k full component draws (measured 1.55 ms against a 0.58 ms portfolio mean), and at 0.14 the single-core sandbox rate fell to 1.05x the reference trainer — close enough that a contended core starves the lane, which DEC-CA-0001 makes a straight token loss. The king affords 0.18 because it runs 4 workers with prefetch depth 3; this is one synchronous process. Pipelining the generator is the change that would unlock the rest.

Files

20 items
  • tests/test_weft.py

    31746442c3c0

    48.3 KB

  • weft/domains.py

    4651e58e6de3

    38.8 KB

  • weft/families.py

    5a783fa9ede3

    25.0 KB

  • weft/observation.py

    9ca60e727806

    19.6 KB

  • README.md

    b1652db2ec2a

    16.2 KB

  • weft/schedule.py

    59c0a2ebce2f

    15.1 KB

  • weft/modifiers.py

    2bc3e1f4671d

    13.6 KB

  • tools/paired_lcb.py

    90fac9214059

    10.8 KB

  • generator.py

    c8a780d99f1d

    9.9 KB

  • tools/local_duel.py

    6cdb77ecd11f

    9.8 KB

  • weft/rebase.py

    8e1bfd13f89f

    9.4 KB

  • weft/spectral.py

    d9d0705daffb

    8.6 KB

  • tools/profile_corpus.py

    a42a90df37cf

    7.2 KB

  • tools/determinism_profile.py

    68c0f681dc7e

    4.4 KB

  • weft/sanitize.py

    d63034aafe0e

    3.3 KB

  • tools/tailrun_profile.py

    104f7fd9208c

    3.2 KB

  • weft/rng.py

    76bf4f6d972d

    2.2 KB

  • config.json

    c1e5394e3fb2

    1.8 KB

  • weft/__init__.py

    b3765b25b897

    439 B

  • requirements.txt

    48ee1272b209

    93 B