# M1-final: patch-response reconstruction and evaluation

[简体中文](m1_final_model.zh_cn.md) · [English dashboard](../m1_results.en.html) · [Reproduce](../models/meta/m1_final/README.md)

The same final release is overwritten without a version bump. Direction and magnitude are closer to requested S16 expectations, but this is not demonstrated S16 accuracy. S15 Worlds was used to accept candidates and is not an independent blind test.

## Architecture

Previously, old momentum could overwhelm nerfs to high-ban champions while near-zero champions struggled to emerge. The reconstruction uses a professional context baseline plus absolute probability changes.

1. **Context:** cloglog regression uses decayed history, reliability-adjusted momentum, prior professional DPM and gold difference at 15 minutes. Patch magnitude attenuates old momentum. Picks are champion × role; bans are global.
2. **Professional response:** nonnegative shared-shrinkage Huber regression learns future probability minus baseline on completed S15 Full Fearless transitions. Absolute/relative ranked changes and semantic changes enter jointly, without automatic high-ban attenuation by `p(1-p)`.
3. **Mechanism prior:** ultimate/basic, area/single/unknown scope, effect type and phases remain distinct. Relative duration changes and nonlinear hinges have no sleep bonus. Strategic ultimate cooldowns (map impact, control, isolation) use a dedicated frequency channel. Damage, clear, resource, growth, durability and items remain active.
4. **Fusion:** response is 50% learned professional change and 50% prior. Observed ranked fusion is 35% inside the prior, not of the whole model. Cold emergence requires favorable dated mechanics, net positive response and ranked usage. Bans can use supported-role pick threat, with no champion floor or fixed ban/pick ratio.
5. **Practice and competition:** 6% role-specific practice precedes Full Fearless availability and series-length projection. Per-game totals are 10 picks, 2 per role and 10 bans; champion BP ≤100%. Released priority reallocates to competitors. Historical Ridge estimates mean pick position.

## Learned values and assumptions

Context and professional-response coefficients are learned. Strategic-ultimate importance, prior amplitude, fusion and emergence are generic assumptions calibrated within the S15 error budget; history does not establish sixfold causal ultimate importance. Equal phase weighting in the prior is also an assumption. Effect types classify events before hard control is pooled into scope/phase channels; sleep/stun receive no special multipliers.

The professional response head uses S15 Full Fearless, without pooling five years of Classic BP targets. Historical 2021–25 ranked/skill archives and audits remain, but this run did not re-estimate every prior weight from five-year ranked labels. Preparation contains 69 forecast windows and 37 transition records, not all independent.

```json
{
  "history_half_life": 14,
  "cold_scale": 2.0,
  "response_ridge": 0.03,
  "response_gain": 2.0,
  "ban_gate": 0.02,
  "strategic_importance": 6.0,
  "area_importance": 1.0,
  "prior_strength": 3.0,
  "prior_rank_fusion": 0.35,
  "prior_mix": 0.5,
  "prior_cold_scale": 2.0,
  "ban_pick_threat": 1.0
}
```

## Selection and time boundaries

Selection compares 64 professional-response, 144 prior and 3 ban-emergence candidates, with April/June/August temporal splits. Final regression uses completed records before 2025-10-14. The 84 S15 Worlds games participate in candidate acceptance, not coefficient regression, so they are no longer independent test data.

Predeclared limits: development/terminal/Worlds BP RMSE ≤1.10× previous; BP MAE ≤previous+0.5 pp; pick and ban RMSE ≤1.15× previous; terminal low-base BP RMSE ≤previous+0.5 pp. Passing candidates favor stronger response on S15 development inputs, then response error. This is a sensitivity/error tradeoff, not minimum Worlds error. S16 hero targets are absent from the loss, but S16 expectations motivated architecture and prior design.

## Errors (percentage points)

| Stage | Metric | Previous | Current | Difference |
|---|---|---:|---:|---:|
| Development | pick_mae_pp | 2.712 | 2.665 | -0.046 |
| Development | pick_rmse_pp | 4.466 | 4.420 | -0.047 |
| Development | ban_mae_pp | 3.949 | 3.787 | -0.162 |
| Development | ban_rmse_pp | 7.902 | 7.700 | -0.202 |
| Development | bp_mae_pp | 5.359 | 5.130 | -0.229 |
| Development | bp_rmse_pp | 9.256 | 8.972 | -0.284 |
| Development | role_rmse_pp | 2.040 | 2.020 | -0.020 |
| Development | low_base_bp_rmse_pp | 2.361 | 2.242 | -0.119 |
| Terminal | pick_mae_pp | 2.580 | 2.589 | +0.010 |
| Terminal | pick_rmse_pp | 4.366 | 4.412 | +0.046 |
| Terminal | ban_mae_pp | 3.441 | 3.458 | +0.017 |
| Terminal | ban_rmse_pp | 7.444 | 7.573 | +0.129 |
| Terminal | bp_mae_pp | 4.395 | 4.434 | +0.039 |
| Terminal | bp_rmse_pp | 8.071 | 8.233 | +0.162 |
| Terminal | role_rmse_pp | 2.005 | 2.023 | +0.018 |
| Terminal | low_base_bp_rmse_pp | 2.388 | 2.534 | +0.145 |
| Worlds acceptance | pick_mae_pp | 2.401 | 2.502 | +0.101 |
| Worlds acceptance | pick_rmse_pp | 3.923 | 4.141 | +0.218 |
| Worlds acceptance | ban_mae_pp | 3.254 | 3.555 | +0.301 |
| Worlds acceptance | ban_rmse_pp | 7.171 | 7.717 | +0.545 |
| Worlds acceptance | bp_mae_pp | 5.199 | 5.642 | +0.442 |
| Worlds acceptance | bp_rmse_pp | 9.273 | 10.134 | +0.862 |
| Worlds acceptance | role_rmse_pp | 1.801 | 1.879 | +0.077 |

Worlds BP RMSE relative change: +9.29%; terminal: +2.00% . All limits pass, but Worlds error did increase.

## S16 monitoring

Reference: the model deployed before this run. Differences are not isolated causal patch effects.

| Champion | Old BP | New pick | New ban | New BP | BP change |
|---|---:|---:|---:|---:|---:|
| Lillia | 3.70% | 5.27% | 6.71% | 11.97% | +8.27 pp |
| Mordekaiser | 0.70% | 2.34% | 2.35% | 4.69% | +3.99 pp |
| Aphelios | 14.98% | 22.20% | 14.79% | 36.99% | +22.01 pp |
| Aurora | 9.39% | 12.38% | 4.81% | 17.19% | +7.81 pp |
| Nocturne | 73.24% | 9.08% | 35.23% | 44.31% | -28.93 pp |
| Vi | 68.79% | 14.08% | 40.07% | 54.15% | -14.65 pp |
| Poppy | 23.27% | 5.33% | 15.77% | 21.09% | -2.17 pp |
| Ryze | 30.92% | 14.21% | 8.68% | 22.89% | -8.04 pp |
| Smolder | 1.53% | 2.14% | 2.02% | 4.16% | +2.64 pp |
| Amumu | 1.87% | 1.67% | 0.00% | 1.67% | -0.20 pp |
| Samira | 0.28% | 0.21% | 0.00% | 0.21% | -0.07 pp |
| Lucian | 43.95% | 12.83% | 13.88% | 26.72% | -17.24 pp |
| Azir | 5.33% | 2.97% | 2.67% | 5.64% | +0.32 pp |

## Verified active corrections

Component-off replay holds known patch-transition context fixed. Ability and item effects are subsets of all mechanics. Differences include allocation competition, are not additive and are not proven causal effects.

| Disabled component | Max role-pick difference | Max ban difference |
|---|---:|---:|
| no_practice | 0.942 pp | 0.341 pp |
| no_mechanistic | 14.166 pp | 15.515 pp |
| no_observed_ranked | 8.651 pp | 35.024 pp |
| no_items | 1.843 pp | 5.430 pp |
| no_targeting | 4.355 pp | 13.175 pp |

S15 has no comparable practice snapshot and its target-window item ablation is zero; S16 activation does not independently validate those effects on S15 Worlds.

## Data, viewer and reproduction

Frozen inputs reused; no fresh collection in this run. Cutoff: 2026-10-11T18:06:15.025015+00:00

Practice includes 2,763 eligible games from 60/95 roster players over 30 days: verified identity, non-China, solo/duo, actual role matching registered role. Missing accounts are not zero practice; the 6% weight lacks comparable S15 validation. The KR Master+ 26.20 snapshot is partial; per-game pick rates are not doubled. Player-champion win rate/KDA remain disabled.

The dark bilingual viewer retains portraits, role stacks, gray/red outlines and historical-day/event filters. Details show practice players, ranked changes, patches, items, ablations and pre-allocation baseline/learned/prior responses. No old-model selector is added.

Validation includes exact S15/S16 replay, role/allocation contracts, sparse-ban gating, 12 controlled monotonicity checks, time boundaries and module activation. UI tests execute actual HTML scripts in a DOM fixture, not screenshot inspection. Artifacts contain candidates, model and checks. Older control_axes_validation.json is a representation audit, not proof of new prior weights; retraining_summary.json records this revision.
