BPLoL BPDRAFT ANALYTICS
← 返回首页
← 返回结果入口 / Back to results

S16 V3: corrected model, reproduction and all errors

中文 · English

Predict per-game pick rate, active ban rate, presence and conditional mean selection position for the entire Worlds event. Roles come from played lineups; every role rate uses the same game denominator. Original V2 outputs/parameters are preserved; V3 is a separate version.

- Inputs: 2556 games / 1003 series / 40 events / 173 champions × 5 roles.

- Cutoff: 2026-10-09T19:20:41Z; patch 26.20; Data Dragon 16.20.1.

- Preserve the match archive, target teams, skill-importance and role-relevance assumptions. Correct the inference structure and the Rocketbelt haste sign uniformly. No champion-specific overrides or S16-based parameter selection.

1. Historical game weighting

w(g) ∝ exp(-ln(2)·age_days/45 - 0.25·patch_gap) × importance(g) × exp(0.8·clip((Elo_before(g)-1500)/400,-1,1)) × event_games_before_cutoff^(-0.25) × first_game_factor

Older games and larger patch gaps receive smaller weights. Importance: MSI1.8, EWC1.4, First Stand1.3, qualifiers/regional finals1.45, playoffs/grand finals1.35, KeSPA/DCGI0.75, promotion0.4, others1. Strength uses the two teams’ mean pre-game Elo, updated with K=20. Initial ratings: LCK1600, LPL1580, LEC1500, LCS1450, LCP/CBLOL1430, outside challengers1350. First games receive factor1.4; known Worlds-team history has mixture0.45, moderating deep Fearless pools.

2. Historical regression layer

historical_candidate = V1_no_patch_baseline + blend × (intercept + standardized_features × coefficients)

Six frequency ridge models cover five roles and bans; five position ridge models estimate conditional mean pick rank. Coefficients are shared across champions within each output; no champion-identity features. These 17 historical features use fitted means/scales (scale floor0.025):

baseline_probability, recent14_rate, recent45_rate, season_rate, recent_minus_baseline, recent_squared, baseline_squared, qualified_rate, champion_ban_rate, champion_pick_rate, early_pick_fraction, early_ban_fraction, role_share, flex_entropy, attack_range, recent_role_partner_strength, qualified_team_coverage

recent14/recent45 add exponential half-life weighting of14/45 days to the main weights; they are not hard windows. The dashboard’s recent45 comparator is a direct count in the hard45-day window. Historical co-selection partner strength is correlational, not a causal synergy claim.

ridge=0.03; residual_blend=0.75; position_blend=0.5

3. Role evidence and structural zeros

support(c,r) = historical_pick_count_before_cutoff(c,r) > 0

States without historical role picks stay exactly zero after regression and throughout projection; epsilon and intercepts cannot revive them. All865 states remain in error denominators. There are277 supported and588 unsupported states. Rare observed roles retain small forecasts: Aurora top has62 historical picks. No hero-name or subjective manual exclusion is used.

Zero means conservative abstention, not impossibility. Champions never picked in the archive receive no invented role pick rates, so novel debuts/flex roles can be missed. Their actual test picks still count as errors. At zero pick rate, conditional positions export/display as null/blank; if a first-time champion appears in a test, position metrics use the V1 conditional prior fallback instead of dropping the champion.

4. Separately applied patch priors

delta(c,r) = clip(Σ favorable_log_change × skill_importance × role_relevance, -1.6, 1.6)

patched_pick(c,r) ∝ supported_historical_candidate(c,r) × exp(delta(c,r))

patched_ban(c) ∝ historical_ban(c) × exp(Σ supported_V1_role_share(c,r) × delta(c,r))

Preserve V1 gain1 for numeric, mechanism and system effects. Numeric effects use favorable-direction log(new/old): cooldowns, costs and damage taken invert sign; attack speed uses total multipliers, mitigation uses remaining damage. Ability importance and role relevance remain declared priors. Patch26.19 is exposure-adjusted for older historical games;26.20 applies directly. Historical regression and support projection precede the patch multiplier and final projection, avoiding additive patch uplifts in unsupported roles.

Remove V2’s numeric_patch, mechanism_patch, system_patch, positive_patch_saturation and patch_x_low_frequency learned features. S15 historical target windows lack complete patch-change history and these columns are mostly near zero: Aurora-top saturation0.962 extrapolated far beyond training mean0.000077. Patch effects remain explicit priors above; their true elasticity has not been learned/validated. Rocketbelt haste20→10 uses negative log cast-frequency denominator120→110, corrected uniformly for13 affected champions.

5. Positions and marginal constraints

Σ_c p(c,r)=2; Σ_(c,r) p(c,r)=10; Σ_c ban(c)=10; Σ_r p(c,r)+ban(c)≤1

E[BP_step] = E[pick_rank] + 6 + 4·P(pick_rank≥7)

Projection preserves support/availability zeros. Pick rank lies in1–10 and full BP step in7–20; role position residuals cap at±2 ranks, then patch priority priors apply with valid moments enforced. Overall positions are weighted by predicted role picks. Bans have no observed final role and are not assigned fictitious lanes.

6. Training and retesting

V3 uses2392 pre-Worlds S15 games, forming39 event/patch target windows with1771 target games. The first27 of34 rolling windows select ridge∈{0.03,0.1,0.3,1,3}, residual blend∈{0.25,0.5,0.75,1} and position blend∈{0,0.25,0.5,0.75,1}:100 combinations. Loss=0.45×presence RMSE+0.30×pick RMSE+0.15×role RMSE+0.10×rank MAE. Every fold fits only target windows ending before its start; final shared coefficients fit all39. S15 Worlds labels are accessed only after selection/fitting for diagnostics.

In-sample errors across all39 S15 target windows (V3 only; not generalization evidence):

MetricV3
Pick MAE / pp2.4404
Pick RMSE / pp4.0968
Ban MAE / pp3.3445
Presence RMSE / pp8.1483
Champion × role RMSE / pp1.8726
Pick-weighted rank MAE1.0736
Pick-weighted BP-step MAE1.7664

Mean errors on the first27 S15 development windows:

MetricV2V3
Pick MAE / pp2.54052.5047
Pick RMSE / pp4.23914.2391
Ban MAE / pp3.42413.4241
Presence RMSE / pp8.41408.3976
Champion × role RMSE / pp1.94161.9366
Pick-weighted rank MAE1.07331.0719
Pick-weighted BP-step MAE1.74321.7398

Last7 S15 chronological checks (previously inspected, not fresh blind tests):

MetricV2V3
Pick MAE / pp2.40942.3951
Pick RMSE / pp3.91203.9287
Ban MAE / pp3.31453.3139
Presence RMSE / pp7.86727.8514
Champion × role RMSE / pp1.78251.7872
Pick-weighted rank MAE1.09331.0919
Pick-weighted BP-step MAE1.81031.8051

S15 Worlds84-game pre-event reconstruction (retrospective diagnostic):

MetricV2V3
Pick MAE / pp2.65322.4051
Pick RMSE / pp4.23593.9278
Ban MAE / pp3.03843.0165
Presence RMSE / pp8.44308.4203
Champion × role RMSE / pp1.91621.8094
Pick-weighted rank MAE0.90620.8500
Pick-weighted BP-step MAE1.49471.4275

Frozen S16 parameter transfer:13 windows/210 games:

MetricV2V3
Pick MAE / pp3.31613.2834
Pick RMSE / pp5.38105.3696
Ban MAE / pp3.74303.7435
Presence RMSE / pp8.61468.5959
Champion × role RMSE / pp2.45992.4542
Pick-weighted rank MAE1.46221.4602
Pick-weighted BP-step MAE2.46652.4612

S16 uses identical whole-series splits and recomputes old V2 metrics to verify agreement. Event/patch groups require≥30 games, earlier65% whole series provide history, later≥10 games test; other events are cut off too, with no overlapping series. Metadata matches the historical patch; new-patch shocks are zero. These checks test within-patch continuation, not unseen26.19/26.20 shocks. Windows are equally averaged, not pooled independent samples.

Actual role picks missed by the support mask: S15 7 windows=8; S15 Worlds=0; S16=10/2100.

S16 gains are small; pick/role RMSE slightly worsen on the last7 S15 windows. Direct repair evidence is exact unsupported-role zeros and removal of extreme patch extrapolation, not universally large accuracy gains. S16 check outcomes do not tune the parameters.

7. Uncertainty, reproduction and handoff

80 bootstrap draws resample historical complete series with a fixed pre-cutoff support mask and fixed coefficients;95% intervals cover sampling variation only. Sensitivity removes all patch effects, scales numeric magnitude±25%, or removes mechanism/system priors. Four toggle combinations yield two-factor Shapley26.19/26.20 contributions: model sensitivity, not observed/causal patch effects. Novel roles, stage mix, future schedules and Fearless availability uncertainty are not fully in these intervals.

```powershell

python -X utf8 research/scripts/train_worlds_model_v3.py

python -X utf8 research/scripts/forecast_s16_worlds_v3.py

python -X utf8 research/scripts/render_s16_forecast_v3.py

python -X utf8 research/scripts/audit_s16_forecast_v3.py

node research/scripts/check_s16_dashboard.cjs s16_preparation/outputs_v3

```

Teammates need V3 source,17-feature order, coefficients/means/scales/hyperparameters, matching V1 parameters, support policy, corrected patch encoding and input hashes. Do not interchange22-feature V2 with17-feature V3 coefficients or copy S15 per-champion posteriors. The new bundle is experiments/s15_candidates/s15-v3-supported-roles; original frozen transfer_parameters remain unchanged.

- Event-specific disabled roster not obtained; all 173 champions assumed eligible.

- CBLOL-only games excluded by requested collection scope: LOS/FURIA team preferences are estimated from sparse cross-region observations.

- No completed 26.19/26.20 games in this dataset; target-patch shocks are extrapolations.

- Future series lengths, stage participation and within-series unavailable pools are not explicitly simulated; this is the same marginal whole-event model as S15. First-game weighting moderates observed Fearless depth.

- Independent game-level source checks cover only part of the archive; no claim of every draft step verified against video.

- Historical-series bootstrap measures sampling variation, not all model, patch or event-format uncertainty.

- Zero unsupported roles are conservative abstentions; novel professional flex roles and unseen champion debuts are not forecast as positive picks. All actual test roles remain in error denominators.

- The patch response is an assumed V1 multiplicative prior. S15 lacks complete historical patch shocks, so the five V2 learned patch features are removed; no claim of empirically validated 26.19/26.20 response.

- Per-window S16 errors

- S15 training and retrospective errors

- Forecast/input hashes

- Shared V3 coefficients

- Corrected patch facts

- Patch 26.19

- Patch 26.20