From 977dc237ede5c1c043bb80843aa9e0d66162bf1f Mon Sep 17 00:00:00 2001 From: devRaGonSa Date: Tue, 24 Mar 2026 08:08:52 +0100 Subject: [PATCH] Add monthly player ranking data audit --- ...K-065-monthly-player-ranking-data-audit.md | 98 +++++++ docs/monthly-player-ranking-data-audit.md | 246 ++++++++++++++++++ 2 files changed, 344 insertions(+) create mode 100644 ai/tasks/done/TASK-065-monthly-player-ranking-data-audit.md create mode 100644 docs/monthly-player-ranking-data-audit.md diff --git a/ai/tasks/done/TASK-065-monthly-player-ranking-data-audit.md b/ai/tasks/done/TASK-065-monthly-player-ranking-data-audit.md new file mode 100644 index 0000000..493c005 --- /dev/null +++ b/ai/tasks/done/TASK-065-monthly-player-ranking-data-audit.md @@ -0,0 +1,98 @@ +# TASK-065-monthly-player-ranking-data-audit + +## Goal +Auditar con precisión los datos disponibles actualmente en el proyecto y en la fuente histórica CRCON/scoreboard para determinar qué métricas pueden alimentar un futuro ranking mensual de mejores jugadores y con qué nivel de fiabilidad. + +## Context +Se quiere diseñar un futuro ranking de “Top 3 mejores jugadores del mes”, idealmente teniendo en cuenta variables como: +- kills +- KPM +- KDA +- apoyo +- garrisons/OPs quitados si estuvieran disponibles +- enfrentamientos directos si fueran extraíbles +- impacto en partida + +Pero antes de diseñar el ranking o asignar pesos, es imprescindible saber exactamente: +1. qué datos ya están persistidos en la base local del proyecto +2. qué datos existen en la fuente CRCON/scoreboard pero aún no guardamos +3. qué datos no están realmente disponibles o no son fiables +4. qué combinación de métricas sería viable para una primera versión realista del ranking + +## Steps +1. Revisar el modelo histórico actual del proyecto, incluyendo: + - tablas SQLite + - modelos históricos + - payloads + - snapshots existentes +2. Inventariar qué métricas están ya disponibles y persistidas hoy por: + - jugador + - partida + - jugador por partida +3. Confirmar con precisión qué campos existen ya en el histórico bruto y qué calidad tienen para un ranking serio. +4. Revisar la fuente histórica CRCON/scoreboard que usa el proyecto para verificar si, además de lo ya persistido, existen datos accesibles sobre: + - garrisons destruidos + - OPs destruidos + - enfrentamientos directos / duelos + - otras métricas útiles de impacto táctico +5. Documentar para cada métrica potencial: + - si ya existe en la base + - si existe en la fuente pero no se persiste aún + - si no existe realmente o no puede extraerse de forma fiable + - si sería recomendable usarla en una futura V1 del ranking +6. Dejar una matriz clara de disponibilidad y fiabilidad, por ejemplo: + - métrica + - disponible hoy + - persistida hoy + - calidad/fiabilidad + - coste de implementación adicional + - recomendable para V1 sí/no +7. Incluir una recomendación final sobre qué conjunto de métricas sería razonable para: + - una V1 del ranking mensual + - una posible V2 más ambiciosa +8. No diseñar todavía la fórmula final del ranking salvo una recomendación muy preliminar si ayuda a contextualizar. +9. No implementar todavía nuevas tablas, nuevas ingestas ni nuevas vistas de UI. +10. Al completar la implementación: + - dejar el repositorio consistente + - hacer commit + - hacer push al remoto si el entorno lo permite + +## Files to Read First +- AGENTS.md +- ai/repo-context.md +- ai/architecture-index.md +- backend/README.md +- backend/app/historical_models.py +- backend/app/historical_storage.py +- backend/app/historical_ingestion.py +- backend/app/historical_snapshots.py +- backend/app/payloads.py +- docs/historical-domain-model.md +- docs/historical-data-quality-notes.md +- docs/historical-crcon-source-discovery.md +- docs/historical-coverage-report.md +- cualquier evidencia de endpoints o JSON reales de la fuente CRCON/scoreboard que ya use el proyecto + +## Expected Files to Modify +- un documento nuevo, por ejemplo: + - docs/monthly-player-ranking-data-audit.md +- opcionalmente ai/architecture-index.md si conviene enlazar esta auditoría +- opcionalmente docs/decisions.md si surge una decisión técnica clara sobre el alcance de métricas para V1/V2 + +## Constraints +- No implementar todavía el ranking mensual. +- No modificar producto visible. +- No crear nuevas rutas, vistas ni snapshots de ranking MVP. +- No hacer cambios destructivos. +- Mantener el trabajo centrado en discovery, inventario y recomendación técnica. + +## Validation +- Existe un documento claro que explica exactamente con qué datos podemos contar hoy. +- Queda claro qué métricas ya están persistidas y cuáles no. +- Queda claro qué métricas existen en la fuente CRCON/scoreboard pero requerirían trabajo adicional. +- Existe una recomendación técnica razonable para una V1 y una V2 del ranking mensual de jugadores. +- Los cambios quedan committeados y se hace push si el entorno lo permite. + +## Change Budget +- Preferir menos de 4 archivos modificados o creados. +- Preferir menos de 220 líneas cambiadas. diff --git a/docs/monthly-player-ranking-data-audit.md b/docs/monthly-player-ranking-data-audit.md new file mode 100644 index 0000000..b39b8d6 --- /dev/null +++ b/docs/monthly-player-ranking-data-audit.md @@ -0,0 +1,246 @@ +# Monthly Player Ranking Data Audit + +## Validation Date + +- 2026-03-24 + +## Scope + +Auditoria tecnica del estado real de datos para un futuro ranking mensual de +"mejores jugadores" usando: + +- codigo y esquema historico del backend +- persistencia local en `backend/data/hll_vietnam_dev.sqlite3` +- snapshots historicos ya generados en `backend/data/snapshots/` +- discovery ya documentada de la fuente CRCON/scoreboard + +No se implementa todavia ninguna formula de ranking, tabla nueva ni cambio de +UI. + +## Evidence Reviewed + +- `backend/app/historical_models.py` +- `backend/app/historical_storage.py` +- `backend/app/historical_ingestion.py` +- `backend/app/historical_snapshots.py` +- `backend/app/historical_snapshot_storage.py` +- `backend/app/payloads.py` +- `docs/historical-domain-model.md` +- `docs/historical-data-quality-notes.md` +- `docs/historical-crcon-source-discovery.md` +- `docs/historical-coverage-report.md` + +## Current Persisted State + +Local SQLite currently contains: + +- `historical_servers`: `3` +- `historical_matches`: `9638` +- `historical_players`: `163506` +- `historical_player_match_stats`: `1062244` +- `historical_ingestion_runs`: `32` + +Coverage visible in the local database today: + +- `comunidad-hispana-01`: `8602` matches, from `2024-05-17T20:48:40Z` to `2026-03-23T16:01:20Z` +- `comunidad-hispana-02`: `753` matches, from `2025-11-04T17:10:19Z` to `2026-03-23T18:58:06Z` +- `comunidad-hispana-03`: `283` matches, from `2026-01-14T22:34:18Z` to `2026-03-08T18:11:52Z` + +Important quality notes from the local dataset: + +- all `historical_player_match_stats` rows have populated values for kills, + deaths, teamkills, time, KPM, KDA, combat, offense, defense, support, level + and team side +- `85,270 / 163,506` players have SteamID; the rest currently depend on + `crcon-player:*` identity, so identity continuity is usable but not equally + strong for every player +- all persisted matches have start/end timestamps, map and game mode +- `7,961 / 9,638` persisted matches currently have both allied/axis score + +## What Is Persisted Today + +### Match level + +Persisted per match: + +- server +- external match id +- creation/start/end timestamps +- map name, pretty name, game mode, image +- allied score +- axis score + +Not persisted at match level: + +- raw full CRCON JSON payload +- derived win/loss per player +- any tactical event ledger + +### Player identity level + +Persisted per player: + +- stable player key +- display name +- SteamID when available +- source player id +- first seen / last seen + +### Player per match level + +Persisted per player-match row: + +- level +- team side +- kills +- deaths +- teamkills +- time seconds +- kills per minute +- deaths per minute +- kill/death ratio +- combat +- offense +- defense +- support + +## What Exists In CRCON Source But Is Not Persisted + +The documented CRCON detail payload already exposes fields that the project does +not currently store: + +- `kills_by_type` +- `kills_streak` +- `longest_life_secs` +- `shortest_life_secs` +- `most_killed` +- `death_by` +- `weapons` +- `death_by_weapons` + +These fields are visible in the source discovery, but the current upsert logic +only persists the smaller normalized subset listed above. + +## What Was Not Confirmed As Available + +The current repository evidence does not confirm any stable source fields for: + +- garrisons destroyed +- outposts destroyed +- direct duel history in a structured reusable form +- tactical actions such as node building, dismantling or commander abilities + +For direct encounters, the source does expose `most_killed` and `death_by`, but +that is not the same thing as a complete duel graph and is not stored today. + +## Availability And Reliability Matrix + +| Metric / signal | Exists in source | Persisted today | Reliability for ranking | Extra work | V1? | +| --- | --- | --- | --- | --- | --- | +| Kills | Yes | Yes | High | None | Yes | +| Deaths | Yes | Yes | High | None | Yes | +| Support | Yes | Yes | High | None | Yes | +| Combat | Yes | Yes | Medium-High | Query only | Maybe | +| Offense | Yes | Yes | Medium-High | Query only | Maybe | +| Defense | Yes | Yes | Medium-High | Query only | Maybe | +| Teamkills | Yes | Yes | High as penalty signal | Query only | Maybe | +| Match count | Yes | Derivable | High | Query only | Yes | +| Time played | Yes | Yes | High | Query only | Yes | +| KPM | Yes | Yes | Medium-High if computed from totals, lower if averaging raw per-match KPM | Query only | Yes | +| KDA / KD ratio | Yes | Yes | Medium-High if computed from totals, lower if averaging raw per-match KDA | Query only | Yes | +| 100+ kill matches | Derivable | Exposed in leaderboard | Medium | None | No | +| Win/loss context | Partially | Derivable from team side + scores when scores exist | Medium | Query and validation | Maybe | +| Weapons profile | Yes | No | Medium-Low for V1 | New persistence/modeling | No | +| Kill streak / life metrics | Yes | No | Medium-Low for V1 | New persistence/modeling | No | +| Direct encounters / duels | Partial only | No | Low today | New extraction plus modeling | No | +| Garrisons destroyed | Not confirmed | No | Unknown | Source validation first | No | +| OPs destroyed | Not confirmed | No | Unknown | Source validation first | No | +| Tactical impact composite | Partial proxies only | Partial | Medium after design work | Query/design | No for strict V1 | + +## Current Product Readiness + +The backend is already able to expose monthly leaderboard snapshots, but only +for these metrics: + +- `kills` +- `deaths` +- `support` +- `matches_over_100_kills` + +This means: + +- the project already supports a monthly ranking surface operationally +- the current ranking surface is narrower than the real data persisted in SQLite +- offense, defense, combat, KPM and KDA are available in the database but not + yet wired as first-class monthly leaderboard metrics + +## Recommendation For Ranking V1 + +A realistic V1 should use only metrics already persisted with strong coverage +and low modeling risk: + +- total kills +- total support +- KPM recomputed from `SUM(kills) / SUM(time_seconds)` +- KDA recomputed from `SUM(kills) / NULLIF(SUM(deaths), 0)` +- minimum participation gate based on matches played and/or minutes played +- optional small penalty for teamkills + +Why this is the safest V1: + +- no new ingestion is required +- all needed raw fields already exist locally +- the ranking can avoid inflated outliers by requiring minimum activity +- KPM and KDA become more defensible when derived from totals, not from average + of precomputed per-match ratios + +## Recommendation For Ranking V2 + +A stronger V2 can expand the model with already persisted but not yet surfaced +signals: + +- offense +- defense +- combat +- win/loss context derived from player side and match result when scores exist + +V2 may also evaluate source-only fields if a later task decides to persist them: + +- weapons-based detail +- kill streak and life-span signals +- partial rivalry/encounter signals from `most_killed` and `death_by` + +## Metrics Not Recommended For Early Use + +Not recommended for V1 and not yet defensible for a serious monthly ranking: + +- garrisons destroyed +- OPs destroyed +- duel ranking +- generic "impact in match" as a single opaque score + +Reason: + +- either the source availability is not confirmed +- or the source exists but the project does not yet persist enough structure to + make the metric auditable and stable + +## Final Conclusion + +The repository already has enough persisted historical data for a credible +monthly Top 3 V1 without touching ingestion: + +- kills +- support +- time played +- deaths +- teamkills +- offense +- defense +- combat + +The most realistic first release is a constrained monthly ranking based on +volume plus efficiency, using only persisted fields and explicit participation +thresholds. Tactical metrics such as garrisons, OPs and real duel graphs should +stay out of scope until the source is revalidated and the missing structures are +persisted deliberately.