How the Data Lab Works
Sources, and what we refuse to do
Every dataset traces to a primary source: professional player settings from publicly published player configs, swept weekly, with career metadata (roster status, years active, social links) from Liquipedia’s wiki (CC BY-SA, queried through its official API with rate limits and a named user-agent); game data extracted directly from Counter-Strike 2’s own files after game updates; marketplace facts read from each operator’s own published pages, never from other comparison sites. We do not backfill history we didn’t record, publish aggregates below sample thresholds, or present community consensus as verified fact — where a number cannot be traced, the report says so instead of filling the gap. Site-wide sources, refresh cadence and derived-stat definitions live on the methodology page.
Snapshots and change detection
Time-series datasets (player settings, site reachability, the market index) are built from dated snapshots recorded by scheduled jobs. A change is only claimed when two snapshots actually differ; a field appearing or disappearing from a source is treated as source noise, not a change. Machine-detected changes are labelled provisional until reviewed. History is append-only — earlier values are never rewritten, and corrections are published as dated entries rather than silent edits.
Publication thresholds
Each report defines a minimum dataset size and history length before it goes live; until then it is listed on the Data Lab hubas “in collection” with no public page. Aggregate statistics within a report are suppressed for any group below its stated minimum sample.
Citation
Reports may be cited freely with attribution. Format: Source: CSDB.gg, [report title], accessed [date], linking the report URL. Corrections and questions: via the contact details.