A sports model should be tested as a measurement system before it is treated as a forecast. I start by freezing the data cut-off, separating training and evaluation periods, and writing down which fields were available before each event.
Then I compare a simple baseline with the model, inspect calibration, and keep a log of missing values and rule changes. If the result cannot be reproduced from the same snapshot, the score is not yet meaningful. This process is about learning from historical data, not promising a winning selection.
I keep methodology notes and public-data experiments at koei-ai.
Top comments (0)