Can AI Predict the Exact Score? What France 4-6 England Taught Us About the Limits of Prediction
On Saturday night our model looked at the World Cup third-place match, weighed France against England, and locked a call: France to win, most likely 2-1, with 3-1 and 3-2 as the near neighbours. The game finished France 4-6 England. Ten goals. A scoreline our model, and every model, and every bookmaker on earth, had priced at somewhere near one chance in a thousand. We missed the winner, missed the margin, missed the entire character of the match. So let us ask the question that miss demands, plainly and technically: can an AI ever actually predict the exact score of a football match — and is there any real hope of getting it right? The honest answer is more interesting than a simple yes or no, and it starts with understanding exactly what kind of problem a scoreline is.
What we predicted, and what the universe delivered
Our locked call put France ahead at roughly a coin-flip-plus, with an expected-goals profile that produced a tight, low-scoring scoreline family — 2-1, 3-1, 3-2 — the shapes a competitive elite match usually takes. That was not a lazy guess; it was the output of a model that has graded 102 predictions in public this tournament, hit 83% of its decisive calls, and reads the base rates of knockout football correctly most of the time. And it was wrong on every axis a scoreline can be wrong. England won. The margin was two. The total was ten. If you had asked the model, before kickoff, for the probability of precisely 4-6, it would have returned a number with two zeros after the decimal point before the first significant digit. This is the uncomfortable, clarifying fact at the centre of the whole question: the exact final score of a football match is one of the genuinely hardest things in applied prediction to call, and the reasons are mathematical, not motivational. No amount of "trying harder" reaches 4-6.
The mathematics of goals: why football scores behave like a Poisson process
Start with the machinery every serious football model is built on. Goals in football are rare, roughly independent events scattered across ninety minutes, and events with that shape are described remarkably well by the Poisson distribution. If you can estimate how many goals a team is expected to score — call it lambda, its scoring rate — the Poisson distribution hands you the probability of them scoring exactly zero, one, two, three, and so on. Multiply the two teams' distributions together (with a correction we will get to) and you have a probability for every possible scoreline: 0-0, 1-0, 2-1, 4-6, all of them.
Here is the part that surprises people. Suppose you had a perfect estimate of both teams' expected goals — not a good estimate, a flawless one, handed to you by an oracle. Feed it into the Poisson machinery and ask for the single most likely exact scoreline. In a typical match between competitive sides, that most likely score — usually 1-0, 1-1, or 2-1 — carries a probability of only about 11 to 13 percent. That is the ceiling. Even with perfect knowledge of the underlying goal rates, the most probable single scoreline is right barely one time in eight. The distribution of possible scores is simply too spread out for any single cell to dominate. This is not a flaw in Poisson; it is the truth Poisson is reporting. Football scorelines are inherently, irreducibly high-variance.
Why the most likely score is still usually wrong
Sit with the consequence for a moment, because it reframes everything. A model that is perfectly calibrated — one whose probabilities exactly match reality — will still, if you force it to name a single exact scoreline, be wrong roughly 85 to 88 percent of the time. Not because it is a bad model. Because the world it is predicting is one where 1-0 and 2-1 and 1-1 and 0-0 and 3-2 all happen with meaningful, overlapping frequency, and no single one of them is the "right answer" in any stable sense. The exact score is a draw from a wide distribution, and drawing the modal value is, by construction, a minority event.
Our own public ledger says exactly this. Across the tournament our top single scoreline pick landed about 15 percent of the time — a hair above the theoretical modal ceiling, which is roughly where an honest, well-tuned model should sit. When people ask us to "predict the score," and imagine a future in which AI nails 4-6 before kickoff, they are imagining a model that beats a limit no model can beat, because the limit is not in the model. It is in the game.
Two kinds of uncertainty — and why football has too much of the wrong kind
To see where hope lives and where it doesn't, you have to split uncertainty into its two species, because they behave completely differently.
- Epistemic uncertainty is the uncertainty of ignorance — things you don't know but could. Who is actually in the starting eleven. Whether a key defender is carrying a knock. The tactical plan. Recent form, travel, rest, motivation. This uncertainty is reducible: better data and better models genuinely shrink it.
- Aleatoric uncertainty is the uncertainty of chance — the genuine randomness baked into the event itself. Whether a well-struck shot beats the keeper by an inch or hits the post. Whether a deflection loops in or loops out. Whether the referee sees the handball. This uncertainty is irreducible: no amount of study removes it, because it is not ignorance, it is noise.
Every prediction problem is a mix of the two, and the mix determines what is possible. Predicting tomorrow's sunrise is almost pure epistemic — near-zero noise, so near-perfect accuracy. Predicting a single coin flip is almost pure aleatoric — no study helps, 50% is the ceiling forever. Football scorelines sit painfully far toward the coin-flip end. Goals are scarce, so each one is enormously consequential; a single deflected shot in the 92nd minute — the kind that decided one of this tournament's semi-finals — flips a result that ninety minutes of skill had left level. When the decisive events are few and each is soaked in chance, the aleatoric floor is high, and the exact scoreline is the single most aleatoric-dominated quantity in the entire sport. You cannot model your way past a coin's randomness by studying the coin harder, and a scoreline is closer to a fistful of coins than most fans want to believe.
Why 4-6 specifically was almost unpredictable
Now the France-England game itself. A ten-goal match is not merely a low-probability scoreline; it is a low-probability regime. It requires several rare things to co-occur: both teams' finishing running hot at once, both defensive structures failing repeatedly, and a game state that keeps opening up rather than closing down. Under any reasonable pre-match model, the probability that a match produces ten or more total goals is a fraction of one percent. The probability of the specific cell 4-6 is smaller still — on the order of one in a thousand or rarer. A responsible model does not, and should not, place meaningful probability there. If it did — if some model had "called" 4-6 — that would not be intelligence; it would be a broken model spraying probability into the tails, and it would look ridiculous across the other ninety-nine games where sanity prevailed.
There is one honest caveat, and it belongs to the epistemic column: this was a third-place play-off, a fixture with a documented tendency to sprawl open. Bronze-medal games average nearly four goals and have not produced a goalless draw in roughly ninety years — nobody defends for their life in a consolation match, both managers rotate, and the whole thing tilts toward chaos. A model that carried a strong "dead-rubber openness" feature should have widened its scoreline distribution and shaded its totals upward. Ours did widen the family somewhat, but nowhere near ten goals — because nowhere near ten goals is the correct pre-match position. Even a model that perfectly understood the dead-rubber effect would have predicted, say, a 3-2 shape with fatter tails, not 4-6. The dead-rubber feature is a real, reducible edge we can sharpen. It moves the distribution. It does not, and cannot, let you call the specific avalanche.
The information ceiling: why ~15% is a law, not a target
There is a deeper way to see the ceiling, borrowed from information theory. The distribution of plausible scorelines in a football match — everything from 0-0 up through the realistic 4-3s and 5-2s — carries an entropy of roughly three-and-a-half to four bits. Entropy is a measure of genuine unpredictability, and that figure says, formally, that the outcome space is wide and flat: many scorelines share the probability mass, none towers over the rest. When you sample a single draw from a distribution that flat, the best any predictor can do at naming the exact value — even a predictor that knows the true distribution perfectly — is to name its mode, and the mode wins only at the modal frequency. For football scorelines, that frequency is the 12-to-17-percent band that both the theory and our live results keep landing on. Academic studies of football forecasting, across many leagues and methods, converge on the same ceiling: the exact-score top-1 hit rate for state-of-the-art models tops out around 15 to 18 percent. It has not moved much in decades, not because the models haven't improved — they have, dramatically — but because the ceiling was never about the models.
So is there hope? Yes — but you have to redefine "accurate"
Here is the constructive turn, and it is the reason this platform exists. The question "can AI predict the exact score" smuggles in a definition of accuracy — name the single correct cell — that is close to a parlor trick and, as we have seen, physically capped near 15%. Change the definition to something both honest and useful, and the hope becomes real and large.
- Predict the distribution, not the point. The genuinely accurate output is not "2-1" but the whole probability surface — 1-0 at 13%, 1-1 at 11%, 2-1 at 9%, and so on. Judge that by calibration: when the model says 30%, does it happen 30% of the time? A model can be beautifully accurate in this sense while being "wrong" on the single score in four games out of five, because it was never claiming to know the single score — it was claiming to know the odds, and it did. Our published Brier score, a proper measure of probabilistic accuracy, is the number we actually stake our name on, not the exact-score hit rate.
- Predict a set, not a point. If the top single score hits ~15%, the top six most likely scores together capture around 50%. That is a genuinely useful object: a hedge that reflects the real uncertainty instead of hiding it. Our tiered product does exactly this — an outcome call, a top-3 set, and a top-6 set — precisely because a single-score claim would be dishonest about what is knowable.
- Climb to the layers that are actually predictable. Scoreline is the hardest layer. The layers above it are far more tractable: the match outcome (win/draw/loss) is callable at 55-65% by a strong model; goal totals (over/under a line) are callable well above chance; who advances in a knockout is more predictable still. Accuracy is not one number — it is a ladder, and the higher rungs bear real weight even when the exact-score rung stays slippery.
- Wait for the game to speak. The single biggest source of scoreline accuracy is time itself. Every minute of play converts aleatoric uncertainty into observed fact. A pre-match model staring at France-England could never call 4-6 — but a live model at the 70th minute, having watched five goals already fly in and the defences visibly disintegrate, can shift its distribution violently toward a high-scoring finish. In-play prediction is a fundamentally easier problem than pre-match prediction, because the noise is resolving in front of you. The hope for "accurate scoreline prediction" is overwhelmingly a live-prediction hope.
What actually moves the needle (and how far)
None of this means pre-match models are stuck. The epistemic frontier is real, and better modelling genuinely pushes the ceiling — just not to the moon. The improvements that matter: richer expected-goals inputs built from shot quality rather than raw historical goals; lineup-aware modelling that reacts to who is actually playing; correlated scoring models — the Dixon-Coles correction and its descendants — that fix Poisson's known underestimation of low scores and draws by acknowledging the two teams' goals are not independent; context features like the dead-rubber effect, weather, congestion, and rest; and blending in market odds, which pack the crowd's information into a single sharp signal. Each of these is worth real accuracy. Stacked together, across a season, they might lift exact-score top-1 from the low teens toward the high teens, and — more valuably — sharpen the whole distribution so the top-6 set and the outcome calls get measurably better. That is the honest size of the prize: not a leap from 15% to 50% on the exact score, which is impossible, but a steady tightening of a distribution that will always, correctly, refuse to promise you the single cell.
The verdict — and why we published our 4-6 miss instead of burying it
So, can AI predict the exact score? For a single match, before kickoff, as a confident point call: no — and any product that claims otherwise is either lying or misunderstanding its own model. The exact scoreline is dominated by irreducible chance; the modal cell caps near 15%; 4-6 lives in a tail no responsible model can inflate. That ceiling is a law of the game, not a temporary limitation of the technology, and no future model, however large, repeals it.
But is there hope for accurate AI football prediction? Emphatically yes — once you let "accurate" mean the true and useful thing rather than the parlor trick. There is deep, compounding hope in well-calibrated probability distributions; in honest hedged sets that capture half the outcomes in six scorelines; in the eminently callable layers of outcome and total and advancement; and in live models that grow sharp as the match resolves the randomness in real time. The frontier of football prediction is not a race to name 4-6 before kickoff. It is a race to quantify uncertainty more honestly than anyone else — and that race is wide open, and worth running.
Which is exactly why, when our France call came back 4-6, we did not quietly delete it. We left it on the card, crossed out, graded a miss in public — right beside the calibration numbers and the top-6 sets that tell the real story of what this model knows and what no model can. A prediction engine that pretended it could have called that game would be a worse product, not a better one. The honesty is not damage control. The honesty is the accuracy — the accuracy about our own limits — and in a field this dominated by chance, that may be the only kind worth trusting. AI predictions are for entertainment; the honesty is the product.
AI 能预测准确比分吗?法国 4-6 英格兰给我们上的一课
周六晚上,我们的模型审视了世界杯季军战,权衡法国与英格兰,锁定了一个判断:法国获胜,最可能 2-1,邻近的 3-1、3-2。这场比赛最终 法国 4-6 英格兰。十个进球。这是一个我们的模型、所有模型、地球上每一家博彩公司都定价在约千分之一的比分。我们判错了胜者、判错了净胜球、判错了整场比赛的性格。所以让我们直接、技术性地问这个失误所要求的问题:AI 究竟能不能预测一场足球赛的准确比分——真的有希望预测对吗?诚实的答案比简单的是或否更有意思,而它始于理解"比分"到底是一个什么样的问题。
我们预测了什么,宇宙又交付了什么
我们锁定的判断把法国放在略高于五五开的位置,其预期进球画像产出一组紧凑、低比分的比分家族——2-1、3-1、3-2——这是一场势均力敌的顶级比赛通常呈现的形状。这不是懒惰的猜测;它是一个本届公开评分 102 次、决胜判断命中 83%、大多数时候正确读出淘汰赛基础概率的模型的输出。而它在比分能错的每一个维度上都错了。英格兰赢了。净胜球是二。总进球是十。如果开球前你问模型精确 4-6 的概率,它会返回一个小数点后先有两个零、才出现第一个有效数字的数。这就是整个问题核心那个令人不适却澄清一切的事实:一场足球赛的准确终场比分是应用预测中真正最难判断的对象之一,而原因是数学的,不是态度的。任何"更努力"都到不了 4-6。
进球的数学:为什么足球比分像一个泊松过程
先从每个严肃足球模型都建立其上的机制说起。足球进球是稀少、大致独立、散布在九十分钟里的事件,而这种形状的事件被泊松分布描述得出奇地好。如果你能估计一支球队预期进多少球——称之为 lambda,即它的进球率——泊松分布就把它进恰好零、一、二、三……球的概率交给你。把两队的分布相乘(还有一个稍后会讲的修正),你就得到了每一个可能比分的概率:0-0、1-0、2-1、4-6,全都有。
这里是让人吃惊的部分。假设你对两队预期进球有一个完美的估计——不是好的估计,而是由神谕交给你的、毫无瑕疵的估计。把它喂进泊松机制,问最可能的那个准确比分。在一场势均力敌两队的典型比赛里,那个最可能的比分——通常是 1-0、1-1 或 2-1——只携带约 11% 到 13% 的概率。这就是天花板。即便完美知道底层进球率,最可能的单一比分也勉强八次里对一次。可能比分的分布实在太分散,没有任何单一格子能占主导。这不是泊松的缺陷;这是泊松在报告的真相。足球比分本质上、不可约地高方差。
为什么最可能的比分仍然通常是错的
请在这个后果上停留片刻,因为它重构了一切。一个完美校准的模型——其概率精确匹配现实——如果你逼它说出单一准确比分,它仍会大约 85% 到 88% 的时候错。不是因为它是个坏模型,而是因为它所预测的世界里,1-0、2-1、1-1、0-0、3-2 都以有意义、彼此重叠的频率发生,其中没有任何单一比分在任何稳定意义上是"正确答案"。准确比分是从一个宽分布里抽出的一个样本,而抽中众数值,按其构造,就是一个少数事件。
我们自己的公开台账正是这么说的。整届比赛我们最高单一比分预测约 15% 的时候命中——比理论众数天花板高出一丝,一个诚实、调校良好的模型大致就该待在这里。当人们要我们"预测比分",并想象一个 AI 在开球前钉死 4-6 的未来时,他们想象的是一个击败了没有模型能击败之极限的模型,因为极限不在模型里。它在比赛里。
两种不确定性——以及为什么足球有太多错的那一种
要看清希望在哪、不在哪,你必须把不确定性拆成它的两个物种,因为它们的行为完全不同。
- 认知不确定性是无知的不确定性——你不知道、但本可以知道的东西。谁真正首发。某位关键后卫是否带伤。战术计划。近期状态、旅途、休息、动机。这种不确定性是可约的:更好的数据和更好的模型真的能缩小它。
- 偶然不确定性是机遇的不确定性——烙进事件本身的真正随机。一记打得很正的射门是差一寸骗过门将,还是击中立柱。一次折射是弹进还是弹出。裁判是否看见那次手球。这种不确定性是不可约的:再多研究也去不掉,因为它不是无知,它是噪声。
每个预测问题都是两者的混合,而混合比例决定了什么是可能的。预测明天日出几乎是纯认知的——近乎零噪声,因而近乎完美的准确。预测一次抛硬币几乎是纯偶然的——研究无助,50% 永远是天花板。足球比分痛苦地偏向抛硬币那一端。进球稀少,所以每一个都极其关键;第 92 分钟一记折射射门——正是决定本届一场半决赛的那种——推翻了九十分钟球技留下的平局。当决定性事件寥寥、且每一个都浸在机遇里,偶然的地板就很高,而准确比分是整项运动中最受偶然主宰的单一量。你无法靠更用力研究硬币来绕过硬币的随机,而一个比分比大多数球迷愿意相信的更接近一把硬币。
为什么偏偏 4-6 几乎无法预测
现在说法国-英格兰这场本身。一场十球的比赛不仅仅是一个低概率比分;它是一个低概率状态。它需要几件稀罕的事同时发生:两队的把握机会同时火热、两条防线反复失守、以及一个不断打开而非关闭的比赛态势。在任何合理的赛前模型下,一场比赛产出总共十球或以上的概率是不到百分之一的零头。具体格子 4-6 的概率更小——量级在千分之一或更罕见。一个负责任的模型不会、也不应该在那里放上有意义的概率。如果它放了——如果某个模型"命中"了 4-6——那不是智能;那是一个把概率喷向长尾的坏模型,而它会在其余九十九场理智占上风的比赛里显得荒唐。
有一个诚实的但书,它属于认知那一栏:这是一场季军附加赛,一种有据可查地倾向于摊开的赛事。铜牌战场均近四球,约九十年没出现过 0-0——安慰性质的比赛没人拼命防守,两位主帅都轮换,整件事都朝混乱倾斜。一个带有强烈"死局开放度"特征的模型本应加宽其比分分布、并把总进球往上调。我们的确实把家族加宽了一些,但远不到十球——因为远不到十球才是正确的赛前立场。即便一个完美理解死局效应的模型,也会预测比如一个尾部更肥的 3-2 形状,而不是 4-6。死局特征是我们能磨利的、真实的、可约的优势。它移动分布。它不会、也不能让你叫出那场具体的雪崩。
信息天花板:为什么约 15% 是定律,不是目标
还有一种更深的方式看这个天花板,借自信息论。一场足球赛里可信比分的分布——从 0-0 一路到现实的 4-3 和 5-2——携带约三个半到四比特的熵。熵是对真正不可预测性的度量,而那个数字正式地说明:结果空间又宽又平——许多比分分享着概率质量,没有任何一个高耸于其余之上。当你从一个如此平的分布里抽一个样本,任何预测者在叫出准确值上能做的最好——即便是一个完美知道真实分布的预测者——就是叫出它的众数,而众数只以众数频率获胜。对足球比分而言,那个频率就是理论与我们的实盘结果反复落在的 12% 到 17% 区间。足球预测的学术研究,跨越许多联赛与方法,收敛到同一个天花板:最先进模型的准确比分 top-1 命中率大致封顶在 15% 到 18%。几十年来它没怎么移动,不是因为模型没进步——它们进步巨大——而是因为天花板从来就不关乎模型。
那么,有希望吗?有——但你必须重新定义"准确"
这里是建设性的转折,也是这个平台存在的理由。"AI 能预测准确比分吗"这个问题偷偷夹带了一个准确的定义——叫出那个唯一正确的格子——而它接近一个杂耍戏法,且如我们所见,物理上封顶在约 15%。把定义换成某种既诚实又有用的东西,希望就变得真实而巨大。
- 预测分布,而非点。真正准确的输出不是"2-1",而是整个概率面——1-0 占 13%,1-1 占 11%,2-1 占 9%,如此等等。用校准来评判它:当模型说 30%,它是否 30% 的时候发生?一个模型在这个意义上可以美妙地准确,同时在五场里有四场"判错"单一比分,因为它从未声称知道那个单一比分——它声称知道赔率,而它确实知道。我们公布的 Brier 分数,一个对概率准确性的正当度量,才是我们真正押上名声的数字,而不是准确比分命中率。
- 预测一个集合,而非一个点。如果最高单一比分命中约 15%,那么最可能的前 六 个比分加起来约捕获 50%。那是一个真正有用的对象:一个反映真实不确定性、而非隐藏它的对冲。我们的分层产品正是这么做——一个胜负判断、一个 top-3 集合、一个 top-6 集合——正因为单一比分的声称会对"什么是可知的"不诚实。
- 爬到真正可预测的层。比分是最难的层。它上面的层要好办得多:比赛胜负(胜/平/负)由一个强模型可在 55% 到 65% 判出;总进球(大/小于某盘口)可远高于随机地判出;淘汰赛谁晋级更可预测。准确不是一个数字——它是一把梯子,即便准确比分那一级仍然滑,更高的级也承受真正的重量。
- 等比赛开口说话。比分准确性最大的单一来源是时间本身。每一分钟的比赛都把偶然不确定性转化为被观测的事实。一个盯着法国-英格兰的赛前模型永远叫不出 4-6——但一个在第 70 分钟、已经看着五个球飞进、防线肉眼可见地瓦解的实时模型,能把它的分布猛烈地推向高比分收尾。实时预测本质上是比赛前预测更容易的问题,因为噪声正在你面前消解。"准确比分预测"的希望压倒性地是一个实时预测的希望。
什么才真正拨动指针(以及能拨多远)
这一切都不意味着赛前模型无路可走。认知的前沿是真实的,更好的建模确实推动天花板——只是推不到月亮。真正要紧的改进:用射门质量而非原始历史进球构建的更丰富的预期进球输入;对谁真正上场作出反应的阵容感知建模;相关进球模型——Dixon-Coles 修正及其后代——通过承认两队进球并非独立,修正泊松对低比分与平局的已知低估;死局效应、天气、赛程拥挤、休息等情境特征;以及混入市场赔率,它把人群的信息打包成一个锐利的单一信号。这每一项都值得真实的准确。跨越一个赛季叠加起来,它们或许能把准确比分 top-1 从十几提到接近二十,而——更有价值地——锐化整个分布,使 top-6 集合与胜负判断可度量地更好。这就是奖赏诚实的大小:不是准确比分从 15% 跃到 50%(那不可能),而是一个分布的稳步收紧,而那个分布将永远、正确地拒绝向你承诺那唯一的格子。
裁决——以及我们为什么公布而非埋葬我们的 4-6 失误
那么,AI 能预测准确比分吗?对一场比赛、开球前、作为一个自信的点判断:不能——任何声称能的产品,要么在撒谎,要么误解了自己的模型。准确比分被不可约的机遇主宰;众数格子封顶在约 15%;4-6 活在没有任何负责任模型能吹大的长尾里。那个天花板是比赛的定律,不是技术的暂时局限,没有任何未来模型,无论多大,能废除它。
但对准确的 AI 足球预测有希望吗?斩钉截铁地有——一旦你让"准确"意味着那个真实而有用的东西,而非杂耍戏法。在校准良好的概率分布里、在用六个比分捕获一半结果的诚实对冲集合里、在胜负与总进球与晋级这些极其可判的层里、在随比赛实时消解随机而变锐的实时模型里,有着深厚而复利的希望。足球预测的前沿不是一场在开球前叫出 4-6 的竞赛。它是一场比任何人都更诚实地量化不确定性的竞赛——而那场竞赛敞开着,值得去跑。
这正是为什么,当我们的法国判断回来是 4-6 时,我们没有悄悄删掉它。我们把它留在卡片上,划着叉,公开评为一次失误——就在那些校准数字与 top-6 集合旁边,它们讲述着这个模型知道什么、以及没有模型能知道什么的真实故事。一个假装自己本可以叫出那场比赛的预测引擎,会是一个更差的产品,不是更好的。诚实不是危机公关。诚实就是准确——关于我们自身极限的准确——而在一个如此被机遇主宰的领域里,那或许是唯一值得信任的那种准确。AI 预测仅供娱乐;诚实才是产品。
Bolehkah AI Meramal Skor Tepat? Apa France 4-6 England Ajar Kita Tentang Had Ramalan
Sabtu malam, model kami menilai perlawanan tempat ketiga Piala Dunia, menimbang France lawan England, dan mengunci panggilan: France menang, paling mungkin 2-1, dengan 3-1 dan 3-2 sebagai jiran terdekat. Perlawanan berakhir France 4-6 England. Sepuluh gol. Satu skor yang model kami, setiap model, dan setiap penjudi di dunia, telah harga pada sekitar satu peluang dalam seribu. Kami tersasar pemenang, tersasar margin, tersasar seluruh sifat perlawanan. Jadi mari tanya soalan yang dituntut oleh kesilapan itu, secara terus dan teknikal: bolehkah AI benar-benar meramal skor tepat sesebuah perlawanan bola sepak — dan adakah harapan sebenar untuk betul? Jawapan jujurnya lebih menarik daripada ya atau tidak yang mudah, dan ia bermula dengan memahami jenis masalah apa sebenarnya sebuah skor itu.
Apa kami ramal, dan apa alam semesta sampaikan
Panggilan berkunci kami meletak France mendahului pada kira-kira lambungan-syiling-lebih-sikit, dengan profil jangkaan gol yang menghasilkan keluarga skor yang ketat dan rendah — 2-1, 3-1, 3-2 — bentuk yang biasa diambil oleh perlawanan elit yang seimbang. Itu bukan tekaan malas; ia hasil sebuah model yang telah digredkan 102 ramalan secara terbuka kejohanan ini, menepati 83% panggilan keputusannya, dan membaca kadar asas bola sepak kalah mati dengan betul kebanyakan masa. Dan ia tersasar pada setiap paksi sebuah skor boleh tersasar. England menang. Margin dua. Jumlah sepuluh. Jika anda tanya model, sebelum sepak mula, kebarangkalian tepat 4-6, ia akan pulangkan nombor dengan dua sifar selepas titik perpuluhan sebelum digit bererti pertama. Inilah fakta janggal namun menjelaskan di tengah seluruh soalan: skor akhir tepat sesebuah perlawanan bola sepak ialah salah satu perkara paling sukar dalam ramalan gunaan untuk dipanggil, dan sebabnya matematik, bukan motivasi. Tiada jumlah "cuba lebih kuat" sampai ke 4-6.
Matematik gol: mengapa skor bola sepak berkelakuan seperti proses Poisson
Mula dengan jentera yang setiap model bola sepak serius dibina di atasnya. Gol dalam bola sepak ialah peristiwa jarang, lebih kurang bebas, bertaburan merentas sembilan puluh minit, dan peristiwa berbentuk begitu digambarkan dengan sangat baik oleh taburan Poisson. Jika anda boleh anggar berapa gol sesebuah pasukan dijangka jaring — panggil ia lambda, kadar jaringannya — taburan Poisson memberi anda kebarangkalian mereka menjaring tepat sifar, satu, dua, tiga, dan seterusnya. Darabkan taburan kedua-dua pasukan (dengan satu pembetulan yang kita akan sampai) dan anda ada kebarangkalian untuk setiap skor mungkin: 0-0, 1-0, 2-1, 4-6, semuanya.
Inilah bahagian yang mengejutkan orang. Andaikan anda ada anggaran sempurna jangkaan gol kedua-dua pasukan — bukan anggaran baik, satu yang tanpa cela, dihulur oleh sang orakel. Suapkan ke jentera Poisson dan tanya skor tepat yang paling mungkin. Dalam perlawanan tipikal antara pihak yang seimbang, skor paling mungkin itu — biasanya 1-0, 1-1, atau 2-1 — membawa kebarangkalian hanya sekitar 11 hingga 13 peratus. Itulah silingnya. Walau dengan pengetahuan sempurna kadar gol asas, skor tunggal paling berkemungkinan betul hampir-hampir sekali dalam lapan. Taburan skor mungkin terlalu tersebar untuk mana-mana sel tunggal mendominasi. Ini bukan kecacatan Poisson; ini kebenaran yang Poisson laporkan. Skor bola sepak secara semula jadi, secara tak boleh dikurang, bervarians tinggi.
Mengapa skor paling mungkin masih biasanya tersasar
Duduk dengan akibatnya seketika, kerana ia membingkai semula segalanya. Sebuah model yang ditentukur sempurna — yang kebarangkaliannya tepat sepadan realiti — masih akan, jika anda paksa ia namakan satu skor tepat, tersasar kira-kira 85 hingga 88 peratus masa. Bukan kerana ia model buruk. Kerana dunia yang ia ramal ialah dunia di mana 1-0 dan 2-1 dan 1-1 dan 0-0 dan 3-2 semua berlaku dengan kekerapan bererti yang bertindih, dan tiada satu pun daripadanya ialah "jawapan betul" dalam apa jua erti stabil. Skor tepat ialah cabutan daripada taburan yang luas, dan mencabut nilai mod, secara binaan, ialah peristiwa minoriti.
Lejar terbuka kami sendiri berkata betul-betul ini. Merentas kejohanan, pilihan skor tunggal teratas kami menepati kira-kira 15 peratus masa — sehelai rambut di atas siling mod teori, kira-kira di mana sebuah model jujur dan ditala baik sepatutnya duduk. Bila orang minta kami "ramal skor," dan bayangkan masa depan di mana AI paku 4-6 sebelum sepak mula, mereka bayangkan model yang mengalahkan had yang tiada model boleh kalahkan, kerana had itu bukan dalam model. Ia dalam permainan.
Dua jenis ketidakpastian — dan mengapa bola sepak ada terlalu banyak jenis yang salah
Untuk lihat di mana harapan hidup dan di mana tidak, anda mesti pecahkan ketidakpastian kepada dua spesiesnya, kerana keduanya berkelakuan sama sekali berbeza.
- Ketidakpastian epistemik ialah ketidakpastian kejahilan — benda yang anda tak tahu tetapi boleh. Siapa sebenarnya sebelas pemula. Sama ada seorang pertahanan utama menanggung kecederaan. Rancangan taktikal. Bentuk terkini, perjalanan, rehat, motivasi. Ketidakpastian ini boleh dikurang: data lebih baik dan model lebih baik benar-benar mengecutkannya.
- Ketidakpastian aleatorik ialah ketidakpastian nasib — kerawakan tulen yang terbakar ke dalam peristiwa itu sendiri. Sama ada rembatan yang kena betul mengalahkan penjaga gol sejengkal atau kena tiang. Sama ada lantunan melengkung masuk atau keluar. Sama ada pengadil nampak sepak tangan itu. Ketidakpastian ini tak boleh dikurang: tiada jumlah kajian membuangnya, kerana ia bukan kejahilan, ia hingar.
Setiap masalah ramalan ialah campuran kedua-duanya, dan campuran itu menentukan apa yang mungkin. Meramal matahari terbit esok hampir tulen epistemik — hampir-sifar hingar, jadi hampir-sempurna ketepatan. Meramal satu lambungan syiling hampir tulen aleatorik — tiada kajian membantu, 50% ialah siling selamanya. Skor bola sepak duduk menyakitkan jauh ke hujung lambungan-syiling. Gol jarang, jadi setiap satu amat berpengaruh; satu rembatan terlantun di minit ke-92 — jenis yang memutuskan salah satu separuh akhir kejohanan ini — menterbalikkan keputusan yang sembilan puluh minit kemahiran biarkan seri. Bila peristiwa penentu sedikit dan setiap satu direndam nasib, lantai aleatorik tinggi, dan skor tepat ialah kuantiti tunggal paling didominasi-aleatorik dalam seluruh sukan itu. Anda tak boleh model laluan melepasi kerawakan syiling dengan mengkaji syiling lebih kuat, dan sebuah skor lebih dekat kepada segenggam syiling daripada kebanyakan peminat mahu percaya.
Mengapa 4-6 khususnya hampir tak boleh diramal
Kini perlawanan France-England itu sendiri. Perlawanan sepuluh gol bukan sekadar skor berkebarangkalian rendah; ia rejim berkebarangkalian rendah. Ia perlukan beberapa perkara jarang berlaku serentak: penamat kedua-dua pasukan panas serentak, kedua-dua struktur pertahanan gagal berulang, dan keadaan permainan yang terus membuka bukan menutup. Di bawah mana-mana model pra-perlawanan munasabah, kebarangkalian sebuah perlawanan hasilkan sepuluh gol atau lebih ialah pecahan satu peratus. Kebarangkalian sel khusus 4-6 lebih kecil lagi — pada urutan satu dalam seribu atau lebih jarang. Model bertanggungjawab tidak, dan tak sepatutnya, letak kebarangkalian bererti di sana. Jika ia buat — jika sesetengah model "memanggil" 4-6 — itu bukan kecerdasan; ia model rosak menyembur kebarangkalian ke ekor, dan ia akan nampak mengarut merentas sembilan puluh sembilan perlawanan lain di mana kewarasan menang.
Ada satu peringatan jujur, dan ia milik lajur epistemik: ini perlawanan tempat ketiga, fixture dengan kecenderungan berdokumen untuk terbuka luas. Perlawanan gangsa purata hampir empat gol dan tidak menghasilkan seri kosong dalam kira-kira sembilan puluh tahun — tiada siapa bertahan mati-matian dalam perlawanan saguhati, kedua-dua pengurus putar, dan seluruhnya condong kepada kekacauan. Sebuah model yang membawa ciri "keterbukaan perlawanan mati" yang kuat sepatutnya luaskan taburan skornya dan condongkan jumlahnya ke atas. Kami memang luaskan keluarga itu sedikit, tetapi jauh dari sepuluh gol — kerana jauh dari sepuluh gol ialah pendirian pra-perlawanan yang betul. Walau model yang faham sempurna kesan perlawanan-mati akan ramal, katakan, bentuk 3-2 dengan ekor lebih gemuk, bukan 4-6. Ciri perlawanan-mati ialah kelebihan sebenar, boleh dikurang, yang kami boleh tajamkan. Ia gerak taburan. Ia tidak, dan tak boleh, benarkan anda panggil runtuhan khusus itu.
Siling maklumat: mengapa ~15% ialah hukum, bukan sasaran
Ada cara lebih dalam melihat siling itu, dipinjam dari teori maklumat. Taburan skor munasabah dalam perlawanan bola sepak — semua dari 0-0 hingga 4-3 dan 5-2 yang realistik — membawa entropi kira-kira tiga setengah hingga empat bit. Entropi ialah ukuran ketidakbolehramalan tulen, dan angka itu berkata, secara formal, bahawa ruang keputusan luas dan rata: banyak skor berkongsi jisim kebarangkalian, tiada menjulang atas yang lain. Bila anda ambil satu cabutan dari taburan serata itu, terbaik yang mana-mana peramal boleh buat pada namakan nilai tepat — walau peramal yang tahu taburan sebenar dengan sempurna — ialah namakan modnya, dan mod menang hanya pada kekerapan mod. Untuk skor bola sepak, kekerapan itu ialah jalur 12-hingga-17-peratus yang kedua-dua teori dan keputusan langsung kami terus mendarat. Kajian akademik ramalan bola sepak, merentas banyak liga dan kaedah, menumpu ke siling sama: kadar tepat skor top-1 untuk model tercanggih memuncak sekitar 15 hingga 18 peratus. Ia tak banyak bergerak dalam dekad, bukan kerana model tak bertambah baik — mereka bertambah baik, secara dramatik — tetapi kerana siling itu tidak pernah tentang model.
Jadi adakah harapan? Ya — tetapi anda mesti takrif semula "tepat"
Inilah pusingan konstruktif, dan ia sebab platform ini wujud. Soalan "bolehkah AI ramal skor tepat" menyeludup satu takrifan ketepatan — namakan satu sel betul — yang hampir helah silap mata dan, seperti kita lihat, terhad secara fizikal sekitar 15%. Tukar takrifan kepada sesuatu yang jujur dan berguna, dan harapan jadi nyata dan besar.
- Ramal taburan, bukan titik. Output yang benar-benar tepat bukan "2-1" tetapi seluruh permukaan kebarangkalian — 1-0 pada 13%, 1-1 pada 11%, 2-1 pada 9%, dan seterusnya. Nilai itu dengan penentukuran: bila model kata 30%, adakah ia berlaku 30% masa? Sebuah model boleh tepat dengan indah dalam erti ini sambil "tersasar" pada skor tunggal empat perlawanan dari lima, kerana ia tak pernah dakwa tahu skor tunggal — ia dakwa tahu odds, dan ia tahu. Skor Brier kami yang diterbitkan, ukuran wajar ketepatan kebarangkalian, ialah nombor yang kami betul-betul pertaruh nama, bukan kadar tepat skor.
- Ramal set, bukan titik. Jika skor tunggal teratas menepati ~15%, enam skor paling mungkin bersama menangkap sekitar 50%. Itu objek benar-benar berguna: lindung nilai yang mencerminkan ketidakpastian sebenar bukannya menyembunyikannya. Produk bertingkat kami buat betul-betul ini — panggilan keputusan, set top-3, dan set top-6 — tepatnya kerana dakwaan skor tunggal akan tidak jujur tentang apa yang boleh diketahui.
- Naik ke lapisan yang benar-benar boleh diramal. Skor ialah lapisan tersukar. Lapisan atasnya jauh lebih tertangani: keputusan perlawanan (menang/seri/kalah) boleh dipanggil pada 55-65% oleh model kuat; jumlah gol (atas/bawah satu garis) boleh dipanggil jauh atas peluang; siapa mara dalam kalah mati lebih boleh diramal lagi. Ketepatan bukan satu nombor — ia tangga, dan anak tangga lebih tinggi menanggung berat sebenar walau anak tangga skor-tepat kekal licin.
- Tunggu perlawanan bercakap. Sumber tunggal terbesar ketepatan skor ialah masa itu sendiri. Setiap minit permainan menukar ketidakpastian aleatorik kepada fakta terperhati. Model pra-perlawanan menatap France-England tak pernah boleh panggil 4-6 — tetapi model langsung di minit ke-70, telah menonton lima gol terbang masuk dan pertahanan jelas berkecai, boleh anjak taburannya ganas ke arah penamat berskor tinggi. Ramalan dalam-permainan ialah masalah asasnya lebih mudah daripada ramalan pra-perlawanan, kerana hingar sedang selesai di depan anda. Harapan untuk "ramalan skor tepat" secara berat sebelah ialah harapan ramalan langsung.
Apa yang benar-benar gerakkan jarum (dan sejauh mana)
Tiada satu pun ini bermaksud model pra-perlawanan terperangkap. Sempadan epistemik nyata, dan pemodelan lebih baik benar-benar tolak siling — cuma bukan ke bulan. Penambahbaikan yang penting: input jangkaan gol lebih kaya dibina dari kualiti rembatan bukan gol sejarah mentah; pemodelan sedar-barisan yang bertindak balas kepada siapa sebenarnya bermain; model jaringan berkorelasi — pembetulan Dixon-Coles dan keturunannya — yang membaiki kekurangan anggar Poisson yang diketahui pada skor rendah dan seri dengan mengakui gol kedua-dua pasukan tidak bebas; ciri konteks seperti kesan perlawanan-mati, cuaca, kesesakan, dan rehat; dan mengadun odds pasaran, yang membungkus maklumat orang ramai ke dalam satu isyarat tajam. Setiap ini bernilai ketepatan sebenar. Ditimbun bersama, merentas semusim, ia mungkin angkat tepat skor top-1 dari belasan rendah ke belasan tinggi, dan — lebih bernilai — tajamkan seluruh taburan supaya set top-6 dan panggilan keputusan jadi lebih baik secara boleh diukur. Itulah saiz jujur hadiahnya: bukan lompatan dari 15% ke 50% pada skor tepat, yang mustahil, tetapi pengetatan mantap sebuah taburan yang akan sentiasa, dengan betul, enggan janji anda sel tunggal.
Keputusan — dan mengapa kami terbitkan kesilapan 4-6 kami bukannya tanam
Jadi, bolehkah AI ramal skor tepat? Untuk satu perlawanan, sebelum sepak mula, sebagai panggilan titik yakin: tidak — dan mana-mana produk yang dakwa sebaliknya sama ada menipu atau salah faham modelnya sendiri. Skor tepat didominasi nasib tak boleh dikurang; sel mod memuncak sekitar 15%; 4-6 hidup dalam ekor yang tiada model bertanggungjawab boleh besarkan. Siling itu ialah hukum permainan, bukan batasan sementara teknologi, dan tiada model masa depan, sebesar mana pun, memansuhkannya.
Tetapi adakah harapan untuk ramalan bola sepak AI yang tepat? Tegas ya — sebaik anda biar "tepat" bermaksud benda benar dan berguna bukannya helah silap mata. Ada harapan dalam, berganda, dalam taburan kebarangkalian ditentukur baik; dalam set lindung nilai jujur yang menangkap separuh keputusan dalam enam skor; dalam lapisan keputusan dan jumlah dan kemaraan yang amat boleh dipanggil; dan dalam model langsung yang jadi tajam ketika perlawanan selesaikan kerawakan masa nyata. Sempadan ramalan bola sepak bukan perlumbaan namakan 4-6 sebelum sepak mula. Ia perlumbaan mengukur ketidakpastian lebih jujur daripada sesiapa — dan perlumbaan itu terbuka luas, dan berbaloi dilari.
Itulah tepatnya mengapa, bila panggilan France kami balik 4-6, kami tidak diam-diam padam. Kami biar ia pada kad, bersilang, digred kesilapan secara terbuka — betul-betul di sebelah nombor penentukuran dan set top-6 yang cerita kisah sebenar apa model ini tahu dan apa tiada model boleh. Sebuah enjin ramalan yang berpura ia boleh panggil perlawanan itu akan jadi produk lebih buruk, bukan lebih baik. Kejujuran bukan kawalan kerosakan. Kejujuran ialah ketepatan — ketepatan tentang had kami sendiri — dan dalam bidang yang begini didominasi nasib, itu mungkin satu-satunya jenis yang berbaloi dipercayai. Ramalan AI untuk hiburan; kejujuran ialah produknya.
AI துல்லியமான ஸ்கோரை கணிக்க முடியுமா? France 4-6 England நமக்கு கற்பித்த பாடம்
சனிக்கிழமை இரவு, எங்கள் மாடல் உலகக் கோப்பை மூன்றாம் இட ஆட்டத்தை ஆய்ந்து, France-ஐயும் England-ஐயும் எடைபோட்டு, ஒரு கணிப்பை பூட்டியது: France வெல்லும், மிக சாத்தியமாக 2-1, அருகில் 3-1, 3-2. ஆட்டம் France 4-6 England என முடிந்தது. பத்து கோல்கள். எங்கள் மாடல், எல்லா மாடல்களும், உலகின் ஒவ்வொரு பந்தய நிறுவனமும் ஆயிரத்தில் ஒன்று என விலையிட்ட ஒரு ஸ்கோர். வெற்றியாளரை தவறவிட்டோம், வித்தியாசத்தை தவறவிட்டோம், ஆட்டத்தின் முழு குணத்தையும் தவறவிட்டோம். எனவே அந்த தவறு கோரும் கேள்வியை நேரடியாக, தொழில்நுட்பமாக கேட்போம்: ஒரு AI உண்மையில் ஒரு கால்பந்து ஆட்டத்தின் துல்லிய ஸ்கோரை கணிக்க முடியுமா — சரியாக கணிக்க ஏதேனும் உண்மையான நம்பிக்கை உண்டா? நேர்மையான பதில் எளிய ஆம் அல்லது இல்லையை விட சுவாரஸ்யமானது; அது ஒரு ஸ்கோர் எந்த வகை பிரச்சனை என்பதை புரிந்துகொள்வதில் தொடங்குகிறது.
நாம் கணித்தது, பிரபஞ்சம் வழங்கியது
எங்கள் பூட்டிய கணிப்பு France-ஐ ஏறத்தாழ நாணய-சுண்டல்-கூட்டல் இடத்தில் வைத்தது, அதன் எதிர்பார்க்கப்பட்ட-கோல் விவரம் ஒரு இறுக்கமான, குறைந்த-ஸ்கோர் ஸ்கோர் குடும்பத்தை உருவாக்கியது — 2-1, 3-1, 3-2 — சமபலமான உயர்தர ஆட்டம் பொதுவாக எடுக்கும் வடிவங்கள். அது சோம்பேறி யூகம் அல்ல; இந்த போட்டியில் பொதுவில் 102 கணிப்புகளை மதிப்பிட்டு, தன் தீர்க்கமான கணிப்புகளில் 83% சரியாகி, நாக்-அவுட் கால்பந்தின் அடிப்படை விகிதங்களை பெரும்பாலான நேரங்களில் சரியாக படிக்கும் ஒரு மாடலின் வெளியீடு. அது ஒரு ஸ்கோர் தவறாகக்கூடிய ஒவ்வொரு அச்சிலும் தவறியது. England வென்றது. வித்தியாசம் இரண்டு. மொத்தம் பத்து. தொடக்கத்திற்கு முன், சரியாக 4-6-க்கான நிகழ்தகவை மாடலிடம் கேட்டிருந்தால், முதல் முக்கிய இலக்கத்திற்கு முன் தசம புள்ளிக்குப் பின் இரண்டு பூஜ்ஜியங்களுடன் ஒரு எண்ணை திருப்பியிருக்கும். இதுவே முழு கேள்வியின் மையத்தில் உள்ள அசௌகரியமான, தெளிவூட்டும் உண்மை: ஒரு கால்பந்து ஆட்டத்தின் துல்லிய இறுதி ஸ்கோர் பயன்பாட்டு கணிப்பில் உண்மையிலேயே கடினமான ஒன்று; காரணங்கள் கணிதமானவை, உந்துதல் சார்ந்தவை அல்ல. எவ்வளவு "அதிக முயற்சி" செய்தாலும் 4-6 எட்டாது.
கோல்களின் கணிதம்: கால்பந்து ஸ்கோர்கள் ஏன் ஒரு பாய்சான் செயல்முறை போல செயல்படுகின்றன
ஒவ்வொரு தீவிர கால்பந்து மாடலும் கட்டப்பட்டுள்ள இயந்திரத்தில் தொடங்குவோம். கால்பந்தில் கோல்கள் அரிதான, தோராயமாக சுதந்திரமான, தொண்ணூறு நிமிடங்களில் சிதறிய நிகழ்வுகள்; அந்த வடிவ நிகழ்வுகளை பாய்சான் பரவல் குறிப்பிடத்தக்க அளவில் நன்றாக விவரிக்கிறது. ஒரு அணி எத்தனை கோல்கள் அடிக்கும் என எதிர்பார்க்கப்படுகிறது என்பதை மதிப்பிட முடிந்தால் — அதை lambda, அதன் கோல் விகிதம் என அழையுங்கள் — பாய்சான் பரவல் அவை சரியாக பூஜ்ஜியம், ஒன்று, இரண்டு, மூன்று… அடிக்கும் நிகழ்தகவை தருகிறது. இரு அணிகளின் பரவல்களை பெருக்கினால் (பின்னர் வரும் ஒரு திருத்தத்துடன்) ஒவ்வொரு சாத்தியமான ஸ்கோருக்கும் நிகழ்தகவு கிடைக்கும்: 0-0, 1-0, 2-1, 4-6, எல்லாமே.
மக்களை ஆச்சரியப்படுத்தும் பகுதி இதுதான். இரு அணிகளின் எதிர்பார்க்கப்பட்ட கோல்களுக்கு உங்களிடம் ஒரு சரியான மதிப்பீடு இருப்பதாக வைத்துக்கொள்ளுங்கள் — நல்ல மதிப்பீடு அல்ல, குறையற்ற ஒன்று, ஒரு தேவவாக்கால் கொடுக்கப்பட்டது. அதை பாய்சான் இயந்திரத்தில் ஊட்டி, மிக சாத்தியமான ஒற்றை துல்லிய ஸ்கோரை கேளுங்கள். சமபலமான அணிகளுக்கிடையேயான வழக்கமான ஆட்டத்தில், அந்த மிக சாத்தியமான ஸ்கோர் — பொதுவாக 1-0, 1-1, அல்லது 2-1 — சுமார் 11 முதல் 13 சதவீதம் மட்டுமே நிகழ்தகவை சுமக்கிறது. அதுவே உச்சவரம்பு. அடிப்படை கோல் விகிதங்களை சரியாக அறிந்தாலும், மிக சாத்தியமான ஒற்றை ஸ்கோர் எட்டில் ஒரு முறை மட்டுமே சரி. சாத்தியமான ஸ்கோர்களின் பரவல் எந்த ஒற்றை கட்டமும் ஆதிக்கம் செலுத்த முடியாத அளவுக்கு பரந்திருக்கிறது. இது பாய்சானின் குறை அல்ல; பாய்சான் அறிவிக்கும் உண்மை இது. கால்பந்து ஸ்கோர்கள் இயல்பாகவே, குறைக்க முடியாத வகையில், அதிக மாறுபாடு கொண்டவை.
மிக சாத்தியமான ஸ்கோர் ஏன் இன்னும் பொதுவாக தவறு
அந்த விளைவுடன் சிறிது தங்குங்கள், ஏனெனில் அது எல்லாவற்றையும் மறுவடிவமைக்கிறது. சரியாக அளவீடு செய்யப்பட்ட ஒரு மாடல் — அதன் நிகழ்தகவுகள் யதார்த்தத்துடன் சரியாக பொருந்தும் — ஒற்றை துல்லிய ஸ்கோரை சொல்ல நீங்கள் கட்டாயப்படுத்தினால், இன்னும் சுமார் 85 முதல் 88 சதவீதம் நேரம் தவறாகும். அது மோசமான மாடல் என்பதால் அல்ல. அது கணிக்கும் உலகில் 1-0, 2-1, 1-1, 0-0, 3-2 எல்லாம் அர்த்தமுள்ள, ஒன்றுடன் ஒன்று மேற்பொருந்தும் அதிர்வெண்ணுடன் நிகழ்கின்றன, அவற்றில் எதுவும் எந்த நிலையான அர்த்தத்திலும் "சரியான பதில்" அல்ல. துல்லிய ஸ்கோர் ஒரு பரந்த பரவலிலிருந்து எடுக்கப்பட்ட ஒரு மாதிரி, மற்றும் அதிக-அடர்த்தி மதிப்பை எடுப்பது, அமைப்பின்படியே, ஒரு சிறுபான்மை நிகழ்வு.
எங்கள் சொந்த பொது லெட்ஜரும் இதையே சொல்கிறது. போட்டி முழுவதும் எங்கள் உயர்ந்த ஒற்றை ஸ்கோர் தேர்வு சுமார் 15 சதவீதம் நேரம் சரியானது — கோட்பாட்டு உச்சவரம்பை விட ஒரு முடி மேலே, ஒரு நேர்மையான, நன்கு அமைக்கப்பட்ட மாடல் இருக்க வேண்டிய இடம் அதுதான். மக்கள் எங்களிடம் "ஸ்கோரை கணி" என கேட்டு, தொடக்கத்திற்கு முன் AI 4-6-ஐ ஆணியடிக்கும் எதிர்காலத்தை கற்பனை செய்யும்போது, எந்த மாடலும் வெல்ல முடியாத ஒரு எல்லையை வெல்லும் மாடலை அவர்கள் கற்பனை செய்கிறார்கள், ஏனெனில் எல்லை மாடலில் இல்லை. அது ஆட்டத்தில் உள்ளது.
இரு வகை நிச்சயமின்மை — கால்பந்தில் ஏன் தவறான வகை மிக அதிகம்
நம்பிக்கை எங்கே வாழ்கிறது, எங்கே இல்லை என்பதை காண, நிச்சயமின்மையை அதன் இரு இனங்களாக பிரிக்க வேண்டும், ஏனெனில் அவை முற்றிலும் வேறுபட்டு செயல்படுகின்றன.
- அறிவுசார் நிச்சயமின்மை அறியாமையின் நிச்சயமின்மை — உங்களுக்கு தெரியாத, ஆனால் தெரிந்திருக்க முடிந்த விஷயங்கள். உண்மையில் யார் தொடக்க பதினொருவர். ஒரு முக்கிய பாதுகாவலர் காயத்துடன் இருக்கிறாரா. தந்திர திட்டம். சமீபத்திய படிவம், பயணம், ஓய்வு, உந்துதல். இந்த நிச்சயமின்மை குறைக்கக்கூடியது: சிறந்த தரவும் சிறந்த மாடலும் அதை உண்மையிலேயே சுருக்குகின்றன.
- சீரற்ற (aleatoric) நிச்சயமின்மை வாய்ப்பின் நிச்சயமின்மை — நிகழ்விலேயே சுட்டெரிக்கப்பட்ட உண்மையான சீரற்ற தன்மை. நன்கு அடித்த ஷாட் கோல்காப்பாளரை ஒரு அங்குலத்தால் வெல்கிறதா அல்லது கம்பத்தில் மோதுகிறதா. ஒரு திசைதிருப்பல் உள்ளே வளைகிறதா வெளியே வளைகிறதா. நடுவர் அந்த கை-பந்தை பார்க்கிறாரா. இந்த நிச்சயமின்மை குறைக்க முடியாதது: எவ்வளவு படித்தாலும் நீங்காது, ஏனெனில் அது அறியாமை அல்ல, அது இரைச்சல்.
ஒவ்வொரு கணிப்பு பிரச்சனையும் இரண்டின் கலவை, அந்த கலவையே சாத்தியமானதை தீர்மானிக்கிறது. நாளைய சூரிய உதயத்தை கணிப்பது கிட்டத்தட்ட தூய அறிவுசார் — கிட்டத்தட்ட-பூஜ்ஜிய இரைச்சல், எனவே கிட்டத்தட்ட-சரியான துல்லியம். ஒரு நாணய சுண்டலை கணிப்பது கிட்டத்தட்ட தூய சீரற்ற — படிப்பு உதவாது, 50% என்றென்றும் உச்சவரம்பு. கால்பந்து ஸ்கோர்கள் வலிநிறைந்த வகையில் நாணய-சுண்டல் முனையை நோக்கி அமர்ந்துள்ளன. கோல்கள் அரிதானவை, எனவே ஒவ்வொன்றும் மகத்தான விளைவு கொண்டது; 92வது நிமிடத்தில் ஒரு திசைதிருப்பப்பட்ட ஷாட் — இந்த போட்டியின் ஒரு அரையிறுதியை தீர்மானித்த வகை — தொண்ணூறு நிமிட திறமை சமமாக விட்ட முடிவை புரட்டிப்போடுகிறது. தீர்க்கமான நிகழ்வுகள் சிலவாக இருந்து, ஒவ்வொன்றும் வாய்ப்பில் நனைந்திருக்கும்போது, சீரற்ற தளம் உயர்ந்தது, மற்றும் துல்லிய ஸ்கோர் முழு விளையாட்டிலும் மிகவும் சீரற்ற-ஆதிக்க அளவு. நாணயத்தை கடினமாக படிப்பதன் மூலம் நாணயத்தின் சீரற்ற தன்மையை நீங்கள் மாடல் செய்து கடக்க முடியாது, மற்றும் ஒரு ஸ்கோர் பெரும்பாலான ரசிகர்கள் நம்ப விரும்புவதை விட ஒரு கைப்பிடி நாணயங்களுக்கு நெருக்கமானது.
ஏன் குறிப்பாக 4-6 கிட்டத்தட்ட கணிக்க முடியாதது
இப்போது France-England ஆட்டமே. பத்து கோல் ஆட்டம் வெறும் குறைந்த-நிகழ்தகவு ஸ்கோர் அல்ல; அது ஒரு குறைந்த-நிகழ்தகவு ஆட்சி. அது பல அரிதான விஷயங்கள் ஒருங்கே நிகழ வேண்டும்: இரு அணிகளின் முடித்தல் ஒரே நேரத்தில் சூடாக ஓடுவது, இரு பாதுகாப்பு அமைப்புகளும் மீண்டும் மீண்டும் தோல்வியடைவது, மற்றும் மூடுவதற்குப் பதிலாக திறந்து கொண்டே இருக்கும் ஒரு ஆட்ட நிலை. எந்த நியாயமான ஆட்டத்திற்கு-முந்தைய மாடலின் கீழும், ஒரு ஆட்டம் மொத்தம் பத்து அல்லது அதற்கு மேற்பட்ட கோல்கள் தரும் நிகழ்தகவு ஒரு சதவீதத்தின் பகுதி. குறிப்பிட்ட கட்டம் 4-6-இன் நிகழ்தகவு இன்னும் சிறியது — ஆயிரத்தில் ஒன்று அல்லது அரிதான வரிசையில். பொறுப்பான மாடல் அங்கே அர்த்தமுள்ள நிகழ்தகவை வைக்காது, வைக்கக்கூடாது. வைத்தால் — ஏதேனும் மாடல் 4-6-ஐ "அழைத்தால்" — அது புத்திசாலித்தனம் அல்ல; அது வால்களுக்கு நிகழ்தகவை தெளிக்கும் ஒரு உடைந்த மாடல், மற்றும் விவேகம் வென்ற மற்ற தொண்ணூற்றொன்பது ஆட்டங்களில் அது கேலிக்குரியதாக தோன்றும்.
ஒரு நேர்மையான எச்சரிக்கை உள்ளது, அது அறிவுசார் நெடுவரிசைக்கு சொந்தமானது: இது ஒரு மூன்றாம் இட ஆட்டம், ஆவணப்படுத்தப்பட்ட வகையில் திறந்து பரவும் போக்கு கொண்ட ஆட்டம். வெண்கல ஆட்டங்கள் சராசரியாக கிட்டத்தட்ட நான்கு கோல்கள், சுமார் தொண்ணூறு ஆண்டுகளில் கோல் இல்லா சமநிலை தரவில்லை — ஆறுதல் ஆட்டத்தில் யாரும் உயிரை பணயம் வைத்து பாதுகாக்கவில்லை, இரு மேலாளர்களும் சுழற்றுகிறார்கள், முழுவதும் குழப்பத்தை நோக்கி சாய்கிறது. ஒரு வலுவான "இறந்த-ஆட்ட திறந்த தன்மை" அம்சம் கொண்ட மாடல் தன் ஸ்கோர் பரவலை அகலப்படுத்தி, மொத்தத்தை மேலே சாய்த்திருக்க வேண்டும். எங்களுடையது குடும்பத்தை சற்று அகலப்படுத்தியது, ஆனால் பத்து கோல்களுக்கு அருகில் இல்லை — ஏனெனில் பத்து கோல்களுக்கு அருகில் இல்லாததுதான் சரியான ஆட்டத்திற்கு-முந்தைய நிலைப்பாடு. இறந்த-ஆட்ட விளைவை சரியாக புரிந்த மாடலும்கூட, சொல்லுங்கள், கொழுத்த வால்களுடன் ஒரு 3-2 வடிவத்தை கணித்திருக்கும், 4-6 அல்ல. இறந்த-ஆட்ட அம்சம் நாம் கூர்மைப்படுத்தக்கூடிய உண்மையான, குறைக்கக்கூடிய நன்மை. அது பரவலை நகர்த்துகிறது. அது அந்த குறிப்பிட்ட பனிச்சரிவை அழைக்க உங்களை அனுமதிக்காது.
தகவல் உச்சவரம்பு: ~15% ஏன் ஒரு விதி, இலக்கு அல்ல
அந்த உச்சவரம்பை காண இன்னும் ஆழமான வழி உள்ளது, தகவல் கோட்பாட்டிலிருந்து கடன் வாங்கியது. ஒரு கால்பந்து ஆட்டத்தில் சாத்தியமான ஸ்கோர்களின் பரவல் — 0-0 முதல் யதார்த்தமான 4-3, 5-2 வரை — சுமார் மூன்றரை முதல் நான்கு பிட் என்ட்ரோபியை சுமக்கிறது. என்ட்ரோபி உண்மையான கணிக்க-முடியாத்தன்மையின் அளவீடு, அந்த எண் முறையாக சொல்கிறது: விளைவு வெளி பரந்து தட்டையானது — பல ஸ்கோர்கள் நிகழ்தகவு நிறையை பகிர்கின்றன, எதுவும் மற்றவற்றை விட உயரவில்லை. அத்தகைய தட்டையான பரவலிலிருந்து ஒற்றை மாதிரியை எடுக்கும்போது, துல்லிய மதிப்பை சொல்வதில் எந்த கணிப்பாளரும் செய்யக்கூடிய சிறந்தது — உண்மையான பரவலை சரியாக அறிந்த கணிப்பாளர் கூட — அதன் அதிக-அடர்த்தியை சொல்வதுதான், மற்றும் அது அதிக-அடர்த்தி அதிர்வெண்ணில் மட்டுமே வெல்கிறது. கால்பந்து ஸ்கோர்களுக்கு, அந்த அதிர்வெண் கோட்பாடும் எங்கள் நேரடி முடிவுகளும் தொடர்ந்து தரையிறங்கும் 12 முதல் 17 சதவீத பட்டை. கால்பந்து கணிப்பின் கல்வி ஆய்வுகள், பல லீக்குகள் மற்றும் முறைகளில், அதே உச்சவரம்புக்கு ஒன்றுசேர்கின்றன: அதிநவீன மாடல்களுக்கான துல்லிய-ஸ்கோர் top-1 சரி விகிதம் சுமார் 15 முதல் 18 சதவீதத்தில் உச்சம். பல தசாப்தங்களாக அது அதிகம் நகரவில்லை, மாடல்கள் மேம்படாததால் அல்ல — அவை வியத்தகு அளவில் மேம்பட்டன — ஆனால் உச்சவரம்பு ஒருபோதும் மாடல்களை பற்றியதல்ல.
அப்படியானால் நம்பிக்கை உண்டா? உண்டு — ஆனால் "துல்லியம்" என்பதை மறுவரையறை செய்ய வேண்டும்
இதுவே ஆக்கபூர்வமான திருப்பம், இந்த தளம் இருப்பதற்கான காரணமும் இதுவே. "AI துல்லிய ஸ்கோரை கணிக்க முடியுமா" என்ற கேள்வி ஒரு துல்லிய வரையறையை — அந்த ஒற்றை சரியான கட்டத்தை சொல் — கடத்துகிறது, அது ஒரு மந்திர வித்தைக்கு அருகில், மற்றும் நாம் பார்த்தபடி, இயற்பியல் ரீதியாக சுமார் 15%-இல் மட்டுப்படுத்தப்பட்டது. வரையறையை நேர்மையான மற்றும் பயனுள்ள ஒன்றாக மாற்றுங்கள், நம்பிக்கை உண்மையானதாகவும் பெரியதாகவும் ஆகிறது.
- புள்ளியை அல்ல, பரவலை கணி. உண்மையிலேயே துல்லியமான வெளியீடு "2-1" அல்ல, முழு நிகழ்தகவு பரப்பு — 1-0 13%, 1-1 11%, 2-1 9%, இப்படியே. அதை அளவீட்டு (calibration) மூலம் மதிப்பிடுங்கள்: மாடல் 30% சொன்னால், அது 30% நேரம் நிகழ்கிறதா? ஒரு மாடல் இந்த அர்த்தத்தில் அழகாக துல்லியமாக இருக்கலாம், அதே நேரம் ஐந்தில் நான்கு ஆட்டங்களில் ஒற்றை ஸ்கோரில் "தவறாக" இருக்கலாம், ஏனெனில் அது ஒற்றை ஸ்கோரை அறிவதாக ஒருபோதும் கூறவில்லை — அது வாய்ப்புகளை அறிவதாக கூறியது, அறிந்தது. எங்கள் வெளியிட்ட Brier மதிப்பெண், நிகழ்தகவு துல்லியத்தின் சரியான அளவீடு, துல்லிய-ஸ்கோர் சரி விகிதம் அல்ல, நாம் உண்மையில் பெயரை பணயம் வைக்கும் எண்.
- புள்ளியை அல்ல, தொகுப்பை கணி. உயர்ந்த ஒற்றை ஸ்கோர் ~15% சரியானால், மிக சாத்தியமான முதல் ஆறு ஸ்கோர்கள் சேர்ந்து சுமார் 50% பிடிக்கின்றன. அது உண்மையிலேயே பயனுள்ள பொருள்: உண்மையான நிச்சயமின்மையை மறைக்காமல் பிரதிபலிக்கும் ஒரு பாதுகாப்பு. எங்கள் அடுக்கு தயாரிப்பு இதையே செய்கிறது — ஒரு முடிவு கணிப்பு, ஒரு top-3 தொகுப்பு, ஒரு top-6 தொகுப்பு — ஒற்றை-ஸ்கோர் கூற்று அறியக்கூடியதைப் பற்றி நேர்மையற்றதாக இருக்கும் என்பதால்தான்.
- உண்மையில் கணிக்கக்கூடிய அடுக்குகளுக்கு ஏறு. ஸ்கோர் மிக கடினமான அடுக்கு. அதற்கு மேலுள்ள அடுக்குகள் மிக எளிதானவை: ஆட்ட முடிவு (வெற்றி/சமன்/தோல்வி) வலுவான மாடலால் 55-65% கணிக்கக்கூடியது; கோல் மொத்தம் (ஒரு கோட்டுக்கு மேல்/கீழ்) வாய்ப்பை விட நன்றாக கணிக்கக்கூடியது; நாக்-அவுட்டில் யார் முன்னேறுகிறார் இன்னும் கணிக்கக்கூடியது. துல்லியம் ஒரு எண் அல்ல — அது ஒரு ஏணி, துல்லிய-ஸ்கோர் படி வழுக்கினாலும் உயர்ந்த படிகள் உண்மையான எடையை தாங்குகின்றன.
- ஆட்டம் பேசும் வரை காத்திரு. ஸ்கோர் துல்லியத்தின் மிகப்பெரிய ஒற்றை மூலம் நேரமே. ஒவ்வொரு நிமிட ஆட்டமும் சீரற்ற நிச்சயமின்மையை கவனிக்கப்பட்ட உண்மையாக மாற்றுகிறது. France-England-ஐ பார்க்கும் ஆட்டத்திற்கு-முந்தைய மாடல் 4-6-ஐ ஒருபோதும் அழைக்க முடியாது — ஆனால் 70வது நிமிடத்தில், ஏற்கனவே ஐந்து கோல்கள் பறந்து, பாதுகாப்புகள் கண்கூடாக சிதைவதை பார்த்த ஒரு நேரடி மாடல், தன் பரவலை உயர்-ஸ்கோர் முடிவை நோக்கி கடுமையாக நகர்த்த முடியும். ஆட்டத்திற்குள் கணிப்பு ஆட்டத்திற்கு-முந்தைய கணிப்பை விட அடிப்படையில் எளிதான பிரச்சனை, ஏனெனில் இரைச்சல் உங்கள் முன்னால் தீர்ந்துகொண்டிருக்கிறது. "துல்லிய ஸ்கோர் கணிப்பு"க்கான நம்பிக்கை பெருமளவில் ஒரு நேரடி-கணிப்பு நம்பிக்கை.
உண்மையில் எது ஊசியை நகர்த்துகிறது (எவ்வளவு தூரம்)
இதில் எதுவும் ஆட்டத்திற்கு-முந்தைய மாடல்கள் சிக்கியுள்ளன என்று அர்த்தமல்ல. அறிவுசார் எல்லை உண்மையானது, சிறந்த மாடலிங் உண்மையிலேயே உச்சவரம்பை தள்ளுகிறது — சந்திரனுக்கு அல்ல. முக்கியமான மேம்பாடுகள்: மூல வரலாற்று கோல்களை விட ஷாட் தரத்திலிருந்து கட்டப்பட்ட செழுமையான எதிர்பார்க்கப்பட்ட-கோல் உள்ளீடுகள்; உண்மையில் யார் விளையாடுகிறார் என்பதற்கு பதிலளிக்கும் அணி-அறிந்த மாடலிங்; தொடர்புடைய கோல் மாடல்கள் — Dixon-Coles திருத்தமும் அதன் வழித்தோன்றல்களும் — இரு அணிகளின் கோல்கள் சுதந்திரமானவை அல்ல என்பதை ஒப்புக்கொண்டு குறைந்த ஸ்கோர்கள் மற்றும் சமன்களை பாய்சான் குறைமதிப்பிடுவதை சரிசெய்கின்றன; இறந்த-ஆட்ட விளைவு, வானிலை, நெரிசல், ஓய்வு போன்ற சூழல் அம்சங்கள்; மற்றும் சந்தை வாய்ப்புகளை கலப்பது, கூட்டத்தின் தகவலை ஒரு கூர்மையான ஒற்றை சமிக்ஞையாக பொதிகிறது. இவை ஒவ்வொன்றும் உண்மையான துல்லியத்திற்கு மதிப்பு. ஒரு பருவம் முழுவதும் அடுக்கினால், அவை துல்லிய-ஸ்கோர் top-1-ஐ பத்துக்குள் இருந்து இருபதுக்கு அருகில் உயர்த்தலாம், மற்றும் — இன்னும் மதிப்புமிக்கதாக — முழு பரவலையும் கூர்மைப்படுத்தி top-6 தொகுப்பும் முடிவு கணிப்புகளும் அளவிடக்கூடிய அளவில் சிறப்பாகின்றன. அதுவே பரிசின் நேர்மையான அளவு: துல்லிய ஸ்கோரில் 15%-இலிருந்து 50%-க்கு தாவல் அல்ல, அது சாத்தியமற்றது, ஆனால் அந்த ஒற்றை கட்டத்தை உங்களுக்கு உறுதியளிக்க எப்போதும், சரியாக, மறுக்கும் ஒரு பரவலின் நிலையான இறுக்கம்.
தீர்ப்பு — நமது 4-6 தவறை புதைக்காமல் ஏன் வெளியிட்டோம்
எனவே, AI துல்லிய ஸ்கோரை கணிக்க முடியுமா? ஒற்றை ஆட்டத்திற்கு, தொடக்கத்திற்கு முன், ஒரு நம்பிக்கையான புள்ளி கணிப்பாக: முடியாது — வேறுவிதமாக கூறும் எந்த தயாரிப்பும் பொய் சொல்கிறது அல்லது தன் சொந்த மாடலை தவறாக புரிந்துகொள்கிறது. துல்லிய ஸ்கோர் குறைக்க முடியாத வாய்ப்பால் ஆதிக்கம் செலுத்தப்படுகிறது; அதிக-அடர்த்தி கட்டம் சுமார் 15%-இல் உச்சம்; 4-6 எந்த பொறுப்பான மாடலும் ஊதிப்பெருக்க முடியாத ஒரு வாலில் வாழ்கிறது. அந்த உச்சவரம்பு ஆட்டத்தின் விதி, தொழில்நுட்பத்தின் தற்காலிக வரம்பல்ல, எந்த எதிர்கால மாடலும், எவ்வளவு பெரியதாக இருந்தாலும், அதை ரத்து செய்யாது.
ஆனால் துல்லியமான AI கால்பந்து கணிப்புக்கு நம்பிக்கை உண்டா? உறுதியாக உண்டு — "துல்லியம்" என்பது மந்திர வித்தையை அல்ல, உண்மையான பயனுள்ள விஷயத்தை குறிக்க அனுமதித்தவுடன். நன்கு-அளவீடு செய்யப்பட்ட நிகழ்தகவு பரவல்களில்; ஆறு ஸ்கோர்களில் பாதி விளைவுகளை பிடிக்கும் நேர்மையான பாதுகாப்பு தொகுப்புகளில்; முடிவு, மொத்தம், முன்னேற்றம் என்ற மிக கணிக்கக்கூடிய அடுக்குகளில்; ஆட்டம் நிகழ்நேரத்தில் சீரற்ற தன்மையை தீர்க்கும்போது கூர்மையாகும் நேரடி மாடல்களில் — ஆழமான, கூட்டு நம்பிக்கை உண்டு. கால்பந்து கணிப்பின் எல்லை தொடக்கத்திற்கு முன் 4-6-ஐ சொல்லும் ஒரு பந்தயம் அல்ல. அது எவரையும் விட நிச்சயமின்மையை நேர்மையாக அளவிடும் ஒரு பந்தயம் — அந்த பந்தயம் திறந்திருக்கிறது, ஓட மதிப்புள்ளது.
எங்கள் France கணிப்பு 4-6 என திரும்பியபோது, நாம் அதை அமைதியாக நீக்கவில்லை என்பதற்கு இதுவே காரணம். அதை அட்டையில் விட்டோம், குறுக்கு கோடிட்டு, பொதுவில் ஒரு தவறாக மதிப்பிட்டோம் — இந்த மாடல் என்ன அறியும், எந்த மாடலும் என்ன அறியாது என்ற உண்மை கதையை சொல்லும் அளவீட்டு எண்கள் மற்றும் top-6 தொகுப்புகளுக்கு அருகிலேயே. அந்த ஆட்டத்தை தன்னால் அழைக்க முடிந்திருக்கும் என பாசாங்கு செய்யும் ஒரு கணிப்பு இயந்திரம் ஒரு மோசமான தயாரிப்பாக இருக்கும், சிறந்ததல்ல. நேர்மை சேத கட்டுப்பாடு அல்ல. நேர்மைதான் துல்லியம் — நமது சொந்த வரம்புகளைப் பற்றிய துல்லியம் — வாய்ப்பால் இவ்வளவு ஆதிக்கம் செலுத்தப்படும் ஒரு துறையில், அதுவே நம்பத்தக்க ஒரே வகையாக இருக்கலாம். AI கணிப்புகள் பொழுதுபோக்கிற்காக; நேர்மைதான் தயாரிப்பு.