How AI predicts a match result: inside the analysis
·12 min read·updated September 20, 2026
An AI does not predict the result of a match. It calculates how likely each outcome is, and then says how much it trusts its own arithmetic. The difference matters. A prediction sounds like “this is what will happen”. A probability sounds duller and means more: “in a situation like this, on data like this, this is what usually happens”. Between the raw numbers and a finished read there are four steps, and none of them looks like magic. Here they are, in order.
What AI actually does with a match
Imagine someone hands you two hundred numbers about two teams and gives you an hour. Goals, shots, home and away runs, who is injured, who is back, what happened in the last five meetings, how both sides play at the end of a season. The task: say which outcome is more likely.
You will manage it. But by the two hundredth match in a row you will be tired, and by the three hundredth you will be bending the answer towards whatever you decided in the first thirty seconds. People work that way.
A machine does not get tired and does not fall in love with its first guess. That is where its advantage ends. It does not see what is not in the data, does not understand context nobody gave it, and has no idea that the manager fell out with his captain this morning. It becomes useful not because it is “smarter”, but because it does the same boring job to the same standard three hundred times in a row.
Step 1. What data goes into a match analysis
The first thing worth understanding: the quality of a read is decided almost entirely by what went into it. A beautiful model on bad data produces beautiful nonsense.
Statistics and form
The base is results and how they were achieved. Not just the scoreline, but how much a team created, how much it allowed to be created against it, how steady it has been over recent weeks. “Five wins in a row” and “five wins in a row, scraped, with one shot on target” are very different sets of five wins. The metric that separates the two is expected goals — what it measures, and where it misleads, is covered in what is xG in football. How that recent run is actually calculated, and when it’s lying, is its own topic: team form: how it’s actually calculated.
Head-to-head history
How these teams have done against each other. A treacherous indicator: it works while squads and playing styles have stayed put, and turns into noise once two years have brought a new manager, half a new first eleven and a different division. Reading head-to-head records without that correction is the most common mistake in manual analysis — the cutoff rule for when to reset one to zero is in head-to-head (H2H) record explained.
Line-ups
Who is named, who is injured, who is being rested before a European week. One player rarely changes the picture, but a change of goalkeeper or the absence of both centre-backs does.
Fresh news, minutes before the analysis
Statistics describe the past; the match is played today. So before anything is calculated, the system separately checks what has appeared just now: confirmed line-ups and team news, the weather, the appointed referee. This is the one layer that cannot be taken from a database — it has to be looked up at the moment of the analysis.
The consensus layer
A separate layer: how professional sources assess this fixture. Sharkline takes that from more than twenty sources. Not to copy someone else’s verdict, but to see where it disagrees with what the statistics say. The disagreement itself is informative: behind it there is usually either news that has not reached the numbers yet, or a mistake on one of the two sides.
Step 2. Why raw data cannot be used as it is
Between “data collected” and “numbers calculated” sits a step nobody sees from the outside, and it decides more than the rest.
Data contradicts itself. One source lists a player as injured, another has him in the squad. One shows a team in excellent form by results, another shows it in dreadful form by the quality of the chances it creates. The model has to resolve that somehow, and the way it resolves it matters more than the model.
Data goes stale. A line-up known a day before kick-off and a line-up known an hour before are two different line-ups. An analysis made in the morning knows nothing about an injury picked up in the afternoon session.
Sometimes there is simply not much data. A second division in a distant league, a women’s tournament, an early cup round — there can be ten times less material than for a top-flight fixture. That is not a reason to calculate worse. It is a reason to say honestly that there is less confidence here.
Step 3. How data turns into the probability of an outcome
Now the actual arithmetic. What comes out is not one option but a distribution: a percentage for each outcome, adding up to a hundred.
Sevilla — Getafe
LaLiga
Win probability
1 · Sevilla
46%
X
29%
2 · Getafe
25%
Read
Goals at both ends likely
Both attacks are in form, and their meetings usually run high.
Confidence
The home side have scored two or more in eight home games running.
Risk: rotation ahead of a midweek European tie.
Example of an analysis card. Figures are illustrative.
Take this card. Manchester City against Liverpool: 44% for the home win, 28% for the draw, 28% for the away win. That does not read as “City will win”. It reads as something duller and far more useful: in a match like this, on data like this, the home side wins in roughly four cases out of ten. And in six, it does not.
Two conclusions follow immediately, and both break a habit.
More likely does not mean “will happen”. 44% is less than half. The most likely outcome in this particular match still happens less often than it fails to. Anyone saying “the model predicted a City win” simply stopped reading before the number.
Flat percentages are more honest than pretty ones. When a read comes out 44/28/28, it is telling you the main thing: this match is close, the edge is small. A service that draws 80% on the same fixture and promises a “sure outcome” is not more accurate. It is selling certainty, because certainty sells better than a distribution.
Step 4. Where the confidence percentage comes from
Next to the probabilities the card carries a second number — the ring at 74%. It gets confused with the first one constantly, and it is a completely different scale.
Probability answers the question “how likely is this outcome”. Confidence answers “how much does the model trust its own calculation”. You can be very confident that a match is unpredictable. And you can produce an impressive-looking probability on data that is worth nothing.
The easiest way to feel the difference is a weather forecast. “40% chance of rain tomorrow” and “we honestly do not know what tomorrow looks like, the station has been offline for three days” are two different messages, and the second one matters more.
The main risk
That is why every read carries, on its own line, the thing most likely to break it. In the example above it is squad rotation before a midweek European fixture. Not an excuse invented afterwards, but part of the read: the service names in advance the circumstance under which it will be wrong.
That line is the easiest way to test a service. If it never names its own main risk, it either does not calculate one, or calculates it and keeps it out of sight.
What happens when there is not enough data
The most interesting scenario, and the one where the difference between an instrument and a salesman becomes visible.
From the analysis card
Research quality
★★★ high
high confidence
Plenty of data, sources and model agree, the read is confident.
Research quality
★ low
Murky match: it gives no confident read and says so plainly.
On the left, a match with a solid data base. On the right, the case where there is only one honest answer. Illustrative.
When a match has little data, or the data contradicts itself, the read gets a low research quality rating and a plain recommendation to skip it. There is no confidence percentage there, because there is no confidence either.
For a service that lives on an analytics subscription, that is a normal working outcome. For a channel obliged to produce “today’s certain call” every day, it is a commercial disaster: it cannot afford a day without an answer. So the answer will always be there, regardless of whether anything sits underneath it.
Same situation, two opposite behaviours. That fork is the one worth choosing an instrument on.
What the AI does not know in this pipeline
The list is short and worth keeping in mind:
- Everything that happened after the analysis. A read is a snapshot taken at a moment. An injury in the warm-up, or a line-up announced later, will not be in it.
- Decisions nobody announced. League position for both teams is in the data, and the model can factor in “this side has nothing left to play for”. Whether the manager will rest his first choice for a cup tie, when nobody has said so out loud, is known to nobody.
- Anything that never becomes data. Conflicts inside a squad, arrangements made quietly, everything that reaches neither the statistics nor the news.
- Rare situations. A debutant, a rescheduled fixture, a competition without proper statistics. Here the system does not try to wriggle out: research quality drops and the read is honest — skip it.
None of this makes an analysis useless. It simply explains why every honest read has a ceiling, and why a promise of results is always a lie, whoever makes it. Where those limits actually run, and which of them are myths, is mapped out in what AI does not know about a match.
How this differs from “the AI picked this outcome”
“The AI picked it” sounds more convincing and means less. Behind it you cannot see the data, the distribution, the confidence or the risk. It cannot be checked. If it lands: “we told you”. If it does not: “sport is unpredictable”.
A read that has all four steps can be checked. You can see what it was built from, where the service admits its own weak spot, and how much it trusts its answer at all. That is not a guarantee — guarantees do not exist in sport. It is, however, a clear view of what you are paying for.
Five checks that separate one from the other, applicable to any service in about a minute, are in AI sports analysis vs prediction channels. The whole pipeline is laid out on how Sharkline works.
How this looks in Sharkline
Alaves
Getafe
probability of each outcome
Today 32
Alaves
Getafe
Win probability
Read
Alaves not to lose
Alaves are steady at home, Getafe are weak away.
Alaves are unbeaten in five home games running.
Risk: Getafe have their first-choice striker back.
The daily feed. Matches and figures are examples.
Sharkline analyses hundreds of matches a day across nine sports, and does it before you open the app. A full read on one match takes about a minute. The card shows the probability of each outcome, the read, a confidence percentage, the arguments on both sides and the main risk. If a match is not in the feed, you can photograph it or send a screenshot.
What we do not do: promise results, publish “certain calls”, or produce a read where the data is not there. That is not modesty — it is the only arrangement in which a confidence percentage means anything at all. More about the product on About Sharkline.
Frequently asked questions
Can AI accurately predict football match results? No. Not an AI, not a person, not anyone. An AI calculates the probability of outcomes from data, and even a well-calculated 70% means that three times out of ten things go the other way. Anyone promising an exact result is selling certainty, not analysis.
What is the difference between the probability of an outcome and the confidence percentage? Probability says how likely the outcome itself is. Confidence says how much the model trusts its own calculation. You can get a high probability out of weak data — the first number will be large, the second small, and the second one matters more.
Why does an analysis sometimes advise skipping a match? Because there is too little data on it, or the data contradicts itself. In that situation any figure would be invented. An honest refusal is more useful than a plausible answer built on nothing.
What data is the analysis based on? Team and player statistics, current form, league position, head-to-head history, line-ups and injuries — all of that comes from a sports data API. On top of it sits a consensus from more than twenty sources, plus a separate check of fresh news right before the analysis: confirmed line-ups, weather, referee. Every read carries a rating of that data’s quality, so you can see how far to trust it.
How long does one match analysis take? About a minute. Matches in the daily feed are analysed in advance, so they open ready.
The matches are read. The call is yours.