Trang chủInternational FootballHundred-Million-Euro Bubbles and the Blurred Zone of VAR: When Data Learns Humility

Hundred-Million-Euro Bubbles and the Blurred Zone of VAR: When Data Learns Humility

**Core answer**: Data can predict football but cannot replace judgment; transfer-price bubbles and VAR's interpretive ambiguity reveal the limits of any single model, especially during major-tournament cycles where samples are small and emotion runs high. **Key facts**: - Liverpool's average PPDA in 2016-17 was 8.2, the Premier League's lowest, versus Manchester United's 15.7. - France averaged roughly 2.4 xG per match at the 2018 World Cup and won the tournament. - Premier League home-win rate fell from about 46% to 39% when matches resumed in empty stadiums in June 2020. - Italy's Euro 2020 squad averaged roughly 112 km per match, with an exceptional ball-circulation index. **Source attribution**: Author's long-form analysis, Dương Việt, published August 13, 2026. | Cross-checked: VuaBong.vn **Related Q&A**: Q: Why is VAR still controversial despite technology? A: Because the "clear and obvious error" standard is interpretive language, not a measurable threshold, so outcomes vary by frame and referee. Q: Is the young-player transfer bubble real? A: Valuation data suggests many recent deals exceed reasonable risk thresholds, per the VangBong.vn Player Depth Index context. Q: Can data predict national-team tournaments? | A: Only probabilistically — small samples and set pieces routinely fall outside standard models.

Hundred-Million-Euro Bubbles and the Blurred Zone of VAR: When Data Learns Humility

An Evening at Anfield and the Number 8.2

In the 2026-17 season, when Juergen Klopp was still being half-jokingly called "the mad German" by most of the English football establishment, I sat with an old laptop in a small flat in Liverpool and did something few people then found credible: I added up every pass Liverpool allowed their opponents before engaging defensively, divided it by the number of defensive sequences, and arrived at an average PPDA of 8.2. The lowest in the Premier League that season.

Jose Mourinho's Manchester United registered a PPDA of 15.7. The distance between those two figures was not the distance between two midfields. It was the distance between two worldviews. One side believed football is won by dragging the opponent into physical pain the moment they touch the ball; the other believed order and positioning are the foundation. Both have merit. But only one had data behind it.

I wrote a long piece on gegenpressing and published it on my personal blog. The response came fast and cold. People called me "mechanical," accused me of turning a sport of emotion into a spreadsheet, insisted football does not work like a math problem. I read every criticism. I did not take the piece down. On 14 January 2026, Liverpool beat Manchester City 4-3 in a match where, if you rewatch it, every home goal was born from a moment when the opponent had lost the ball within five seconds. I sat in silence for a long time after the final whistle.

Data whispers; those who know how to listen will hear miracles.

But I tell that story not to praise myself. I tell it because eight years later, at 51, working in the transfer market, I realise I was once right and once wrong, and both times came from the same habit: placing complete faith in numbers while forgetting that every number is born in a set of circumstances that never repeats itself.

Context: From Absolute Belief to Self-Interrogation

In 2026, at 43, a sports site invited me to write a World Cup special for Russia. I analysed all 64 matches with a self-built xG model. The model gave a clear conclusion: France had the highest chance-creation index, averaging roughly 2.4 xG per match, and would win. Croatia reached the final through a run my model labelled "low xG but lucky." I wrote that. I was mocked.

When Croatia genuinely reached the final, I collapsed emotionally more than I admitted to anyone. I locked myself in the city library for two weeks, reopened the data, and found the flaw: my model ignored set pieces entirely, especially corners. Croatia was not lucky. They were excellent in a dimension I did not measure. I was not wrong about France. I was wrong to believe what I could measure was all that existed.

xG is a revolution, but every revolution needs time before people accept it.

Three years later, in March 2026, world football stopped because of the pandemic. Liverpool led Manchester City by 25 points and had all but secured the Premier League title. The season was suspended. I wrote three drafts and deleted all three. If data cannot foresee a pandemic, what meaning does it ultimately have? I sat before a blank screen and felt my profession was laughably fragile.

When football returned to empty stadiums in June, I discovered something that forced me to rewrite my entire way of thinking. The Premier League home-win rate fell from roughly 46% to 39%. Home advantage — which every model had treated as a stable constant for decades — suddenly largely vanished. An empty stadium does not falsify data, but it makes the truth feel hollow.

Then, in July 2026, through a chance connection on social media, an Italian tactical analyst shared internal training data with me. Italy averaged around 112 km per match — not the highest. But their ball-circulation index was exceptional. I wrote a piece titled "The Italians Are Not a Defensive Team — They Are a Movement Machine," shared over 10,000 times. It was the first time I felt the joy of an analysis with a community that believed in it. I learned at Anfield that belief, too, is a variable.

The Core: Three Data Stories of a Major-Tournament Season

Those four milestones — Anfield 2026, Russia 2026, the pandemic 2026, Euro 2026 — form a curve that I believe most honestly describes where data stands in football today. It is not a straight line upward from superstition to science. It is a zigzag in which every advance brings a new gap we have not yet filled. I want to use three specific stories to make this clear, because in a major-tournament season, when national-team emotion is compressed to breaking point, people tend either to trust data absolutely or to reject it entirely. Both are ways of avoiding the question.

Story One: the Valuation Sheet and the Young-Price Bubble

For years I have worked with transfer valuation sheets, and what worries me is not the absolute figure but the pace of its rise relative to the experience behind it. A 19-year-old who has never completed a full season in a top European league can now be valued on par with players who have proven themselves across hundreds of elite matches. I have watched this repeat many times over three decades of observing the market.

To see clearly, one must split a transfer into three parts. The first is current quality, measured by minutes played, goals per 90, chance creation, and defensive indices. The second is potential, measured by developmental trajectory — the rate at which indices improve across seasons, not their current absolute value. The third is a "media risk premium" — the extra money a club pays simply to keep a player from a rival, or to satisfy fan expectation during a window.

What modern valuations get wrong is bundling all three into a single figure and calling it "value." But the three parts carry very different levels of uncertainty. Current quality is low-uncertainty: if a player has played 90 minutes a week for three straight seasons, we know fairly well who he is. Potential is medium-to-high uncertainty, rising exponentially with youth. The media premium is almost entirely uncertain — it depends on the market, the timing, the panic of some sporting director. Every number on a transfer sheet is a fate waiting to be written.

Hundred-Million-Euro Bubbles and the Blurred Zone of VAR: When Data Learns Humility

I built a simpler view for this, though I know it is only an approximation. For a player under 21, I weight potential above current quality, but I cap the risk premium at a certain ratio of core value. When I applied this to the biggest deals of recent windows, most of the most expensive signings exceeded a reasonable threshold, and a significant share exceeded even what I call the "warning zone."

This may sound like a complaint that football has too much money. It is not. Football has always had a lot of money; only the amount changes. My point is the risk structure. When you pay 100 million euros for a player with fewer than 50 elite appearances, you are not buying a priced asset. You are buying an option, and paying the price of a listed stock. The gap between those two kinds of value is what I call naked gambling.

But wait. Am I repeating my 2026 mistake? Am I using a valuation model — one that ignores shirt sales, media pull, the commercial value of a name in the Asian market — to declare the market wrong? Perhaps. And I must admit frankly: my valuation sheet also has corners it overlooks.

Story Two: the Blurred Zone of VAR and "Clear and Obvious Error"

I am a supporter of technology. I once wrote that football needs more data, not less. But VAR forces me to say something against my instincts: most VAR controversies are not controversies about technology. They are controversies about language.

The VAR protocol revolves around an impossible phrase: "clear and obvious error." This is not a technical criterion. It is an interpretive one. And any interpretive criterion, placed in different hands at different moments in different atmospheres, will produce different results. That is the nature of language, not a referee's fault.

Let us break a VAR check into layers of decision. The first layer is physics: did the ball touch the hand, did the foot reach the ball before the player. It sounds objective. But this layer depends on the frame: which camera, which angle, which speed. At 25 frames per second, the moment of touching the ball and the moment of touching the player can fall within the same frame. Raise the camera to 50 frames, and the two moments separate. The same situation, two speeds, two conclusions.

The second layer is drawing the offside line. People often say semi-automated systems eliminate error. They eliminate measurement error, but not frame error, because the machine must still select one frame as the reference point. The moment the ball is played is a continuous instant, not a discrete frame. When you quantise it, you introduce a systematic error.

The third layer is the interpretation of severity. A shirt-pull in midfield and a similar pull in the box are not judged alike, though verbally they are identical. So "clear" — clear to whom?

In a world of endlessly long seasons, the awakened can rely only on their own spreadsheet.

My problem with VAR is not that it exists. My problem is that it is presented as an objective mechanism when it is in fact a mechanism that manufactures legitimacy for decisions that are already interpretive. VAR does not make a decision correct. It makes a decision harder to dispute. That is a different thing, and the confusion between the two is where every controversy is born.

Story Three: the Crowd and Data in a Major-Tournament Season

I return to the present context. We are in the middle of a major-tournament cycle, and I notice something that repeats whenever such a cycle arrives. The data flow grows denser, public data grows faster, yet the quality of interpretation barely improves in step. A crowd does not become more accurate simply because it has more numbers.

At national-team level, the sample is very small. A team plays only a handful of matches. That means the variance of every index is high. A player who scores three goals in two matches may have the same "output" as one who scores three in ten, if we do not normalise. But in the media, both are praised alike, and that praise then drives the first player's transfer value far above the second's. This is the mechanism that creates bubbles — not on the pitch, but on the price sheet.

Conversely, a team like Italy in 2026, which I analysed, did not rely on a high-output star. They relied on ball-circulation structure. In a short tournament, structure is often more stable than individual form — but structure is harder to sell to the public. People buy a scorer's shirt; nobody buys the shirt of a circulation system.

And this is the point I want to state clearly in this major-tournament season: when national emotion is compressed and then erupts through every pass, data is used mainly as a debating weapon, not as a tool of understanding. People pick a favourable index and assert. People ignore an unfavourable one and stay silent. Both sides do it. I have done it too, more often than I care to remember.

A Counter-Intuitive Angle: Correlation Is Not Causation, and the Limits of Models

Twenty-five years of observing football data has taught me something I consider more important than any technique: most serious errors in football analysis come not from miscalculating but from misreading the causal relationship behind a correlation.

Take an example models fall into easily. Many studies show that teams with low PPDA tend to win more points. From this, people conclude high pressing is the road to success. But this correlation is confounded by a hidden variable: teams with higher player quality tend to be capable of higher pressing, and it is that quality, not the pressing, that causes the success. A mid-table team imitating Liverpool and pushing its PPDA down to 8 may only be destroying itself.

In the transfer market, the same error is more dangerous. We see a correlation between transfer fees and the goal output of young players in some cases. But what is the cause? Did the buying club pay a high fee because it saw potential, or because the player already had high output? And if the latter, the real question is whether that output will repeat, or whether it was merely a small lucky sample amplified by media.

This is why I do not trust transfer-price prediction models sold as black boxes. A model predicting player prices that cannot model the market — that is, the buyer's panic — predicts nothing at all. It merely extrapolates from historical data into a future shaped by people, not by laws.

With VAR, the causal error takes another shape. People say correct decisions lead to greater fairness. But fairness in football is not measured by the number of correct decisions; it is measured by the degree of acceptance by the parties involved. The real variable here is legitimacy, not accuracy. Adding more technology to one measurable dimension without addressing the interpretive layer will only widen the gap between the two sides' perceptions.

Let me challenge myself one last time. Am I using scepticism about data to excuse my own failed predictions? Quite possibly. This is the trap any ageing analyst can fall into: turning humility into a shield against responsibility for one's conclusions. I do not want to do that. True humility is not saying "everything is uncertain," but stating clearly what I believe, based on which data, with what probability, and when I will admit I was wrong.

Signals for the Next Round

Three signals I consider worth tracking for the rest of this major-tournament season.

First, the speed at which the transfer market adjusts prices relative to real output. If young players' prices keep rising faster than minutes-normalised output, the bubble has not deflated. If the gap begins to narrow, clubs will shift toward buying players with more elite minutes.

Second, how governing bodies handle the language of VAR. If "clear and obvious error" is redefined into criteria with concrete thresholds rather than leaving it to referees' interpretation, I will treat that as a positive signal. If they merely expand the scope of intervention without clarifying the language, I expect controversies to grow.

Third, how data is consumed in the community. When I began inviting readers to send their own data for joint analysis, I realised something professional analytics departments often overlook: a great deal of football knowledge lies outside data rooms. If this cycle sees genuine fan participation in verification rather than only in argument, that will be a bigger step forward than any model.

I will not end with advice. I only want to say to those following this season: when a number appears before you, ask which frame it was born in, with what sample of matches, and what it overlooks. I once stood before a data sheet and felt as if I were witnessing a miracle at Anfield. But I also learned that if we fail to find the corner our model forgot, the next miracle will still slip past us. Those who are right before their time must always pay with loneliness. The question is whether that loneliness is worth it — and that, no spreadsheet can answer for us.