The Reep Register

Case study

First official data partner

SkillCorner asked for an outside view of its identity metadata, then acted on it.

In July, SkillCorner did something most data companies leave in-house. It opened its player and team registry to an independent check. Not the tracking data or physical outputs it sells, but the identity metadata underneath: the names, dates of birth and identifiers that decide which player a row of tracking belongs to. Every provider resolves that layer on its own, and none has an outside view of where it drifts.

Four weeks later we pulled their registry again and measured what had changed.

  • 102 duplicate name-and-birthday groups, down from 853
  • 811 dates of birth corrected
  • No new duplicates created

What Reep brought was cross-provider triangulation: because their register resolves the same entities from several independent sources, they could point us at specific records rather than general suspicions.

Mohamed Rahmouni, Data Product Manager, SkillCorner

Why an outside check is hard to come by

Companies that produce football data know a second opinion on the quality of their metadata is worth having. Getting one properly, and regularly, is a different problem from their core business. It means entity resolution at scale and access to multiple independent sources, and both of these come with real opportunity costs.

This is where the Reep Register comes in. It carries 1.8 million football entities (people, teams, competitions, matches), each resolved from several independent sources and mapped to the provider IDs the industry already uses. It exists to take down the silos between providers: one stable ID per player, team, competition, season, match, coach and referee, with every provider's own identifier mapped back to it. Because the register can contradict any single source with evidence from the others, a disagreement becomes a specific, checkable claim against one record rather than an opinion. Running that across SkillCorner's whole identity layer was an afternoon's work.

What the first report said

Three kinds of finding:

  • Suspected duplicates. The same person or team holding two or more IDs
  • Contested birthdays. Dates of birth that two or more genuinely independent collections disagreed with (flipped day and month, off-by-one, wrong year)
  • Retired identifiers. De-duplicated or removed IDs that a customer might still be using internally

What SkillCorner did about it

SkillCorner’s data team worked through the report, and the results showed. A second pass against the same registry a month later found duplicate groups down from 853 to 102. 811 birthdays had changed, and the corrections landed in exactly the three patterns the first report named: 100 first-of-January placeholders replaced with real dates, 20 year-level corrections and 18 day-month transpositions.

One finding was harder to prove than the rest. Counting duplicates is easy but knowing whether a fix actually worked is not, if you are constantly cleaning up after it. Here the faulty import left a signature on every record it touched, so we could watch for that instead. Looking at new duplicates created since our audit, none carried the tell-tale signature. Being able to confirm an issue has been addressed – with evidence – is comfort for any busy data team.

duplicate name-and-birthday groups: July 853, of which 91 with the import's signature; August 102, of which 9 with the import's signature; Created since 42, of which 0 with the import's signature. July first pull 853 91 with signature August four weeks later 102 9 with signature Created since after the fix 42 none with signature duplicate name-and-birthday groups with the import's signature
The cause is closed, not just cleaned up after. The faulty import left a signature on every record it touched. Its share of the duplicate groups fell from 91 to 9 between the two pulls, and none of the 42 groups created since carry it.

In SkillCorner's words

We gave Reep access to our metadata deliberately. Identifiers, names, dates of birth — this isn't what we sell, and we think the industry as a whole has far more to gain from that layer being open than from each provider guarding its own version of it.

The reality of this work is that entity resolution never closes. New matches, new players and new competitions arrive continuously, so duplicates and errors will appear no matter how good your process is. We have an operations team dedicated to cleaning our registry for exactly that reason. What Reep brought was cross-provider triangulation: because their register resolves the same entities from several independent sources, they could point us at specific records rather than general suspicions.

That let our team prioritise a set of fixes immediately instead of working towards them over a longer horizon, and it independently confirmed a few issues we were already aware of, which reinforced the case for addressing them in the short to medium term. The less time this industry spends reconciling who is who, the more of it goes into the analysis itself.

Mohamed Rahmouni, Data Product Manager, SkillCorner

Why any of this matters

Identity is the layer everything else in football data sits on. A club running tracking data next to event data next to a valuation dataset is joining three systems on a single assumption: that "the same player" means the same player in all three. It should ‘just work’ whether they’re in the Women’s Super League or in the Ecuadorian Segunda Categoría. When an identifier is duplicated, or points at a twin or a namesake in an internal mapping, we’ve all seen what can happen. The numbers still add up. They are just about the wrong person or team, and discovering that is yet another papercut for a club’s internal teams.

SkillCorner is the Reep Register's First Official Data Partner, and the first data company we’ve seen to treat an outside check as something to invite. Their customers' trust is ultimately the thing they are protecting, and we’re glad to be part of what that looks like in practice.

Nothing proprietary from SkillCorner's systems enters the public register, we treat their data the same as any other provider’s, and we name our partners publicly so readers can judge that independence for themselves.

This collaboration is ultimately about our shared goal of fomenting an open analytics community, without unnecessary barriers between different providers' datasets. We trust that our next report for SkillCorner, already in progress, will be as valuable as the first - for them and for the industry.