← The Fine Print

The Fine Print

Footnotes

Every place a filing departed from the harmonised template this cycle, in the form or structure of its data. We classify these divergences as follows: (1) departure from regulation, (2) permitted by regulation, (3) not mandated in regulation, (4) other changes. Each entry records what the filing showed and the step we took to put it on equal terms. It concerns form, not substance, and assesses no provider's figures or practices.

1. Departure from regulation

The regulation stated a requirement and the filing did not meet it.

Non-UTF-8 file encoding

The regulation binds filings to UTF-8. AliExpress filed file 11 in ISO-8859-1, and Temu filed files 3, 8, 9 and 11 in it, inconsistently within an otherwise UTF-8 bundle.

AliExpressTemu

We attempt UTF-8 first, then fall back to Latin-1 and Windows-1252, and record the encoding each file used. UTF-8 is one of the regulation's two binding format requirements.

Codes outside the closed taxonomy

Annex II defines a closed code list, with every moderation action classified against the codes declared in the categories file. Loading the filings grew that list from 86 declared codes to 135. Meta, across its Facebook and Instagram filings, added the most, roughly 41 finer codes, and LinkedIn five, while TikTok and X added none.

declaredKEYWORD_OTHER
addedKEYWORD_OTHER_FRAUD_AND_DECEPTION
FacebookInstagramLinkedIn

We load every code as filed and record the codes that fall outside the declared list. The taxonomy is closed by the regulation, so a code beyond it is an addition to it.

2. Permitted by regulation

The regulation explicitly allowed the option the provider took.

Omitted optional columns

Annex II permits a provider to omit columns for which it has no data. Snapchat omitted four columns in file 4 and the category-of-incompatibility column in file 6. The file 6 omission removes the dimension on which its terms-and-conditions actions could otherwise be broken down by category.

Snapchat

We store the omitted columns as null and record the omission. The template allows it, so the missing dimension is permitted rather than withheld.

3. Not mandated in regulation

The regulation did not require or address the point, and the choices varied across providers.

Non-conforming scope ranges

File 3 expects one value per row, either TOTAL or a single member-state code. To record no orders across the member states, three providers wrote a bracketed range between two anchor codes, each rendering the placeholder differently. The rows carried zero across every data column.

AliExpress, 86 rowsAT, [_], SE
Shein, 85 rowsAT, […], SE
Temu, 85 rowsAT, [¡­], SE
AliExpressSheinTemu

We store these under a synthetic _NON_CONFORMING_RANGE marker and keep the original text, so downstream views can include or exclude them. The template's single-value structure would express the same as a zero in each member state's own row.

Duplicate rows under the template's row identity

File 8 is unique by indicator, language and scope. Four providers filed it with empty duplicate rows under that identity: Pinterest 5, Zalando 6, Amazon Store 6, Temu 156. Most likely a template-export artefact.

PinterestZalandoAmazon StoreTemu

We keep the first occurrence of each row and drop and count the rest. The duplicates carried no data, so nothing reported is lost.

Restatement signalled three ways

The regulation does not say how to flag a restatement. Three providers restated during the cycle and each signalled it differently. Amazon Store republished file 6 with a contextual note, its total moving from 150,504,153 to 402,704,488. Shein annotated its file 1 date field. Snapchat republished the whole report, recording the original date in file 1.

Shein, file 1 date2026/2/28[Updated 2026/04/30]
Amazon StoreSheinSnapchat

We show the latest version filed. Amazon Store's and Snapchat's originals are held alongside their restatements, but Shein's was not separately retrievable, so for that one only the restated figures exist. The regulation fixed the figures to report but not how to mark a correction to them.

Detection accuracy expressed several ways

File 8 records the accuracy, precision and recall of automated detection. The regulation does not fix the form these figures take, and the filings varied. Most providers gave them as decimal fractions, while Booking, AliExpress, Temu and Snapchat gave them as percentages. Snapchat went further and left its aggregate figures as a placeholder of 1, reporting the real values only per country, where accuracy ran from 0.69 to 0.92.

as a decimal0.84
as a percentage95.31%
Snapchat aggregate1
BookingAliExpressTemuSnapchat

We read percentages and decimals onto the same scale of 0 to 1 so the figures compare, and we suppress Snapchat's placeholder aggregate, showing its per-country figures instead. The regulation set out these measures but not the form they take.

AMAR reported over a shifted window

The cycle runs 1 July to 31 December 2025. Shein reported its AMAR over 1 August 2025 to 31 January 2026 while its other metrics used the standard window, disclosed in a free-text annotation in the reporting-period field.

Shein, file 1 period2025/12/31 (except for AMR figures which is reported from 2025/8/1 to 2026/1/31)
Shein

We load the figure as filed and attach a caveat wherever a per-AMAR comparison involves Shein. AMAR is the denominator for per-user comparison, so a shifted window changes how its rates compare.

Audience as a band rather than a figure

Some providers did not report audience as a single exact number. Booking reported the EU total as a threshold, more than 45 million, confirming only that it clears the very-large-platform line. Google gave no figure in the audience file and pointed to a separate report, whose per-country figures are rounded bands, and which reports two separate bases, signed-in accounts and signed-out sessions, that count different things and have no single total.

Booking.comGoogle MapsGoogle PlayGoogle ShoppingYouTube

We store banded and threshold values exactly, flag them as non-exact, and suppress per-user rates wherever the audience figure is banded or absent, so a provider that disclosed only a band does not sit on a rate comparison as though it had given a number. Google's two bases are kept separate and no single total is invented. The Commission's guidance on publishing user numbers addresses how active recipients are counted rather than the form the figure takes, and the filings show it still reported in non-comparable ways this cycle.

Two rows under one code

In file 4, X filed two rows carrying the same category code under the same parent category, distinguished only by their free-text description. None of the other providers did this.

filed twice, one parentKEYWORD_OTHER
X

We widened the row identity to include the description text, so neither row is dropped as a duplicate. The template's row structure did not anticipate two distinct entries under one code.

Spreadsheet-export merged cells

Instagram's filing was exported from a spreadsheet with merged category cells, leaving the category code present only on the first row of each group and blank on the continuation rows.

Instagram

We forward-fill the code down each group so every row carries its category. The merge is an export artefact, not a value the template defines.

Inconsistent date formats

The regulation does not specify a date format. Providers used at least three conventions in file 1, and Shein embedded annotations within the field.

ISO2026-02-27
ISO with time2026-02-27 00:00:00
European27/02/2026
TikTokAmazon StoreLinkedInSnapchatShein

We parse against common formats and strip annotations, preserving the original value. The template fixed the fields but left their notation free.

Different number notation

The same magnitude appears written three ways across filings, with thousands marked by a comma, by a dot, or replaced with a magnitude suffix.

Instagram1,234,567
Zalando1.234.567
Temu1.2M
InstagramZalandoTemu

We parse all three to a plain integer, distinguishing a decimal point from a thousands dot. The regulation fixed the figures but not how to write them.

Divergent country codes

Greece appears under EL from Instagram and LinkedIn and under GR from TikTok, the EU institutional code against the ISO code. A comparable split affects the Czech and Irish codes.

Instagram, LinkedInEL
TikTokGR
InstagramLinkedInTikTok

We map each to one canonical code, so a country's figures reconcile across providers. The regulation names the indicators but not which code set to use.

How absence is expressed

For own-initiative action against illegal content, the providers that addressed it expressed absence differently. TikTok filed the form blank, LinkedIn filed explicit zeros, X filed nulls with no usable total, and Instagram filed figures. For Article 23 misuse suspensions, three filed zero and TikTok filed null.

TikTokLinkedInXInstagram

We preserve the distinction between blank, zero and null, and treat a blank as not reported rather than as a zero. A reported zero is a measurement, and a blank is its absence.

Underspecified indicators

The template named an indicator but not how to compute it. Out-of-court dispute outcomes are reported as a percentage, but a value like 0.403 could mean 40.3 percent or 0.403 percent, and the data cannot resolve which. The Article 9 and Article 10 order-compliance medians were filed without a defined moment for the clock to start.

out-of-court, share implemented0.403
TikTokInstagramLinkedInX

We report both as filed and flag that they are not comparable across providers. The indicator is named by the regulation, but its unit and its clock are not.

Parts that do not reconcile to the whole

Some totals do not equal the sum of their parts as filed. The four restriction-type columns sum to the reported total exactly for most providers, past it for AliExpress and TikTok, and short of it for Pornhub and Snapchat. The per-language moderator figures overlap, since a multilingual moderator is counted under each language, and TikTok's declared moderator total exceeds its language sum by 472.

TikTok, restriction types vs totaltype sum 261,778,094 / reported total 258,691,522
AliExpressTikTokPornhubSnapchat

We report the figures as filed and do not force them to reconcile. This is an observed property of the data as published, not a step we took, and the parts relate to the whole differently across providers.

4. Other changes

Changes to the form of the submission itself, in how it was packaged, named, or when it arrived, rather than to the data inside it.

Advertising filed as separate services

Google filed advertising as a distinct service for four of its products, each shipping a separate ads bundle alongside the main filing.

Google MapsGoogle PlayGoogle ShoppingYouTube

We load each ads bundle as its own service filing and box its figures off, never blended into a platform's headline totals. The filing arrived as more than one submission per provider.

A repeated file in the bundle

X's bundle contained file 7, appeals and recidivism, twice.

X

We load a single copy. The repetition is a packaging artefact, not additional data.

Per-tool rather than aggregate automation reporting

For the accuracy of automated detection, most providers reported a single aggregate figure. XVideos and XNXX instead reported the accuracy of each third-party classifier they run, one row per detector.

XVideosXNXX

We widened the row identity so these distinct per-detector rows are kept rather than collapsed as duplicates. The same indicator was filed at a finer grain than the aggregate the others reported.

Single-spreadsheet bundles

Where most providers filed eleven separate CSVs, Shein, Apple App Store and Wikipedia each filed a single spreadsheet with the files as tabs.

SheinApple App StoreWikipedia

We split each into the eleven files before loading. This is a packaging choice the regulation does not constrain.

The harmonised template narrowed the ways these reports could differ, but it did not remove them. The entries above are every structural divergence we found this cycle and the step each one required. They are why the figures elsewhere on the site carry the caveats they do, and why comparing two providers' raw numbers means first knowing how each was filed.

This page concerns how faithfully and how differently each provider followed the template. It does not assess the accuracy of any figure or the substance of any provider's moderation. The loaders and numbered migrations behind every transform are in thedata repository.