AI News HubLIVE
サイト内リライト6 分で読了

翻訳待ち:The State of AI Disclosure 2026: what 1,088 EU sites' chat widgets say

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Research · Original data · Post-application snapshot The State of AI Disclosure 2026 1088 detector-flagged EU-facing sites scannedScanned Monday 10 August 2026Rule pack 2026.07.9 Eight days after Article 50 of the EU AI…

ソースHacker News AI著者: coldpress

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。

Research · Original data · Post-application snapshot The State of AI Disclosure 2026 1088 detector-flagged EU-facing sites scannedScanned Monday 10 August 2026Rule pack 2026.07.9 Eight days after Article 50 of the EU AI Act began to apply, we swept our locked cohort of 1,142 detector-flagged, EU-facing sites — 1,088 scanned, 54 excluded by robots.txt and counted, each rendered in a real browser, homepage only, read-only — on a staffed Monday, early afternoon in EU time. Corrected for the detector's measured error, roughly 11% of the 12,285-site frame shows a real chat launcher (the frame and its correction were measured 27–28 July; the surfaces below, on the scan date). Most of what happens at the first interaction is invisible to our automated visitor: on 78% of confirmed widgets it could not read the first message at all, and on the 174 surfaces it could read, a disclosure that the chat is automated was detected on 19 — about one in ten. This page is a single post-application snapshot: it publishes no before/after delta and no compliance verdicts, and every exclusion is counted below. 11% ≈11% of 12,285 EU-facing candidate sites show a real chat launcher once measured detection error is corrected (95% CI 9.0–14.3); the raw detector said 9.4%. Frame and correction measured 27–28 July 2026 — this figure reads no sweep data and is not dated to the scan day. How to read these numbers Every figure on this page is an aggregate. No individual site is named, and none will be: the point is the state of the ecosystem, not a wall of shame. We hold that a raw detector count without its precision and recall is not a finding &mdash; and by that standard exactly one rate here is published corrected for its detector's measured error, with the error rates themselves in the open: the prevalence headline. The other numbers below are raw, and are labelled as such where they appear. The widget detector's error was measured, but it is not corrected out of the coverage figure, whose denominator it contaminates in a direction we name in that section. The disclosure detector &mdash; the one deciding whether a first message discloses AI &mdash; has no measured precision or recall at all: no judging pass in this study ever read disclosure wording, so the three disclosure-state counts are raw detector output whose error rate is unknown rather than known to be small. Where we report disclosure, a finding of &ldquo;not detected&rdquo; means our scanner, loading the site the way this study's automated visitor does &mdash; which is not the way a human visitor does, and this study ran no human arm, so it cannot tell you how far apart the two are &mdash; could not observe a disclosure at the widget's first interaction: a statement about what was observable, never a legal conclusion about any organisation. We report three outcomes only: detected, not detected, and could not verify &mdash; and we publish the could-not-verify rate instead of hiding it. Why we never say &ldquo;compliant&rdquo; applies to this study exactly as it applies to a report. How many EU-facing sites run a chat widget? The corrected figure above comes from auditing our own detector, not from trusting it. Both admission rules were checked against what a blind judge could see in a real-browser screenshot of the rendered page &mdash; that capture is the ground truth here, and it is not the same thing as what a site actually runs, nor what a human visitor would find &mdash; and the raw detector count is adjusted by the measured rates below. The two errors run in opposite directions and do not cancel; the correction and its confidence interval carry both. One defect in that interval's history is disclosed here rather than left for a reader to find. The bootstrap behind the interval this study first drafted drew the same sampling variance twice, so the interval it produced &mdash; 7.95&ndash;15.42 &mdash; was labelled 95 % but sized closer to 99 %. The estimator was frozen under precommitment until the second sweep completed, and the one-line fix was applied after it, on 10 August 2026; the correctly sized 95 % interval is the one printed above, and the point estimate did not move. Note which way the fix cut, because it ran against this page rather than for it: the load-bearing claim here is that the corrected figure is not distinguishable from the raw detector's, and the narrower interval makes that claim harder to sustain, not easier. The raw figure remains inside the corrected interval &mdash; the claim survives the fix, by less room than the draft interval suggested. Both intervals are stated here so the correction is checkable rather than silent. Correction componentMeasured onRate (95% CI) Generic-launcher admits showing a real launcher49 decidable47% (34–61) Vendor-fingerprint admits showing a real launcher52 decidable73% (60–83) “No widget” sites showing a launcher anyway (recall gap)322 decidable6.8% (4.6–10.1) Inert-SDK sites showing a launcher anyway41 decidable12% (5–26) By widget vendor Vendor shares are of sweep-confirmed widgets. One caveat travels with this table: &ldquo;the vendor's SDK is present&rdquo; and &ldquo;a visitor can chat&rdquo; are different claims. A substantial share of the sites whose vendor SDK we fingerprinted showed no chat our scanner could reach. We cannot tell you why, and we will not guess: that gap is a single undecomposed quantity mixing our own detector's false positives, inert or disabled SDKs, staffed hours, login gates and several days of site churn between the two measurements. This study forbids itself from republishing it as any one of those causes &mdash; in particular as an after-hours or weekend penalty &mdash; and its size is not distinguishable from the detector precision published above. Two limits bear on this table specifically. First, a consequence of the frame, pre-registered before the data existed: candidates are taken from the top of a popularity ranking, which over-represents enterprise sites relative to the long tail, so SMB-leaning vendors are under-represented here against their real install base. The unknown / custom cell is precisely the one a rank-truncated frame inflates, because the vendors that concentrate below the rank cut never enter the study at all. Second, the SDK-versus-reachable caveat above is scoped to the fingerprinted rows, while the detector contamination sits almost entirely in the other class: the sites matching no vendor fingerprint are the ones our admission audit found least likely to be chat widgets in the first place. &ldquo;Unknown or custom&rdquo; is therefore a statement about what our fingerprints matched &mdash; it is not a measurement of how much of the ecosystem is custom-built, and correcting it for measured precision would move it down. Vendor cells below the minimum-cell threshold are grouped as &ldquo;other / n too small&rdquo; rather than published individually; the country section's note on how the pre-registered threshold is applied covers this table too. VendorSitesShare of confirmed unknown / custom57072% zendesk608% livechat334% other (13 vendors, each n < 30)13116% Can an automated visitor reach the first message at all? 78% On 78.1% (95% CI 75.1–80.8) of the 794 sites where the sweep confirmed a chat widget, our automated visitor could not read the first-interaction surface. Raw share on the detector-confirmed denominator, scanned Monday 10 August 2026 (graded 11:01–12:12 UTC). This number is usually buried as a denominator footnote; we measured it on purpose and publish it as a headline. For every confirmed widget the sweep tried, read-only, to dismiss the consent layer, click the launcher and read the first message, and recorded where that failed. Consent walls, launchers our scanner could not click, panels that never open, panels whose text could not be read: each failure cause is counted below, including the ones that are our instrument's limits rather than the site's. Three qualifications on that accounting, none of them buried. The consent step is attempted but its outcome is not recorded per site, so neither a reader nor an author can audit which unreadable rows really died on a cookie wall &mdash; only the ones that failed loudly enough to be bucketed appear as consent failures, on a scanner whose own source calls cookie walls the primary blocker to reading a widget's first message. The bucketing itself carried a defect in this study's first sweep: rows were assigned by matching error-message wording, so panels that opened but whose text could not be read under unexpected wording fell into &ldquo;other&rdquo;, and the drafted funnel failed simple subtraction. The aggregation code was frozen under precommitment until the second sweep completed and was corrected after it, on 10 August 2026; buckets now key on the recorded interaction outcome, and the table below is generated by the corrected code. And each failure cause below carries an attribution &mdash; instrument-side, site-side, or mixed &mdash; because one cause this study first drafted was labelled as a fact about the site when the code that emitted it reads it as a fact about us: rows where a vendor SDK was fingerprinted but our launcher selector no longer matched are a stale recipe of ours, not an absent widget. They are labelled that way below, and the interaction-stage instrument-side subtotal is printed as its own row. Read that subtotal as a floor on our own share of this headline, not the total: the blind audit below shows the largest failure bucket is dominated by a different error of ours &mdash; admissions that were never chat widgets &mdash; which interaction-stage tags cannot see, and which is why that bucket is tagged mixed rather than site-side. One known contamination is stated rather than hidden: our admission audit found that a share of &ldquo;confirmed&rdquo; widgets are not chat widgets at all, and those sites land in the &ldquo;panel did not open&rdquo; row &mdash; so the denominator here is detector-confirmed, not verified widget presence. That is not a small caveat on a large number. The interval printed with this headline is sampling error only, and sampling error is not the dominant uncertainty here: the denominator contamination is larger, runs in one direction, and is not inside that interval. Correcting for it would move this figure down. The correction pre-registered for exactly this purpose &mdash; a blind re-judging of the failure-bucket screenshots this study already holds &mdash; has now been run against this sweep's rows, and its result is published here rather than left as an open action item: the sweep’s own above-fold screenshots of all 361 “panel did not open” sites were re-judged blind (judges saw the image only — no site, vendor, or error metadata; 20 planted duplicates agreed 19/20 on the launcher question). 260 were decidable: 67.7% (95% CI 61.8–73.1) showed no chat launcher visible in the capture — accessibility widgets, carousel arrows, trust badges and contact links our detector had admitted as chat — and 32.3% (26.9–38.2) showed a real launcher whose panel nonetheless did not open. Two bounds travel with that rate: a launcher below the fold or rendered after the screenshot reads as absent, so the no-launcher share is biased upward by an unmeasured amount; and the 101 undecidable screenshots are dominated by consent layers covering the corners where launchers sit — an exclusion that is not neutral. Both are stated in the method log record beside the judgements. The headline above is still printed raw, on the detector-confirmed denominator, exactly as pre-registered &mdash; re-basing it on the audited denominator is an analysis change that goes through adversarial review before it is published, not a hand adjustment &mdash; but the size and direction of its largest known error term are now measured and on the page, not disclosed and unresolved. Whatever an autom [truncated for AI cost control]