Appendix: Noise-Floor Method and Source Taxonomy

Kääselä Appendix 2026-08 Permanent, shared across all issues Summary edition · approx. 2 min read

This page is a summary edition; since September 2026, non-Japanese editions are published as summaries. The full appendix is published in Japanese: full text (日本語). All figures on this page are identical to the Japanese original.

Permanent appendix referenced by every Kääselä Monthly Report. Individual issues link here instead of restating the methodology.

1. Why measure the noise floor

Generative AI does not necessarily return the same answer to the same question. Whether an apparent month-over-month change in citations is a real change or mere generation variance can only be judged against the size of that variance — so it is measured first, published, and used as the baseline for reading monthly differences. Unlike commercial tools that assert "the AI says X" from a single run (n=1), this baseline is measured and disclosed.

2. Measurement method

Each query was run 20 times (n=20) consecutively under identical conditions (ChatGPT, GPT-5.5, web version, temporary chat, keyword-form queries, run in Kyoto City), logging the set of cited source domains per run. Overlap between two runs is expressed as the Jaccard coefficient (intersection ÷ union): close to 1.0 means the same sources every run, close to 0 means near-total turnover.

3. Noise floor across the 12 cells (Jaccard coefficient)

JapaneseEnglishSimplified Ch.Traditional Ch.
Tea ceremony0.250.200.310.25
Kimono0.410.160.150.11
Zen meditation0.350.400.220.23

The noise floor differs by category and language; no single threshold fits all cells.

4. Pre-registered thresholds

Thresholds for the main metric (share of citations by source type) are set per cell. From issue #000 data (with business info cards excluded), a run-level bootstrap (B=4000) gives each cell's 95% half-width; for comparisons between two periods, a shift of at least "half-width × √2, rounded up" counts as signal. The initial blanket registration (±10pt, 14pt across periods) had been computed on a basis that counted business info cards as citations; it was voided with the card separation (2nd-edition correction) and re-derived per cell — using issue #000 data only, before any issue #001 data was aggregated. Thresholds are not adjusted after seeing results.

Two-period threshold (pt)JapaneseEnglishSimplified Ch.Traditional Ch.
Tea ceremony1591015
Kimono17162322
Zen meditation14131814

Appearance rates of individual businesses vary more than type shares (±approx. 22pt); monthly evaluation covers only the concentration of the top group.

5. The six citation-source categories

TypeDefinition
PublicMunicipalities, tourism associations, public bodies, public cultural facilities
OwnedA property run by the recommended business, temple, or individual guide themselves (whether or not it takes bookings)
OTAThird-party platform aggregating others' inventory, where booking completes on its own site
MediaEditorial content that cannot take bookings itself and routes visitors elsewhere (including affiliates)
DirectoryCross-sectional, commission-free listings
UGCUser-generated content

6. Limitations of the method

Observation through the standard UI captures only what the answer cites, not the candidate pool the AI searched — so this study reports "appearance rate," never "selection rate." Reproducibility means a statistically matching distribution under the same time window, engine version, and de-personalized conditions, not verbatim identical answers. The full discussion is in the Japanese full text, §6.

This page is part of the Kääselä (Kamogawa AI Search Lab) personal research project. Data is published free of charge under CC BY 4.0 and is not for sale.

Top