
What was collected—and what was not
The 8 October 2026 local catalogue contained 15,436 conference records across 24 subject categories. It combined 15,430 directory-derived candidates with six retained official-link records from an AI pilot. The prior catalogue report describes coverage across 124 countries. This was an operational collection, not a random sample of all conferences. [1]
The screening exercise did not conduct a representative audit of search-engine results, run a chatbot recommendation benchmark or complete a full SCVS assessment on every record. Its question was narrower: which collected records met the local rules for a shortlist supported by an official event link?
All figures below refer to records, not unique organiser companies. A repeated event series can contribute many dated records. Nothing here establishes the number of fake sites worldwide or the proportion of harmful results in a search engine.
The measured screening breakdown
The rules assigned mutually exclusive exclusion reasons: 10,612 records had a repeated-title series pattern, 4,817 remained unresolved aggregator-only listings, and one matched an audit-listed organiser brand attributed by a directory. Six records were retained with six distinct canonical official links. [1]
The repeated-title flag used an existing threshold of at least ten listings with the same normalised title across dates or locations. That may identify a pattern worth reviewing, but a legitimate recurring event can also repeat its title. Numeric IDs in listing URLs were not, on their own, an exclusion reason.
The original equals excluded plus retained: 15,436 = 15,430 + 6. The three exclusion reasons sum to 15,430. Category totals reconcile independently in the published aggregate data.
Measured snapshot / 8 October 2026
How the 15,436 records were classified
Common scale: 0–15,436 records. The one-record and six-record bars are nearly invisible at this scale; the labels preserve their exact counts.
| Exclusive reason | Records | Share of sample |
|---|---|---|
| Repeated-title pattern | 10612 | 68.748% |
| Unresolved aggregator listing | 4817 | 31.206% |
| Audit-listed brand attribution | 1 | 0.006% |
| Retained official-link record | 6 | 0.039% |
Sample composition / six largest categories
Coverage was uneven across fields
Why the six links are not a global quality ranking
All six retained records were in AI. The other 23 categories had zero retained records under this filter. That says something about the collection and its evidence rules; it does not mean those fields have no trustworthy conferences.
The six official links came from earlier field checks. No new comprehensive legitimacy audit was run for this screening. The retained set should therefore be described as a small official-source shortlist, not as the only legitimate conferences or a certified league table.
The dataset is strongly shaped by directory collection and by uneven official-source enrichment. Broader collection from scholarly societies, university event calendars and discipline-specific organisers could change the retained count substantially. That next collection step is necessary before generalising beyond this snapshot.
What the case suggests for search and AI
Our interpretation is that discovery and verification need separate stages. A listing can be useful for finding a possible event while still lacking evidence sufficient for a payment or submission decision. A polished website, repeated title or fluent recommendation does not resolve those missing facts.
Independent research has documented fabricated citations in generated literature reviews. That is a different task and dataset from this conference pilot. It supports checking generated references, but cannot be used to claim that these 15,430 excluded records were recommended by current AI services. [2]
A defensible next study would specify the search queries, countries, dates, result positions and recommendation prompts in advance, then have reviewers check the same evidence standard for each result. Report agreement, unknowns and false positives. Until such a benchmark exists, claims about search or chatbot prevalence remain unmeasured here.
How ScholarVault applies the lesson
ScholarVault’s approach is to keep official sources, evidence coverage and assessment limits visible in the research decision. SCVS is an automated risk assessment, not a guarantee; a candidate list should remain distinguishable from a human-reviewed conclusion. [3]
For a university library or research office, this case offers a workshop exercise: take a small shortlist, record the source of each claim and mark what remains unknown. Do not publicly label organisers as fraudulent from a screening flag. Preserve the initial record and the reason for any later decision.
Reproduce the figures: download the aggregate JSON or category CSV. These contain no organiser names or candidate URLs. The snapshot, thresholds and limitations are included with the data.
Sources & reading notes
Sources reviewed for this draft on 9 October 2026. Institutional examples illustrate a source trail; they do not imply a ScholarVault partnership or endorsement. Workflow recommendations are ScholarVault’s editorial interpretation.
- ScholarVault local screening snapshot — aggregate data and methodology
- Walters & Wilder, 2023 — study of generated bibliographic citations
- ScholarVault — SCVS methodology and assessment limits
Have a correction or a newer institutional source? Send it to the editorial team.