Abstract
Benchmarks such as MMLU suggest flagship language models approach factuality saturation, with scores above 90%. We show this picture is incomplete. *LLMpedia* generates encyclopedic articles entirely from parametric memory, producing 1M articles across three model families without retrieval. For gpt-5-mini, the verifiable true rate on Wikipedia-covered subjects is only 74.7% -- more th