u/KauravaLivesMatter

ELI5 - Why do LLMs hallucinate?

I have seen videos about the transformer architecture etc., and I get that large language models generate responses based on some statistical likelihood of words and terms. However, I still don't get how they can completely make up facts and even references.

Why can't they state facts that they have come across in their training as they are? What is it, either from a mathematical standpoint or from an architectural standpoint of large language models that causes them to hallucinate?

reddit.com
u/KauravaLivesMatter — 10 days ago