I was deep in thought about how much of a model hallucination is RLHF-based assumption in the absence of training data or context. Pretraining gives the model the ability to confabulate. Post-training often influences whether it chooses to confabulate rather than say "I don't know." A base language model is trained to predict plausible continuations. If the evidence needed to answer is absent from its weights or context, there's no fundamental mechanism in next-token prediction that says "st...