Inside language models like LLaMA lie tangled patterns, their strength depending on the training method. But to an outside observer, the model always yields the same result—much like a black hole that 'forgets' every detail of the matter it swallows. This explains why simple fine-tuning methods are
arXiv:2601.06788 · 2026-01-11