Self-adaptive fuzzing exposes the hallucination cracks in multimodal LLMs
A new evaluation framework pairs a unified taxonomy benchmark with self-adaptive multimodal fuzzing (SAMF) and shows that state-of-the-art MLLMs degrade under stress, revealing a gap between reasoning and factual grounding. RL alignment even worsens ...