As Multimodal Language Models (MLMs) become essential to how we access and interpret online content and world knowledge, their inability to master multimodal pragmatic reasoning remains a critical safety flaw. When faced with cross-modal incongruity—such as an image of people standing in a disastrous flood captioned with "living the dream"—MLMs struggle to bridge the inference gap. Crucially, this is not merely a perceptual flaw of missing visual cues; it is a fundamental reasoning failure. MLMs frequently fail to synthesize interacting modalities to deduce emergent, unspoken meanings like irony, implicit bias, or contextual reframing. When an MLM fails to reason pragmatically, the result is not just a technical error—it is a direct safety risk that can propagate harmful stereotypes or provide dangerous guidance, particularly to vulnerable users. Current safety measures rely heavily on post-hoc filters, attempting to muzzle harmful outputs at the surface level without addressing the underlying cognitive deficit. Consequently, models remain highly vulnerable to adversarial inputs where benign individual elements are combined to induce a harmful interpretation. This talk proposes a more fundamental approach: Gaze-Guided Reasoning. By integrating human cognitive signals, we strive to provide MLMs with a blueprint for cross-modal reasoning. Because human eye movement is inherently tied to cognitive processing, it offers a faithful behavioral signal to teach AI how to integrate complex information. By moving beyond brittle output filters, we aim to show how human gaze can alleviate current reasoning bottlenecks, paving the way for multimodal AI systems that are more interpretable, aligned, and intrinsically safer.
Slides from presentation
Slides from the presentation will be visible on this site if the speaker in question wishes to share them.
Please note that you need to be signed in in order to see them.