VLM幻覺可提前攔截

2026年9月18日 00:00
站內 AI 整理稿

Computer Science > Computer Vision and Pattern Recognition arXiv:2606.00435 (cs) [Submitted on 29 May 2026 (v1), last revised 16 Sep 2026 (this version, v4)] Title:Detect Before You Leap: Mirage Detection in Vision-Language Models Authors:Md.Shaown Miah, S.M.

Taiabul Haque, Syed Ishtiaque Ahmed, Sayeed Shafayet Chowdhury View a PDF of the paper titled Detect Before You Leap: Mirage Detection in Vision-Language Models, by Md.

Shaown Miah and 3 other authors View PDF HTML (experimental) Abstract:Vision-language models (VLMs) can produce confident answers without relevant visual evidence, a failure mode known as mirage reasoning (Asadi et al., 2026).

To that end, we study pre-release mirage detection: deciding whether a VLM answer should be released or withheld.

Our model-agnostic method, Text-Conditioned Layer-wise Internal Alignment (TC-LIA), tracks question-image alignment across the layers of a frozen CLIP ViT-H/14 encoder, summarizing patch-text alignment by final similarity, late-layer top-k alignment, early-to-late gain, and slope.

TC-LIA is purely unsupervised (fixed projections, fixed scoring weights, no labels, no training) and already delivers strong detection independently.

Additionally, when combined with blank/noise detection, domain routing, and VLM self-assessment, it forms an ensemble whose supervised training improves performance but is an optional add-on.On 19,004 samples spanning ten VQA domains, fourteen state-of-the-art VLMs exhibit 57.3-75.

0% base mirage rates.Our proposed TC-LIA alone cuts this to 7.5% with 83.5% Related/Unrelated/Blank-Noise classification accuracy, and the ensemble reaches 84.3-88.4% accuracy with 5.9-7.2% mirage rates (best joint result: 88.4% accuracy, 6.4% mirage rate).

Notably, an ensemble trained on a single backbone transfers well to unseen backbones, with the best-transferring source staying within 1.2% accuracy points of per-backbone training across thirteen held-out VLMs.Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.

AI) Cite as: arXiv:2606.00435 [cs.CV] (or arXiv:2606.00435v4 [cs.CV] for this version) https://doi.org/10.48550/arXiv.2606.00435 Focus to learn more arXiv-issued DOI via DataCite Submission history From: Md.

Shaown Miah [view email] [v1] Fri, 29 May 2026 23:51:35 UTC (11,848 KB) [v2] Mon, 15 Jun 2026 15:09:01 UTC (13,483 KB) [v3] Mon, 27 Jul 2026 23:31:43 UTC (16,234 KB) [v4] Wed, 16 Sep 2026 06:47:00 UTC (16,286 KB) Full-text links: Access Paper: View a PDF of the paper titled Detect Before You Leap: Mirage Detection in Vision-Language Models, by Md.

Shaown Miah and 3 other authorsView PDFHTML (experimental)TeX Source view license Current browse context: cs.CV < prev | next > new | recent | 2026-06 Change to browse by: cs cs.AI References & Citations NASA ADSGoogle Scholar Semantic Scholar export BibTeX citation Loading...

BibTeX formatted citation × loading...Data provided by: Bookmark Bibliographic Tools Bibliographic and Citation Tools Bibliographic Explorer Toggle Bibliographic Explorer (What is the Explorer?) Connected Papers Toggle Connected Papers (What is Connected Papers?

) Litmaps Toggle Litmaps (What is Litmaps?) scite.ai Toggle scite Smart Citations (What are Smart Citations?) Code, Data, Media Code, Data and Media Associated with this Article alphaXiv Toggle alphaXiv (What is alphaXiv?) Links to Code Toggle CatalyzeX Code Finder for Papers (What is CatalyzeX?

) DagsHub Toggle DagsHub (What is DagsHub?) GotitPub Toggle Gotit.pub (What is GotitPub?) Huggingface Toggle Hugging Face (What is Huggingface?) ScienceCast Toggle ScienceCast (What is ScienceCast?) Demos Demos Replicate Toggle Replicate (What is Replicate?

) Spaces Toggle Hugging Face Spaces (What is Spaces?) Spaces Toggle TXYZ.AI (What is TXYZ.AI?) Related Papers Recommenders and Search Tools Link to Influence Flower Influence Flower (What are Influence Flowers?) Core recommender toggle CORE Recommender (What is CORE?

) Author Venue Institution Topic About arXivLabs arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.

Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy.arXiv is committed to these values and only works with partners that adhere to them.

Have an idea for a project that will add value for arXiv's community?Learn more about arXivLabs.Which authors of this paper are endorsers?| Disable MathJax (What is MathJax?)

Related

相關文章

AI 負責創造,人來幹髒活,這事兒能否停一下?

AI 负责天马行空的创意,人类却在为它收拾数据、清洗标注、调试参数——这种“AI 负责创造,人来干脏活”的倒挂分工,正在成为许多从业者的日常。所谓“脏活”,指的是那些低创造性、高重复性的劳动:整理混乱的原始数据、手动修正模型输出的幻觉、把格式错乱的回答重新排版。讽刺的是,AI 被吹捧为解放生产力的工具,结果它把最需要智力的部分留给了自己,把最机械的部分甩给了人。

3 天前

剛剛,阿迪王呼籲為AI緊急降速,馬斯克奧特曼罕見完全支持:遞歸自我進化太危險了

Anthropic 執行長 Dario Amodei 近日公開呼籲各界緊急為人工智慧發展踩煞車,直言「遞歸自我進化」的風險已經大到不容忽視。這項發言意外獲得向來意見分歧的 Tesla 創辦人馬斯克與 OpenAI 執行長奧特曼一致表態支持,三人在 AI 安全議題上形成罕見共識。Amodei 認為,當 AI 系統具備自我改寫與持續升級的能力,人類將可能失去控制權,導致不可逆的災難性後果。

4 天前
IT之家其他AI

Anthropic CEO 阿莫迪稱 AI 行業應當放緩發展速度

作者:沁滄(實習) 責編:沁滄 評論: 9 月 12 日消息,Anthropic CEO 達里奧 · 阿莫迪(Dario Amodei)今日發文,稱人工智能行業應當放緩發展速度。阿莫迪表示,從 2026 年夏季開始,AI 協助構建下一代 AI 的能力快速演進,若不加節制,技術突破速度將徹底超越人類的理解與掌控能力。

6 天前