VLM幻覺可提前攔截
Computer Science > Computer Vision and Pattern Recognition arXiv:2606.00435 (cs) [Submitted on 29 May 2026 (v1), last revised 16 Sep 2026 (this version, v4)] Title:Detect Before You Leap: Mirage Detection in Vision-Language Models Authors:Md.Shaown Miah, S.M.
Taiabul Haque, Syed Ishtiaque Ahmed, Sayeed Shafayet Chowdhury View a PDF of the paper titled Detect Before You Leap: Mirage Detection in Vision-Language Models, by Md.
Shaown Miah and 3 other authors View PDF HTML (experimental) Abstract:Vision-language models (VLMs) can produce confident answers without relevant visual evidence, a failure mode known as mirage reasoning (Asadi et al., 2026).
To that end, we study pre-release mirage detection: deciding whether a VLM answer should be released or withheld.
Our model-agnostic method, Text-Conditioned Layer-wise Internal Alignment (TC-LIA), tracks question-image alignment across the layers of a frozen CLIP ViT-H/14 encoder, summarizing patch-text alignment by final similarity, late-layer top-k alignment, early-to-late gain, and slope.
TC-LIA is purely unsupervised (fixed projections, fixed scoring weights, no labels, no training) and already delivers strong detection independently.
Additionally, when combined with blank/noise detection, domain routing, and VLM self-assessment, it forms an ensemble whose supervised training improves performance but is an optional add-on.On 19,004 samples spanning ten VQA domains, fourteen state-of-the-art VLMs exhibit 57.3-75.
0% base mirage rates.Our proposed TC-LIA alone cuts this to 7.5% with 83.5% Related/Unrelated/Blank-Noise classification accuracy, and the ensemble reaches 84.3-88.4% accuracy with 5.9-7.2% mirage rates (best joint result: 88.4% accuracy, 6.4% mirage rate).
Notably, an ensemble trained on a single backbone transfers well to unseen backbones, with the best-transferring source staying within 1.2% accuracy points of per-backbone training across thirteen held-out VLMs.Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.
AI) Cite as: arXiv:2606.00435 [cs.CV] (or arXiv:2606.00435v4 [cs.CV] for this version) https://doi.org/10.48550/arXiv.2606.00435 Focus to learn more arXiv-issued DOI via DataCite Submission history From: Md.
Shaown Miah [view email] [v1] Fri, 29 May 2026 23:51:35 UTC (11,848 KB) [v2] Mon, 15 Jun 2026 15:09:01 UTC (13,483 KB) [v3] Mon, 27 Jul 2026 23:31:43 UTC (16,234 KB) [v4] Wed, 16 Sep 2026 06:47:00 UTC (16,286 KB) Full-text links: Access Paper: View a PDF of the paper titled Detect Before You Leap: Mirage Detection in Vision-Language Models, by Md.
Shaown Miah and 3 other authorsView PDFHTML (experimental)TeX Source view license Current browse context: cs.CV < prev | next > new | recent | 2026-06 Change to browse by: cs cs.AI References & Citations NASA ADSGoogle Scholar Semantic Scholar export BibTeX citation Loading...
BibTeX formatted citation × loading...Data provided by: Bookmark Bibliographic Tools Bibliographic and Citation Tools Bibliographic Explorer Toggle Bibliographic Explorer (What is the Explorer?) Connected Papers Toggle Connected Papers (What is Connected Papers?
) Litmaps Toggle Litmaps (What is Litmaps?) scite.ai Toggle scite Smart Citations (What are Smart Citations?) Code, Data, Media Code, Data and Media Associated with this Article alphaXiv Toggle alphaXiv (What is alphaXiv?) Links to Code Toggle CatalyzeX Code Finder for Papers (What is CatalyzeX?
) DagsHub Toggle DagsHub (What is DagsHub?) GotitPub Toggle Gotit.pub (What is GotitPub?) Huggingface Toggle Hugging Face (What is Huggingface?) ScienceCast Toggle ScienceCast (What is ScienceCast?) Demos Demos Replicate Toggle Replicate (What is Replicate?
) Spaces Toggle Hugging Face Spaces (What is Spaces?) Spaces Toggle TXYZ.AI (What is TXYZ.AI?) Related Papers Recommenders and Search Tools Link to Influence Flower Influence Flower (What are Influence Flowers?) Core recommender toggle CORE Recommender (What is CORE?
) Author Venue Institution Topic About arXivLabs arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy.arXiv is committed to these values and only works with partners that adhere to them.
Have an idea for a project that will add value for arXiv's community?Learn more about arXivLabs.Which authors of this paper are endorsers?| Disable MathJax (What is MathJax?)
Related
相關文章

美國聯邦法官駁回 OpenAI 請求,蘋果無需公佈與 SpaceXAI 和解協議
儘管此前有公佈保密和解協議的先例,但這種情況的前提是其包含與訴訟具體帶決爭議點相關的信息,對於 SpaceXAI 訴 OpenAI 案而言,SpaceXAI 與蘋果的和解協議不屬於這類情形。

單個分子沒有溫度,單個神經元也沒有智能 | 關於“湧現是什麼”、“智能是什麼”的一些思考
一隻分子有溫度嗎?直覺上似乎有。水是熱的,組成水的分子當然也應該是熱的。但嚴格來說,單個分子只有質量、速度、方向和動能。它會運動,會與其他分子碰撞,卻沒有我們通常所說的“溫度”。

AI 負責創造,人來幹髒活,這事兒能否停一下?
AI 负责天马行空的创意,人类却在为它收拾数据、清洗标注、调试参数——这种“AI 负责创造,人来干脏活”的倒挂分工,正在成为许多从业者的日常。所谓“脏活”,指的是那些低创造性、高重复性的劳动:整理混乱的原始数据、手动修正模型输出的幻觉、把格式错乱的回答重新排版。讽刺的是,AI 被吹捧为解放生产力的工具,结果它把最需要智力的部分留给了自己,把最机械的部分甩给了人。

全國網絡安全標準化技術委員會發布《人工智能安全治理框架 3.0》
網安標委在 2026 國家網絡安全宣傳週開幕式上發佈《人工智能安全治理框架 3.0》,該框架更新了風險分類,優化了治理措施,旨在提升人工智能安全治理能力,保障人工智能造福人類。#人工智能安全治理#

剛剛,阿迪王呼籲為AI緊急降速,馬斯克奧特曼罕見完全支持:遞歸自我進化太危險了
Anthropic 執行長 Dario Amodei 近日公開呼籲各界緊急為人工智慧發展踩煞車,直言「遞歸自我進化」的風險已經大到不容忽視。這項發言意外獲得向來意見分歧的 Tesla 創辦人馬斯克與 OpenAI 執行長奧特曼一致表態支持,三人在 AI 安全議題上形成罕見共識。Amodei 認為,當 AI 系統具備自我改寫與持續升級的能力,人類將可能失去控制權,導致不可逆的災難性後果。

Anthropic CEO 阿莫迪稱 AI 行業應當放緩發展速度
作者:沁滄(實習) 責編:沁滄 評論: 9 月 12 日消息,Anthropic CEO 達里奧 · 阿莫迪(Dario Amodei)今日發文,稱人工智能行業應當放緩發展速度。阿莫迪表示,從 2026 年夏季開始,AI 協助構建下一代 AI 的能力快速演進,若不加節制,技術突破速度將徹底超越人類的理解與掌控能力。