Liqiang Jing Yue Zhang Xinya Du Jing Recent Advances in Multimodal Hallucination

Recent Advances in Multimodal Hallucination

von Liqiang Jing Yue Zhang Xinya Du

Preis unbekannt

Buch in deiner Nähe kaufen


...oder deine aktuelle Postleitzahl eingeben:
oder

Beschreibung

Hallucination in Multimodal Models is the first comprehensive research monograph dedicated to the growing challenge of hallucinations in large-scale multimodal AI systems, particularly vision-language models (VLMs) and multimodal large language models (MLLMs). The book systematically defines, categorizes, evaluates, and mitigates hallucinations — cases where models generate content that is factually inconsistent, visually unsupported, or commonsensically implausible. These hallucinations have become increasingly problematic in real-world applications of AI, including robotics, autonomous systems, and AI-generated media, where the consequences of inaccurate outputs can be severe.

The purpose of this book is threefold:

(1) to formalize the types and causes of hallucination in multimodal models;

(2) to present state-of-the-art evaluation frameworks, such as FaithScore and FIHA, for quantifying hallucination at a fine-grained level; and

(3) to introduce a unified mitigation framework.

The book presents new research results, including multiple benchmark datasets (QA-VisualGenome, QA-FB15k, FIHA-v1) and empirical studies across leading MLLMs like LLaVA, InstructBLIP, and VisualGLM. Of particular interest are the analytical insights into which components of vision-language architectures (e.g., LLM backbones, vision encoders, visual-textual connectors) are most responsible for hallucination behaviors. The book also highlights the relationship between model scale, prompt design, and hallucination frequency, offering practical tools for both researchers and engineers.

This book complements and extends the existing literature on multimodal model evaluation (e.g., MME, SEED-Bench, LAMM) by moving beyond surface-level metrics to offer deeper, interpretable, and automated hallucination analysis. Unlike survey papers or benchmarks that only diagnose the problem, our monograph provides a cohesive solution path from diagnosis to mitigation, built upon novel technical contributions and real-world implementations.

As this is the first edition, it introduces original theoretical frameworks, algorithms, benchmarks, and design paradigms. It is intended to serve as a reference for graduate students, academic researchers, and industry practitioners working in natural language processing, computer vision, embodied AI, and trustworthy AI systems.


Hallucination in Multimodal Models is the first comprehensive research monograph dedicated to the growing challenge of hallucinations in large-scale multimodal AI systems, particularly vision-language models (VLMs) and multimodal large language models (MLLMs). The book systematically defines, categorizes, evaluates, and mitigates hallucinations — cases where models generate content that is factually inconsistent, visually unsupported, or commonsensically implausible. These hallucinations have become increasingly problematic in real-world applications of AI, including robotics, autonomous systems, and AI-generated media, where the consequences of inaccurate outputs can be severe.

The purpose of this book is threefold:

(1) to formalize the types and causes of hallucination in multimodal models;

(2) to present state-of-the-art evaluation frameworks, such as FaithScore and FIHA, for quantifying hallucination at a fine-grained level; and

(3) to introduce a unified mitigation framework.

The book presents recent research results, including FaithScore, FIHA, FIFA, FGAIF, and Dentist, together with empirical studies across leading large vision-language models. Of particular interest are its fine-grained approaches to atomic fact verification, semantic dependency modeling, unified text-video hallucination evaluation, reward-based alignment, and training-free hallucination mitigation. The book also examines how these methods can improve the faithfulness and reliability of multimodal model outputs, offering practical tools for both researchers and engineers.

This book complements and extends the existing literature on multimodal model evaluation (e.g., MME, SEED-Bench, LAMM) by moving beyond surface-level metrics to offer deeper, interpretable, and automated hallucination analysis. Unlike survey papers or benchmarks that only diagnose the problem, our monograph provides a cohesive solution path from diagnosis to mitigation, built upon novel technical contributions and real-world implementations.

As this is the first edition, it introduces original theoretical frameworks, algorithms, benchmarks, and design paradigms. It is intended to serve as a reference for graduate students, academic researchers, and industry practitioners working in natural language processing, computer vision, embodied AI, and trustworthy AI systems.


Fine-Grained, Automated Hallucination Evaluation Frameworks (FaithScore & FIHA) Unified Hallucination Mitigation Framework Across Models and Modalities Release of Public Benchmarks, Toolkits, and Interactive Demos

Autor*in

Liqiang Jing

Themen in »Recent Advances in Multimodal Hallucination«

Multimodal AI Vision-language models Hallucination in AI Large language models MLLM evaluation Faithful image captioning Visual hallucination detection AI hallucination mitigation World models Embodied AI Trustworthy AI VLM benchmarking Visual grounding Multimodal reasoning Generative AI evaluation

Stimmen zu »Recent Advances in Multimodal Hallucination«

Details

ISBN: 9783032363213
Verlag: Springer International Publishing
Erscheinung: 07.11.2026

Link teilen


Über buchnah.de | Die Buchhandlungen | Die Verlage | Impressum & Kontakt | Datenschutz | Presse


Auf dieser Seite kannst Du Buchhandlungen in der Nähe finden