<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Medical Hallucinations | Xi Zhang</title><link>https://x-izhang.github.io/tags/medical-hallucinations/</link><atom:link href="https://x-izhang.github.io/tags/medical-hallucinations/index.xml" rel="self" type="application/rss+xml"/><description>Medical Hallucinations</description><generator>Hugo Blox Builder (https://hugoblox.com)</generator><language>en-us</language><lastBuildDate>Wed, 17 Jun 2026 00:00:00 +0000</lastBuildDate><image><url>https://x-izhang.github.io/media/icon_hu134860076176174952.png</url><title>Medical Hallucinations</title><link>https://x-izhang.github.io/tags/medical-hallucinations/</link></image><item><title>🎤 Talk at Glasgow AI4BioMed Lab!</title><link>https://x-izhang.github.io/post/2026ai4bioccs/</link><pubDate>Wed, 17 Jun 2026 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/post/2026ai4bioccs/</guid><description>&lt;h3 id="-talk-title">💡 Talk Title&lt;/h3>
&lt;p>&lt;a href="https://x-izhang.github.io/publication/zhang-2026-ccsclinicalconsensusselection/">&lt;strong>&amp;ldquo;CCS: Clinical Consensus Selection for Radiology Report Generation&amp;rdquo;&lt;/strong>&lt;/a>&lt;/p>
&lt;figure>&lt;img src="https://x-izhang.github.io/post/2026ai4bioccs/talk.png">
&lt;/figure>
&lt;h5 id="-abstract">🖇️ Abstract&lt;/h5>
&lt;p>Radiology report generation (RRG) is commonly framed as a single-path task, where a multimodal large language model (MLLM) produces one decoded report as the final output. While progress has largely come from scaling training data, model capacity, and retrieval mechanisms, improving report quality at inference time remains underexplored. We observe that fixed radiology MLLMs often generate clinically stronger reports elsewhere in their candidate pool than the one selected by default decoding, suggesting that inference-time decision making is an overlooked bottleneck. To address this, we propose &lt;strong>C&lt;/strong>linical &lt;strong>C&lt;/strong>onsensus &lt;strong>S&lt;/strong>election (&lt;strong>CCS&lt;/strong>), a decoder-agnostic inference-time selection framework that samples multiple candidate reports and selects the one with the highest clinical consensus across the rollout pool. CCS combines text-based utilities with a radiology-adapted utility computed by an image&amp;ndash;report-trained multimodal embedder, measuring agreement beyond surface-level textual similarity. Across three datasets and multiple radiology MLLMs, CCS consistently improves inference-time performance over single-path decoding and generic Best-of-N baselines, with especially clear gains on clinical metrics. Further analysis shows that image-grounded utility forms a selection axis distinct from textual consensus, and that substantial headroom remains for improving RRG at inference time.&lt;/p>
&lt;p>🔗 &lt;strong>Project Website:&lt;/strong> &lt;a href="https://x-izhang.github.io/CCS/" target="_blank" rel="noopener">https://x-izhang.github.io/CCS/&lt;/a>&lt;/p>
&lt;h3 id="-slides">📺 Slides&lt;/h3>
&lt;div style="text-align: center;">
&lt;iframe src="https://docs.google.com/presentation/d/e/2PACX-1vRuYdC5UkgKIXrivSKfN_ULo2nUagGypgN27be1LM0cAQgtlmbgI7WbneXLSkSi9Wp9cboKrg02PaU8/pubembed?start=true&amp;loop=true&amp;delayms=3000" frameborder="0" width="700" height="422" allowfullscreen="true" mozallowfullscreen="true" webkitallowfullscreen="true">&lt;/iframe>
&lt;/div>
&lt;p>📍 &lt;a href="https://ai4biomed.org/" target="_blank" rel="noopener">&lt;strong>Glasgow AI4BioMed Lab&lt;/strong>&lt;/a> — An interdisciplinary research lab focusing on AI applications in biomedical sciences, based in Glasgow.&lt;/p>
&lt;h3 id="heading">📅&lt;/h3>
&lt;p>&lt;strong>When:&lt;/strong> Wednesday, June 17, 2026 at 2pm&lt;br>
&lt;strong>Where:&lt;/strong> F121&lt;/p>
&lt;p>See you there!&lt;/p></description></item><item><title>🎉 Paper Accepted — See You at ACL 2026!</title><link>https://x-izhang.github.io/post/2026acl/</link><pubDate>Mon, 06 Apr 2026 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/post/2026acl/</guid><description>&lt;p>&lt;a href="https://x-izhang.github.io/publication/zhang-2025-ccdmitigatinghallucinationsradiology/">&lt;em>&amp;ldquo;CCD: Mitigating Hallucinations in Radiology MLLMs via Clinical Contrastive Decoding&amp;rdquo;&lt;/em>&lt;/a> has been accepted to &lt;a href="https://2026.aclweb.org/" target="_blank" rel="noopener">ACL 2026&lt;/a>!&lt;/p>
&lt;p>📍 &lt;strong>ACL 2026&lt;/strong> — The 64th Annual Meeting of the Association for Computational Linguistics will take place in &lt;strong>San Diego, California&lt;/strong>, July 2026&lt;/p>
&lt;p>See you there!&lt;/p>
&lt;!-- ### 📺 Talk
&lt;div style="position: relative; padding-bottom: 56.25%; height: 0; overflow: hidden;">
&lt;iframe allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" allowfullscreen="allowfullscreen" loading="eager" referrerpolicy="strict-origin-when-cross-origin" src="https://www.youtube.com/embed/_R8XUaaAU3g?autoplay=0&amp;controls=1&amp;end=0&amp;loop=0&amp;mute=0&amp;start=0" style="position: absolute; top: 0; left: 0; width: 100%; height: 100%; border:0;" title="YouTube video"
>&lt;/iframe>
&lt;/div>
-->
&lt;h3 id="-poster">🪧 Poster&lt;/h3>
&lt;figure>&lt;img src="https://x-izhang.github.io/post/2026acl/2026acl_poster.jpg">
&lt;/figure></description></item><item><title>CCD: Mitigating Hallucinations in Radiology MLLMs via Clinical Contrastive Decoding</title><link>https://x-izhang.github.io/publication/zhang-2025-ccdmitigatinghallucinationsradiology/</link><pubDate>Mon, 06 Apr 2026 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/publication/zhang-2025-ccdmitigatinghallucinationsradiology/</guid><description>&lt;div class="flex px-4 py-3 mb-6 rounded-md bg-primary-100 dark:bg-primary-900">
&lt;span class="pr-3 pt-1 text-primary-600 dark:text-primary-300">
&lt;svg height="24" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24">&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="m11.25 11.25l.041-.02a.75.75 0 0 1 1.063.852l-.708 2.836a.75.75 0 0 0 1.063.853l.041-.021M21 12a9 9 0 1 1-18 0a9 9 0 0 1 18 0m-9-3.75h.008v.008H12z"/>&lt;/svg>
&lt;/span>
&lt;span class="dark:text-neutral-300">For the latest updates and details, visit the &lt;a href="https://x-izhang.github.io/CCD/" target="_blank" rel="noopener">CCD Project Website&lt;/a>.&lt;/span>
&lt;/div>
&lt;h2 id="why-ccd-">Why CCD? 🔬&lt;/h2>
&lt;figure>&lt;img src="https://x-izhang.github.io/publication/zhang-2025-ccdmitigatinghallucinationsradiology/figure_1.png">
&lt;/figure>
&lt;p>Radiology MLLMs remain vulnerable to prompt-induced hallucinations when clinical sections contain counterfactual details or ambiguous guidance. The figure above contrasts baseline predictions with &lt;strong>CCD-enabled&lt;/strong> outputs across report generation and question answering tasks. &lt;strong>Red&lt;/strong> highlights mark unsupported findings that can compromise patient care, while &lt;strong>blue&lt;/strong> text reflects misleading prompt context the model must resist.&lt;/p>
&lt;p>Clinical Contrastive Decoding (CCD) mitigates these risks at inference time. Rather than retraining models or relying on retrieval corpora, &lt;strong>CCD&lt;/strong> injects trustworthy, image-grounded signals distilled from specialist expert models. The result is a decoding policy that maintains fluency yet stays faithful to the radiograph.&lt;/p>
&lt;h2 id="ccd-at-a-glance-">CCD at a Glance ⚙️&lt;/h2>
&lt;p>&lt;strong>Clinical Contrastive Decoding (CCD)&lt;/strong> is a plug-and-play inference framework designed to reduce medical hallucinations in radiology MLLMs. It introduces structured clinical supervision from expert models (e.g., DenseNet or MedSigLIP, or others) at decoding time, without modifying model weights or requiring external retrieval.&lt;/p>
&lt;p>Given a chest radiograph, the expert model predicts symptom-level probabilities across 14 CheXpert categories. CCD integrates this signal through a dual-stage logit refinement strategy:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>Symptom-grounded Contrastive Decoding (SCD)&lt;/strong>: constructs an anchor prompt using high-confidence findings (e.g., “Atelectasis, Cardiomegaly”) and generates a contrastive logits path conditioned on this prompt. The final logits are a weighted interpolation between the anchor-conditioned and original paths, encouraging the model to mention supported findings and suppress false negatives.&lt;/li>
&lt;li>&lt;strong>Expert-informed Contrastive Decoding (ECD)&lt;/strong>: transforms expert probabilities into token-level logit biases via log-odds conversion. These biases are injected into the logits from the first stage, softly penalising unsupported findings and reducing false positives.&lt;/li>
&lt;/ul>
&lt;p>This two-stage mechanism enables &lt;strong>CCD&lt;/strong> to progressively guide generation with both symbolic supervision (via anchor prompts) and probabilistic constraints (via confidence scores), achieving robust improvements in both report generation and VQA. &lt;strong>CCD&lt;/strong> is fully model-agnostic and integrates seamlessly with state-of-the-art radiology MLLMs such as MAIRA-2, Libra, LLaVA-Rad, and LLaVA-Med.&lt;/p>
&lt;p>Unlike prior contrastive decoding approaches that rely on perturbed visual or textual inputs, &lt;strong>CCD&lt;/strong> leverages clinically grounded signals from expert models to provide task-specific and symptom-level control during generation.&lt;/p>
&lt;h2 id="key-contributions-">Key Contributions ✨&lt;/h2>
&lt;ul>
&lt;li>
&lt;p>&lt;strong>Empirical Insight&lt;/strong>: We conduct a systematic analysis of &lt;strong>prompt-induced hallucinations&lt;/strong> in radiology MLLMs, revealing that noisy clinical sections (such as &lt;em>&lt;strong>irrelevant&lt;/strong>&lt;/em> or &lt;em>&lt;strong>contradictory clinical&lt;/strong>&lt;/em> details) can trigger &lt;strong>unsupported findings&lt;/strong> across multiple datasets.&lt;/p>
&lt;/li>
&lt;li>
&lt;p>&lt;strong>Inference-time Framework&lt;/strong>: We introduce &lt;strong>Clinical Contrastive Decoding (CCD)&lt;/strong> — a &lt;strong>dual-stage inference-time strategy&lt;/strong> that leverages expert-derived labels as &lt;strong>anchor prompts&lt;/strong> and performs &lt;strong>probabilistic logit adjustments&lt;/strong>, requiring &lt;strong>no retraining or architectural modifications&lt;/strong>.&lt;/p>
&lt;/li>
&lt;li>
&lt;p>&lt;strong>Consistent Gains&lt;/strong>: Extensive experiments on &lt;strong>MIMIC-CXR&lt;/strong>, &lt;strong>IU-Xray&lt;/strong>, and &lt;strong>CheXpert Plus&lt;/strong> demonstrate that &lt;strong>CCD&lt;/strong> achieves up to &lt;strong>+17% improvement in RadGraph-F1&lt;/strong> with SOTA radiology MLLMs (MAIRA-2), along with higher &lt;strong>VQA accuracy&lt;/strong>, all &lt;strong>without altering model weights or structures&lt;/strong>.&lt;/p>
&lt;/li>
&lt;/ul>
&lt;h2 id="evaluation-highlights-">Evaluation Highlights 📈&lt;/h2>
&lt;figure>&lt;img src="https://x-izhang.github.io/publication/zhang-2025-ccdmitigatinghallucinationsradiology/figure_2.png">
&lt;/figure>
&lt;ul>
&lt;li>Achieves up to &lt;strong>+17% RadGraph-F1&lt;/strong> on MIMIC-CXR when attached to state-of-the-art report generators.&lt;/li>
&lt;li>Reduces mention-level hallucinations on IU X-Ray while maintaining &lt;strong>BLEU / ROUGE&lt;/strong> scores, confirming that &lt;strong>CCD&lt;/strong> improves factuality without harming surface-level metrics.&lt;/li>
&lt;/ul>
&lt;h2 id="ablation-studies-">Ablation Studies 🧪&lt;/h2>
&lt;figure>&lt;img src="https://x-izhang.github.io/publication/zhang-2025-ccdmitigatinghallucinationsradiology/figure_3.png">
&lt;/figure>
&lt;p>Ablations show both contrastive stages matter: removing the anchor projection (SCD) or the logit sharpening (ECD) degrades factual metrics by &lt;strong>~5–8%&lt;/strong>.&lt;/p>
&lt;h2 id="my-musings-">My Musings ⏳&lt;/h2>
&lt;p>As I reflect on the development and impact of &lt;strong>Clinical Contrastive Decoding (CCD)&lt;/strong>,&lt;/p>
&lt;blockquote>
&lt;p>It’s better to be roughly right than precisely wrong.&lt;br>
— &lt;em>Carveth Read, Logic: Deductive and Inductive&lt;/em>&lt;/p>
&lt;/blockquote>
&lt;p>This quote perfectly captures the spirit of our approach. In the high-stakes field of radiology, the cost of hallucinations can be severe. By integrating expert-derived signals at inference time, &lt;strong>CCD&lt;/strong> prioritises &lt;strong>clinical accuracy&lt;/strong> over mere &lt;strong>linguistic fluency&lt;/strong>. It’s a reminder that in medical AI, being &lt;em>approximately correct and grounded in reality&lt;/em> is far more valuable than generating &lt;em>fluent but misleading&lt;/em> text.&lt;/p>
&lt;h2 id="bibtex-">BibTeX 📚&lt;/h2>
&lt;div class="highlight">&lt;pre tabindex="0" class="chroma">&lt;code class="language-bibtex" data-lang="bibtex">&lt;span class="line">&lt;span class="cl">&lt;span class="nc">@article&lt;/span>&lt;span class="p">{&lt;/span>&lt;span class="nl">zhang2025ccd&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">title&lt;/span>&lt;span class="p">=&lt;/span>&lt;span class="s">{CCD: Mitigating Hallucinations in Radiology MLLMs via Clinical Contrastive Decoding}&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">author&lt;/span>&lt;span class="p">=&lt;/span>&lt;span class="s">{Zhang, Xi and Meng, Zaiqiao and Lever, Jake and Ho, Edmond SL}&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">journal&lt;/span>&lt;span class="p">=&lt;/span>&lt;span class="s">{arXiv preprint arXiv:2509.23379}&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">year&lt;/span>&lt;span class="p">=&lt;/span>&lt;span class="s">{2025}&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="p">}&lt;/span>
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div></description></item><item><title>🎤 Talk at Glasgow AI4BioMed Lab!</title><link>https://x-izhang.github.io/post/2026ai4bioccd/</link><pubDate>Wed, 28 Jan 2026 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/post/2026ai4bioccd/</guid><description>&lt;h3 id="-talk-title">💡 Talk Title&lt;/h3>
&lt;p>&lt;a href="https://x-izhang.github.io/publication/zhang-2025-ccdmitigatinghallucinationsradiology/">&lt;strong>&amp;ldquo;CCD: Mitigating Hallucinations in Radiology MLLMs via Clinical Contrastive Decoding&amp;rdquo;&lt;/strong>&lt;/a>&lt;/p>
&lt;figure>&lt;img src="https://x-izhang.github.io/post/2026ai4bioccd/talk.jpeg">
&lt;/figure>
&lt;h5 id="-abstract">🖇️ Abstract&lt;/h5>
&lt;p>Radiology multimodal large language models (MLLMs) often generate clinically unsupported descriptions, posing serious risks in medical applications. We introduce &lt;strong>Clinical Contrastive Decoding (CCD)&lt;/strong>, a training-free inference framework that integrates structured clinical signals from radiology expert models to mitigate hallucinations. &lt;strong>CCD&lt;/strong> refines token-level logits during generation through a dual-stage contrastive mechanism, enhancing clinical fidelity without modifying the base MLLM. On the MIMIC-CXR dataset,&lt;strong>CCD&lt;/strong> yields up to &lt;strong>17%&lt;/strong> improvement in RadGraph-F1, providing a lightweight solution for bridging expert models and MLLMs in radiology.&lt;/p>
&lt;p>🔗 &lt;strong>Project Website:&lt;/strong> &lt;a href="https://x-izhang.github.io/CCD/" target="_blank" rel="noopener">https://x-izhang.github.io/CCD/&lt;/a>&lt;/p>
&lt;h3 id="-slides">📺 Slides&lt;/h3>
&lt;div style="text-align: center;">
&lt;iframe src="https://docs.google.com/presentation/d/e/2PACX-1vQ-1I0vwB28WW7VlHrKmvKYRx6-SIRF5uqEmvbooXQC_wteF2Pb_j9B9Ob9VQUiyTUQs3eGkBYcqj2G/pubembed?start=true&amp;loop=true&amp;delayms=3000" frameborder="0" width="700" height="422" allowfullscreen="true" mozallowfullscreen="true" webkitallowfullscreen="true">&lt;/iframe>
&lt;/div>
&lt;p>📍 &lt;a href="https://ai4biomed.org/" target="_blank" rel="noopener">&lt;strong>Glasgow AI4BioMed Lab&lt;/strong>&lt;/a> — An interdisciplinary research lab focusing on AI applications in biomedical sciences, based in Glasgow.&lt;/p>
&lt;h3 id="heading">📅&lt;/h3>
&lt;p>&lt;strong>When:&lt;/strong> Wednesday, January 28, 2026 at 2pm&lt;br>
&lt;strong>Where:&lt;/strong> F121&lt;/p>
&lt;p>See you there!&lt;/p></description></item><item><title>🎤 Invited Talk at DICTA 2025!</title><link>https://x-izhang.github.io/post/dicta2025/</link><pubDate>Tue, 25 Nov 2025 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/post/dicta2025/</guid><description>&lt;h3 id="-talk-title">💡 Talk Title&lt;/h3>
&lt;p>&lt;strong>Multimodal Medical Models: Cross-modal Alignment and Consistency&lt;/strong>&lt;/p>
&lt;h5 id="-abstract">🖇️ Abstract&lt;/h5>
&lt;p>This talk provides an overview of recent developments in medical vision–language modelling and examines key sources of misalignment that lead to hallucinations. Emerging strategies to enhance visual, semantic, and temporal consistency will also be discussed, highlighting pathways toward safer and more trustworthy clinical AI systems.&lt;/p>
&lt;h3 id="-slides">📺 Slides&lt;/h3>
&lt;div style="text-align: center;">
&lt;iframe src="https://docs.google.com/presentation/d/e/2PACX-1vSMAkxuWzc2RErcGE-iC5tXpdOgpOoq5KHyzXfuXa5g296Djm80oc8cl-aGNSgpKo0bFndCikieTPL9/pubembed?start=true&amp;loop=true&amp;delayms=5000" frameborder="0" width="700" height="422" allowfullscreen="true" mozallowfullscreen="true" webkitallowfullscreen="true">&lt;/iframe>
&lt;/div>
&lt;p>📍 &lt;strong>DICTA 2025&lt;/strong> — &lt;a href="https://dicta2025.dictaconference.org/" target="_blank" rel="noopener">The 26th International Conference on Digital Image Computing: Techniques and Applications&lt;/a> will take place in &lt;strong>Adelaide, Australia&lt;/strong>, from 3–5 December 2025.&lt;/p>
&lt;p>🏥 &lt;a href="https://sites.google.com/view/medai-chas/medai-chas" target="_blank" rel="noopener">&lt;strong>MedAI-CHAS&lt;/strong>&lt;/a> — A workshop focusing on Challenges, Hallucinations, and Solutions for Advancing Clinical Utility in Medical AI.&lt;/p>
&lt;p>See you there!&lt;/p></description></item></channel></rss>