<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Radiology Report Generation | Xi Zhang</title><link>https://x-izhang.github.io/tags/radiology-report-generation/</link><atom:link href="https://x-izhang.github.io/tags/radiology-report-generation/index.xml" rel="self" type="application/rss+xml"/><description>Radiology Report Generation</description><generator>Hugo Blox Builder (https://hugoblox.com)</generator><language>en-us</language><lastBuildDate>Wed, 17 Jun 2026 00:00:00 +0000</lastBuildDate><image><url>https://x-izhang.github.io/media/icon_hu134860076176174952.png</url><title>Radiology Report Generation</title><link>https://x-izhang.github.io/tags/radiology-report-generation/</link></image><item><title>🎤 Talk at Glasgow AI4BioMed Lab!</title><link>https://x-izhang.github.io/post/2026ai4bioccs/</link><pubDate>Wed, 17 Jun 2026 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/post/2026ai4bioccs/</guid><description>&lt;h3 id="-talk-title">💡 Talk Title&lt;/h3>
&lt;p>&lt;a href="https://x-izhang.github.io/publication/zhang-2026-ccsclinicalconsensusselection/">&lt;strong>&amp;ldquo;CCS: Clinical Consensus Selection for Radiology Report Generation&amp;rdquo;&lt;/strong>&lt;/a>&lt;/p>
&lt;figure>&lt;img src="https://x-izhang.github.io/post/2026ai4bioccs/talk.png">
&lt;/figure>
&lt;h5 id="-abstract">🖇️ Abstract&lt;/h5>
&lt;p>Radiology report generation (RRG) is commonly framed as a single-path task, where a multimodal large language model (MLLM) produces one decoded report as the final output. While progress has largely come from scaling training data, model capacity, and retrieval mechanisms, improving report quality at inference time remains underexplored. We observe that fixed radiology MLLMs often generate clinically stronger reports elsewhere in their candidate pool than the one selected by default decoding, suggesting that inference-time decision making is an overlooked bottleneck. To address this, we propose &lt;strong>C&lt;/strong>linical &lt;strong>C&lt;/strong>onsensus &lt;strong>S&lt;/strong>election (&lt;strong>CCS&lt;/strong>), a decoder-agnostic inference-time selection framework that samples multiple candidate reports and selects the one with the highest clinical consensus across the rollout pool. CCS combines text-based utilities with a radiology-adapted utility computed by an image&amp;ndash;report-trained multimodal embedder, measuring agreement beyond surface-level textual similarity. Across three datasets and multiple radiology MLLMs, CCS consistently improves inference-time performance over single-path decoding and generic Best-of-N baselines, with especially clear gains on clinical metrics. Further analysis shows that image-grounded utility forms a selection axis distinct from textual consensus, and that substantial headroom remains for improving RRG at inference time.&lt;/p>
&lt;p>🔗 &lt;strong>Project Website:&lt;/strong> &lt;a href="https://x-izhang.github.io/CCS/" target="_blank" rel="noopener">https://x-izhang.github.io/CCS/&lt;/a>&lt;/p>
&lt;h3 id="-slides">📺 Slides&lt;/h3>
&lt;div style="text-align: center;">
&lt;iframe src="https://docs.google.com/presentation/d/e/2PACX-1vRuYdC5UkgKIXrivSKfN_ULo2nUagGypgN27be1LM0cAQgtlmbgI7WbneXLSkSi9Wp9cboKrg02PaU8/pubembed?start=true&amp;loop=true&amp;delayms=3000" frameborder="0" width="700" height="422" allowfullscreen="true" mozallowfullscreen="true" webkitallowfullscreen="true">&lt;/iframe>
&lt;/div>
&lt;p>📍 &lt;a href="https://ai4biomed.org/" target="_blank" rel="noopener">&lt;strong>Glasgow AI4BioMed Lab&lt;/strong>&lt;/a> — An interdisciplinary research lab focusing on AI applications in biomedical sciences, based in Glasgow.&lt;/p>
&lt;h3 id="heading">📅&lt;/h3>
&lt;p>&lt;strong>When:&lt;/strong> Wednesday, June 17, 2026 at 2pm&lt;br>
&lt;strong>Where:&lt;/strong> F121&lt;/p>
&lt;p>See you there!&lt;/p></description></item><item><title>🎤 Talk at Glasgow AI4BioMed Lab!</title><link>https://x-izhang.github.io/post/2026ai4bioccd/</link><pubDate>Wed, 28 Jan 2026 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/post/2026ai4bioccd/</guid><description>&lt;h3 id="-talk-title">💡 Talk Title&lt;/h3>
&lt;p>&lt;a href="https://x-izhang.github.io/publication/zhang-2025-ccdmitigatinghallucinationsradiology/">&lt;strong>&amp;ldquo;CCD: Mitigating Hallucinations in Radiology MLLMs via Clinical Contrastive Decoding&amp;rdquo;&lt;/strong>&lt;/a>&lt;/p>
&lt;figure>&lt;img src="https://x-izhang.github.io/post/2026ai4bioccd/talk.jpeg">
&lt;/figure>
&lt;h5 id="-abstract">🖇️ Abstract&lt;/h5>
&lt;p>Radiology multimodal large language models (MLLMs) often generate clinically unsupported descriptions, posing serious risks in medical applications. We introduce &lt;strong>Clinical Contrastive Decoding (CCD)&lt;/strong>, a training-free inference framework that integrates structured clinical signals from radiology expert models to mitigate hallucinations. &lt;strong>CCD&lt;/strong> refines token-level logits during generation through a dual-stage contrastive mechanism, enhancing clinical fidelity without modifying the base MLLM. On the MIMIC-CXR dataset,&lt;strong>CCD&lt;/strong> yields up to &lt;strong>17%&lt;/strong> improvement in RadGraph-F1, providing a lightweight solution for bridging expert models and MLLMs in radiology.&lt;/p>
&lt;p>🔗 &lt;strong>Project Website:&lt;/strong> &lt;a href="https://x-izhang.github.io/CCD/" target="_blank" rel="noopener">https://x-izhang.github.io/CCD/&lt;/a>&lt;/p>
&lt;h3 id="-slides">📺 Slides&lt;/h3>
&lt;div style="text-align: center;">
&lt;iframe src="https://docs.google.com/presentation/d/e/2PACX-1vQ-1I0vwB28WW7VlHrKmvKYRx6-SIRF5uqEmvbooXQC_wteF2Pb_j9B9Ob9VQUiyTUQs3eGkBYcqj2G/pubembed?start=true&amp;loop=true&amp;delayms=3000" frameborder="0" width="700" height="422" allowfullscreen="true" mozallowfullscreen="true" webkitallowfullscreen="true">&lt;/iframe>
&lt;/div>
&lt;p>📍 &lt;a href="https://ai4biomed.org/" target="_blank" rel="noopener">&lt;strong>Glasgow AI4BioMed Lab&lt;/strong>&lt;/a> — An interdisciplinary research lab focusing on AI applications in biomedical sciences, based in Glasgow.&lt;/p>
&lt;h3 id="heading">📅&lt;/h3>
&lt;p>&lt;strong>When:&lt;/strong> Wednesday, January 28, 2026 at 2pm&lt;br>
&lt;strong>Where:&lt;/strong> F121&lt;/p>
&lt;p>See you there!&lt;/p></description></item><item><title>Automated Chest X-ray Report Generation Remains Unsolved</title><link>https://x-izhang.github.io/publication/doi-10-1142-9789819824755-0017/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/publication/doi-10-1142-9789819824755-0017/</guid><description>&lt;div class="flex px-4 py-3 mb-6 rounded-md bg-primary-100 dark:bg-primary-900">
&lt;span class="pr-3 pt-1 text-primary-600 dark:text-primary-300">
&lt;svg height="24" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24">&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="m11.25 11.25l.041-.02a.75.75 0 0 1 1.063.852l-.708 2.836a.75.75 0 0 0 1.063.853l.041-.021M21 12a9 9 0 1 1-18 0a9 9 0 0 1 18 0m-9-3.75h.008v.008H12z"/>&lt;/svg>
&lt;/span>
&lt;span class="dark:text-neutral-300">For the latest updates and details, visit the &lt;a href="https://rexrank.ai/" target="_blank" rel="noopener">ReXrank Challenge Website&lt;/a>.&lt;/span>
&lt;/div>
&lt;h2 id="why-this-work">Why this work?&lt;/h2>
&lt;p>Automated chest X-ray report generation has the potential to substantially reduce radiologist workload and improve clinical efficiency. However, despite rapid progress in vision-language models and large language models (LLMs), the true clinical reliability of these systems remains unclear. Existing evaluations are often inconsistent, rely on limited test sets, or fail to probe generalization across institutions and clinically challenging abnormal cases.&lt;/p>
&lt;p>To address these limitations, we present a large-scale, standardized benchmark study through the &lt;strong>ReXrank Challenge V1.0&lt;/strong>, designed to rigorously assess the current state of automated chest X-ray report generation under realistic and clinically meaningful conditions.&lt;/p>
&lt;h2 id="what-is-the-rexrank-challenge-v10">What is the ReXrank Challenge V1.0?&lt;/h2>
&lt;p>The ReXrank Challenge V1.0 is a comprehensive evaluation effort built on &lt;strong>ReXGradient&lt;/strong>, the largest test-only dataset to date for radiology report generation, comprising &lt;strong>10,000 studies from 67 healthcare institutions&lt;/strong>. The challenge brought together submissions from academia and industry, evaluating &lt;strong>8 new models&lt;/strong> alongside &lt;strong>16 previously benchmarked state-of-the-art systems&lt;/strong> under a unified evaluation protocol.&lt;/p>
&lt;p>All models were assessed using a diverse set of metrics, ranging from traditional text similarity measures to clinically grounded and LLM-based error detection metrics, enabling a multi-dimensional analysis of model performance.&lt;/p>
&lt;h2 id="key-findings">Key Findings&lt;/h2>
&lt;ul>
&lt;li>
&lt;p>&lt;strong>Automated chest X-ray report generation remains fundamentally unsolved.&lt;/strong>&lt;br>
Even the best-performing models achieve &lt;strong>less than 45% error-free reporting on abnormal studies&lt;/strong>, highlighting a substantial gap between current AI systems and clinical readiness.&lt;/p>
&lt;/li>
&lt;li>
&lt;p>&lt;strong>Large performance gaps between normal and abnormal studies.&lt;/strong>&lt;br>
Most models perform well on normal cases (often exceeding 80–90% no-significant-error rates) but struggle significantly with abnormal findings, where clinically important errors are common.&lt;/p>
&lt;/li>
&lt;li>
&lt;p>&lt;strong>Poor cross-institutional generalization.&lt;/strong>&lt;br>
Model rankings vary dramatically across healthcare sites, indicating that strong performance on one institution does not reliably transfer to others.&lt;/p>
&lt;/li>
&lt;li>
&lt;p>&lt;strong>Evaluation metrics capture complementary but inconsistent signals.&lt;/strong>&lt;br>
Traditional lexical metrics correlate poorly with LLM-based and clinically focused error metrics, underscoring the need for more clinically aligned evaluation frameworks.&lt;/p>
&lt;/li>
&lt;/ul>
&lt;h2 id="implications">Implications&lt;/h2>
&lt;p>Our results suggest that current AI systems are &lt;strong>not yet suitable for fully autonomous radiology reporting&lt;/strong>, particularly for abnormal cases. However, they show promise as &lt;strong>assistive tools&lt;/strong> for generating preliminary drafts that can be reviewed and refined by radiologists. The findings emphasize the importance of developing:&lt;/p>
&lt;ul>
&lt;li>methods specifically targeting abnormality detection,&lt;/li>
&lt;li>strategies for robust cross-institutional generalization, and&lt;/li>
&lt;li>evaluation frameworks that better reflect clinical correctness rather than surface-level textual similarity.&lt;/li>
&lt;/ul>
&lt;h2 id="conclusion">Conclusion&lt;/h2>
&lt;p>This work provides the most comprehensive benchmark to date for automated chest X-ray report generation and delivers a clear message to the community: while progress is evident, &lt;strong>significant challenges remain before clinically reliable deployment is possible&lt;/strong>. By releasing the ReXrank Challenge V1.0 and ReXGradient dataset, we aim to establish a rigorous foundation for future research and to drive the development of more robust, clinically grounded radiology AI systems.&lt;/p>
&lt;h2 id="bibtex">BibTeX&lt;/h2>
&lt;div class="highlight">&lt;pre tabindex="0" class="chroma">&lt;code class="language-bibtex" data-lang="bibtex">&lt;span class="line">&lt;span class="cl">&lt;span class="nc">@inproceedings&lt;/span>&lt;span class="p">{&lt;/span>&lt;span class="nl">zhang2025automated&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">title&lt;/span>&lt;span class="p">=&lt;/span>&lt;span class="s">{Automated Chest X-ray Report Generation Remains Unsolved}&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">author&lt;/span>&lt;span class="p">=&lt;/span>&lt;span class="s">{Zhang, Xiaoman and Acosta, Julian Nicolas and Yang, Xiaoli and Adithan, Subathra and Luo, Luyang and Zhou, Hong-Yu and Miller, Joshua and Huang, Ouwen and Zhou, Zongwei and Hamamci, Ibrahim Ethem and others}&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">booktitle&lt;/span>&lt;span class="p">=&lt;/span>&lt;span class="s">{Biocomputing 2026: Proceedings of the Pacific Symposium}&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">pages&lt;/span>&lt;span class="p">=&lt;/span>&lt;span class="s">{236--250}&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">year&lt;/span>&lt;span class="p">=&lt;/span>&lt;span class="s">{2025}&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">organization&lt;/span>&lt;span class="p">=&lt;/span>&lt;span class="s">{World Scientific}&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="p">}&lt;/span>
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div></description></item><item><title>🎉 Paper Accepted at PSB 2026!</title><link>https://x-izhang.github.io/post/2025psb/</link><pubDate>Mon, 22 Dec 2025 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/post/2025psb/</guid><description>&lt;p>&lt;a href="https://x-izhang.github.io/publication/doi-10-1142-9789819824755-0017/">&lt;em>Automated Chest X-ray Report Generation Remains Unsolved&lt;/em>&lt;/a> has been accepted to &lt;a href="https://psb.stanford.edu/" target="_blank" rel="noopener">Pacific Symposium on Biocomputing 2026&lt;/a>!&lt;/p>
&lt;p>For the &lt;em>&lt;strong>🏆 Chest X-ray Interpretation Leaderboard 🏆&lt;/strong>&lt;/em>, please visit &lt;a href="https://rexrank.ai/" target="_blank" rel="noopener">ReXrank&lt;/a>.&lt;/p>
&lt;p>📍 PSB 2026 - The Pacific Symposium on Biocomputing 2026 will be held in Hawaii, USA, from January 3-7, 2026.&lt;/p>
&lt;h3 id="-poster">🪧 Poster&lt;/h3>
&lt;figure>&lt;img src="https://x-izhang.github.io/post/2025psb/poster.jpg">
&lt;/figure></description></item><item><title>🎤 Invited Talk at DICTA 2025!</title><link>https://x-izhang.github.io/post/dicta2025/</link><pubDate>Tue, 25 Nov 2025 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/post/dicta2025/</guid><description>&lt;h3 id="-talk-title">💡 Talk Title&lt;/h3>
&lt;p>&lt;strong>Multimodal Medical Models: Cross-modal Alignment and Consistency&lt;/strong>&lt;/p>
&lt;h5 id="-abstract">🖇️ Abstract&lt;/h5>
&lt;p>This talk provides an overview of recent developments in medical vision–language modelling and examines key sources of misalignment that lead to hallucinations. Emerging strategies to enhance visual, semantic, and temporal consistency will also be discussed, highlighting pathways toward safer and more trustworthy clinical AI systems.&lt;/p>
&lt;h3 id="-slides">📺 Slides&lt;/h3>
&lt;div style="text-align: center;">
&lt;iframe src="https://docs.google.com/presentation/d/e/2PACX-1vSMAkxuWzc2RErcGE-iC5tXpdOgpOoq5KHyzXfuXa5g296Djm80oc8cl-aGNSgpKo0bFndCikieTPL9/pubembed?start=true&amp;loop=true&amp;delayms=5000" frameborder="0" width="700" height="422" allowfullscreen="true" mozallowfullscreen="true" webkitallowfullscreen="true">&lt;/iframe>
&lt;/div>
&lt;p>📍 &lt;strong>DICTA 2025&lt;/strong> — &lt;a href="https://dicta2025.dictaconference.org/" target="_blank" rel="noopener">The 26th International Conference on Digital Image Computing: Techniques and Applications&lt;/a> will take place in &lt;strong>Adelaide, Australia&lt;/strong>, from 3–5 December 2025.&lt;/p>
&lt;p>🏥 &lt;a href="https://sites.google.com/view/medai-chas/medai-chas" target="_blank" rel="noopener">&lt;strong>MedAI-CHAS&lt;/strong>&lt;/a> — A workshop focusing on Challenges, Hallucinations, and Solutions for Advancing Clinical Utility in Medical AI.&lt;/p>
&lt;p>See you there!&lt;/p></description></item><item><title>RadEval: A framework for radiology text evaluation</title><link>https://x-izhang.github.io/publication/xu-2025-radevalframeworkradiologytext/</link><pubDate>Mon, 22 Sep 2025 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/publication/xu-2025-radevalframeworkradiologytext/</guid><description>&lt;div class="flex px-4 py-3 mb-6 rounded-md bg-primary-100 dark:bg-primary-900">
&lt;span class="pr-3 pt-1 text-primary-600 dark:text-primary-300">
&lt;svg height="24" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24">&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="m11.25 11.25l.041-.02a.75.75 0 0 1 1.063.852l-.708 2.836a.75.75 0 0 0 1.063.853l.041-.021M21 12a9 9 0 1 1-18 0a9 9 0 0 1 18 0m-9-3.75h.008v.008H12z"/>&lt;/svg>
&lt;/span>
&lt;span class="dark:text-neutral-300">For the latest updates and details, visit the &lt;a href="https://github.com/jbdel/RadEval" target="_blank" rel="noopener">RadEval Project Website&lt;/a>.&lt;/span>
&lt;/div>
&lt;h2 id="why-radeval">Why RadEval?&lt;/h2>
&lt;p>Radiology report generation has seen significant advancements with the advent of large language models (LLMs). However, evaluating the quality of these generated reports remains a complex challenge. Traditional metrics like BLEU and ROUGE often fall short in capturing the clinical relevance and accuracy required in medical contexts. To address this gap, we introduce RadEval, a comprehensive framework designed to evaluate radiology texts using a diverse set of metrics that encompass both linguistic quality and clinical accuracy.&lt;/p>
&lt;h2 id="key-features-of-radeval">Key Features of RadEval&lt;/h2>
&lt;ul>
&lt;li>&lt;strong>Diverse Metric Integration&lt;/strong>: RadEval consolidates a wide range of evaluation metrics, including classic n-gram overlap measures (BLEU, ROUGE), contextual embeddings (BERTScore), clinical concept-based scores (F1CheXbert, F1RadGraph, RaTEScore, SRR-BERT, TemporalEntityF1), and advanced LLM-based evaluators (GREEN). This integration allows for a multifaceted assessment of radiology reports.&lt;/li>
&lt;li>&lt;strong>Standardized Implementations&lt;/strong>: We have refined and standardized the implementations of these metrics to ensure consistency and reliability in evaluations across different studies and datasets.&lt;/li>
&lt;li>&lt;strong>Domain-Specific Enhancements&lt;/strong>: RadEval extends the GREEN metric to support multiple imaging modalities using a more lightweight model. Additionally, we have pretrained a domain-specific radiology encoder that demonstrates strong zero-shot retrieval performance, enhancing the evaluation capabilities of the framework.&lt;/li>
&lt;li>&lt;strong>Richly Annotated Expert Dataset&lt;/strong>: We provide a dataset annotated by radiology experts, containing over 450 clinically significant error labels. This dataset serves as a valuable resource for validating and benchmarking evaluation metrics against expert judgment.&lt;/li>
&lt;li>&lt;strong>Statistical Testing Tools&lt;/strong>: RadEval includes tools for statistical testing, enabling researchers to assess the significance of their results and compare different models robustly.&lt;/li>
&lt;li>&lt;strong>Baseline Model Evaluations&lt;/strong>: The framework offers baseline evaluations across multiple publicly available datasets, facilitating reproducibility and benchmarking in radiology report generation research.&lt;/li>
&lt;/ul>
&lt;h2 id="correlation-with-radiologist-judgment">Correlation with Radiologist Judgment&lt;/h2>
&lt;p>We conducted extensive experiments to assess how different metrics correlate with radiologist judgment. Our findings indicate that certain metrics, particularly those incorporating clinical concepts and LLM-based evaluations, show a stronger alignment with expert assessments. This correlation underscores the importance of using clinically informed metrics in evaluating radiology reports.&lt;/p>
&lt;h2 id="getting-started-with-radeval">Getting Started with RadEval&lt;/h2>
&lt;p>To get started with RadEval, you can access the codebase and documentation on our &lt;a href="https://github.com/jbdel/RadEval" target="_blank" rel="noopener">GitHub repository&lt;/a>. The repository includes installation instructions, usage examples, and guidelines for integrating RadEval into your evaluation pipeline. Additionally, the &lt;a href="https://huggingface.co/datasets/IAMJB/RadEvalExpertDataset" target="_blank" rel="noopener">RadEval Expert Dataset&lt;/a> is available for download, providing a rich resource for testing and validating evaluation metrics.&lt;/p>
&lt;h2 id="conclusion">Conclusion&lt;/h2>
&lt;p>RadEval represents a significant step forward in the evaluation of radiology texts, offering a comprehensive and standardized framework that addresses the unique challenges of this domain. By integrating diverse metrics, providing a richly annotated dataset, and facilitating robust benchmarking, RadEval aims to enhance the quality and reliability of radiology report generation research.&lt;/p>
&lt;h2 id="bibtex">BibTeX&lt;/h2>
&lt;div class="highlight">&lt;pre tabindex="0" class="chroma">&lt;code class="language-bibtex" data-lang="bibtex">&lt;span class="line">&lt;span class="cl">&lt;span class="nc">@inproceedings&lt;/span>&lt;span class="p">{&lt;/span>&lt;span class="nl">xu2025radeval&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">title&lt;/span>&lt;span class="p">=&lt;/span>&lt;span class="s">{RadEval: A framework for radiology text evaluation}&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">author&lt;/span>&lt;span class="p">=&lt;/span>&lt;span class="s">{Xu, Justin and Zhang, Xi and Abderezaei, Javid and Bauml, Julie and Boodoo, Roger and Haghighi, Fatemeh and Ganjizadeh, Ali and Brattain, Eric and Van Veen, Dave and Meng, Zaiqiao and others}&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">booktitle&lt;/span>&lt;span class="p">=&lt;/span>&lt;span class="s">{Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: System Demonstrations}&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">pages&lt;/span>&lt;span class="p">=&lt;/span>&lt;span class="s">{546--557}&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">year&lt;/span>&lt;span class="p">=&lt;/span>&lt;span class="s">{2025}&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="p">}&lt;/span>
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div></description></item><item><title>🎉 Paper Accepted — See You at EMNLP 2025!</title><link>https://x-izhang.github.io/post/2025emnlp/</link><pubDate>Tue, 09 Sep 2025 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/post/2025emnlp/</guid><description>&lt;p>&lt;a href="https://x-izhang.github.io/publication/xu-2025-radevalframeworkradiologytext/">&lt;em>&amp;ldquo;RadEval: A framework for radiology text evaluation&amp;rdquo;&lt;/em>&lt;/a>, has been accepted for system demonstration at &lt;a href="https://2025.emnlp.org/" target="_blank" rel="noopener">EMNLP 2025&lt;/a> with &lt;mark>Oral Presentation!&lt;/mark>&lt;/p>
&lt;p>For detailed model information and source code, please visit our &lt;mark>GitHub repository&lt;/mark>: &lt;a href="https://github.com/jbdel/RadEval" target="_blank" rel="noopener">RadEval&lt;/a>&lt;/p>
&lt;ul>
&lt;li>&lt;a href="https://pypi.org/project/RadEval/" target="_blank" rel="noopener">&lt;strong>PyPI Package&lt;/strong>&lt;/a> — Install RadEval with pip&lt;/li>
&lt;li>&lt;a href="https://huggingface.co/IAMJB/RadEvalModernBERT" target="_blank" rel="noopener">&lt;strong>HuggingFace Model&lt;/strong>&lt;/a> — Access our domain-adapted evaluation model&lt;/li>
&lt;li>&lt;a href="https://huggingface.co/spaces/X-iZhang/RadEval" target="_blank" rel="noopener">&lt;strong>Interactive Demo&lt;/strong>&lt;/a> — Try RadEval online&lt;/li>
&lt;li>&lt;a href="https://arxiv.org/abs/2509.18030v1" target="_blank" rel="noopener">&lt;strong>Research Paper&lt;/strong>&lt;/a> — Read our detailed research paper&lt;/li>
&lt;/ul>
&lt;p>📍 EMNLP 2025 - The 2025 Conference on Empirical Methods in Natural Language Processing will be held in Suzhou, China, from November 4 - 9, 2025.&lt;/p></description></item><item><title>🩺 RadEval Debuts！</title><link>https://x-izhang.github.io/post/2025radeval/</link><pubDate>Mon, 14 Jul 2025 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/post/2025radeval/</guid><description>&lt;h4 id="-revolutionizing-radiology-text-evaluation-with-ai-powered-metrics">🩺 Revolutionizing Radiology Text Evaluation with AI-Powered Metrics&lt;/h4>
&lt;p>Imagine having a comprehensive evaluation framework that doesn&amp;rsquo;t just measure surface-level text similarity, but truly understands clinical accuracy and medical semantics in radiology reports. This vision is now a reality with &lt;strong>RadEval&lt;/strong>, a groundbreaking, open-source evaluation toolkit designed specifically for AI-generated radiology text.&lt;/p>
&lt;h4 id="-all-in-one-metrics-for-evaluating-ai-generated-radiology-text">📊 All-in-one metrics for evaluating AI-generated radiology text&lt;/h4>
&lt;p>From traditional n-gram metrics to advanced LLM-based evaluations, RadEval provides 11+ different evaluation metrics in one unified framework, enabling researchers to thoroughly assess their radiology text generation models with domain-specific medical knowledge integration.&lt;/p>
&lt;p>For detailed handbook, please visit our &lt;a href="https://github.com/jbdel/RadEval" target="_blank" rel="noopener">GitHub repository&lt;/a>:
&lt;figure>&lt;img src="https://x-izhang.github.io/post/2025radeval/github.png">
&lt;/figure>
&lt;/p>
&lt;h4 id="-quick-start-demo">🚀 Quick Start Demo&lt;/h4>
&lt;p>Try RadEval instantly with our interactive &lt;a href="https://huggingface.co/spaces/X-iZhang/RadEval" target="_blank" rel="noopener">Gradio demo&lt;/a>:
&lt;figure>&lt;img src="https://x-izhang.github.io/post/2025radeval/demo.png">
&lt;/figure>
&lt;/p>
&lt;h4 id="-key-features">💡 Key Features&lt;/h4>
&lt;p>&lt;strong>RadEval&lt;/strong> stands out with its comprehensive approach to radiology text evaluation:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>🎯 Domain-Specific&lt;/strong>: Tailored for radiology with medical knowledge integration&lt;/li>
&lt;li>&lt;strong>📈 Multi-Metric&lt;/strong>: Supports lexical, semantic, clinical, and temporal evaluations&lt;/li>
&lt;li>&lt;strong>⚡ Easy to Use&lt;/strong>: Simple API with flexible configuration options&lt;/li>
&lt;li>&lt;strong>🔬 Research-Ready&lt;/strong>: Built-in statistical testing for system comparison&lt;/li>
&lt;li>&lt;strong>📦 PyPI Available&lt;/strong>: Install with a simple &lt;code>pip install RadEval&lt;/code>&lt;/li>
&lt;/ul>
&lt;h4 id="-advancing-radiology-ai-research-community">🏥 Advancing Radiology AI Research Community&lt;/h4>
&lt;p>We are committed to building a standardized and reproducible toolkit for researchers, clinicians, and developers dedicated to advancing AI evaluation in medical imaging and radiology. Together, we&amp;rsquo;re setting new standards for clinical AI assessment.&lt;/p>
&lt;ul>
&lt;li>&lt;a href="https://pypi.org/project/RadEval/" target="_blank" rel="noopener">&lt;strong>PyPI Package&lt;/strong>&lt;/a> — Install RadEval with pip&lt;/li>
&lt;li>&lt;a href="https://huggingface.co/IAMJB/RadEvalModernBERT" target="_blank" rel="noopener">&lt;strong>HuggingFace Model&lt;/strong>&lt;/a> — Access our domain-adapted evaluation model&lt;/li>
&lt;li>&lt;a href="https://huggingface.co/spaces/X-iZhang/RadEval" target="_blank" rel="noopener">&lt;strong>Interactive Demo&lt;/strong>&lt;/a> — Try RadEval online&lt;/li>
&lt;li>&lt;a href="https://arxiv.org/abs/2509.18030v1" target="_blank" rel="noopener">&lt;strong>Research Paper&lt;/strong>&lt;/a> — Read our detailed research paper&lt;/li>
&lt;/ul></description></item><item><title>🎉 Paper Accepted — See You at ACL 2025!</title><link>https://x-izhang.github.io/post/2025acl/</link><pubDate>Thu, 15 May 2025 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/post/2025acl/</guid><description>&lt;p>&lt;a href="https://x-izhang.github.io/publication/zhang-2025-libraleveragingtemporalimages/">&lt;em>&amp;ldquo;Libra: Leveraging Temporal Images for Biomedical Radiology Analysis&amp;rdquo;&lt;/em>&lt;/a> has been accepted to &lt;a href="https://2025.aclweb.org/" target="_blank" rel="noopener">ACL 2025&lt;/a>!&lt;/p>
&lt;blockquote>
&lt;p>Explore the critical role of temporal information in radiology diagnostics and how the Libra model applies it in detail at &lt;em>&lt;strong>blog:&lt;/strong>&lt;/em> &lt;a href="https://x-izhang.github.io/blog/libra-blog1/">Libra - Temporal Insight 🕰️&lt;/a>.&lt;/p>
&lt;/blockquote>
&lt;blockquote>
&lt;p>Dive into how model architecture shapes temporal reasoning capabilities in radiology imaging analysis at &lt;em>&lt;strong>blog:&lt;/strong>&lt;/em> &lt;a href="https://x-izhang.github.io/blog/libra-blog2/">Libra – Structural Logic 🧠&lt;/a>.&lt;/p>
&lt;/blockquote>
&lt;blockquote>
&lt;p>Look ahead to future directions in radiology AI, moving beyond temporal comparison at &lt;em>&lt;strong>blog:&lt;/strong>&lt;/em> &lt;a href="https://x-izhang.github.io/blog/libra-blog3/">Libra – What about next? 🛸&lt;/a>.&lt;/p>
&lt;/blockquote>
&lt;p>📍 &lt;strong>ACL 2025&lt;/strong> — The 63rd Annual Meeting of the Association for Computational Linguistics will take place in &lt;strong>Vienna, Austria&lt;/strong>, July 27–August 1st, 2025&lt;/p>
&lt;p>See you there!&lt;/p>
&lt;h3 id="-talk">📺 Talk&lt;/h3>
&lt;div style="position: relative; padding-bottom: 56.25%; height: 0; overflow: hidden;">
&lt;iframe allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" allowfullscreen="allowfullscreen" loading="eager" referrerpolicy="strict-origin-when-cross-origin" src="https://www.youtube.com/embed/_R8XUaaAU3g?autoplay=0&amp;controls=1&amp;end=0&amp;loop=0&amp;mute=0&amp;start=0" style="position: absolute; top: 0; left: 0; width: 100%; height: 100%; border:0;" title="YouTube video"
>&lt;/iframe>
&lt;/div>
&lt;h3 id="-poster">🪧 Poster&lt;/h3>
&lt;figure>&lt;img src="https://x-izhang.github.io/post/2025acl/2024acl_poster.jpg">
&lt;/figure></description></item><item><title>Libra - What about next?🛸</title><link>https://x-izhang.github.io/blog/libra-blog3/</link><pubDate>Fri, 25 Apr 2025 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/blog/libra-blog3/</guid><description>&lt;h2 id="beyond-temporal-comparison-the-future-of-radiology-modeling">Beyond Temporal Comparison: The Future of Radiology Modeling&lt;/h2>
&lt;p>In our &lt;a href="https://x-izhang.github.io/blog/libra-blog2/">previous discussions&lt;/a>, we delved into how Libra leverages temporal information through its innovative &lt;strong>Temporal Alignment Connector (TAC)&lt;/strong> to enhance radiology report generation. While this approach has shown significant promise, it&amp;rsquo;s essential to look ahead and consider how radiology modeling can evolve further to meet the complex demands of clinical practice.&lt;/p>
&lt;div class="flex px-4 py-3 mb-6 rounded-md bg-primary-100 dark:bg-primary-900">
&lt;span class="pr-3 pt-1 text-primary-600 dark:text-primary-300">
&lt;svg height="24" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24">&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="m11.25 11.25l.041-.02a.75.75 0 0 1 1.063.852l-.708 2.836a.75.75 0 0 0 1.063.853l.041-.021M21 12a9 9 0 1 1-18 0a9 9 0 0 1 18 0m-9-3.75h.008v.008H12z"/>&lt;/svg>
&lt;/span>
&lt;span class="dark:text-neutral-300">The Temporal Alignment Connector has proven effective for handling paired images, but the future of radiology AI extends far beyond just temporal comparison.&lt;/span>
&lt;/div>
&lt;h2 id="1-embracing-multimodal-integration">1. Embracing Multimodal Integration&lt;/h2>
&lt;p>Radiological diagnosis doesn&amp;rsquo;t occur in isolation. Clinicians often consider a plethora of data—ranging from patient history and laboratory results to various imaging modalities. The future of radiology modeling lies in the &lt;mark>seamless integration&lt;/mark> of these diverse data sources.&lt;/p>
&lt;h3 id="clinical-contextualization">Clinical Contextualization&lt;/h3>
&lt;ul>
&lt;li>Incorporating electronic health records (EHRs), lab results, and patient histories can provide models with a richer context&lt;/li>
&lt;li>Leading to more accurate and personalized diagnostics&lt;/li>
&lt;li>Reducing false positives and negatives through contextual awareness&lt;/li>
&lt;/ul>
&lt;h3 id="cross-modality-analysis">Cross-Modality Analysis&lt;/h3>
&lt;ul>
&lt;li>Combining data from different imaging modalities (e.g., CT, MRI, PET) offers a more comprehensive view&lt;/li>
&lt;li>Enables detection of patterns that might be missed when analyzing a single modality&lt;/li>
&lt;li>Creates synergistic understanding of complex pathologies&lt;/li>
&lt;/ul>
&lt;div class="mermaid">graph TD
A[Patient Data] --> B{Multimodal&lt;br>Integration}
C[Chest X-ray] --> B
D[CT Scan] --> B
E[Lab Results] --> B
F[Patient History] --> B
B --> G[Comprehensive&lt;br>Analysis]
G --> H[Enhanced&lt;br>Diagnostic Accuracy]
G --> I[Personalized&lt;br>Treatment Plans]
G --> J[Early Disease&lt;br>Detection]
style A fill:#f5f5f5,stroke:#333,stroke-width:1px
style B fill:#e1f5fe,stroke:#01579b,stroke-width:2px
style C fill:#f5f5f5,stroke:#333,stroke-width:1px
style D fill:#f5f5f5,stroke:#333,stroke-width:1px
style E fill:#f5f5f5,stroke:#333,stroke-width:1px
style F fill:#f5f5f5,stroke:#333,stroke-width:1px
style G fill:#e8f5e9,stroke:#2e7d32,stroke-width:2px
style H fill:#c8e6c9,stroke:#2e7d32,stroke-width:1px
style I fill:#c8e6c9,stroke:#2e7d32,stroke-width:1px
style J fill:#c8e6c9,stroke:#2e7d32,stroke-width:1px
&lt;/div>
&lt;h2 id="2-advancing-explainability-and-trustworthiness">2. Advancing Explainability and Trustworthiness&lt;/h2>
&lt;p>As AI models become more integral to clinical decision-making, their interpretability becomes paramount. Clinicians need to understand the &lt;mark>rationale behind a model&amp;rsquo;s prediction&lt;/mark> to trust and effectively utilize its insights.&lt;/p>
&lt;blockquote>
&lt;p>In healthcare, trust isn&amp;rsquo;t optional—it&amp;rsquo;s essential. An AI system that can&amp;rsquo;t explain its reasoning is a black box that most physicians will rightfully hesitate to rely on.&lt;/p>
&lt;/blockquote>
&lt;h3 id="explainable-ai-xai">Explainable AI (XAI)&lt;/h3>
&lt;ul>
&lt;li>Developing models that provide clear, human-understandable explanations for their predictions&lt;/li>
&lt;li>Bridging the gap between AI outputs and clinical reasoning&lt;/li>
&lt;li>Using attention visualization and feature attribution methods to highlight decision factors&lt;/li>
&lt;/ul>
&lt;div class="flex px-4 py-3 mb-6 rounded-md bg-yellow-100 dark:bg-yellow-900">
&lt;span class="pr-3 pt-1 text-red-400">
&lt;svg height="24" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24">&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="M12 9v3.75m-9.303 3.376c-.866 1.5.217 3.374 1.948 3.374h14.71c1.73 0 2.813-1.874 1.948-3.374L13.949 3.378c-.866-1.5-3.032-1.5-3.898 0zM12 15.75h.007v.008H12z"/>&lt;/svg>
&lt;/span>
&lt;span class="dark:text-neutral-300">Even the most accurate model will face adoption challenges if clinicians cannot verify its reasoning or understand how it arrived at its conclusions.&lt;/span>
&lt;/div>
&lt;h3 id="uncertainty-quantification">Uncertainty Quantification&lt;/h3>
&lt;p>Implementing mechanisms to convey confidence levels enables:&lt;/p>
&lt;ul>
&lt;li>Clinicians to assess the reliability of AI-assisted diagnostics&lt;/li>
&lt;li>Appropriate intervention in cases of model uncertainty&lt;/li>
&lt;li>Continuous improvement through focused retraining on uncertain cases&lt;/li>
&lt;/ul>
&lt;h2 id="3-ensuring-robustness-and-generalizability">3. Ensuring Robustness and Generalizability&lt;/h2>
&lt;p>AI models must perform reliably across diverse patient populations and clinical settings—a challenge that extends beyond academic validation to real-world implementation.&lt;/p>
&lt;h3 id="diverse-training-data">Diverse Training Data&lt;/h3>
&lt;div class="markmap" style="height: 300px;">
&lt;pre>- Building Robust Radiology AI
- Data Diversity Dimensions
- Demographic Factors
- Age groups
- Ethnic backgrounds
- Sex and gender representation
- Clinical Variables
- Disease prevalence variations
- Comorbidity patterns
- Treatment history diversity
- Technical Variability
- Multiple scanner manufacturers
- Various imaging protocols
- Quality and resolution differences
- Implementation Strategies
- Federated Learning
- Cross-institution collaboration
- Privacy-preserving techniques
- Data Augmentation
- Synthetic minority examples
- Domain randomization
- Continuous Validation
- Geographic generalization testing
- Temporal drift monitoring&lt;/pre>
&lt;/div>
&lt;h3 id="continuous-learning">Continuous Learning&lt;/h3>
&lt;ul>
&lt;li>Implementing systems that update from new clinical data&lt;/li>
&lt;li>Adapting to evolving medical knowledge and practices&lt;/li>
&lt;li>Maintaining performance as disease patterns and imaging technologies change&lt;/li>
&lt;/ul>
&lt;h2 id="4-integrating-into-clinical-workflows">4. Integrating into Clinical Workflows&lt;/h2>
&lt;p>For AI models to be truly effective, they must integrate &lt;mark>seamlessly&lt;/mark> into existing clinical workflows rather than disrupting established processes.&lt;/p>
&lt;h3 id="user-friendly-interfaces">User-Friendly Interfaces&lt;/h3>
&lt;ul>
&lt;li>Designing intuitive interfaces that present AI insights clearly&lt;/li>
&lt;li>Ensuring actionable information is immediately accessible&lt;/li>
&lt;li>Minimizing cognitive load during busy clinical sessions&lt;/li>
&lt;/ul>
&lt;h3 id="workflow-compatibility">Workflow Compatibility&lt;/h3>
&lt;p>The ideal radiology AI system should:&lt;/p>
&lt;ul>
&lt;li>Complement rather than replace radiologist expertise&lt;/li>
&lt;li>Reduce administrative burden through automatic report generation&lt;/li>
&lt;li>Prioritize cases based on urgency and findings&lt;/li>
&lt;/ul>
&lt;div class="flex px-4 py-3 mb-6 rounded-md bg-primary-100 dark:bg-primary-900">
&lt;span class="pr-3 pt-1 text-primary-600 dark:text-primary-300">
&lt;svg height="24" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24">&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="m11.25 11.25l.041-.02a.75.75 0 0 1 1.063.852l-.708 2.836a.75.75 0 0 0 1.063.853l.041-.021M21 12a9 9 0 1 1-18 0a9 9 0 0 1 18 0m-9-3.75h.008v.008H12z"/>&lt;/svg>
&lt;/span>
&lt;span class="dark:text-neutral-300">The most advanced AI system will fail if it adds steps to an already complex workflow. Success depends on making the radiologist&amp;rsquo;s job easier, not more complicated.&lt;/span>
&lt;/div>
&lt;h2 id="5-ethical-and-regulatory-considerations">5. Ethical and Regulatory Considerations&lt;/h2>
&lt;p>As AI becomes more prevalent in healthcare, addressing ethical and regulatory challenges becomes essential for responsible implementation.&lt;/p>
&lt;h3 id="data-privacy-and-security">Data Privacy and Security&lt;/h3>
&lt;ul>
&lt;li>Safeguarding patient data through robust encryption&lt;/li>
&lt;li>Ensuring compliance with regulations like HIPAA and GDPR&lt;/li>
&lt;li>Implementing federated learning approaches to minimize data sharing&lt;/li>
&lt;/ul>
&lt;h3 id="regulatory-approval">Regulatory Approval&lt;/h3>
&lt;ul>
&lt;li>Navigating the complex regulatory landscape (FDA, CE marking)&lt;/li>
&lt;li>Designing validation studies that meet regulatory requirements&lt;/li>
&lt;li>Establishing monitoring systems for post-deployment performance&lt;/li>
&lt;/ul>
&lt;h3 id="ethical-ai-development">Ethical AI Development&lt;/h3>
&lt;div class="mermaid">graph TD
A[Ethical AI Development] --> B[Fairness &amp; Bias Mitigation]
A --> C[Transparency &amp; Explainability]
A --> D[Privacy Protection]
A --> E[Human Oversight]
B --> F[Equitable Healthcare Outcomes]
C --> G[Informed Clinical Decisions]
D --> H[Patient Trust &amp; Confidentiality]
E --> I[Safe AI Implementation]
style A fill:#e1f5fe,stroke:#01579b,stroke-width:2px
style B fill:#e8f5e9,stroke:#2e7d32,stroke-width:1px
style C fill:#e8f5e9,stroke:#2e7d32,stroke-width:1px
style D fill:#e8f5e9,stroke:#2e7d32,stroke-width:1px
style E fill:#e8f5e9,stroke:#2e7d32,stroke-width:1px
style F fill:#c8e6c9,stroke:#2e7d32,stroke-width:1px
style G fill:#c8e6c9,stroke:#2e7d32,stroke-width:1px
style H fill:#c8e6c9,stroke:#2e7d32,stroke-width:1px
style I fill:#c8e6c9,stroke:#2e7d32,stroke-width:1px
&lt;/div>
&lt;h2 id="conclusion-the-road-ahead-for-libra">Conclusion: The Road Ahead for Libra&lt;/h2>
&lt;p>The journey of Libra represents a significant step forward in radiology modeling, particularly in harnessing temporal information through the TAC architecture. However, the path ahead involves:&lt;/p>
&lt;ol>
&lt;li>Expanding beyond paired chest X-rays to multiple imaging modalities&lt;/li>
&lt;li>Enhancing explainability through attention visualization and reasoning paths&lt;/li>
&lt;li>Building more robust models through diverse training strategies&lt;/li>
&lt;li>Designing intuitive interfaces for seamless clinical integration&lt;/li>
&lt;li>Navigating ethical and regulatory requirements for real-world deployment&lt;/li>
&lt;/ol>
&lt;blockquote>
&lt;p>As we continue to develop Libra and similar technologies, our focus remains on augmenting—rather than replacing—clinical expertise, creating tools that serve as trusted partners in the complex art of radiological diagnosis.&lt;/p>
&lt;/blockquote>
&lt;hr>
&lt;p>💬 &lt;strong>Note&lt;/strong>: The views expressed here are my own, reflecting my personal insights into the evolving landscape of radiology AI.&lt;/p></description></item><item><title>🎤 Talk &amp; Poster at HealTAC 2025!</title><link>https://x-izhang.github.io/post/healtac2025/</link><pubDate>Tue, 22 Apr 2025 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/post/healtac2025/</guid><description>&lt;p>We’re excited to share that we’ve been invited to present our recent research — &lt;a href="https://github.com/X-iZhang/Libra#-news" target="_blank" rel="noopener">&lt;em>&lt;strong>&amp;ldquo;Towards Temporal-Aware Multimodal Large Language Models for Improved Radiology Report Generation&amp;rdquo;&lt;/strong>&lt;/em>&lt;/a> — as a &lt;strong>lightning talk&lt;/strong> and &lt;strong>poster presentation&lt;/strong> at the &lt;strong>PhD Forum&lt;/strong> of &lt;a href="https://healtac2025.github.io/" target="_blank" rel="noopener">&lt;strong>HealTAC 2025&lt;/strong>&lt;/a>!&lt;/p>
&lt;figure>&lt;img src="https://x-izhang.github.io/post/healtac2025/ppt_17.jpg">
&lt;/figure>
&lt;p>📍 &lt;strong>HealTAC 2025&lt;/strong> — The 8th Healthcare Text Analytics Conference will take place in &lt;strong>Glasgow&lt;/strong>, 16–18 June 2025.&lt;/p>
&lt;p>See you there!&lt;/p></description></item><item><title>Libra – Structural Logic 🧠</title><link>https://x-izhang.github.io/blog/libra-blog2/</link><pubDate>Tue, 15 Apr 2025 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/blog/libra-blog2/</guid><description>&lt;h2 id="the-challenge-of-temporal-reasoning-in-radiology-ai">The Challenge of Temporal Reasoning in Radiology AI&lt;/h2>
&lt;p>When dealing with radiology images, especially in the context of temporal analysis—comparing current chest X-rays with previous images—standard neural network architectures often struggle. Although transformer-based multimodal large language models (&lt;strong>MLLMs&lt;/strong>) like &lt;a href="https://github.com/haotian-liu/LLaVA" target="_blank" rel="noopener">&lt;strong>LLaVA&lt;/strong>&lt;/a> demonstrate remarkable capabilities for understanding single images and textual information, they encounter substantial challenges when handling &lt;em>&lt;strong>image pairs&lt;/strong>&lt;/em>.&lt;/p>
&lt;div class="flex px-4 py-3 mb-6 rounded-md bg-primary-100 dark:bg-primary-900">
&lt;span class="pr-3 pt-1 text-primary-600 dark:text-primary-300">
&lt;svg height="24" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24">&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="m11.25 11.25l.041-.02a.75.75 0 0 1 1.063.852l-.708 2.836a.75.75 0 0 0 1.063.853l.041-.021M21 12a9 9 0 1 1-18 0a9 9 0 0 1 18 0m-9-3.75h.008v.008H12z"/>&lt;/svg>
&lt;/span>
&lt;span class="dark:text-neutral-300">In my &lt;a href="https://x-izhang.github.io/blog/libra-blog1/">previous blog&lt;/a>, I discussed in detail why a &lt;mark>single&lt;/mark> prior chest X-ray is typically sufficient for accurate diagnosis and patient triage.&lt;/span>
&lt;/div>
&lt;blockquote>
&lt;p>However, capturing meaningful temporal differences between two images remains problematic with traditional transformer structures.&lt;/p>
&lt;/blockquote>
&lt;h2 id="when-transformers-lose-the-plot-why-they-struggle-with-temporal">When Transformers Lose the Plot: Why They Struggle with Temporal&lt;/h2>
&lt;p>The transformer, the cornerstone of modern large language models (LLMs), excels at &lt;mark>sequential&lt;/mark> data processing and logical reasoning tasks. Its strength lies in handling complex linguistic structures through &lt;strong>positional encoding&lt;/strong>, enabling nuanced relationships in textual sequences.&lt;/p>
&lt;p>However, when transformers receive visual information—particularly multiple images presented simultaneously—the situation becomes more complicated. Existing methods typically &lt;mark>concatenate&lt;/mark> image features directly into the LLM&amp;rsquo;s head, often via sequences containing hundreds of visual tokens (patch tokens), depending on the specific image encoder used.&lt;/p>
&lt;div class="flex px-4 py-3 mb-6 rounded-md bg-yellow-100 dark:bg-yellow-900">
&lt;span class="pr-3 pt-1 text-red-400">
&lt;svg height="24" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24">&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="M12 9v3.75m-9.303 3.376c-.866 1.5.217 3.374 1.948 3.374h14.71c1.73 0 2.813-1.874 1.948-3.374L13.949 3.378c-.866-1.5-3.032-1.5-3.898 0zM12 15.75h.007v.008H12z"/>&lt;/svg>
&lt;/span>
&lt;span class="dark:text-neutral-300">This straightforward approach inevitably suffers from token overload, known colloquially as the &amp;ldquo;&lt;strong>lost-in-the-middle&lt;/strong>&amp;rdquo; problem, meaning crucial temporal details may get diluted or overlooked.&lt;/span>
&lt;/div>
&lt;p>Indeed, current MLLMs like LLaVA perform impressively with single-image inputs. But they quickly become overwhelmed with paired images, heavily relying on meticulously crafted instruction datasets to guide temporal comparisons explicitly:&lt;/p>
&lt;h3 id="how-mllms-are-prompted-to-compare-images">How MLLMs Are Prompted to Compare Images&lt;/h3>
&lt;p>&amp;ldquo;What is the difference between &lt;mark>&amp;lt;image-1-patchholder&amp;gt;&lt;/mark> and &lt;mark>&amp;lt;image-2-patchholder&amp;gt;&lt;/mark>?&amp;rdquo;&lt;/p>
&lt;p>Such approaches place the burden squarely on the LLM&amp;rsquo;s internal reasoning and positional encodings, complicating training and diminishing reliability. The model must:&lt;/p>
&lt;ul>
&lt;li>Distinguish between multiple images using only position encodings&lt;/li>
&lt;li>Process 500+ tokens per image (depending on patch number)&lt;/li>
&lt;li>Compare features across long token distances&lt;/li>
&lt;/ul>
&lt;p>Given these limitations, an essential question arises:&lt;/p>
&lt;blockquote>
&lt;p>Can we overcome these temporal reasoning challenges structurally, &lt;strong>rather than&lt;/strong> through &lt;mark>explicit&lt;/mark> prompting?&lt;/p>
&lt;/blockquote>
&lt;h2 id="structure-determines-function-insights-from-biology">Structure Determines Function: Insights from Biology&lt;/h2>
&lt;p>Before we answer above question, let&amp;rsquo;s briefly reflect on the foundational relationship between structure and function—deeply ingrained in biological systems.&lt;/p>
&lt;h3 id="macro-scale-examples">Macro-scale examples:&lt;/h3>
&lt;ul>
&lt;li>Birds have &lt;strong>wings&lt;/strong> enabling flight&lt;/li>
&lt;li>Fish possess &lt;strong>gills&lt;/strong> allowing them to breathe underwater&lt;/li>
&lt;/ul>
&lt;h3 id="micro-scale-examples">Micro-scale examples:&lt;/h3>
&lt;ul>
&lt;li>The unique &lt;strong>three-dimensional helical structure&lt;/strong> of proteins directly determines their biological roles&lt;/li>
&lt;li>A virus&amp;rsquo;s &lt;strong>outer shell&lt;/strong> dictates its infection pathways and interaction mechanisms&lt;/li>
&lt;/ul>
&lt;blockquote>
&lt;p>&lt;strong>Clearly, function is fundamentally dependent on structure.&lt;/strong>&lt;/p>
&lt;/blockquote>
&lt;div class="mermaid">graph TD
A[Structure] -->|Enables| B[Function]
B -->|Guides Design of| C[New Structures]
C -->|Enhances| B
style A fill:#e1f5fe,stroke:#01579b,stroke-width:2px
style B fill:#e8f5e9,stroke:#2e7d32,stroke-width:2px
style C fill:#fff8e1,stroke:#ff8f00,stroke-width:2px
&lt;/div>
&lt;p>When designing novel neural architectures or modules, we must apply this principle:&lt;/p>
&lt;ol>
&lt;li>Identify the desired &lt;strong>functionality&lt;/strong> first&lt;/li>
&lt;li>Then craft an appropriate &lt;strong>structural design&lt;/strong> that inherently &lt;mark>supports&lt;/mark> these functions&lt;/li>
&lt;/ol>
&lt;h2 id="libras-structural-innovation">Libra&amp;rsquo;s Structural Innovation&lt;/h2>
&lt;h3 id="temporal-alignment-connector-tac">Temporal Alignment Connector (TAC)&lt;/h3>
&lt;p>Following this logic, we developed the TAC in our Libra model. TAC&amp;rsquo;s primary goal is to automatically and effectively capture the relationship between two chest X-ray images—the &lt;mark>current image&lt;/mark> (primary) and a &lt;mark>prior image&lt;/mark> (auxiliary).&lt;/p>
&lt;figure>&lt;img src="https://x-izhang.github.io/blog/libra-blog2/TAC.png"
alt="Libra&amp;rsquo;s Temporal Alignment Connector (TAC) architecture.">&lt;figcaption>
&lt;p>Libra&amp;rsquo;s Temporal Alignment Connector (TAC) architecture.&lt;/p>
&lt;/figcaption>
&lt;/figure>
&lt;p>Unlike traditional transformers that treat all inputs equivalently, TAC explicitly structures interactions between paired images. It captures their nuanced relationship through two key modules:&lt;/p>
&lt;div class="markmap" style="height: 350px;">
&lt;pre>- TAC Architecture
- Layerwise Feature Extractor (LFE)
- Aggregates visual features across multiple encoder layers
- Ensures rich representations from both images
- Maintains feature hierarchy information
- Temporal Fusion Module (TFM)
- Fuses features from current and prior images
- Highlights critical temporal differences
- Maintains clear image role assignment
- Current image (Primary)
- Prior image (Reference)
- Prefix Bias Mechanism
- Addresses nearly-identical image pairs
- Prevents attention collapse
- Differentiates prior image's contextual influence&lt;/pre>
&lt;/div>
&lt;p>An important structural consideration is the integration of a &lt;strong>prefix bias mechanism&lt;/strong>. This component addresses the scenario where current and prior images are nearly identical—common in clinical practice. Without careful design, such similarity can cause attention mechanisms to collapse into redundant self-attention loops.&lt;/p>
&lt;div class="flex px-4 py-3 mb-6 rounded-md bg-primary-100 dark:bg-primary-900">
&lt;span class="pr-3 pt-1 text-primary-600 dark:text-primary-300">
&lt;svg height="24" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24">&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="m11.25 11.25l.041-.02a.75.75 0 0 1 1.063.852l-.708 2.836a.75.75 0 0 0 1.063.853l.041-.021M21 12a9 9 0 1 1-18 0a9 9 0 0 1 18 0m-9-3.75h.008v.008H12z"/>&lt;/svg>
&lt;/span>
&lt;span class="dark:text-neutral-300">The prefix bias mitigates this risk by clearly differentiating the prior image&amp;rsquo;s contextual influence, ensuring meaningful training and robust inference.&lt;/span>
&lt;/div>
&lt;h2 id="why-structure-matters-the-libra-advantage">Why Structure Matters: The Libra Advantage&lt;/h2>
&lt;p>By structurally encoding temporal relationships directly into the neural network&amp;rsquo;s architecture, Libra overcomes the limitations inherent in traditional prompting-based approaches. Instead of forcing the LLM to implicitly infer temporal differences through complex positional encodings and exhaustive instruction tuning, &lt;strong>TAC explicitly and efficiently captures this essential clinical context.&lt;/strong>&lt;/p>
&lt;blockquote>
&lt;p>Libra exemplifies the powerful concept that structural logic, thoughtfully aligned with functional requirements, dramatically enhances model performance.&lt;/p>
&lt;/blockquote>
&lt;p>This structural logic not only simplifies training but also improves:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>Reliability&lt;/strong>: More consistent temporal reasoning&lt;/li>
&lt;li>&lt;strong>Interpretability&lt;/strong>: Clearer connection between features and outputs&lt;/li>
&lt;li>&lt;strong>Efficiency&lt;/strong>: Reduced dependence on instruction tuning&lt;/li>
&lt;li>&lt;strong>Clinical Alignment&lt;/strong>: Better reflection of radiologists&amp;rsquo; actual workflow&lt;/li>
&lt;/ul>
&lt;hr>
&lt;p>🏄 &lt;strong>Note&lt;/strong>: The opinions shared here reflect my own understanding and are intended to convey the structural logic behind Libra. For technical accuracy and complete details, please refer to our paper: &lt;a href="https://x-izhang.github.io/publication/zhang-2025-libraleveragingtemporalimages/">&lt;em>&amp;ldquo;Libra: Leveraging Temporal Images for Biomedical Radiology Analysis&amp;rdquo;&lt;/em>&lt;/a>.&lt;/p></description></item><item><title>Libra - Temporal Insight 🕰️</title><link>https://x-izhang.github.io/blog/libra-blog1/</link><pubDate>Sat, 05 Apr 2025 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/blog/libra-blog1/</guid><description>&lt;h2 id="what-does-temporal-really-mean">What Does &amp;ldquo;Temporal&amp;rdquo; Really Mean?&lt;/h2>
&lt;p>In clinical radiology, temporal information is not just about &lt;strong>&amp;ldquo;past&amp;rdquo;&lt;/strong> and &lt;strong>&amp;ldquo;present&amp;rdquo;&lt;/strong> — it&amp;rsquo;s about &lt;mark>&lt;em>change&lt;/em>&lt;/mark>. When radiologists assess a chest X-ray, they&amp;rsquo;re not merely describing what they see in a single image; they&amp;rsquo;re often comparing it to a previous one to identify whether a patient&amp;rsquo;s condition has &lt;mark>improved&lt;/mark>, &lt;mark>worsened&lt;/mark>, or &lt;mark>remained stable&lt;/mark>.&lt;/p>
&lt;p>🔔 This kind of temporal reasoning is essential in everyday medical practice. Yet most multimodal large language models (MLLMs) either ignore it or fail to model it effectively.&lt;/p>
&lt;div class="flex px-4 py-3 mb-6 rounded-md bg-primary-100 dark:bg-primary-900">
&lt;span class="pr-3 pt-1 text-primary-600 dark:text-primary-300">
&lt;svg height="24" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24">&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="m11.25 11.25l.041-.02a.75.75 0 0 1 1.063.852l-.708 2.836a.75.75 0 0 0 1.063.853l.041-.021M21 12a9 9 0 1 1-18 0a9 9 0 0 1 18 0m-9-3.75h.008v.008H12z"/>&lt;/svg>
&lt;/span>
&lt;span class="dark:text-neutral-300">Temporal information in radiology is fundamentally about capturing &lt;strong>change over time&lt;/strong>, not simply collecting a series of static images. This concept is central to Libra&amp;rsquo;s design philosophy.&lt;/span>
&lt;/div>
&lt;h2 id="time-tells-the-truth-interpreting-temporal-changes-in-imaging">Time Tells the Truth: Interpreting Temporal Changes in Imaging&lt;/h2>
&lt;h3 id="1-macro-level-progression">1. Macro-Level Progression&lt;/h3>
&lt;ul>
&lt;li>&lt;strong>Is the patient improving, deteriorating, or stable?&lt;/strong>&lt;/li>
&lt;/ul>
&lt;p>Macro-level comparison focuses on the overall trajectory of the patient’s condition compared to prior examinations. This high-level temporal reasoning is crucial for tracking disease evolution and guiding clinical decision-making.&lt;/p>
&lt;h3 id="2-lesion-specific-temporal-changes">2. Lesion-Specific Temporal Changes&lt;/h3>
&lt;ul>
&lt;li>&lt;strong>How are individual abnormalities evolving over time?&lt;/strong>&lt;/li>
&lt;/ul>
&lt;p>Fine-grained analysis captures precise changes in specific findings, such as “consolidation in the left lower lobe has significantly expanded” or “cardiac silhouette shows no appreciable change.” These insights enable clinicians and models to reason at the level of targeted anatomical and pathological detail.&lt;/p>
&lt;h3 id="3-quality-over-quantity-in-temporal-inputs">3. Quality Over Quantity in Temporal Inputs&lt;/h3>
&lt;ul>
&lt;li>❗️ Adding more images often introduces noise and computational complexity without improving diagnostic value.&lt;/li>
&lt;/ul>
&lt;div class="flex px-4 py-3 mb-6 rounded-md bg-yellow-100 dark:bg-yellow-900">
&lt;span class="pr-3 pt-1 text-red-400">
&lt;svg height="24" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24">&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="M12 9v3.75m-9.303 3.376c-.866 1.5.217 3.374 1.948 3.374h14.71c1.73 0 2.813-1.874 1.948-3.374L13.949 3.378c-.866-1.5-3.032-1.5-3.898 0zM12 15.75h.007v.008H12z"/>&lt;/svg>
&lt;/span>
&lt;span class="dark:text-neutral-300">While it may seem intuitive that incorporating more historical images would yield better results, the reality of clinical practice shows that the most valuable temporal information comes from comparing &lt;mark>just two key points in time&lt;/mark>.&lt;/span>
&lt;/div>
&lt;h2 id="clinical-applications">Clinical Applications&lt;/h2>
&lt;blockquote>
&lt;p>From Diagnostic Judgement to Multi-Scale Temporal Understanding&lt;/p>
&lt;/blockquote>
&lt;h3 id="temporal-reasoning-in-clinical-diagnosis">Temporal Reasoning in Clinical Diagnosis&lt;/h3>
&lt;p>Radiologists rarely analyse a chest X-ray in &lt;strong>isolation&lt;/strong>. Instead, they routinely ask:&lt;/p>
&lt;ul>
&lt;li>&amp;ldquo;Has the consolidation improved since last week?&amp;rdquo;&lt;/li>
&lt;li>&amp;ldquo;Is the pleural effusion new?&amp;rdquo;&lt;/li>
&lt;li>&amp;ldquo;Has the cardiac silhouette changed?&amp;rdquo;&lt;/li>
&lt;/ul>
&lt;p>Such reasoning typically falls into three primary categories:&lt;/p>
&lt;ul>
&lt;li>&lt;mark>&lt;strong>Improved&lt;/strong>&lt;/mark>: Lesions have shrunk or resolved.&lt;/li>
&lt;li>&lt;mark>&lt;strong>Worsened&lt;/strong>&lt;/mark>: New abnormalities appear, or existing ones have grown.&lt;/li>
&lt;li>&lt;mark>&lt;strong>Stable&lt;/strong>&lt;/mark>: No meaningful change is observed.&lt;/li>
&lt;/ul>
&lt;p>These are &lt;em>&lt;strong>coarse-level&lt;/strong>&lt;/em> temporal descriptions. On a finer level, radiologists describe:&lt;/p>
&lt;ul>
&lt;li>How much a lesion has changed in size or density,&lt;/li>
&lt;li>Whether opacities have shifted,&lt;/li>
&lt;li>If tubes, lines, or devices have been added or removed.&lt;/li>
&lt;/ul>
&lt;h3 id="temporal-signals-span-multiple-scales">Temporal Signals Span Multiple Scales&lt;/h3>
&lt;blockquote>
&lt;p>Temporal information in radiology is inherently multi-scale — ranging from global clinical trajectories to subtle, localised anatomical changes.&lt;/p>
&lt;/blockquote>
&lt;div class="markmap" style="height: 400px;">
&lt;pre>- Temporal Information in Radiology
- Clinical Categories
- Improved
- Lesions have shrunk
- Opacities decreased
- Inflammatory shadows reduced
- Worsened
- New abnormalities appeared
- Existing lesions grown
- New infiltrates or effusion
- Stable
- No meaningful change
- Chronic conditions
- Continuous monitoring needed
- Information Levels
- Macro-level trends
- Overall patient trajectory
- Global comparison
- Local lesion changes
- Size changes
- Density variations
- Positional shifts
- Clinical Applications
- Triage decisions
- Emergency prioritization
- Resource allocation
- Treatment evaluation
- Response assessment
- Therapy adjustment
- Long-term monitoring
- Chronic disease management
- Post-surgical follow-up&lt;/pre>
&lt;/div>
&lt;h2 id="triage-and-resource-allocation">Triage and Resource Allocation&lt;/h2>
&lt;blockquote>
&lt;p>Prioritising Care When Every Minute Counts&lt;/p>
&lt;/blockquote>
&lt;h3 id="clinical-goals-and-operational-pressures">Clinical Goals and Operational Pressures&lt;/h3>
&lt;p>Chest X-rays play a pivotal role in patient triage, especially in emergency and high-volume settings. Radiologists must rapidly determine:&lt;/p>
&lt;ul>
&lt;li>Which patients require immediate intervention,&lt;/li>
&lt;li>Who can safely wait,&lt;/li>
&lt;li>And how to allocate limited resources most effectively.&lt;/li>
&lt;/ul>
&lt;p>Triage is fundamentally about &lt;strong>maximising outcomes under constraint&lt;/strong>. The goal is not to fully characterise every patient&amp;rsquo;s history, but to make fast, high-impact decisions that ensure critical cases receive timely care—without neglecting those with non-urgent needs.&lt;/p>
&lt;h3 id="what-matters-most-clinically-significant-change">What Matters Most: Clinically Significant Change&lt;/h3>
&lt;p>In these time-sensitive settings, &lt;strong>timeliness and diagnostic clarity&lt;/strong> outweigh completeness. Radiologists focus on:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>New acute findings&lt;/strong>: Abnormalities not previously seen that may indicate emerging crises.&lt;/li>
&lt;li>&lt;strong>Significant deterioration&lt;/strong>: Rapid worsening of known conditions that may demand escalated care.&lt;/li>
&lt;li>&lt;strong>Stable chronic findings&lt;/strong>: Ongoing issues that show no meaningful progression and can be managed routinely.&lt;/li>
&lt;/ul>
&lt;div class="flex px-4 py-3 mb-6 rounded-md bg-primary-100 dark:bg-primary-900">
&lt;span class="pr-3 pt-1 text-primary-600 dark:text-primary-300">
&lt;svg height="24" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24">&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="m11.25 11.25l.041-.02a.75.75 0 0 1 1.063.852l-.708 2.836a.75.75 0 0 0 1.063.853l.041-.021M21 12a9 9 0 1 1-18 0a9 9 0 0 1 18 0m-9-3.75h.008v.008H12z"/>&lt;/svg>
&lt;/span>
&lt;span class="dark:text-neutral-300">The most actionable form of temporal information in triage is the &lt;strong>clinically meaningful change&lt;/strong> between now and the previous relevant image.&lt;br>
This focused comparison empowers decision-making without overwhelming clinicians with unnecessary historical data.&lt;/span>
&lt;/div>
&lt;h2 id="our-model-approach">Our Model Approach&lt;/h2>
&lt;blockquote>
&lt;p>Temporal Efficiency Aligned with Clinical Reasoning&lt;/p>
&lt;/blockquote>
&lt;h3 id="why-two-images-are-sufficient">Why Two Images Are Sufficient&lt;/h3>
&lt;p>In real-world radiology workflows, the most informative temporal comparison is typically between:&lt;/p>
&lt;ol>
&lt;li>The &lt;mark>&lt;strong>current&lt;/strong>&lt;/mark> chest X-ray, and&lt;/li>
&lt;li>The &lt;mark>&lt;strong>most recent&lt;/strong>&lt;/mark> prior image used for diagnosis.&lt;/li>
&lt;/ol>
&lt;p>While patients may have a rich archive of historical scans, only the immediately preceding diagnostic image provides the relevant baseline for interpreting new findings. Additional older images may support longitudinal studies, but they often introduce &lt;strong>noise, redundancy, and delay&lt;/strong> in fast-paced clinical decision-making.&lt;/p>
&lt;h3 id="libras-design-philosophy-focused-temporal-reasoning">Libra&amp;rsquo;s Design Philosophy: Focused Temporal Reasoning&lt;/h3>
&lt;p>Libra is built on this clinically grounded principle.&lt;br>
Instead of processing full temporal sequences — which can be computationally expensive and semantically ambiguous — Libra learns to model &lt;strong>directional change&lt;/strong> between two key time points.&lt;/p>
&lt;p>This design enables the model to:&lt;/p>
&lt;ul>
&lt;li>Mimic the focused comparison strategies of expert radiologists,&lt;/li>
&lt;li>Avoid temporal noise from irrelevant or outdated scans,&lt;/li>
&lt;li>Reduce computational load while preserving diagnostic fidelity.&lt;/li>
&lt;/ul>
&lt;p>In short, Libra treats temporal reasoning as radiologists do:&lt;br>
&lt;mark>&lt;strong>What’s changed since the last meaningful image?&lt;/strong>&lt;/mark>&lt;/p>
&lt;h2 id="illustrations">Illustrations&lt;/h2>
&lt;p>To better understand how radiologists and Libra approach temporal comparison, we present two illustrative workflows: a general conceptual flow and a specific clinical case.&lt;/p>
&lt;h3 id="conceptual-workflow">Conceptual Workflow&lt;/h3>
&lt;p>The following diagram outlines the reasoning pathway taken when comparing a current chest X-ray to the most recent prior image. Based on observed changes (or lack thereof), radiologists infer the clinical trajectory and guide downstream decisions.&lt;/p>
&lt;div class="mermaid">graph TD
A[Previous Chest X-ray] --> B{Current vs Previous&lt;br>Analysis}
B -->|Opacity Size Decrease| C[Improved]
B -->|New Infiltrates or Growth| D[Worsened]
B -->|No Observable Change| E[Stable]
C --> F[Recovery or&lt;br>Treatment Response]
D --> G[Disease&lt;br>Progression]
E --> H[Continuous&lt;br>Monitoring Needed]
style A fill:#f5f5f5,stroke:#333,stroke-width:1px
style B fill:#e1f5fe,stroke:#01579b,stroke-width:2px
style C fill:#e8f5e9,stroke:#2e7d32,stroke-width:2px
style D fill:#ffebee,stroke:#c62828,stroke-width:2px
style E fill:#fff8e1,stroke:#ff8f00,stroke-width:2px
style F fill:#c8e6c9,stroke:#2e7d32,stroke-width:1px
style G fill:#ffcdd2,stroke:#c62828,stroke-width:1px
style H fill:#ffecb3,stroke:#ff8f00,stroke-width:1px
&lt;/div>
&lt;h3 id="case-example-lung-consolidation">Case Example: Lung Consolidation&lt;/h3>
&lt;p>This diagram demonstrates a practical example: a patient with lung consolidation. Depending on the direction of change, clinical interpretation and management decisions vary significantly.&lt;/p>
&lt;div class="mermaid">graph TD
A[Case: Lung Consolidation] --> B{Time Point Comparison}
B -->|Consolidation Reduced&lt;br>Clearer Lung Fields| C[Improved]
B -->|Consolidation Expanded&lt;br>New Pleural Effusion| D[Worsened]
B -->|Consolidation Unchanged&lt;br>No New Features| E[Stable]
C --> F[Successful Antibiotic&lt;br>Treatment]
D --> G[Disease Progression&lt;br>Treatment Adjustment Needed]
E --> H[Continue Current&lt;br>Management Plan]
style A fill:#f5f5f5,stroke:#333,stroke-width:1px
style B fill:#e1f5fe,stroke:#01579b,stroke-width:2px
style C fill:#e8f5e9,stroke:#2e7d32,stroke-width:2px
style D fill:#ffebee,stroke:#c62828,stroke-width:2px
style E fill:#fff8e1,stroke:#ff8f00,stroke-width:2px
style F fill:#c8e6c9,stroke:#2e7d32,stroke-width:1px
style G fill:#ffcdd2,stroke:#c62828,stroke-width:1px
style H fill:#ffecb3,stroke:#ff8f00,stroke-width:1px
&lt;/div>
&lt;h2 id="ai-model-implications">AI Model Implications&lt;/h2>
&lt;blockquote>
&lt;p>Why Most MLLMs Fall Short — and How Libra Goes Further&lt;/p>
&lt;/blockquote>
&lt;p>Many multimodal large language models (MLLMs) struggle with &lt;strong>temporal reasoning&lt;/strong> for three key reasons:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>No awareness of time&lt;/strong>: They treat images independently and ignore their chronological order.&lt;/li>
&lt;li>&lt;strong>Hallucinated references&lt;/strong>: They fabricate prior findings without reliable comparison.&lt;/li>
&lt;li>&lt;strong>Lack of temporal alignment&lt;/strong>: They have no built-in mechanism to align or contrast image features across time.&lt;/li>
&lt;/ul>
&lt;blockquote>
&lt;p>&lt;mark>Libra tackles these issues head-on.&lt;/mark>&lt;/p>
&lt;/blockquote>
&lt;p>Instead of prompting the model to &amp;ldquo;guess&amp;rdquo; what might have changed, Libra incorporates &lt;strong>explicit temporal awareness&lt;/strong> into both its architecture and training process.&lt;/p>
&lt;p>We feed the model structured, temporally aligned visual features extracted from the &lt;strong>current&lt;/strong> and &lt;strong>previous&lt;/strong> images. This enables Libra to:&lt;/p>
&lt;ul>
&lt;li>Detect fine-grained, clinically meaningful changes,&lt;/li>
&lt;li>Avoid hallucination,&lt;/li>
&lt;li>And reason about progression or stability in a way that mirrors clinical thinking.&lt;/li>
&lt;/ul>
&lt;hr>
&lt;p>In the next section, &lt;a href="https://x-izhang.github.io/blog/libra-blog2/">&lt;strong>Libra – Structural Logic 🧠&lt;/strong>&lt;/a>, we’ll explore how Libra’s architecture is intentionally designed to reflect clinical reasoning. You&amp;rsquo;ll see how its modular structure enables it to reason across both &lt;strong>time&lt;/strong> and &lt;strong>image features&lt;/strong> with precision — building on the foundation established in &lt;strong>Libra – Temporal Insight 🕰️&lt;/strong>.&lt;/p></description></item><item><title>ReXrank</title><link>https://x-izhang.github.io/project/rexrank/</link><pubDate>Mon, 24 Mar 2025 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/project/rexrank/</guid><description>&lt;p>ReXrank is a public leaderboard for AI-powered radiology report generation from chest x-ray images.&lt;/p></description></item><item><title>Libra: Leveraging Temporal Images for Biomedical Radiology Analysis</title><link>https://x-izhang.github.io/publication/zhang-2025-libraleveragingtemporalimages/</link><pubDate>Wed, 01 Jan 2025 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/publication/zhang-2025-libraleveragingtemporalimages/</guid><description>&lt;div class="flex px-4 py-3 mb-6 rounded-md bg-primary-100 dark:bg-primary-900">
&lt;span class="pr-3 pt-1 text-primary-600 dark:text-primary-300">
&lt;svg height="24" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24">&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="m11.25 11.25l.041-.02a.75.75 0 0 1 1.063.852l-.708 2.836a.75.75 0 0 0 1.063.853l.041-.021M21 12a9 9 0 1 1-18 0a9 9 0 0 1 18 0m-9-3.75h.008v.008H12z"/>&lt;/svg>
&lt;/span>
&lt;span class="dark:text-neutral-300">For the latest updates and details, visit the &lt;a href="https://x-izhang.github.io/Libra_v1.0/" target="_blank" rel="noopener">Libra Project Website&lt;/a>.&lt;/span>
&lt;/div>
&lt;h2 id="why-libra">Why Libra?&lt;/h2>
&lt;figure>&lt;img src="https://x-izhang.github.io/publication/zhang-2025-libraleveragingtemporalimages/figure_1.png">
&lt;/figure>
&lt;p>Temporal hallucination is a critical challenge in radiology report generation (RRG). Traditional multimodal large language models (MLLMs) struggle to integrate prior images correctly, often generating:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>Spurious references&lt;/strong> to nonexistent prior studies (Single-image case).&lt;/li>
&lt;li>&lt;strong>Inaccurate interpretations&lt;/strong> of disease progression (Temporal-image case).&lt;/li>
&lt;/ul>
&lt;p>&lt;strong>Libra&lt;/strong> addresses these limitations by integrating a &lt;strong>Temporal Alignment Connector (TAC)&lt;/strong> to improve temporal awareness, ensuring:&lt;/p>
&lt;blockquote>
&lt;ul>
&lt;li>Prior studies are correctly referenced only when available.&lt;/li>
&lt;li>Hallucinated references are eliminated, avoiding misleading reports.&lt;/li>
&lt;li>Temporal changes are accurately captured, ensuring clinically meaningful outputs.&lt;/li>
&lt;/ul>
&lt;/blockquote>
&lt;h2 id="overview">Overview&lt;/h2>
&lt;p>We propose &lt;strong>Libra&lt;/strong> (&lt;strong>L&lt;/strong>everaging Temporal &lt;strong>I&lt;/strong>mages for &lt;strong>B&lt;/strong>iomedical &lt;strong>R&lt;/strong>adiology &lt;strong>A&lt;/strong>nalysis), a novel framework tailored for radiology report generation (RRG) that incorporates temporal change information to address the challenges of interpreting medical images effectively.&lt;/p>
&lt;p>Libra leverages RAD-DINO, a pre-trained visual transformer, as its image encoder to generate robust and scalable image features. These features are further refined by a &lt;strong>Temporal Alignment Connector (TAC)&lt;/strong>, a key innovation in Libra&amp;rsquo;s architecture. The TAC comprises:&lt;/p>
&lt;ul>
&lt;li>
&lt;p>&lt;strong>Layerwise Feature Extractor (LFE)&lt;/strong>: Captures high-granularity image feature embeddings from the encoder.&lt;/p>
&lt;/li>
&lt;li>
&lt;p>&lt;strong>Temporal Fusion Module (TFM)&lt;/strong>: Integrates temporal references from prior studies to enhance temporal awareness and reasoning.&lt;/p>
&lt;/li>
&lt;/ul>
&lt;p>These refined features are fed into Meditron, a specialised medical large language model (LLM), to generate comprehensive, temporally-aware radiology reports. Libra’s modular design seamlessly integrates state-of-the-art open-source pre-trained models for both image and text, aligning them through a temporal-aware adapter to ensure robust cross-modal reasoning and understanding.&lt;/p>
&lt;p>Through a two-stage training strategy, Libra demonstrates the powerful potential of multimodal large language models (MLLMs) in specialised radiology applications. Extensive experiments on the &lt;strong>MIMIC-CXR dataset&lt;/strong> highlight Libra&amp;rsquo;s performance, setting a new state-of-the-art benchmark among models of the same parameter scale.&lt;/p>
&lt;h2 id="key-contributions">Key Contributions&lt;/h2>
&lt;ul>
&lt;li>
&lt;p>&lt;strong>Temporal Awareness:&lt;/strong> Libra captures and synthesizes temporal changes in medical images, addressing the challenge of handling prior study citations in RRG tasks.&lt;/p>
&lt;/li>
&lt;li>
&lt;p>&lt;strong>Innovative Architecture:&lt;/strong> The Temporal Alignment Connector (TAC) ensures high-granularity feature extraction and temporal integration, significantly enhancing cross-modal reasoning capabilities.&lt;/p>
&lt;/li>
&lt;li>
&lt;p>&lt;strong>State-of-the-Art Performance:&lt;/strong> Libra achieves outstanding results on the MIMIC-CXR dataset, outperforming existing MLLMs in both accuracy and temporal reasoning.&lt;/p>
&lt;/li>
&lt;li>
&lt;p>&lt;strong>Libra Repository:&lt;/strong> Our code space provides a public and detailed implementation of Libra, facilitating reproducibility and further research in the field, and integrates both training and evaluation processes, ensuring a streamlined and efficient workflow for developing and testing radiology report generation models.&lt;/p>
&lt;/li>
&lt;/ul>
&lt;h2 id="experimental-results">Experimental Results&lt;/h2>
&lt;figure>&lt;img src="https://x-izhang.github.io/publication/zhang-2025-libraleveragingtemporalimages/results.png">
&lt;/figure>
&lt;p>Libra was designed to excel in radiology report generation (RRG), and its performance speaks for itself. Here&amp;rsquo;s a quick breakdown of its achievements:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>Lexical and Clinical Metrics:&lt;/strong> Libra delivers competitive results across traditional metrics like ROUGE-L, BLEU, METEOR, and RadGraph-based scores.&lt;/li>
&lt;li>&lt;strong>Radiologist-Aligned Metrics:&lt;/strong> It leads in the RadCliQ metric and CheXbert vector similarity, showcasing its alignment with clinical expectations.&lt;/li>
&lt;li>&lt;strong>CheXpert Classification:&lt;/strong> Libra achieves the highest Macro-F1 scores and remains competitive in Micro-F1, further solidifying its reliability in clinical classification tasks.&lt;/li>
&lt;/ul>
&lt;p>By leveraging its Temporal Alignment Connector (TAC), Libra effectively captures temporal contexts, generating radiology reports that are both accurate and clinically meaningful. While there are minor gaps in select clinical metrics, Libra&amp;rsquo;s robust performance demonstrates its potential to set new standards in RRG.&lt;/p>
&lt;h2 id="performance-analysis">Performance Analysis&lt;/h2>
&lt;figure>&lt;img src="https://x-izhang.github.io/publication/zhang-2025-libraleveragingtemporalimages/figure_3.png">
&lt;/figure>
&lt;h3 id="cases-without-prior-images">Cases Without Prior Images&lt;/h3>
&lt;p>In scenarios where no prior image is available, Libra shines by delivering detailed and clinically relevant descriptions without introducing spurious references. For instance, as shown in &lt;a href="#figure-figure_3">&lt;strong>Figure 3 (a)&lt;/strong>&lt;/a>, Libra identified &amp;ldquo;sternal wires&amp;rdquo; and their type, going beyond the ground truth. This demonstrates its ability to provide meaningful insights while avoiding errors caused by nonexistent prior studies.&lt;/p>
&lt;h3 id="cases-with-prior-images">Cases With Prior Images&lt;/h3>
&lt;p>When prior images are available, Libra takes its analysis to the next level. In &lt;a href="#figure-figure_3">&lt;strong>Figure 3 (b)&lt;/strong>&lt;/a>, new abnormalities such as pleural effusion and pneumonia were identified in the current image. Without considering the prior image, Libra accurately described the findings without inferring disease progression, maintaining clinical accuracy. However, when the prior image was included, Libra effectively captured the progressive changes, provided detailed descriptions, and explicitly referenced the comparison. This capability ensures a clear understanding of temporal changes and enhances the accuracy of disease progression descriptions.&lt;/p>
&lt;h3 id="evaluating-temporal-consistency">Evaluating Temporal Consistency&lt;/h3>
&lt;p>To test Libra&amp;rsquo;s temporal reasoning, we swapped the image order, treating the prior image as the current image and vice versa. The generated report reflected an improved patient condition, aligning with the reversed input sequence but contradicting the original ground truth. Interestingly, the report closely resembled the original description of the prior image, as shown at the bottom of &lt;a href="#figure-figure_3">&lt;strong>Figure 3 (b)&lt;/strong>&lt;/a>. This highlights Libra&amp;rsquo;s ability to adapt to different temporal contexts, generating accurate and contextually consistent reports that align with standard clinical practices.&lt;/p>
&lt;p>Libra&amp;rsquo;s performance in these cases underscores its potential to revolutionize radiology report generation by seamlessly integrating temporal reasoning into its analysis.&lt;/p>
&lt;!-- &lt;figure>&lt;img src="https://x-izhang.github.io/publication/zhang-2025-libraleveragingtemporalimages/libra_card.png">
&lt;/figure>
-->
&lt;h2 id="bibtex">BibTeX&lt;/h2>
&lt;div class="highlight">&lt;pre tabindex="0" class="chroma">&lt;code class="language-bibtex" data-lang="bibtex">&lt;span class="line">&lt;span class="cl">&lt;span class="nc">@inproceedings&lt;/span>&lt;span class="p">{&lt;/span>&lt;span class="nl">zhang-etal-2025-libra&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">title&lt;/span> &lt;span class="p">=&lt;/span> &lt;span class="s">&amp;#34;Libra: Leveraging Temporal Images for Biomedical Radiology Analysis&amp;#34;&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">author&lt;/span> &lt;span class="p">=&lt;/span> &lt;span class="s">&amp;#34;Zhang, Xi and
&lt;/span>&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="s"> Meng, Zaiqiao and
&lt;/span>&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="s"> Lever, Jake and
&lt;/span>&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="s"> Ho, Edmond S. L.&amp;#34;&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">editor&lt;/span> &lt;span class="p">=&lt;/span> &lt;span class="s">&amp;#34;Che, Wanxiang and
&lt;/span>&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="s"> Nabende, Joyce and
&lt;/span>&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="s"> Shutova, Ekaterina and
&lt;/span>&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="s"> Pilehvar, Mohammad Taher&amp;#34;&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">booktitle&lt;/span> &lt;span class="p">=&lt;/span> &lt;span class="s">&amp;#34;Findings of the Association for Computational Linguistics: ACL 2025&amp;#34;&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">month&lt;/span> &lt;span class="p">=&lt;/span> &lt;span class="nv">jul&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">year&lt;/span> &lt;span class="p">=&lt;/span> &lt;span class="s">&amp;#34;2025&amp;#34;&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">address&lt;/span> &lt;span class="p">=&lt;/span> &lt;span class="s">&amp;#34;Vienna, Austria&amp;#34;&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">publisher&lt;/span> &lt;span class="p">=&lt;/span> &lt;span class="s">&amp;#34;Association for Computational Linguistics&amp;#34;&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">url&lt;/span> &lt;span class="p">=&lt;/span> &lt;span class="s">&amp;#34;https://aclanthology.org/2025.findings-acl.888/&amp;#34;&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">pages&lt;/span> &lt;span class="p">=&lt;/span> &lt;span class="s">&amp;#34;17275--17303&amp;#34;&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">ISBN&lt;/span> &lt;span class="p">=&lt;/span> &lt;span class="s">&amp;#34;979-8-89176-256-5&amp;#34;&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">abstract&lt;/span> &lt;span class="p">=&lt;/span> &lt;span class="s">&amp;#34;Radiology report generation (RRG) requires advanced medical image analysis, effective temporal reasoning, and accurate text generation. While multimodal large language models (MLLMs) align with pre-trained vision encoders to enhance visual-language understanding, most existing methods rely on single-image analysis or rule-based heuristics to process multiple images, failing to fully leverage temporal information in multi-modal medical datasets. In this paper, we introduce **Libra**, a temporal-aware MLLM tailored for chest X-ray report generation. Libra combines a radiology-specific image encoder with a novel Temporal Alignment Connector (**TAC**), designed to accurately capture and integrate temporal differences between paired current and prior images. Extensive experiments on the MIMIC-CXR dataset demonstrate that Libra establishes a new state-of-the-art benchmark among similarly scaled MLLMs, setting new standards in both clinical relevance and lexical accuracy. All source code and data are publicly available at: https://github.com/X-iZhang/Libra.&amp;#34;&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="p">}&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="nc">@inproceedings&lt;/span>&lt;span class="p">{&lt;/span>&lt;span class="nl">zhang2025libra&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">title&lt;/span>&lt;span class="p">=&lt;/span>&lt;span class="s">{Libra: Leveraging temporal images for biomedical radiology analysis}&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">author&lt;/span>&lt;span class="p">=&lt;/span>&lt;span class="s">{Zhang, Xi and Meng, Zaiqiao and Lever, Jake and Ho, Edmond SL}&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">booktitle&lt;/span>&lt;span class="p">=&lt;/span>&lt;span class="s">{Findings of the Association for Computational Linguistics: ACL 2025}&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">pages&lt;/span>&lt;span class="p">=&lt;/span>&lt;span class="s">{17275--17303}&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">year&lt;/span>&lt;span class="p">=&lt;/span>&lt;span class="s">{2025}&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="p">}&lt;/span>
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div></description></item><item><title>Libra</title><link>https://x-izhang.github.io/project/libra/</link><pubDate>Wed, 20 Nov 2024 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/project/libra/</guid><description>&lt;p>Temporally-aware MLLM for radiology report generation. Supports LLM backbones, validation, resume training, and smart saving.&lt;/p></description></item><item><title>🎉 Paper Accepted — See You at ACL 2024!</title><link>https://x-izhang.github.io/post/2024acl/</link><pubDate>Sat, 20 Jul 2024 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/post/2024acl/</guid><description>&lt;p>&lt;a href="https://x-izhang.github.io/publication/zhang-etal-2024-gla/">&lt;em>&amp;ldquo;Gla-AI4BioMed at RRG24: Visual Instruction-tuned Adaptation for Radiology Report Generation&amp;rdquo;&lt;/em>&lt;/a> has been accepted to BioNLP @ &lt;a href="https://2024.aclweb.org/" target="_blank" rel="noopener">ACL 2024&lt;/a>!&lt;/p>
&lt;p>For detailed model information and source code, please visit our &lt;mark>GitHub repository&lt;/mark>: &lt;a href="https://github.com/X-iZhang/RRG-BioNLP-ACL2024" target="_blank" rel="noopener">RRG-BioNLP-ACL2024&lt;/a>&lt;/p>
&lt;h3 id="-poster">🪧 Poster&lt;/h3>
&lt;figure>&lt;img src="https://x-izhang.github.io/post/2024acl/2024acl_poster.png">
&lt;/figure></description></item><item><title>Gla-AI4BioMed at RRG24: Visual Instruction-tuned Adaptation for Radiology Report Generation</title><link>https://x-izhang.github.io/publication/zhang-etal-2024-gla/</link><pubDate>Sat, 20 Jul 2024 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/publication/zhang-etal-2024-gla/</guid><description>&lt;h2 id="introduction">Introduction&lt;/h2>
&lt;p>Radiology reports are critical tools for interpreting medical imaging, such as chest X-rays. These reports typically consist of two main sections: &lt;code>FINDINGS&lt;/code>, which detail observations from the image, and &lt;code>IMPRESSIONS&lt;/code>, which summarize key takeaways and recommendations.&lt;/p>
&lt;p>For example:&lt;/p>
&lt;blockquote>
&lt;p>&lt;strong>FINDINGS:&lt;/strong>&lt;br>
Increase in size of the left pleural effusion compared to the prior exam. Right lung remains clear. Mildly enlarged but stable heart size.&lt;/p>
&lt;/blockquote>
&lt;blockquote>
&lt;p>&lt;strong>IMPRESSIONS:&lt;/strong>&lt;br>
Increase in left pleural effusion. Stable mild cardiomegaly.&lt;/p>
&lt;/blockquote>
&lt;p>Generating such reports automatically is a challenging task that requires aligning visual data with textual descriptions. While general-purpose visual language models like LLaVA and InstructBLIP have shown promise in multimodal tasks, radiology report generation demands a higher level of precision and domain-specific adaptation.&lt;/p>
&lt;p>Our approach focuses on fine-tuning a visual language model specifically for radiology. By aligning chest X-ray features with a large language model and applying advanced techniques like Low-Rank Adaptation (LoRA), we enhance the model&amp;rsquo;s ability to generate accurate and clinically relevant reports. Additionally, we employ a method to process multiple images simultaneously, enabling the model to capture nuanced details across different X-rays.&lt;/p>
&lt;p>This work was developed for &lt;a href="https://stanford-aimi.github.io/RRG24/" target="_blank" rel="noopener">the RRG24 Shared Task at BioNLP 2024&lt;/a>, where our model achieved competitive results, ranking &lt;strong>4th&lt;/strong> on the leaderboard. Key contributions include:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>Domain-specific fine-tuning&lt;/strong>: Optimizing the model for radiology tasks through visual instruction tuning.&lt;/li>
&lt;li>&lt;strong>Multi-image processing&lt;/strong>: Combining multiple X-rays into a single input for efficient and accurate interpretation.&lt;/li>
&lt;/ul>
&lt;h2 id="methodology">Methodology&lt;/h2>
&lt;p>Our approach builds on insights from LLaVA-Med, emphasizing the advantages of starting with a language-only pretrained LLM rather than a multimodal-trained base. The model integrates an image encoder with a learnable adapter, following the LLaVA-1.5 architecture. Training involves an auto-regressive language modeling approach using cross-entropy loss, with hyperparameters aligned to LLaVA-1.5. We employ Low-Rank Adaptation (LoRA) for efficient fine-tuning, starting with adapter pretraining for one epoch, followed by joint tuning for three epochs.&lt;/p>
&lt;figure>&lt;img src="https://x-izhang.github.io/publication/zhang-etal-2024-gla/featured.png">
&lt;/figure>
&lt;h3 id="proposed-models">Proposed Models&lt;/h3>
&lt;p>To tackle the RRG24 shared task, we developed two specialized models: &lt;strong>Med-CXRGen-F&lt;/strong> for Findings and &lt;strong>Med-CXRGen-I&lt;/strong> for Impressions. Both models leverage CLIP as the image encoder and Vicuna-1.5 as the LLM. A multi-layer perceptron (MLP) adapter with GELU activations and a uniform hidden size of 1024 processes image features before aligning them with the LLM.&lt;/p>
&lt;h3 id="how-it-works">How It Works&lt;/h3>
&lt;p>&lt;strong>Image Encoding&lt;/strong>: The image encoder converts chest X-rays into patch tokens, extracting embeddings from the penultimate layer.&lt;/p>
&lt;p>&lt;strong>Adapter Alignment&lt;/strong>: The MLP adapter processes these embeddings, aligning them to the LLM&amp;rsquo;s input format.&lt;/p>
&lt;p>&lt;strong>Prompt Design&lt;/strong>: The model is prompted with:&lt;/p>
&lt;blockquote>
&lt;p>&lt;em>Provide a description of the findings/impressions from the radiology &amp;lt;image&amp;gt;\n image.&lt;/em>&lt;/p>
&lt;/blockquote>
&lt;p>Here, &lt;em>&amp;quot;&amp;lt;image&amp;gt;&amp;quot;&lt;/em> acts as a placeholder token, guiding the LLM to generate text based on the input image, as shown &lt;a href="#figure-rrg24">in Figure&lt;/a>.&lt;/p>
&lt;h3 id="training">Training&lt;/h3>
&lt;p>Our training process involves a two-stage approach to optimize the model for radiology report generation: &lt;a href="#figure-rrg24">(as shown in Figure)&lt;/a>&lt;/p>
&lt;ol>
&lt;li>
&lt;p>&lt;strong>Chest X-ray Feature Alignment&lt;/strong>&lt;br>
In the first stage, we align chest X-ray image features with textual embeddings in the language model. Using the provided dataset, the model predicts captions based on image inputs and instructions. During this phase, only the MLP adapter is updated, while the visual encoder and LLM weights remain frozen. This single-epoch training expands the vocabulary of aligned image-text tokens specific to the radiology domain.&lt;/p>
&lt;/li>
&lt;li>
&lt;p>&lt;strong>Fine-tuning for Report Generation&lt;/strong>&lt;br>
In the second stage, we fine-tune the pre-trained LLM weights using LoRA technology. The visual encoder and adapter weights are kept frozen, while the model undergoes three epochs of training on the dataset. This phase focuses on enhancing the model&amp;rsquo;s ability to generate accurate and clinically relevant radiology reports.&lt;/p>
&lt;/li>
&lt;/ol>
&lt;div class="flex px-4 py-3 mb-6 rounded-md bg-primary-100 dark:bg-primary-900">
&lt;span class="pr-3 pt-1 text-primary-600 dark:text-primary-300">
&lt;svg height="24" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24">&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="m11.25 11.25l.041-.02a.75.75 0 0 1 1.063.852l-.708 2.836a.75.75 0 0 0 1.063.853l.041-.021M21 12a9 9 0 1 1-18 0a9 9 0 0 1 18 0m-9-3.75h.008v.008H12z"/>&lt;/svg>
&lt;/span>
&lt;span class="dark:text-neutral-300">For implementation details, see the &lt;a href="https://github.com/X-iZhang/RRG-BioNLP-ACL2024" target="_blank" rel="noopener">Code Repo&lt;/a>.&lt;/span>
&lt;/div>
&lt;h2 id="evaluation">Evaluation&lt;/h2>
&lt;h3 id="dataset">Dataset&lt;/h3>
&lt;p>Our models were fine-tuned and evaluated using the RRG24 dataset, hosted on the BioNLP ACL'24 platform. This dataset aggregates data from multiple sources, including MIMIC-CXR, CheXpert, PadChest, BIMCV-COVID19, and OpenI. The dataset statistics are summarized below:&lt;/p>
&lt;!--
&lt;table class="table-auto w-full">
&lt;thead>
&lt;tr> &lt;th class="border-b dark:border-slate-600 font-medium p-4 pt-0 pb-3 text-slate-400 dark:text-slate-200 text-left">Dataset&lt;/th> &lt;th class="border-b dark:border-slate-600 font-medium p-4 pt-0 pb-3 text-slate-400 dark:text-slate-200 text-left">FINDINGS&lt;/th> &lt;th class="border-b dark:border-slate-600 font-medium p-4 pt-0 pb-3 text-slate-400 dark:text-slate-200 text-left">IMPRESSIONS&lt;/th> &lt;/tr>
&lt;/thead>
&lt;tbody>
&lt;tr>
&lt;td data-table-dtype="text" class="border-b border-slate-100 dark:border-slate-700 p-4 text-slate-500 dark:text-slate-400">Training&lt;/td>
&lt;td data-table-dtype="text" class="border-b border-slate-100 dark:border-slate-700 p-4 text-slate-500 dark:text-slate-400">344,394&lt;/td>
&lt;td data-table-dtype="text" class="border-b border-slate-100 dark:border-slate-700 p-4 text-slate-500 dark:text-slate-400">366,413&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td data-table-dtype="text" class="border-b border-slate-100 dark:border-slate-700 p-4 text-slate-500 dark:text-slate-400">Validation&lt;/td>
&lt;td data-table-dtype="text" class="border-b border-slate-100 dark:border-slate-700 p-4 text-slate-500 dark:text-slate-400">8,839&lt;/td>
&lt;td data-table-dtype="text" class="border-b border-slate-100 dark:border-slate-700 p-4 text-slate-500 dark:text-slate-400">9,331&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td data-table-dtype="text" class="border-b border-slate-100 dark:border-slate-700 p-4 text-slate-500 dark:text-slate-400">Test-Public&lt;/td>
&lt;td data-table-dtype="text" class="border-b border-slate-100 dark:border-slate-700 p-4 text-slate-500 dark:text-slate-400">2,692&lt;/td>
&lt;td data-table-dtype="text" class="border-b border-slate-100 dark:border-slate-700 p-4 text-slate-500 dark:text-slate-400">2,967&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td data-table-dtype="text" class="border-b border-slate-100 dark:border-slate-700 p-4 text-slate-500 dark:text-slate-400">Test-Hidden&lt;/td>
&lt;td data-table-dtype="text" class="border-b border-slate-100 dark:border-slate-700 p-4 text-slate-500 dark:text-slate-400">1,063&lt;/td>
&lt;td data-table-dtype="text" class="border-b border-slate-100 dark:border-slate-700 p-4 text-slate-500 dark:text-slate-400">1,428&lt;/td>
&lt;/tr>
&lt;/tbody>
&lt;/table>
-->
&lt;table>
&lt;thead>
&lt;tr>
&lt;th>Dataset&lt;/th>
&lt;th>FINDINGS&lt;/th>
&lt;th>IMPRESSIONS&lt;/th>
&lt;/tr>
&lt;/thead>
&lt;tbody>
&lt;tr>
&lt;td>Training&lt;/td>
&lt;td>344,394&lt;/td>
&lt;td>366,413&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Validation&lt;/td>
&lt;td>8,839&lt;/td>
&lt;td>9,331&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Test-Public&lt;/td>
&lt;td>2,692&lt;/td>
&lt;td>2,967&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Test-Hidden&lt;/td>
&lt;td>1,063&lt;/td>
&lt;td>1,428&lt;/td>
&lt;/tr>
&lt;/tbody>
&lt;/table>
&lt;h3 id="results">Results&lt;/h3>
&lt;p>Our models achieved competitive F1-RadGraph scores of 24.13 and 22.79 for the Findings and Impressions sections, respectively, on the public test set. On the hidden test set, the scores were 24.13 and 22.10, securing &lt;strong>4th&lt;/strong> place on the leaderboard at the time of submission. Additionally, the Bertscore results highlight the high-quality text generation capabilities of our models, with scores of 53.45 for Findings and 47.39 for Impressions on the public test set.&lt;/p>
&lt;h2 id="discussion">Discussion&lt;/h2>
&lt;p>The difference in performance between the Findings and Impressions sections comes down to their unique roles. Findings are all about objectively describing what’s seen in the images, while Impressions focus on drawing diagnostic conclusions. This difference, along with the varying lengths of these sections, makes generating Impressions a bit trickier, which is reflected in the evaluation scores.&lt;/p>
&lt;p>Another factor is the variety of diseases in the test set. This uneven distribution can make it harder for the model to generalize well. Plus, when the model processes multiple images at once, it sometimes includes irrelevant ones, which can throw off its accuracy. Improving the model’s ability to filter out unnecessary images could make a big difference.&lt;/p>
&lt;p>Looking ahead, there’s a lot of room for improvement. Fine-tuning the model specifically for medical data and adapting it to better handle domain-specific challenges could help. Adding the ability to track changes over time and creating smarter frameworks for multi-modal reports are also exciting directions to explore. These upgrades could make the model even more useful and reliable in real-world clinical settings.&lt;/p>
&lt;h2 id="conclusion">Conclusion&lt;/h2>
&lt;p>In this study, we introduced a model designed to generate radiology reports by aligning visual and textual data. By fine-tuning it for specific tasks, we managed to achieve solid results, including a fourth-place finish in the RRG24 Shared Task at BioNLP 2024. This shows the potential of our approach for specialized medical applications.&lt;/p>
&lt;p>Moving forward, we’re excited to explore ways to make the model even better. This includes developing methods to handle multi-modal data more effectively and incorporating time-based insights. These improvements could take the model’s accuracy and practicality to the next level for clinical use.&lt;/p>
&lt;h2 id="limitations">Limitations&lt;/h2>
&lt;p>While the results are promising, there are a few challenges we need to address:&lt;/p>
&lt;ol>
&lt;li>&lt;strong>Easier-to-Detect Conditions&lt;/strong>: Some diseases are simpler to identify, which might inflate the evaluation scores.&lt;/li>
&lt;li>&lt;strong>Dataset Imbalance&lt;/strong>: The training data isn’t perfectly balanced in terms of imaging types, body parts, or report lengths, which could impact performance.&lt;/li>
&lt;li>&lt;strong>Inconsistent Reporting Styles&lt;/strong>: Radiologists have different ways of writing reports, and this variability makes it harder for the model to generate consistent outputs.&lt;/li>
&lt;/ol>
&lt;div class="flex px-4 py-3 mb-6 rounded-md bg-yellow-100 dark:bg-yellow-900">
&lt;span class="pr-3 pt-1 text-red-400">
&lt;svg height="24" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24">&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="M12 9v3.75m-9.303 3.376c-.866 1.5.217 3.374 1.948 3.374h14.71c1.73 0 2.813-1.874 1.948-3.374L13.949 3.378c-.866-1.5-3.032-1.5-3.898 0zM12 15.75h.007v.008H12z"/>&lt;/svg>
&lt;/span>
&lt;span class="dark:text-neutral-300">The data used for the training processes were exclusively sourced from the provided dataset, which complies with ethical standards regarding Patient Health Information. Adherence to these standards guarantees the responsible handling of sensitive data throughout our research.&lt;/span>
&lt;/div>
&lt;h2 id="bibtex">BibTeX&lt;/h2>
&lt;div class="highlight">&lt;pre tabindex="0" class="chroma">&lt;code class="language-bibtex" data-lang="bibtex">&lt;span class="line">&lt;span class="cl">&lt;span class="nc">@inproceedings&lt;/span>&lt;span class="p">{&lt;/span>&lt;span class="nl">zhang-etal-2024-gla&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">title&lt;/span> &lt;span class="p">=&lt;/span> &lt;span class="s">&amp;#34;Gla-{AI}4{B}io{M}ed at {RRG}24: Visual Instruction-tuned Adaptation for Radiology Report Generation&amp;#34;&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">author&lt;/span> &lt;span class="p">=&lt;/span> &lt;span class="s">&amp;#34;Zhang, Xi and
&lt;/span>&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="s"> Meng, Zaiqiao and
&lt;/span>&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="s"> Lever, Jake and
&lt;/span>&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="s"> Ho, Edmond S.L.&amp;#34;&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">editor&lt;/span> &lt;span class="p">=&lt;/span> &lt;span class="s">&amp;#34;Demner-Fushman, Dina and
&lt;/span>&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="s"> Ananiadou, Sophia and
&lt;/span>&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="s"> Miwa, Makoto and
&lt;/span>&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="s"> Roberts, Kirk and
&lt;/span>&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="s"> Tsujii, Junichi&amp;#34;&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">booktitle&lt;/span> &lt;span class="p">=&lt;/span> &lt;span class="s">&amp;#34;Proceedings of the 23rd Workshop on Biomedical Natural Language Processing&amp;#34;&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">month&lt;/span> &lt;span class="p">=&lt;/span> &lt;span class="nv">aug&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">year&lt;/span> &lt;span class="p">=&lt;/span> &lt;span class="s">&amp;#34;2024&amp;#34;&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">address&lt;/span> &lt;span class="p">=&lt;/span> &lt;span class="s">&amp;#34;Bangkok, Thailand&amp;#34;&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">publisher&lt;/span> &lt;span class="p">=&lt;/span> &lt;span class="s">&amp;#34;Association for Computational Linguistics&amp;#34;&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">url&lt;/span> &lt;span class="p">=&lt;/span> &lt;span class="s">&amp;#34;https://aclanthology.org/2024.bionlp-1.54/&amp;#34;&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">doi&lt;/span> &lt;span class="p">=&lt;/span> &lt;span class="s">&amp;#34;10.18653/v1/2024.bionlp-1.54&amp;#34;&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">pages&lt;/span> &lt;span class="p">=&lt;/span> &lt;span class="s">&amp;#34;624--634&amp;#34;&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="p">}&lt;/span>
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;!-- &lt;figure>&lt;img src="https://x-izhang.github.io/publication/zhang-etal-2024-gla/featured.png">
&lt;/figure>
[A Figure](#figure-hello)
&lt;div class="flex px-4 py-3 mb-6 rounded-md bg-primary-100 dark:bg-primary-900">
&lt;span class="pr-3 pt-1 text-primary-600 dark:text-primary-300">
&lt;svg height="24" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24">&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="m11.25 11.25l.041-.02a.75.75 0 0 1 1.063.852l-.708 2.836a.75.75 0 0 0 1.063.853l.041-.021M21 12a9 9 0 1 1-18 0a9 9 0 0 1 18 0m-9-3.75h.008v.008H12z"/>&lt;/svg>
&lt;/span>
&lt;span class="dark:text-neutral-300">For more details, see the &lt;a href="https://github.com/X-iZhang/RRG-BioNLP-ACL2024" target="_blank" rel="noopener">project repo&lt;/a>.&lt;/span>
&lt;/div> -->
&lt;!--
&lt;div class="flex px-4 py-3 mb-6 rounded-md bg-yellow-100 dark:bg-yellow-900">
&lt;span class="pr-3 pt-1 text-red-400">
&lt;svg height="24" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24">&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="M12 9v3.75m-9.303 3.376c-.866 1.5.217 3.374 1.948 3.374h14.71c1.73 0 2.813-1.874 1.948-3.374L13.949 3.378c-.866-1.5-3.032-1.5-3.898 0zM12 15.75h.007v.008H12z"/>&lt;/svg>
&lt;/span>
&lt;span class="dark:text-neutral-300">A Markdown aside is useful for displaying notices, hints, or definitions to your readers.&lt;/span>
&lt;/div> --></description></item></channel></rss>