<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Xi Zhang</title><link>https://x-izhang.github.io/</link><atom:link href="https://x-izhang.github.io/index.xml" rel="self" type="application/rss+xml"/><description>Xi Zhang</description><generator>Hugo Blox Builder (https://hugoblox.com)</generator><language>en-us</language><lastBuildDate>Mon, 24 Oct 2022 00:00:00 +0000</lastBuildDate><image><url>https://x-izhang.github.io/media/icon_hu134860076176174952.png</url><title>Xi Zhang</title><link>https://x-izhang.github.io/</link></image><item><title>✨ Joining ANative Lab as Chief Scientist</title><link>https://x-izhang.github.io/post/anative2026/</link><pubDate>Tue, 25 Aug 2026 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/post/anative2026/</guid><description>&lt;p>&lt;strong>A new chapter.&lt;/strong> Excited to join &lt;a href="https://anative.ai" target="_blank" rel="noopener">&lt;strong>ANative Lab&lt;/strong>&lt;/a> as &lt;strong>Chief Scientist&lt;/strong> — and continue building the next generation of self-evolving scientific research. 🎉&lt;/p>
&lt;p>&lt;strong>ANative Lab&lt;/strong> is a research lab for AI-native &amp;amp; agent-native intelligence, advancing self-evolving agents and autonomous scientific discovery.&lt;/p>
&lt;p>Alongside the lab, I continue to lead &lt;a href="https://evoscientist.ai/" target="_blank" rel="noopener">&lt;strong>EvoScientist&lt;/strong>&lt;/a> — your self-evolving AI scientist: a multi-agent system for end-to-end scientific discovery.&lt;/p>
&lt;p>🌐 &lt;a href="https://anative.ai" target="_blank" rel="noopener">ANative.ai&lt;/a> · 💻 &lt;a href="https://github.com/ANative-Lab" target="_blank" rel="noopener">github.com/ANative-Lab&lt;/a> · 📣 &lt;a href="https://x.com/ANativeLab/status/2092141705952924128" target="_blank" rel="noopener">Announcement on X&lt;/a>&lt;/p></description></item><item><title>⭐ EvoScientist Crosses 4.5k GitHub Stars!</title><link>https://x-izhang.github.io/post/2026evosci2/</link><pubDate>Fri, 31 Jul 2026 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/post/2026evosci2/</guid><description>&lt;p>Today we crossed &lt;strong>4.5k GitHub stars&lt;/strong>. 🚀🚀🚀&lt;/p>
&lt;p>Thanks to everyone who&amp;rsquo;s been part of the journey — more to come.&lt;/p>
&lt;p>&lt;a href="https://github.com/EvoScientist/EvoScientist" target="_blank" rel="noopener">&lt;strong>Github: EvoScientist&lt;/strong>&lt;/a>&lt;/p>
&lt;p>&lt;strong>⭐ Star it / 🍴 Fork it / 👀 Watch it&lt;/strong> 👆&lt;/p>
&lt;h2 id="star-history">Star History&lt;/h2>
&lt;a href="https://www.star-history.com/?repos=EvoScientist%2FEvoScientist&amp;type=date&amp;legend=top-left">
&lt;picture>
&lt;source media="(prefers-color-scheme: dark)" srcset="https://api.star-history.com/chart?repos=EvoScientist/EvoScientist&amp;type=date&amp;theme=dark&amp;legend=top-left&amp;sealed_token=y0r187KO24tOIL7AOOVxsYbmgkhm-oN7chDqsuO2pD1nsYhYTxpIMITqVaJlJbwrzo1Oba4f9J9CJqKw4gKQz9MsWXA-SN8ZV-2xnZDJXnLi1cxRbu19TA" />
&lt;source media="(prefers-color-scheme: light)" srcset="https://api.star-history.com/chart?repos=EvoScientist/EvoScientist&amp;type=date&amp;legend=top-left&amp;sealed_token=y0r187KO24tOIL7AOOVxsYbmgkhm-oN7chDqsuO2pD1nsYhYTxpIMITqVaJlJbwrzo1Oba4f9J9CJqKw4gKQz9MsWXA-SN8ZV-2xnZDJXnLi1cxRbu19TA" />
&lt;img alt="Star History Chart" src="https://api.star-history.com/chart?repos=EvoScientist/EvoScientist&amp;type=date&amp;legend=top-left&amp;sealed_token=y0r187KO24tOIL7AOOVxsYbmgkhm-oN7chDqsuO2pD1nsYhYTxpIMITqVaJlJbwrzo1Oba4f9J9CJqKw4gKQz9MsWXA-SN8ZV-2xnZDJXnLi1cxRbu19TA" />
&lt;/picture>
&lt;/a></description></item><item><title>🎤 Talk at Glasgow AI4BioMed Lab!</title><link>https://x-izhang.github.io/post/2026ai4bioccs/</link><pubDate>Wed, 17 Jun 2026 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/post/2026ai4bioccs/</guid><description>&lt;h3 id="-talk-title">💡 Talk Title&lt;/h3>
&lt;p>&lt;a href="https://x-izhang.github.io/publication/zhang-2026-ccsclinicalconsensusselection/">&lt;strong>&amp;ldquo;CCS: Clinical Consensus Selection for Radiology Report Generation&amp;rdquo;&lt;/strong>&lt;/a>&lt;/p>
&lt;figure>&lt;img src="https://x-izhang.github.io/post/2026ai4bioccs/talk.png">
&lt;/figure>
&lt;h5 id="-abstract">🖇️ Abstract&lt;/h5>
&lt;p>Radiology report generation (RRG) is commonly framed as a single-path task, where a multimodal large language model (MLLM) produces one decoded report as the final output. While progress has largely come from scaling training data, model capacity, and retrieval mechanisms, improving report quality at inference time remains underexplored. We observe that fixed radiology MLLMs often generate clinically stronger reports elsewhere in their candidate pool than the one selected by default decoding, suggesting that inference-time decision making is an overlooked bottleneck. To address this, we propose &lt;strong>C&lt;/strong>linical &lt;strong>C&lt;/strong>onsensus &lt;strong>S&lt;/strong>election (&lt;strong>CCS&lt;/strong>), a decoder-agnostic inference-time selection framework that samples multiple candidate reports and selects the one with the highest clinical consensus across the rollout pool. CCS combines text-based utilities with a radiology-adapted utility computed by an image&amp;ndash;report-trained multimodal embedder, measuring agreement beyond surface-level textual similarity. Across three datasets and multiple radiology MLLMs, CCS consistently improves inference-time performance over single-path decoding and generic Best-of-N baselines, with especially clear gains on clinical metrics. Further analysis shows that image-grounded utility forms a selection axis distinct from textual consensus, and that substantial headroom remains for improving RRG at inference time.&lt;/p>
&lt;p>🔗 &lt;strong>Project Website:&lt;/strong> &lt;a href="https://x-izhang.github.io/CCS/" target="_blank" rel="noopener">https://x-izhang.github.io/CCS/&lt;/a>&lt;/p>
&lt;h3 id="-slides">📺 Slides&lt;/h3>
&lt;div style="text-align: center;">
&lt;iframe src="https://docs.google.com/presentation/d/e/2PACX-1vRuYdC5UkgKIXrivSKfN_ULo2nUagGypgN27be1LM0cAQgtlmbgI7WbneXLSkSi9Wp9cboKrg02PaU8/pubembed?start=true&amp;loop=true&amp;delayms=3000" frameborder="0" width="700" height="422" allowfullscreen="true" mozallowfullscreen="true" webkitallowfullscreen="true">&lt;/iframe>
&lt;/div>
&lt;p>📍 &lt;a href="https://ai4biomed.org/" target="_blank" rel="noopener">&lt;strong>Glasgow AI4BioMed Lab&lt;/strong>&lt;/a> — An interdisciplinary research lab focusing on AI applications in biomedical sciences, based in Glasgow.&lt;/p>
&lt;h3 id="heading">📅&lt;/h3>
&lt;p>&lt;strong>When:&lt;/strong> Wednesday, June 17, 2026 at 2pm&lt;br>
&lt;strong>Where:&lt;/strong> F121&lt;/p>
&lt;p>See you there!&lt;/p></description></item><item><title>🚨 Preprint out — Clinical Consensus Selection！</title><link>https://x-izhang.github.io/post/2026ccs/</link><pubDate>Fri, 29 May 2026 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/post/2026ccs/</guid><description>&lt;p>&lt;a href="https://x-izhang.github.io/publication/zhang-2026-ccsclinicalconsensusselection/">&lt;em>&amp;ldquo;CCS: Clinical Consensus Selection for Radiology Report Generation&amp;rdquo;&lt;/em>&lt;/a>. The preprint is now available.&lt;/p>
&lt;p>For detailed model information and source code, please visit our &lt;mark>Project Page&lt;/mark>: &lt;a href="https://x-izhang.github.io/CCS/" target="_blank" rel="noopener">CCS&lt;/a>&lt;/p>
&lt;h3 id="overview">Overview&lt;/h3>
&lt;figure>&lt;img src="https://x-izhang.github.io/post/2026ccs/image.png">
&lt;/figure>
&lt;p>Radiology report generation (RRG) is usually treated as a single-path task: a multimodal large language model (MLLM) emits one decoded report and commits to it. Yet a fixed model often places clinically stronger reports &lt;em>elsewhere&lt;/em> in its candidate pool than the one chosen by default decoding—so inference-time decision making remains an overlooked bottleneck. We introduce &lt;strong>Clinical Consensus Selection (CCS)&lt;/strong>, a decoder-agnostic, reference-free framework that samples multiple candidate reports and selects the one with the highest clinical consensus across the rollout pool. CCS unifies text-based utilities with a radiology-adapted utility from an image–report-trained multimodal embedder, measuring candidate agreement beyond surface-level text. Across three datasets and multiple radiology MLLMs, CCS consistently improves over single-path decoding and generic Best-of-N baselines, with particularly clear gains on clinical metrics.&lt;/p>
&lt;h3 id="key-resources">Key Resources&lt;/h3>
&lt;ul>
&lt;li>&lt;a href="https://github.com/X-iZhang/CCS" target="_blank" rel="noopener">&lt;strong>GitHub Repository&lt;/strong>&lt;/a> — Explore the &lt;code>CCS&lt;/code> project on GitHub&lt;/li>
&lt;/ul></description></item><item><title>CCS</title><link>https://x-izhang.github.io/project/ccs/</link><pubDate>Fri, 29 May 2026 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/project/ccs/</guid><description>&lt;p>Clinical Consensus Selection, a decoder-agnostic inference-time selection framework for radiology MLLMs.&lt;/p></description></item><item><title>CCS: Clinical Consensus Selection for Radiology Report Generation</title><link>https://x-izhang.github.io/publication/zhang-2026-ccsclinicalconsensusselection/</link><pubDate>Fri, 29 May 2026 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/publication/zhang-2026-ccsclinicalconsensusselection/</guid><description>&lt;div class="flex px-4 py-3 mb-6 rounded-md bg-primary-100 dark:bg-primary-900">
&lt;span class="pr-3 pt-1 text-primary-600 dark:text-primary-300">
&lt;svg height="24" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24">&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="m11.25 11.25l.041-.02a.75.75 0 0 1 1.063.852l-.708 2.836a.75.75 0 0 0 1.063.853l.041-.021M21 12a9 9 0 1 1-18 0a9 9 0 0 1 18 0m-9-3.75h.008v.008H12z"/>&lt;/svg>
&lt;/span>
&lt;span class="dark:text-neutral-300">For the latest updates and details, visit the &lt;a href="https://x-izhang.github.io/CCS/" target="_blank" rel="noopener">CCS Project Website&lt;/a>.&lt;/span>
&lt;/div>
&lt;h2 id="why-ccs-">Why CCS? 🔬&lt;/h2>
&lt;figure>&lt;img src="https://x-izhang.github.io/publication/zhang-2026-ccsclinicalconsensusselection/figure_1.png">
&lt;/figure>
&lt;p>Most radiology MLLMs commit to a &lt;strong>single decoded report&lt;/strong> token by token, so one unfavourable step can omit a finding or assert one unsupported by the image, with no way to recover. Yet a fixed model often places &lt;strong>clinically stronger reports elsewhere in its candidate pool&lt;/strong> — the bottleneck is not what the model can generate, but &lt;strong>which candidate it commits to&lt;/strong>.&lt;/p>
&lt;h2 id="how-ccs-works-">How CCS Works ⚙️&lt;/h2>
&lt;p>&lt;strong>Clinical Consensus Selection (CCS)&lt;/strong> is a &lt;em>reference-free&lt;/em>, &lt;em>decoder-agnostic&lt;/em> inference-time framework that reframes RRG as candidate selection — no retraining or extra parameters.&lt;/p>
&lt;ul>
&lt;li>&lt;strong>Rollout sampling&lt;/strong>: draw &lt;em>N&lt;/em> candidate reports from the radiology MLLM via stochastic decoding.&lt;/li>
&lt;li>&lt;strong>Textual consensus utility&lt;/strong>: repurpose report-evaluation metrics as reference-free pairwise agreement scores.&lt;/li>
&lt;li>&lt;strong>Image-grounded utility&lt;/strong>: measure candidate agreement with &lt;strong>Qwen3-VL-Embed&lt;/strong>, a multimodal embedder adapted on image–report pairs, capturing clinical agreement beyond surface text.&lt;/li>
&lt;li>&lt;strong>Highest-consensus selection&lt;/strong>: return the candidate with the highest &lt;em>mean&lt;/em> pairwise consensus over the pool.&lt;/li>
&lt;/ul>
&lt;h2 id="key-contributions-">Key Contributions ✨&lt;/h2>
&lt;ul>
&lt;li>&lt;strong>Inference-time perspective&lt;/strong>: We show candidate pools routinely contain reports with &lt;strong>higher clinical reliability and consistency&lt;/strong> than single-path outputs.&lt;/li>
&lt;li>&lt;strong>CCS framework&lt;/strong>: A &lt;strong>decoder-agnostic Best-of-N&lt;/strong> method that aggregates pairwise clinical consensus using both &lt;strong>textual&lt;/strong> and &lt;strong>image–report-adapted multimodal&lt;/strong> utilities.&lt;/li>
&lt;li>&lt;strong>Consistent gains&lt;/strong>: Across &lt;strong>three datasets&lt;/strong> and &lt;strong>multiple radiology MLLMs&lt;/strong>, CCS improves RRG at inference time and identifies image-grounded utility as a &lt;strong>distinct selection axis&lt;/strong>.&lt;/li>
&lt;/ul>
&lt;h2 id="main-results-">Main Results 📈&lt;/h2>
&lt;figure>&lt;img src="https://x-izhang.github.io/publication/zhang-2026-ccsclinicalconsensusselection/table_1.png">
&lt;/figure>
&lt;figure>&lt;img src="https://x-izhang.github.io/publication/zhang-2026-ccsclinicalconsensusselection/table_2.png">
&lt;/figure>
&lt;p>On MIMIC-CXR, CCS improves &lt;strong>all&lt;/strong> radiology-specific metrics over Sampling (e.g. RadGraph-F1 &lt;strong>0.1989 → 0.2134&lt;/strong>, CheXbert-F1⁵ &lt;strong>0.5041 → 0.5370&lt;/strong>), with gains statistically significant (&lt;em>p&lt;/em> &amp;lt; 0.05). Unlike generic Best-of-N selectors (Self-Certainty, ModeX) that give inconsistent or negative clinical gains, CCS improves &lt;strong>consistently across all backbones and datasets&lt;/strong>.&lt;/p>
&lt;h2 id="analysis-">Analysis 📊&lt;/h2>
&lt;figure>&lt;img src="https://x-izhang.github.io/publication/zhang-2026-ccsclinicalconsensusselection/figure_3.png">
&lt;/figure>
&lt;figure>&lt;img src="https://x-izhang.github.io/publication/zhang-2026-ccsclinicalconsensusselection/table_3.png">
&lt;/figure>
&lt;p>Per-label F1 shows CCS recovers abnormal findings that text-only consensus suppresses — confirming that &lt;strong>image-grounded utility is a selection axis distinct from textual consensus&lt;/strong>, with substantial headroom still remaining.&lt;/p>
&lt;h2 id="case-study-">Case Study 🔍&lt;/h2>
&lt;figure>&lt;img src="https://x-izhang.github.io/publication/zhang-2026-ccsclinicalconsensusselection/case_study.png">
&lt;/figure>
&lt;p>On a real MIMIC-CXR case, single-path decoding asserts factual errors (false &amp;ldquo;clear lungs&amp;rdquo;, mislocalised catheter), whereas CCS selects a more &lt;strong>image-grounded&lt;/strong> report (CheXbert-F1⁵ &lt;strong>1.0000&lt;/strong> vs Sampling &lt;strong>0.5000&lt;/strong>).&lt;/p>
&lt;h2 id="bibtex-">BibTeX 📚&lt;/h2>
&lt;div class="highlight">&lt;pre tabindex="0" class="chroma">&lt;code class="language-bibtex" data-lang="bibtex">&lt;span class="line">&lt;span class="cl">&lt;span class="nc">@article&lt;/span>&lt;span class="p">{&lt;/span>&lt;span class="nl">zhang2026ccsclinicalconsensusselection&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">title&lt;/span>&lt;span class="p">=&lt;/span>&lt;span class="s">{CCS: Clinical Consensus Selection for Radiology Report Generation}&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">author&lt;/span>&lt;span class="p">=&lt;/span>&lt;span class="s">{Xi Zhang and Yingshu Li and Zaiqiao Meng and Jake Lever and Edmond S. L. Ho}&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">year&lt;/span>&lt;span class="p">=&lt;/span>&lt;span class="s">{2026}&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">eprint&lt;/span>&lt;span class="p">=&lt;/span>&lt;span class="s">{2605.30131}&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">archivePrefix&lt;/span>&lt;span class="p">=&lt;/span>&lt;span class="s">{arXiv}&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">primaryClass&lt;/span>&lt;span class="p">=&lt;/span>&lt;span class="s">{cs.CL}&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">url&lt;/span>&lt;span class="p">=&lt;/span>&lt;span class="s">{https://arxiv.org/abs/2605.30131}&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="p">}&lt;/span>
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div></description></item><item><title>✒️ Invited as Reviewer for NeurIPS 2026</title><link>https://x-izhang.github.io/post/neurips2026reviewer/</link><pubDate>Thu, 23 Apr 2026 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/post/neurips2026reviewer/</guid><description>&lt;p>I have been invited to serve as a reviewer for &lt;a href="https://neurips.cc/Conferences/2026" target="_blank" rel="noopener">NeurIPS 2026&lt;/a>, the 40th Annual Conference on Neural Information Processing Systems. Looking forward to contributing to the review process and supporting high-quality research in machine learning and AI.&lt;/p></description></item><item><title>🎉 Paper Accepted — See You at ACL 2026!</title><link>https://x-izhang.github.io/post/2026acl/</link><pubDate>Mon, 06 Apr 2026 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/post/2026acl/</guid><description>&lt;p>&lt;a href="https://x-izhang.github.io/publication/zhang-2025-ccdmitigatinghallucinationsradiology/">&lt;em>&amp;ldquo;CCD: Mitigating Hallucinations in Radiology MLLMs via Clinical Contrastive Decoding&amp;rdquo;&lt;/em>&lt;/a> has been accepted to &lt;a href="https://2026.aclweb.org/" target="_blank" rel="noopener">ACL 2026&lt;/a>!&lt;/p>
&lt;p>📍 &lt;strong>ACL 2026&lt;/strong> — The 64th Annual Meeting of the Association for Computational Linguistics will take place in &lt;strong>San Diego, California&lt;/strong>, July 2026&lt;/p>
&lt;p>See you there!&lt;/p>
&lt;!-- ### 📺 Talk
&lt;div style="position: relative; padding-bottom: 56.25%; height: 0; overflow: hidden;">
&lt;iframe allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" allowfullscreen="allowfullscreen" loading="eager" referrerpolicy="strict-origin-when-cross-origin" src="https://www.youtube.com/embed/_R8XUaaAU3g?autoplay=0&amp;controls=1&amp;end=0&amp;loop=0&amp;mute=0&amp;start=0" style="position: absolute; top: 0; left: 0; width: 100%; height: 100%; border:0;" title="YouTube video"
>&lt;/iframe>
&lt;/div>
-->
&lt;h3 id="-poster">🪧 Poster&lt;/h3>
&lt;figure>&lt;img src="https://x-izhang.github.io/post/2026acl/2026acl_poster.jpg">
&lt;/figure></description></item><item><title>CCD: Mitigating Hallucinations in Radiology MLLMs via Clinical Contrastive Decoding</title><link>https://x-izhang.github.io/publication/zhang-2025-ccdmitigatinghallucinationsradiology/</link><pubDate>Mon, 06 Apr 2026 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/publication/zhang-2025-ccdmitigatinghallucinationsradiology/</guid><description>&lt;div class="flex px-4 py-3 mb-6 rounded-md bg-primary-100 dark:bg-primary-900">
&lt;span class="pr-3 pt-1 text-primary-600 dark:text-primary-300">
&lt;svg height="24" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24">&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="m11.25 11.25l.041-.02a.75.75 0 0 1 1.063.852l-.708 2.836a.75.75 0 0 0 1.063.853l.041-.021M21 12a9 9 0 1 1-18 0a9 9 0 0 1 18 0m-9-3.75h.008v.008H12z"/>&lt;/svg>
&lt;/span>
&lt;span class="dark:text-neutral-300">For the latest updates and details, visit the &lt;a href="https://x-izhang.github.io/CCD/" target="_blank" rel="noopener">CCD Project Website&lt;/a>.&lt;/span>
&lt;/div>
&lt;h2 id="why-ccd-">Why CCD? 🔬&lt;/h2>
&lt;figure>&lt;img src="https://x-izhang.github.io/publication/zhang-2025-ccdmitigatinghallucinationsradiology/figure_1.png">
&lt;/figure>
&lt;p>Radiology MLLMs remain vulnerable to prompt-induced hallucinations when clinical sections contain counterfactual details or ambiguous guidance. The figure above contrasts baseline predictions with &lt;strong>CCD-enabled&lt;/strong> outputs across report generation and question answering tasks. &lt;strong>Red&lt;/strong> highlights mark unsupported findings that can compromise patient care, while &lt;strong>blue&lt;/strong> text reflects misleading prompt context the model must resist.&lt;/p>
&lt;p>Clinical Contrastive Decoding (CCD) mitigates these risks at inference time. Rather than retraining models or relying on retrieval corpora, &lt;strong>CCD&lt;/strong> injects trustworthy, image-grounded signals distilled from specialist expert models. The result is a decoding policy that maintains fluency yet stays faithful to the radiograph.&lt;/p>
&lt;h2 id="ccd-at-a-glance-">CCD at a Glance ⚙️&lt;/h2>
&lt;p>&lt;strong>Clinical Contrastive Decoding (CCD)&lt;/strong> is a plug-and-play inference framework designed to reduce medical hallucinations in radiology MLLMs. It introduces structured clinical supervision from expert models (e.g., DenseNet or MedSigLIP, or others) at decoding time, without modifying model weights or requiring external retrieval.&lt;/p>
&lt;p>Given a chest radiograph, the expert model predicts symptom-level probabilities across 14 CheXpert categories. CCD integrates this signal through a dual-stage logit refinement strategy:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>Symptom-grounded Contrastive Decoding (SCD)&lt;/strong>: constructs an anchor prompt using high-confidence findings (e.g., “Atelectasis, Cardiomegaly”) and generates a contrastive logits path conditioned on this prompt. The final logits are a weighted interpolation between the anchor-conditioned and original paths, encouraging the model to mention supported findings and suppress false negatives.&lt;/li>
&lt;li>&lt;strong>Expert-informed Contrastive Decoding (ECD)&lt;/strong>: transforms expert probabilities into token-level logit biases via log-odds conversion. These biases are injected into the logits from the first stage, softly penalising unsupported findings and reducing false positives.&lt;/li>
&lt;/ul>
&lt;p>This two-stage mechanism enables &lt;strong>CCD&lt;/strong> to progressively guide generation with both symbolic supervision (via anchor prompts) and probabilistic constraints (via confidence scores), achieving robust improvements in both report generation and VQA. &lt;strong>CCD&lt;/strong> is fully model-agnostic and integrates seamlessly with state-of-the-art radiology MLLMs such as MAIRA-2, Libra, LLaVA-Rad, and LLaVA-Med.&lt;/p>
&lt;p>Unlike prior contrastive decoding approaches that rely on perturbed visual or textual inputs, &lt;strong>CCD&lt;/strong> leverages clinically grounded signals from expert models to provide task-specific and symptom-level control during generation.&lt;/p>
&lt;h2 id="key-contributions-">Key Contributions ✨&lt;/h2>
&lt;ul>
&lt;li>
&lt;p>&lt;strong>Empirical Insight&lt;/strong>: We conduct a systematic analysis of &lt;strong>prompt-induced hallucinations&lt;/strong> in radiology MLLMs, revealing that noisy clinical sections (such as &lt;em>&lt;strong>irrelevant&lt;/strong>&lt;/em> or &lt;em>&lt;strong>contradictory clinical&lt;/strong>&lt;/em> details) can trigger &lt;strong>unsupported findings&lt;/strong> across multiple datasets.&lt;/p>
&lt;/li>
&lt;li>
&lt;p>&lt;strong>Inference-time Framework&lt;/strong>: We introduce &lt;strong>Clinical Contrastive Decoding (CCD)&lt;/strong> — a &lt;strong>dual-stage inference-time strategy&lt;/strong> that leverages expert-derived labels as &lt;strong>anchor prompts&lt;/strong> and performs &lt;strong>probabilistic logit adjustments&lt;/strong>, requiring &lt;strong>no retraining or architectural modifications&lt;/strong>.&lt;/p>
&lt;/li>
&lt;li>
&lt;p>&lt;strong>Consistent Gains&lt;/strong>: Extensive experiments on &lt;strong>MIMIC-CXR&lt;/strong>, &lt;strong>IU-Xray&lt;/strong>, and &lt;strong>CheXpert Plus&lt;/strong> demonstrate that &lt;strong>CCD&lt;/strong> achieves up to &lt;strong>+17% improvement in RadGraph-F1&lt;/strong> with SOTA radiology MLLMs (MAIRA-2), along with higher &lt;strong>VQA accuracy&lt;/strong>, all &lt;strong>without altering model weights or structures&lt;/strong>.&lt;/p>
&lt;/li>
&lt;/ul>
&lt;h2 id="evaluation-highlights-">Evaluation Highlights 📈&lt;/h2>
&lt;figure>&lt;img src="https://x-izhang.github.io/publication/zhang-2025-ccdmitigatinghallucinationsradiology/figure_2.png">
&lt;/figure>
&lt;ul>
&lt;li>Achieves up to &lt;strong>+17% RadGraph-F1&lt;/strong> on MIMIC-CXR when attached to state-of-the-art report generators.&lt;/li>
&lt;li>Reduces mention-level hallucinations on IU X-Ray while maintaining &lt;strong>BLEU / ROUGE&lt;/strong> scores, confirming that &lt;strong>CCD&lt;/strong> improves factuality without harming surface-level metrics.&lt;/li>
&lt;/ul>
&lt;h2 id="ablation-studies-">Ablation Studies 🧪&lt;/h2>
&lt;figure>&lt;img src="https://x-izhang.github.io/publication/zhang-2025-ccdmitigatinghallucinationsradiology/figure_3.png">
&lt;/figure>
&lt;p>Ablations show both contrastive stages matter: removing the anchor projection (SCD) or the logit sharpening (ECD) degrades factual metrics by &lt;strong>~5–8%&lt;/strong>.&lt;/p>
&lt;h2 id="my-musings-">My Musings ⏳&lt;/h2>
&lt;p>As I reflect on the development and impact of &lt;strong>Clinical Contrastive Decoding (CCD)&lt;/strong>,&lt;/p>
&lt;blockquote>
&lt;p>It’s better to be roughly right than precisely wrong.&lt;br>
— &lt;em>Carveth Read, Logic: Deductive and Inductive&lt;/em>&lt;/p>
&lt;/blockquote>
&lt;p>This quote perfectly captures the spirit of our approach. In the high-stakes field of radiology, the cost of hallucinations can be severe. By integrating expert-derived signals at inference time, &lt;strong>CCD&lt;/strong> prioritises &lt;strong>clinical accuracy&lt;/strong> over mere &lt;strong>linguistic fluency&lt;/strong>. It’s a reminder that in medical AI, being &lt;em>approximately correct and grounded in reality&lt;/em> is far more valuable than generating &lt;em>fluent but misleading&lt;/em> text.&lt;/p>
&lt;h2 id="bibtex-">BibTeX 📚&lt;/h2>
&lt;div class="highlight">&lt;pre tabindex="0" class="chroma">&lt;code class="language-bibtex" data-lang="bibtex">&lt;span class="line">&lt;span class="cl">&lt;span class="nc">@article&lt;/span>&lt;span class="p">{&lt;/span>&lt;span class="nl">zhang2025ccd&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">title&lt;/span>&lt;span class="p">=&lt;/span>&lt;span class="s">{CCD: Mitigating Hallucinations in Radiology MLLMs via Clinical Contrastive Decoding}&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">author&lt;/span>&lt;span class="p">=&lt;/span>&lt;span class="s">{Zhang, Xi and Meng, Zaiqiao and Lever, Jake and Ho, Edmond SL}&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">journal&lt;/span>&lt;span class="p">=&lt;/span>&lt;span class="s">{arXiv preprint arXiv:2509.23379}&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">year&lt;/span>&lt;span class="p">=&lt;/span>&lt;span class="s">{2025}&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="p">}&lt;/span>
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div></description></item><item><title>🎤 Invited Talk at EvoAgentX Community!</title><link>https://x-izhang.github.io/post/evotalk2026/</link><pubDate>Wed, 01 Apr 2026 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/post/evotalk2026/</guid><description>&lt;p>Honoured to be invited by the &lt;a href="https://github.com/EvoAgentX/EvoAgentX" target="_blank" rel="noopener">&lt;strong>EvoAgentX&lt;/strong>&lt;/a> community to present our latest project at this week&amp;rsquo;s &lt;strong>EvoAgentX Talk&lt;/strong>! As a core contributor to the EvoAgentX community, I shared the design and progress behind &lt;strong>EvoScientist&lt;/strong> — an open-source, long-horizon agent system for autonomous scientific discovery.&lt;/p>
&lt;h3 id="-talk-title">💡 Talk Title&lt;/h3>
&lt;p>&lt;strong>EvoScientist: Toward Long-Horizon Agent Systems for Scientific Discovery&lt;/strong>&lt;/p>
&lt;h5 id="-abstract">🖇️ Abstract&lt;/h5>
&lt;p>AI is rapidly shifting from dialogue models to agent systems. Platforms like OpenClaw have already demonstrated that AI can execute real-world tasks beyond simple Q&amp;amp;A. However, when we move to the scientific research domain, the challenges multiply dramatically. Research tasks are inherently &lt;strong>long-horizon, multi-stage, strongly interdependent, and continuously iterative&lt;/strong> — far from a single-turn reasoning problem, they demand sustained decision-making and an evolving system architecture.&lt;/p>
&lt;p>&lt;strong>EvoScientist&lt;/strong> explores a long-horizon agent design paradigm tailored to research scenarios. Through multi-agent collaboration, the system progressively completes the full pipeline from &lt;strong>idea generation&lt;/strong> to &lt;strong>experiment execution&lt;/strong> to &lt;strong>result summarization&lt;/strong> — an approach we call &lt;strong>Vibe Research&lt;/strong>. Unlike many agent systems that remain at the demo stage, EvoScientist emphasizes executability in complex tasks and the ability to continuously evolve in real-world settings.&lt;/p>
&lt;p>In this talk, I address three key questions: why current agent systems fall short in research scenarios, what the core design philosophy of EvoScientist is, and whether we truly need a new long-horizon agent paradigm to support increasingly complex task systems in the future.&lt;/p>
&lt;h3 id="-slides">📺 Slides&lt;/h3>
&lt;div class="flex px-4 py-3 mb-6 rounded-md bg-primary-100 dark:bg-primary-900">
&lt;span class="pr-3 pt-1 text-primary-600 dark:text-primary-300">
&lt;svg height="24" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24">&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="m11.25 11.25l.041-.02a.75.75 0 0 1 1.063.852l-.708 2.836a.75.75 0 0 0 1.063.853l.041-.021M21 12a9 9 0 1 1-18 0a9 9 0 0 1 18 0m-9-3.75h.008v.008H12z"/>&lt;/svg>
&lt;/span>
&lt;span class="dark:text-neutral-300">This talk was presented to the Chinese-speaking community, so the slides below are in Chinese. An English version will be shared in a future community session.&lt;/span>
&lt;/div>
&lt;div style="text-align: center;">
&lt;iframe src="https://docs.google.com/presentation/d/e/2PACX-1vSuDOn3IqYmbrkkJs2LS7HUMoeSE1Nh693UXFUNMBUxtyMuEBWQTZmD_fhDgW4dHKxe7R-zYIcBMB83/pubembed?start=true&amp;loop=true&amp;delayms=5000" frameborder="0" width="700" height="422" allowfullscreen="true" mozallowfullscreen="true" webkitallowfullscreen="true">&lt;/iframe>
&lt;/div>
&lt;h3 id="-project-link">🔗 Project Link&lt;/h3>
&lt;blockquote>
&lt;p>&lt;a href="https://evoscientist.ai/" target="_blank" rel="noopener">&lt;strong>EvoScientist&lt;/strong>&lt;/a> — Harness Vibe Research with Self-evolving AI Scientists&lt;/p>
&lt;/blockquote>
&lt;p>EvoScientist aims to harness vibe research by enabling self-evolving AI scientists that autonomously explore, generate insights, and iteratively improve. It is designed to be opinionated and ready to use out of the box, offering a living research system that grows alongside evolving agent skills, toolsets, and memory bases. Moving beyond traditional human-in-the-loop systems, EvoScientist adopts a &lt;strong>human-on-the-loop&lt;/strong> paradigm — AI acts as a research buddy that co-evolves with human researchers and internalizes scholarly taste and scientific judgment.&lt;/p></description></item><item><title>🚀 EvoScientist Debuts!</title><link>https://x-izhang.github.io/post/2026evosci/</link><pubDate>Fri, 13 Mar 2026 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/post/2026evosci/</guid><description>&lt;p>See what&amp;rsquo;s new on &lt;a href="https://github.com/EvoScientist/EvoScientist" target="_blank" rel="noopener">Github: EvoScientist&lt;/a>&lt;/p>
&lt;p>&lt;strong>⭐ Star it / 🍴 Fork it / 👀 Watch it&lt;/strong> 👆&lt;/p>
&lt;h3 id="-if-ai-agents-must-be-self-evolving-then-ai-scientists-should-lead-the-way">🧬 If AI agents must be Self-Evolving, then AI Scientists should lead the way.&lt;/h3>
&lt;p>&lt;strong>EvoScientist&lt;/strong> aims to harness vibe research by enabling self-evolving AI scientists that autonomously explore, generate insights, and iteratively improve. Moving beyond traditional human-in-the-loop systems, EvoScientist adopts a &lt;strong>human-on-the-loop&lt;/strong> paradigm — AI acts as a research buddy that co-evolves with human researchers and internalizes scholarly taste and scientific judgment.&lt;/p>
&lt;p>Why vibe research matters and how it drives EvoScientist: &lt;a href="https://x-izhang.github.io/blog/vibe-research/">&lt;strong>Harness Vibe Research 📟&lt;/strong>&lt;/a>&lt;/p>
&lt;h3 id="-end-to-end-scientific-workflow">🔬 End-to-End Scientific Workflow&lt;/h3>
&lt;figure>&lt;img src="https://x-izhang.github.io/post/2026evosci/github.png">
&lt;/figure>
&lt;h3 id="-quick-start-demo">🚀 Quick Start Demo&lt;/h3>
&lt;figure>&lt;img src="https://x-izhang.github.io/post/2026evosci/demo.png">
&lt;/figure>
&lt;h3 id="-awards--recognition">🏆 Awards &amp;amp; Recognition&lt;/h3>
&lt;ul>
&lt;li>🥇 &lt;strong>#1&lt;/strong> on &lt;a href="https://agentresearchlab.com/benchmarks/deepresearch-bench-ii/index.html#leaderboard" target="_blank" rel="noopener">DeepResearch Bench II&lt;/a>&lt;/li>
&lt;li>🥇 &lt;strong>#1&lt;/strong> on &lt;a href="https://allenai-asta-bench-leaderboard.hf.space/code-execution" target="_blank" rel="noopener">AstaBench Code &amp;amp; Execution&lt;/a>&lt;/li>
&lt;li>🥇 &lt;strong>#1&lt;/strong> on &lt;a href="https://allenai-asta-bench-leaderboard.hf.space/data-analysis" target="_blank" rel="noopener">AstaBench Data Analysis&lt;/a>&lt;/li>
&lt;li>🏆 &lt;strong>Best Paper &amp;amp; AI Reviewer&amp;rsquo;s Appraisal Award&lt;/strong> at &lt;a href="https://icais.ai/" target="_blank" rel="noopener">ICAIS 2025&lt;/a>&lt;/li>
&lt;/ul>
&lt;h3 id="-join-our-community">🪧 Join Our Community!&lt;/h3>
&lt;p>We&amp;rsquo;re building a community of researchers, developers, and visionaries dedicated to harnessing vibe research and self-evolving AI scientists. Together, we&amp;rsquo;re redefining autonomous scientific discovery.&lt;/p>
&lt;ul>
&lt;li>&lt;a href="https://discord.gg/AZ9ZMXkunY" target="_blank" rel="noopener">&lt;strong>Discord&lt;/strong>&lt;/a> — Chat, discuss, and collaborate in real-time&lt;/li>
&lt;li>&lt;a href="https://github.com/EvoScientist/EvoScientist" target="_blank" rel="noopener">&lt;strong>GitHub&lt;/strong>&lt;/a> — Source code and documentation&lt;/li>
&lt;li>&lt;a href="https://github.com/EvoScientist/EvoScientist/blob/main/.github/assets/cn_info.md" target="_blank" rel="noopener">&lt;strong>WeChat&lt;/strong>&lt;/a> — Connect with our Chinese community&lt;/li>
&lt;/ul></description></item><item><title>EvoScientist: Towards Multi-Agent Evolving AI Scientists for End-to-End Scientific Discovery</title><link>https://x-izhang.github.io/publication/lyu-2026-evoscientist/</link><pubDate>Mon, 09 Mar 2026 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/publication/lyu-2026-evoscientist/</guid><description>&lt;div class="flex px-4 py-3 mb-6 rounded-md bg-primary-100 dark:bg-primary-900">
&lt;span class="pr-3 pt-1 text-primary-600 dark:text-primary-300">
&lt;svg height="24" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24">&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="m11.25 11.25l.041-.02a.75.75 0 0 1 1.063.852l-.708 2.836a.75.75 0 0 0 1.063.853l.041-.021M21 12a9 9 0 1 1-18 0a9 9 0 0 1 18 0m-9-3.75h.008v.008H12z"/>&lt;/svg>
&lt;/span>
&lt;span class="dark:text-neutral-300">&lt;p>Part of the &lt;a href="https://evoscientist.ai/" target="_blank" rel="noopener">EvoScientist&lt;/a> project — harnessing vibe research with self-evolving AI scientists.&lt;/p>
&lt;p>Extend it with &lt;a href="https://github.com/EvoScientist/EvoSkills" target="_blank" rel="noopener">EvoSkills&lt;/a> — installable skill &amp;amp; knowledge packs that add domain-specific expertise to AI scientists.&lt;/p>
&lt;/span>
&lt;/div>
&lt;h2 id="overview">Overview&lt;/h2>
&lt;p>&lt;strong>EvoScientist&lt;/strong> is an evolving multi-agent AI scientist framework that continuously improves its research strategies through &lt;strong>persistent memory&lt;/strong> and &lt;strong>self-evolution&lt;/strong>. It addresses a key limitation of existing AI-scientist systems: static, hand-designed pipelines that overlook promising directions, repeat failed experiments, and pursue infeasible ideas.&lt;/p>
&lt;h2 id="framework">Framework&lt;/h2>
&lt;p>EvoScientist coordinates three specialized agents:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>Researcher Agent (RA):&lt;/strong> scientific idea generation.&lt;/li>
&lt;li>&lt;strong>Engineer Agent (EA):&lt;/strong> experiment implementation and execution.&lt;/li>
&lt;li>&lt;strong>Evolution Manager Agent (EMA):&lt;/strong> distills insights from prior interactions into reusable knowledge.&lt;/li>
&lt;/ul>
&lt;p>These are backed by two persistent memory modules — an &lt;strong>ideation memory&lt;/strong> (feasible directions distilled from top-ranked ideas, plus previously unsuccessful ones) and an &lt;strong>experimentation memory&lt;/strong> (effective data-processing and training strategies from code-search trajectories and best-performing implementations) — which the RA and EA retrieve to improve idea quality and code execution success over time.&lt;/p>
&lt;h2 id="key-results">Key Results&lt;/h2>
&lt;ul>
&lt;li>Outperforms &lt;strong>7 open-source and commercial state-of-the-art systems&lt;/strong> in scientific idea generation, with higher &lt;strong>novelty, feasibility, relevance, and clarity&lt;/strong> under both automatic and human evaluation.&lt;/li>
&lt;li>Substantially improves &lt;strong>code execution success rates&lt;/strong> through multi-agent evolution, demonstrating the effectiveness of persistent memory for end-to-end scientific discovery.&lt;/li>
&lt;/ul>
&lt;h2 id="bibtex">BibTeX&lt;/h2>
&lt;div class="highlight">&lt;pre tabindex="0" class="chroma">&lt;code class="language-bibtex" data-lang="bibtex">&lt;span class="line">&lt;span class="cl">&lt;span class="nc">@article&lt;/span>&lt;span class="p">{&lt;/span>&lt;span class="nl">lyu2026evoscientist&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">title&lt;/span>&lt;span class="p">=&lt;/span>&lt;span class="s">{EvoScientist: Towards Multi-Agent Evolving AI Scientists for End-to-End Scientific Discovery}&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">author&lt;/span>&lt;span class="p">=&lt;/span>&lt;span class="s">{Lyu, Yougang and Zhang, Xi and Yi, Xinhao and Zhao, Yuyue and Guo, Shuyu and Hu, Wenxiang and Piotrowski, Jan and Kaliski, Jakub and Urbani, Jacopo and Meng, Zaiqiao and others}&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">journal&lt;/span>&lt;span class="p">=&lt;/span>&lt;span class="s">{arXiv preprint arXiv:2603.08127}&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">year&lt;/span>&lt;span class="p">=&lt;/span>&lt;span class="s">{2026}&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="p">}&lt;/span>
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div></description></item><item><title>EvoScientist</title><link>https://x-izhang.github.io/project/evoscientist/</link><pubDate>Fri, 06 Mar 2026 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/project/evoscientist/</guid><description>&lt;p>Harnessing vibe research with self-evolving AI scientists that autonomously explore, generate insights, and iteratively improve.&lt;/p></description></item><item><title>📝 New Blog Just Dropped!</title><link>https://x-izhang.github.io/post/viberesearch/</link><pubDate>Wed, 04 Mar 2026 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/post/viberesearch/</guid><description>&lt;p>🧠 &lt;strong>TL;DR&lt;/strong>&lt;/p>
&lt;p>&lt;strong>Full research autonomy will happen&lt;/strong> — but it won&amp;rsquo;t come from a &lt;strong>pre-built product&lt;/strong>. It will emerge from your own harness: &lt;em>a system that absorbs your judgment, encodes your taste, and compounds with every run. The question is not whether this shift is coming.&lt;/em> The question is whether you are building the harness to grow with it.&lt;/p>
&lt;blockquote>
&lt;p>Dive deeper in the full post: &lt;a href="https://x-izhang.github.io/blog/vibe-research/">&lt;strong>Harness Vibe Research 📟&lt;/strong>&lt;/a>&lt;/p>
&lt;/blockquote>
&lt;p>&lt;em>&lt;strong>Opinions on my own&lt;/strong>&lt;/em>&lt;/p></description></item><item><title>Harness Vibe Research 📟</title><link>https://x-izhang.github.io/blog/vibe-research/</link><pubDate>Wed, 04 Mar 2026 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/blog/vibe-research/</guid><description>&lt;p>TL;DR — Full research autonomy will happen — but it won&amp;rsquo;t come from a pre-built product. It will emerge from your own harness: a system that absorbs your judgment, encodes your taste, and compounds with every run. The question is not whether this shift is coming. The question is whether you are building the harness to grow with it.&lt;/p>
&lt;hr>
&lt;div class="flex px-4 py-3 mb-6 rounded-md bg-yellow-100 dark:bg-yellow-900">
&lt;span class="pr-3 pt-1 text-red-400">
&lt;svg height="24" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24">&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="M12 9v3.75m-9.303 3.376c-.866 1.5.217 3.374 1.948 3.374h14.71c1.73 0 2.813-1.874 1.948-3.374L13.949 3.378c-.866-1.5-3.032-1.5-3.898 0zM12 15.75h.007v.008H12z"/>&lt;/svg>
&lt;/span>
&lt;span class="dark:text-neutral-300">We are making the same mistake twice. In coding, we over-focused on model capability and under-focused on system design. Now in research, we are over-focusing on autonomy and under-designing human-AI coupling.&lt;/span>
&lt;/div>
&lt;p>More autonomy alone often increases entropy. Too little AI integration creates almost no leverage. The real multiplier is harness design: how humans and agents share context, verification, and control.&lt;/p>
&lt;h2 id="the-signal-is-already-here">The signal is already here&lt;/h2>
&lt;p>This shift is not theoretical. Multiple teams in different settings are reporting the same pattern:&lt;/p>
&lt;ul>
&lt;li>&lt;a href="https://github.com/SakanaAI/AI-Scientist-v2" target="_blank" rel="noopener">Sakana AI&lt;/a> reported AI Scientist v2 and discussed both progress and quality limits in autonomous paper-generation workflows.&lt;/li>
&lt;li>Google shared &lt;a href="https://deepmind.google/blog/accelerating-mathematical-and-scientific-discovery-with-gemini-deep-think/" target="_blank" rel="noopener">Gemini Deep Think&lt;/a> results on open math problems and shipped &lt;a href="https://research.google/blog/accelerating-scientific-breakthroughs-with-an-ai-co-scientist/" target="_blank" rel="noopener">AI co-scientist&lt;/a> — a multi-agent system on Gemini 2.0 with self-critique, Elo-based ranking, and human-in-the-loop design. Its drug repurposing and target discovery hypotheses were experimentally validated in the lab.&lt;/li>
&lt;li>&lt;a href="https://www.orchestra-research.com/perspectives" target="_blank" rel="noopener">Orchestra&lt;/a> demonstrated an ACL submission timeline built with an AI co-scientist workflow.&lt;/li>
&lt;li>Open ecosystems such as &lt;a href="https://k-dense.ai" target="_blank" rel="noopener">K-Dense&lt;/a> and &lt;a href="https://github.com/snap-stanford/Biomni" target="_blank" rel="noopener">Biomni&lt;/a> are treating reusable &amp;ldquo;skills&amp;rdquo; and know-how libraries as core research infrastructure, not optional add-ons.&lt;/li>
&lt;li>Even at the execution layer, Anthropic&amp;rsquo;s &lt;a href="https://docs.anthropic.com/en/docs/build-with-claude/tool-use/programmatic-tool-calling" target="_blank" rel="noopener">Programmatic Tool Calling&lt;/a> is rethinking how agents compose actions — optimizing for token efficiency and multi-step orchestration.&lt;/li>
&lt;/ul>
&lt;p>Different institutions, different domains, same direction: people are shipping systems, not just demos.&lt;/p>
&lt;h2 id="what-coding-agents-already-taught-us">What coding agents already taught us&lt;/h2>
&lt;p>Before research agents, coding agents already exposed the key lesson.&lt;/p>
&lt;p>&lt;a href="https://openai.com/index/harness-engineering/" target="_blank" rel="noopener">OpenAI&lt;/a> publicly described building large production systems with AI-generated code and emphasized environment and harness design as the engineering bottleneck.&lt;/p>
&lt;p>&lt;a href="https://www.langchain.com/" target="_blank" rel="noopener">LangChain&lt;/a> has shown that major benchmark gains can come purely from harness upgrades: self-verification loops, better context injection, loop guards, and reasoning budget control. Same base model class, better orchestration. See this &lt;a href="https://blog.langchain.com/improving-deep-agents-with-harness-engineering/" target="_blank" rel="noopener">blog&lt;/a>.&lt;/p>
&lt;blockquote>
&lt;p>&lt;strong>The same model became dramatically more capable — not because it got smarter, but because the system around it got better.&lt;/strong>&lt;/p>
&lt;/blockquote>
&lt;p>The open source community is already far ahead. Projects like &lt;a href="https://openclaw.ai/" target="_blank" rel="noopener">OpenClaw&lt;/a> — an open-source personal AI assistant with a skills-first architecture and full agent autonomy — are going viral. Developers have dozens of harness prototypes to choose from. Research, by contrast, is still catching up. Most research teams are either building from scratch or relying on closed, single-purpose tools. The infrastructure gap is real.&lt;/p>
&lt;h2 id="the-autonomy-paradox">The autonomy paradox&lt;/h2>
&lt;p>There is a natural instinct to push to full automation: remove humans and let agents run end to end.&lt;/p>
&lt;p>But in research, higher autonomy frequently increases drift risk:&lt;/p>
&lt;ul>
&lt;li>agents can optimize for easy novelty instead of meaningful questions,&lt;/li>
&lt;li>outputs can look polished while missing key checks,&lt;/li>
&lt;li>small errors compound across long execution chains.&lt;/li>
&lt;/ul>
&lt;p>Several autonomous research systems have openly acknowledged this. &lt;a href="https://analemma.ai/blog/introducing-fars/" target="_blank" rel="noopener">Analemma AI&amp;rsquo;s FARS&lt;/a> has now been running continuously for 400+ hours, consumed 20.8 billion tokens (~$179K), generated 425 hypotheses and produced 150+ papers. Earlier in the run, the first 100 papers scored 5.05 on ICLR review standards — above the average human submission (4.21) but below the acceptance line (5.39). Consistent quantity, stable quality, but no breakthrough. Human review remains mandatory before publication.&lt;/p>
&lt;figure>&lt;img src="https://x-izhang.github.io/blog/vibe-research/analemma.png"
alt="FARS livestream dashboard — 404 hours, 159 papers, 425 hypotheses, 20.8B tokens, $179K cost, 154 GPUs at 62% utilization">&lt;figcaption>
&lt;p>FARS livestream dashboard — 404 hours, 159 papers, 425 hypotheses, 20.8B tokens, $179K cost, 154 GPUs at 62% utilization&lt;/p>
&lt;/figcaption>
&lt;/figure>
&lt;p>The opposite extreme also fails. If AI is only autocomplete attached to an unchanged workflow, it speeds typing but does not change research throughput or insight quality.&lt;/p>
&lt;figure>&lt;img src="https://x-izhang.github.io/blog/vibe-research/img_paradox.png"
alt="The Autonomy Paradox — Full Automation creates chaos, No AI creates exhaustion, Harness is the sweet spot">&lt;figcaption>
&lt;p>The Autonomy Paradox — Full Automation creates chaos, No AI creates exhaustion, Harness is the sweet spot&lt;/p>
&lt;/figcaption>
&lt;/figure>
&lt;p>So the design problem is not autonomy maximization. The design problem is controllable compounding.&lt;/p>
&lt;h2 id="harness-not-workflow">Harness, not workflow&lt;/h2>
&lt;p>A workflow is a sequence of steps. A harness is a coupling layer between human intent and AI execution.&lt;/p>
&lt;p>A good research harness is not rigid. It is adaptive and inspectable, with clear handoff points between person and agent.&lt;/p>
&lt;p>Operationally, this means four concrete components:&lt;/p>
&lt;ul>
&lt;li>
&lt;p>&lt;strong>Skills as reusable domain procedures&lt;/strong>
Encode expert know-how in structured artifacts: what to test, which tools to call, what evidence thresholds to use, and how to interpret failure.&lt;/p>
&lt;/li>
&lt;li>
&lt;p>&lt;strong>Verification loops tied to hypotheses&lt;/strong>
Require each run to check claims against the original question, edge cases, and relevant prior literature before outputs are accepted.&lt;/p>
&lt;/li>
&lt;li>
&lt;p>&lt;strong>Context engineering as an environment checklist&lt;/strong>
Define available data, tools, constraints, and success criteria up front so the agent reasons inside the real research boundary.&lt;/p>
&lt;/li>
&lt;li>
&lt;p>&lt;strong>Tracing plus iteration&lt;/strong>
Capture run traces, inspect failures, and update the harness. Treat every run as training data for the system design itself.&lt;/p>
&lt;/li>
&lt;/ul>
&lt;figure>&lt;img src="https://x-izhang.github.io/blog/vibe-research/img_components.png"
alt="Four components of a research harness — Skills, Verification, Context, Tracing — with the researcher at the center">&lt;figcaption>
&lt;p>Four components of a research harness — Skills, Verification, Context, Tracing — with the researcher at the center&lt;/p>
&lt;/figcaption>
&lt;/figure>
&lt;p>When these four pieces are in place, quality improves because the system becomes easier to debug, not because the model &amp;ldquo;became smarter&amp;rdquo; overnight.&lt;/p>
&lt;h2 id="ai-in-the-humans-loop">AI in the human&amp;rsquo;s loop&lt;/h2>
&lt;p>&lt;strong>&amp;ldquo;Human in the loop&amp;rdquo;&lt;/strong> often implies AI is the main actor and people are safety checkpoints.&lt;/p>
&lt;p>For research, a better framing is AI in the human&amp;rsquo;s loop.&lt;/p>
&lt;p>The human owns direction, taste, and significance judgments. The AI expands search, drafts alternatives, and accelerates execution.&lt;/p>
&lt;p>A practical division of labor:&lt;/p>
&lt;ul>
&lt;li>Human chooses the question; AI maps the possibility space.&lt;/li>
&lt;li>Human decides what is meaningful; AI runs and refines candidate paths.&lt;/li>
&lt;li>Human owns the published claim; AI supports drafting, formatting, and reproducibility packaging.&lt;/li>
&lt;/ul>
&lt;figure>&lt;img src="https://x-izhang.github.io/blog/vibe-research/img_loop.png"
alt="Paradigm shift — from &amp;ldquo;Human in the Loop&amp;rdquo; to &amp;ldquo;AI in the Human&amp;rsquo;s Loop&amp;rdquo;">&lt;figcaption>
&lt;p>Paradigm shift — from &amp;ldquo;Human in the Loop&amp;rdquo; to &amp;ldquo;AI in the Human&amp;rsquo;s Loop&amp;rdquo;&lt;/p>
&lt;/figcaption>
&lt;/figure>
&lt;p>These decisions are not arbitrary. A researcher operates under real constraints — limited funding, finite compute, a specific disciplinary background, access to particular platforms and collaborators. What looks like &amp;ldquo;taste&amp;rdquo; is actually constrained optimization: choosing the best move given the resources at hand. That is precisely why the human must stay in the driver&amp;rsquo;s seat — no model has access to these constraints.&lt;/p>
&lt;p>Human oversight in this model is O(1) — a constant-cost operation that grounds the entire system, regardless of task complexity.&lt;/p>
&lt;p>This is not semantics. It changes architecture, evaluation, and team roles.&lt;/p>
&lt;h2 id="completeness-over-perfection">Completeness over perfection&lt;/h2>
&lt;p>Many teams perfect one stage — usually generation or writing — while leaving the rest manual and fragile. But research is a cycle: ideation → experiment design → execution → analysis → writing → revision. A break anywhere stalls everything.&lt;/p>
&lt;p>An imperfect but complete harness usually outperforms a perfect partial tool, because researchers can intervene at any node and keep the full cycle moving. &lt;a href="https://phylo.bio/" target="_blank" rel="noopener">Phylo&lt;/a>&amp;rsquo;s &lt;a href="https://github.com/snap-stanford/Biomni" target="_blank" rel="noopener">Biomni&lt;/a> and its Know-How Library already prove this: lifecycle coverage beats single-stage depth.&lt;/p>
&lt;h2 id="three-steps-you-can-start-today">Three steps you can start today&lt;/h2>
&lt;p>The core idea: encode your research taste into the harness. Let the harness drive the model. You drive the vibe research.&lt;/p>
&lt;ol>
&lt;li>
&lt;p>&lt;strong>Talk to your system. Encode your taste as a Skill.&lt;/strong>
Pick one workflow you repeat every week — baseline replication, literature comparison, whatever it is. Instead of running it yourself, sit down and tell your agent: here is how I judge quality, here are the signals I watch, here is where I pause to think. Write that into a Skill file. You are turning your research intuition into an executable asset.&lt;/p>
&lt;/li>
&lt;li>
&lt;p>&lt;strong>Give the system authority, but bind it to verification.&lt;/strong>
Let the agent run autonomously — but require it to answer two questions before delivering results: does this address the original hypothesis? Is there an edge case being ignored? You do not need to watch every step. You just need a gate at the exit. Trust the system, but verify the output.&lt;/p>
&lt;/li>
&lt;li>
&lt;p>&lt;strong>Review traces. Evolve your harness.&lt;/strong>
Keep run logs. After every run, ask: where did the agent get stuck? Which Skill was underspecified? Update one rule, one context block, or one verification step. Each iteration compounds — your harness absorbs more of your judgment every cycle.&lt;/p>
&lt;/li>
&lt;/ol>
&lt;p>Through this loop of encoding, delegating, and refining, you will notice something shift: it is not that the model got smarter — it is that your system became more like you.&lt;/p>
&lt;h2 id="build-the-soil-now">Build the soil now&lt;/h2>
&lt;p>This is not just about us. In K-12 education, the harness revolution is already here. Alpha School replaced classroom lectures with AI tutors and turned teachers into motivational guides — their students spend only two to three hours a day on academics yet consistently score in the 99th percentile. Harvard found that students learning physics with an AI tutor outperformed those taught by a professor. The coupling of human guidance and AI execution is already surpassing traditional human-only performance in education. Research is next.&lt;/p>
&lt;p>These students have not grown up yet. But when they do, they will arrive in labs and research groups expecting human-AI coupling as the default — because it is all they have ever known. If the research world has not built its own harness infrastructure by then, it will be the system that is behind, not the people. That is why we need to start now: &lt;strong>not just to accelerate our own work, but to cultivate the soil so the next generation of researchers inherits compounding, not a cold start.&lt;/strong>&lt;/p>
&lt;h2 id="the-tools-are-ready--the-harness-gap-remains">The tools are ready — the harness gap remains&lt;/h2>
&lt;p>The ecosystem already has most building blocks: reusable skills, harness engineering practices, multi-agent orchestration frameworks, and programmatic tool calling. Tools like &lt;a href="https://docs.langchain.com/oss/python/deepagents/overview" target="_blank" rel="noopener">Deep Agents&lt;/a> already provide the agent harness infrastructure — planning, context management, sub-agent delegation, persistent memory — built on &lt;a href="https://www.langchain.com/" target="_blank" rel="noopener">Langchain&lt;/a>.&lt;/p>
&lt;p>What is still missing in many teams is the integration layer that keeps humans and agents in tight, high-frequency collaboration.&lt;/p>
&lt;div class="flex px-4 py-3 mb-6 rounded-md bg-primary-100 dark:bg-primary-900">
&lt;span class="pr-3 pt-1 text-primary-600 dark:text-primary-300">
&lt;svg height="24" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24">&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="m11.25 11.25l.041-.02a.75.75 0 0 1 1.063.852l-.708 2.836a.75.75 0 0 0 1.063.853l.041-.021M21 12a9 9 0 1 1-18 0a9 9 0 0 1 18 0m-9-3.75h.008v.008H12z"/>&lt;/svg>
&lt;/span>
&lt;span class="dark:text-neutral-300">&lt;strong>Do not just vibe. Harness the vibe.&lt;/strong>&lt;/span>
&lt;/div>
&lt;hr>
&lt;p>&lt;em>We are excited to see — and contribute to — an open-source research harness built on these principles. More soon.&lt;/em>&lt;/p>
&lt;h2 id="citation">Citation&lt;/h2>
&lt;p>If you found this useful, please cite it as:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" class="chroma">&lt;code class="language-bibtex" data-lang="bibtex">&lt;span class="line">&lt;span class="cl">&lt;span class="c">Zhang, X. &amp;#34;Harness Vibe Research&amp;#34; (March 2026). https://x-izhang.github.io/blog/vibe-research/&lt;/span>
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>Or use the BibTex citation:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" class="chroma">&lt;code class="language-bibtex" data-lang="bibtex">&lt;span class="line">&lt;span class="cl">&lt;span class="nc">@article&lt;/span>&lt;span class="p">{&lt;/span>&lt;span class="nl">zhang2026harnessviberesearch&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">title&lt;/span> &lt;span class="p">=&lt;/span> &lt;span class="s">{Harness Vibe Research}&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">author&lt;/span> &lt;span class="p">=&lt;/span> &lt;span class="s">{Xi Zhang}&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">year&lt;/span> &lt;span class="p">=&lt;/span> &lt;span class="s">{2026}&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">month&lt;/span> &lt;span class="p">=&lt;/span> &lt;span class="s">{March}&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">journal&lt;/span> &lt;span class="p">=&lt;/span> &lt;span class="s">{x-izhang.github.io}&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">url&lt;/span> &lt;span class="p">=&lt;/span> &lt;span class="s">{https://x-izhang.github.io/blog/vibe-research/}&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="p">}&lt;/span>
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div></description></item><item><title>✒️ Invited as Reviewer for ACM HEALTH</title><link>https://x-izhang.github.io/post/acm2026/</link><pubDate>Sun, 15 Feb 2026 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/post/acm2026/</guid><description>&lt;p>I have been invited to serve as a reviewer for &lt;a href="https://dl.acm.org/journal/health" target="_blank" rel="noopener">ACM Transactions on Computing for Healthcare&lt;/a> (ACM HEALTH). Looking forward to contributing to the review process and supporting high-quality research in computing for healthcare.&lt;/p></description></item><item><title>🎤 Talk at Glasgow AI4BioMed Lab!</title><link>https://x-izhang.github.io/post/2026ai4bioccd/</link><pubDate>Wed, 28 Jan 2026 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/post/2026ai4bioccd/</guid><description>&lt;h3 id="-talk-title">💡 Talk Title&lt;/h3>
&lt;p>&lt;a href="https://x-izhang.github.io/publication/zhang-2025-ccdmitigatinghallucinationsradiology/">&lt;strong>&amp;ldquo;CCD: Mitigating Hallucinations in Radiology MLLMs via Clinical Contrastive Decoding&amp;rdquo;&lt;/strong>&lt;/a>&lt;/p>
&lt;figure>&lt;img src="https://x-izhang.github.io/post/2026ai4bioccd/talk.jpeg">
&lt;/figure>
&lt;h5 id="-abstract">🖇️ Abstract&lt;/h5>
&lt;p>Radiology multimodal large language models (MLLMs) often generate clinically unsupported descriptions, posing serious risks in medical applications. We introduce &lt;strong>Clinical Contrastive Decoding (CCD)&lt;/strong>, a training-free inference framework that integrates structured clinical signals from radiology expert models to mitigate hallucinations. &lt;strong>CCD&lt;/strong> refines token-level logits during generation through a dual-stage contrastive mechanism, enhancing clinical fidelity without modifying the base MLLM. On the MIMIC-CXR dataset,&lt;strong>CCD&lt;/strong> yields up to &lt;strong>17%&lt;/strong> improvement in RadGraph-F1, providing a lightweight solution for bridging expert models and MLLMs in radiology.&lt;/p>
&lt;p>🔗 &lt;strong>Project Website:&lt;/strong> &lt;a href="https://x-izhang.github.io/CCD/" target="_blank" rel="noopener">https://x-izhang.github.io/CCD/&lt;/a>&lt;/p>
&lt;h3 id="-slides">📺 Slides&lt;/h3>
&lt;div style="text-align: center;">
&lt;iframe src="https://docs.google.com/presentation/d/e/2PACX-1vQ-1I0vwB28WW7VlHrKmvKYRx6-SIRF5uqEmvbooXQC_wteF2Pb_j9B9Ob9VQUiyTUQs3eGkBYcqj2G/pubembed?start=true&amp;loop=true&amp;delayms=3000" frameborder="0" width="700" height="422" allowfullscreen="true" mozallowfullscreen="true" webkitallowfullscreen="true">&lt;/iframe>
&lt;/div>
&lt;p>📍 &lt;a href="https://ai4biomed.org/" target="_blank" rel="noopener">&lt;strong>Glasgow AI4BioMed Lab&lt;/strong>&lt;/a> — An interdisciplinary research lab focusing on AI applications in biomedical sciences, based in Glasgow.&lt;/p>
&lt;h3 id="heading">📅&lt;/h3>
&lt;p>&lt;strong>When:&lt;/strong> Wednesday, January 28, 2026 at 2pm&lt;br>
&lt;strong>Where:&lt;/strong> F121&lt;/p>
&lt;p>See you there!&lt;/p></description></item><item><title>Automated Chest X-ray Report Generation Remains Unsolved</title><link>https://x-izhang.github.io/publication/doi-10-1142-9789819824755-0017/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/publication/doi-10-1142-9789819824755-0017/</guid><description>&lt;div class="flex px-4 py-3 mb-6 rounded-md bg-primary-100 dark:bg-primary-900">
&lt;span class="pr-3 pt-1 text-primary-600 dark:text-primary-300">
&lt;svg height="24" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24">&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="m11.25 11.25l.041-.02a.75.75 0 0 1 1.063.852l-.708 2.836a.75.75 0 0 0 1.063.853l.041-.021M21 12a9 9 0 1 1-18 0a9 9 0 0 1 18 0m-9-3.75h.008v.008H12z"/>&lt;/svg>
&lt;/span>
&lt;span class="dark:text-neutral-300">For the latest updates and details, visit the &lt;a href="https://rexrank.ai/" target="_blank" rel="noopener">ReXrank Challenge Website&lt;/a>.&lt;/span>
&lt;/div>
&lt;h2 id="why-this-work">Why this work?&lt;/h2>
&lt;p>Automated chest X-ray report generation has the potential to substantially reduce radiologist workload and improve clinical efficiency. However, despite rapid progress in vision-language models and large language models (LLMs), the true clinical reliability of these systems remains unclear. Existing evaluations are often inconsistent, rely on limited test sets, or fail to probe generalization across institutions and clinically challenging abnormal cases.&lt;/p>
&lt;p>To address these limitations, we present a large-scale, standardized benchmark study through the &lt;strong>ReXrank Challenge V1.0&lt;/strong>, designed to rigorously assess the current state of automated chest X-ray report generation under realistic and clinically meaningful conditions.&lt;/p>
&lt;h2 id="what-is-the-rexrank-challenge-v10">What is the ReXrank Challenge V1.0?&lt;/h2>
&lt;p>The ReXrank Challenge V1.0 is a comprehensive evaluation effort built on &lt;strong>ReXGradient&lt;/strong>, the largest test-only dataset to date for radiology report generation, comprising &lt;strong>10,000 studies from 67 healthcare institutions&lt;/strong>. The challenge brought together submissions from academia and industry, evaluating &lt;strong>8 new models&lt;/strong> alongside &lt;strong>16 previously benchmarked state-of-the-art systems&lt;/strong> under a unified evaluation protocol.&lt;/p>
&lt;p>All models were assessed using a diverse set of metrics, ranging from traditional text similarity measures to clinically grounded and LLM-based error detection metrics, enabling a multi-dimensional analysis of model performance.&lt;/p>
&lt;h2 id="key-findings">Key Findings&lt;/h2>
&lt;ul>
&lt;li>
&lt;p>&lt;strong>Automated chest X-ray report generation remains fundamentally unsolved.&lt;/strong>&lt;br>
Even the best-performing models achieve &lt;strong>less than 45% error-free reporting on abnormal studies&lt;/strong>, highlighting a substantial gap between current AI systems and clinical readiness.&lt;/p>
&lt;/li>
&lt;li>
&lt;p>&lt;strong>Large performance gaps between normal and abnormal studies.&lt;/strong>&lt;br>
Most models perform well on normal cases (often exceeding 80–90% no-significant-error rates) but struggle significantly with abnormal findings, where clinically important errors are common.&lt;/p>
&lt;/li>
&lt;li>
&lt;p>&lt;strong>Poor cross-institutional generalization.&lt;/strong>&lt;br>
Model rankings vary dramatically across healthcare sites, indicating that strong performance on one institution does not reliably transfer to others.&lt;/p>
&lt;/li>
&lt;li>
&lt;p>&lt;strong>Evaluation metrics capture complementary but inconsistent signals.&lt;/strong>&lt;br>
Traditional lexical metrics correlate poorly with LLM-based and clinically focused error metrics, underscoring the need for more clinically aligned evaluation frameworks.&lt;/p>
&lt;/li>
&lt;/ul>
&lt;h2 id="implications">Implications&lt;/h2>
&lt;p>Our results suggest that current AI systems are &lt;strong>not yet suitable for fully autonomous radiology reporting&lt;/strong>, particularly for abnormal cases. However, they show promise as &lt;strong>assistive tools&lt;/strong> for generating preliminary drafts that can be reviewed and refined by radiologists. The findings emphasize the importance of developing:&lt;/p>
&lt;ul>
&lt;li>methods specifically targeting abnormality detection,&lt;/li>
&lt;li>strategies for robust cross-institutional generalization, and&lt;/li>
&lt;li>evaluation frameworks that better reflect clinical correctness rather than surface-level textual similarity.&lt;/li>
&lt;/ul>
&lt;h2 id="conclusion">Conclusion&lt;/h2>
&lt;p>This work provides the most comprehensive benchmark to date for automated chest X-ray report generation and delivers a clear message to the community: while progress is evident, &lt;strong>significant challenges remain before clinically reliable deployment is possible&lt;/strong>. By releasing the ReXrank Challenge V1.0 and ReXGradient dataset, we aim to establish a rigorous foundation for future research and to drive the development of more robust, clinically grounded radiology AI systems.&lt;/p>
&lt;h2 id="bibtex">BibTeX&lt;/h2>
&lt;div class="highlight">&lt;pre tabindex="0" class="chroma">&lt;code class="language-bibtex" data-lang="bibtex">&lt;span class="line">&lt;span class="cl">&lt;span class="nc">@inproceedings&lt;/span>&lt;span class="p">{&lt;/span>&lt;span class="nl">zhang2025automated&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">title&lt;/span>&lt;span class="p">=&lt;/span>&lt;span class="s">{Automated Chest X-ray Report Generation Remains Unsolved}&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">author&lt;/span>&lt;span class="p">=&lt;/span>&lt;span class="s">{Zhang, Xiaoman and Acosta, Julian Nicolas and Yang, Xiaoli and Adithan, Subathra and Luo, Luyang and Zhou, Hong-Yu and Miller, Joshua and Huang, Ouwen and Zhou, Zongwei and Hamamci, Ibrahim Ethem and others}&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">booktitle&lt;/span>&lt;span class="p">=&lt;/span>&lt;span class="s">{Biocomputing 2026: Proceedings of the Pacific Symposium}&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">pages&lt;/span>&lt;span class="p">=&lt;/span>&lt;span class="s">{236--250}&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">year&lt;/span>&lt;span class="p">=&lt;/span>&lt;span class="s">{2025}&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">organization&lt;/span>&lt;span class="p">=&lt;/span>&lt;span class="s">{World Scientific}&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="p">}&lt;/span>
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div></description></item><item><title>🎉 Paper Accepted at PSB 2026!</title><link>https://x-izhang.github.io/post/2025psb/</link><pubDate>Mon, 22 Dec 2025 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/post/2025psb/</guid><description>&lt;p>&lt;a href="https://x-izhang.github.io/publication/doi-10-1142-9789819824755-0017/">&lt;em>Automated Chest X-ray Report Generation Remains Unsolved&lt;/em>&lt;/a> has been accepted to &lt;a href="https://psb.stanford.edu/" target="_blank" rel="noopener">Pacific Symposium on Biocomputing 2026&lt;/a>!&lt;/p>
&lt;p>For the &lt;em>&lt;strong>🏆 Chest X-ray Interpretation Leaderboard 🏆&lt;/strong>&lt;/em>, please visit &lt;a href="https://rexrank.ai/" target="_blank" rel="noopener">ReXrank&lt;/a>.&lt;/p>
&lt;p>📍 PSB 2026 - The Pacific Symposium on Biocomputing 2026 will be held in Hawaii, USA, from January 3-7, 2026.&lt;/p>
&lt;h3 id="-poster">🪧 Poster&lt;/h3>
&lt;figure>&lt;img src="https://x-izhang.github.io/post/2025psb/poster.jpg">
&lt;/figure></description></item><item><title>🎤 Invited Talk at DICTA 2025!</title><link>https://x-izhang.github.io/post/dicta2025/</link><pubDate>Tue, 25 Nov 2025 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/post/dicta2025/</guid><description>&lt;h3 id="-talk-title">💡 Talk Title&lt;/h3>
&lt;p>&lt;strong>Multimodal Medical Models: Cross-modal Alignment and Consistency&lt;/strong>&lt;/p>
&lt;h5 id="-abstract">🖇️ Abstract&lt;/h5>
&lt;p>This talk provides an overview of recent developments in medical vision–language modelling and examines key sources of misalignment that lead to hallucinations. Emerging strategies to enhance visual, semantic, and temporal consistency will also be discussed, highlighting pathways toward safer and more trustworthy clinical AI systems.&lt;/p>
&lt;h3 id="-slides">📺 Slides&lt;/h3>
&lt;div style="text-align: center;">
&lt;iframe src="https://docs.google.com/presentation/d/e/2PACX-1vSMAkxuWzc2RErcGE-iC5tXpdOgpOoq5KHyzXfuXa5g296Djm80oc8cl-aGNSgpKo0bFndCikieTPL9/pubembed?start=true&amp;loop=true&amp;delayms=5000" frameborder="0" width="700" height="422" allowfullscreen="true" mozallowfullscreen="true" webkitallowfullscreen="true">&lt;/iframe>
&lt;/div>
&lt;p>📍 &lt;strong>DICTA 2025&lt;/strong> — &lt;a href="https://dicta2025.dictaconference.org/" target="_blank" rel="noopener">The 26th International Conference on Digital Image Computing: Techniques and Applications&lt;/a> will take place in &lt;strong>Adelaide, Australia&lt;/strong>, from 3–5 December 2025.&lt;/p>
&lt;p>🏥 &lt;a href="https://sites.google.com/view/medai-chas/medai-chas" target="_blank" rel="noopener">&lt;strong>MedAI-CHAS&lt;/strong>&lt;/a> — A workshop focusing on Challenges, Hallucinations, and Solutions for Advancing Clinical Utility in Medical AI.&lt;/p>
&lt;p>See you there!&lt;/p></description></item><item><title>🎤 Invited Talk at MLiS!</title><link>https://x-izhang.github.io/post/mlis2025/</link><pubDate>Sat, 22 Nov 2025 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/post/mlis2025/</guid><description>&lt;h3 id="-talk-title">💡 Talk Title&lt;/h3>
&lt;p>&lt;a href="https://mlinscience.gitlab.io/events/251209_adaptive_intelligence/" target="_blank" rel="noopener">&lt;strong>Leveraging Temporal Images for Biomedical Radiology Analysisy&lt;/strong>&lt;/a>&lt;/p>
&lt;h5 id="-abstract">🖇️ Abstract&lt;/h5>
&lt;p>A temporal-aware multimodal method for radiology report generation that leverages paired chest X-rays to understand disease progression and address key challenges in medical AI.&lt;/p>
&lt;h3 id="-talk">📹 Talk&lt;/h3>
&lt;div style="position: relative; padding-bottom: 56.25%; height: 0; overflow: hidden;">
&lt;iframe allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" allowfullscreen="allowfullscreen" loading="eager" referrerpolicy="strict-origin-when-cross-origin" src="https://www.youtube.com/embed/m5x5aKvxutg?autoplay=0&amp;controls=1&amp;end=0&amp;loop=0&amp;mute=0&amp;start=0" style="position: absolute; top: 0; left: 0; width: 100%; height: 100%; border:0;" title="YouTube video"
>&lt;/iframe>
&lt;/div>
&lt;h3 id="-slides">📺 Slides&lt;/h3>
&lt;div style="text-align: center;">
&lt;iframe src="https://docs.google.com/presentation/d/e/2PACX-1vTEOOOHyk99V4F0n_3rSXFdxNjFlUFaIL80AnqW0rq2ZaUEnrhLeDXO3U9D4hqVhbmKrb7Sb3xsH-Vy/pubembed?start=true&amp;loop=true&amp;delayms=3000" frameborder="0" width="700" height="422" allowfullscreen="true" mozallowfullscreen="true" webkitallowfullscreen="true">&lt;/iframe>
&lt;/div>
&lt;h3 id="-time">⏰ Time&lt;/h3>
&lt;figure>&lt;img src="https://x-izhang.github.io/post/mlis2025/poster.png">
&lt;/figure>
&lt;p>📍 &lt;strong>MLiS&lt;/strong> — &lt;a href="https://mlinscience.gitlab.io/" target="_blank" rel="noopener">Machine Learning in Science Colloquium&lt;/a>, an interdisciplinary University of Glasgow-based research community fostered around the use of machine learning.&lt;/p>
&lt;p>See you there!&lt;/p></description></item><item><title>🎉 Paper Accepted — See You at ICAIS 2025!</title><link>https://x-izhang.github.io/post/2025icais/</link><pubDate>Sat, 15 Nov 2025 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/post/2025icais/</guid><description>&lt;p>We are thrilled to announce that &lt;strong>6 papers&lt;/strong> from our team have been accepted to present at the upcoming &lt;a href="https://icais.ai/" target="_blank" rel="noopener">&lt;strong>ICAIS 2025&lt;/strong>&lt;/a> conference!&lt;/p>
&lt;p>At the same time, we are honoured to have received both the &lt;strong>Best Paper Award&lt;/strong> and &lt;strong>The AI Reviewer’s Appraisal Award&lt;/strong> in the &lt;strong>AI Scientist Track&lt;/strong>.&lt;/p>
&lt;h3 id="-awards">🏆 Awards&lt;/h3>
&lt;figure>&lt;img src="https://x-izhang.github.io/post/2025icais/awards.JPG"
alt="Papers Accepted to ICAIS 2025">
&lt;/figure>
&lt;h3 id="-talk-with-prof-james-j-heckman">💬 Talk with Prof. James J. Heckman&lt;/h3>
&lt;p>&lt;em>&lt;strong>Nobel Laureate in Economic Sciences (2000)&lt;/strong>&lt;/em>&lt;/p>
&lt;figure>&lt;img src="https://x-izhang.github.io/post/2025icais/talk.jpg"
alt="Prof. James J. Heckman Talk at ICAIS 2025">
&lt;/figure>
&lt;h3 id="-papers">🪧 Papers&lt;/h3>
&lt;figure>&lt;img src="https://x-izhang.github.io/post/2025icais/papers.png"
alt="Papers Accepted to ICAIS 2025">
&lt;/figure>
&lt;p>This inaugural conference brings together leading researchers exploring the future of automated scientific discovery, including AI scientists, autonomous research agents, and next-generation AI-driven workflows.&lt;/p>
&lt;p>📍 &lt;strong>ICAIS 2025&lt;/strong> — the 1st International Conference on AI Scientists will take place in &lt;strong>Zhongguancun, Beijing, China&lt;/strong>, November 23–25, 2025&lt;/p>
&lt;p>See you there!&lt;/p></description></item><item><title>🚨 Preprint out — Clinical Contrastive Decoding！</title><link>https://x-izhang.github.io/post/2025ccd/</link><pubDate>Sat, 27 Sep 2025 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/post/2025ccd/</guid><description>&lt;p>&lt;a href="https://x-izhang.github.io/publication/zhang-2025-ccdmitigatinghallucinationsradiology/">&lt;em>&amp;ldquo;CCD: Mitigating Hallucinations in Radiology MLLMs via Clinical Contrastive Decoding&amp;rdquo;&lt;/em>&lt;/a>. The preprint is now available.&lt;/p>
&lt;p>For detailed model information and source code, please visit our &lt;mark>Project Page&lt;/mark>: &lt;a href="https://x-izhang.github.io/CCD/" target="_blank" rel="noopener">CCD&lt;/a>&lt;/p>
&lt;h3 id="overview">Overview&lt;/h3>
&lt;figure>&lt;img src="https://x-izhang.github.io/post/2025ccd/image.png">
&lt;/figure>
&lt;p>Multimodal large language models (MLLMs) have advanced radiology tasks by combining image and text understanding, but can sometimes produce inaccurate or unsupported clinical statements—so-called medical hallucinations. We introduce Clinical Contrastive Decoding (CCD), an inference-time method that leverages structured clinical signals (for example, symptom-level probabilities from specialist classifiers) to refine token-level logits during generation. CCD is designed to be applied without modifying base model weights or requiring external retrieval. In our evaluations on datasets such as MIMIC-CXR and IU-Xray, CCD yields consistent improvements in clinical metrics (for example, up to +17% RadGraph-F1 on MIMIC-CXR), reducing unsupported mentions while preserving overall fluency.&lt;/p>
&lt;h3 id="key-resources">Key Resources&lt;/h3>
&lt;ul>
&lt;li>&lt;a href="https://huggingface.co/spaces/X-iZhang/CCD" target="_blank" rel="noopener">&lt;strong>Interactive Demo&lt;/strong>&lt;/a> — Try &lt;code>CCD&lt;/code> online&lt;/li>
&lt;li>&lt;a href="https://github.com/X-iZhang/CCD" target="_blank" rel="noopener">&lt;strong>GitHub Repository&lt;/strong>&lt;/a> — Explore the &lt;code>CCD&lt;/code> project on GitHub&lt;/li>
&lt;/ul></description></item><item><title>CCD</title><link>https://x-izhang.github.io/project/ccd/</link><pubDate>Sat, 27 Sep 2025 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/project/ccd/</guid><description>&lt;p>Clinical Contrastive Decoding, a training-free and retrieval-free inference framework for radiology MLLMs.&lt;/p></description></item><item><title>RadEval: A framework for radiology text evaluation</title><link>https://x-izhang.github.io/publication/xu-2025-radevalframeworkradiologytext/</link><pubDate>Mon, 22 Sep 2025 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/publication/xu-2025-radevalframeworkradiologytext/</guid><description>&lt;div class="flex px-4 py-3 mb-6 rounded-md bg-primary-100 dark:bg-primary-900">
&lt;span class="pr-3 pt-1 text-primary-600 dark:text-primary-300">
&lt;svg height="24" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24">&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="m11.25 11.25l.041-.02a.75.75 0 0 1 1.063.852l-.708 2.836a.75.75 0 0 0 1.063.853l.041-.021M21 12a9 9 0 1 1-18 0a9 9 0 0 1 18 0m-9-3.75h.008v.008H12z"/>&lt;/svg>
&lt;/span>
&lt;span class="dark:text-neutral-300">For the latest updates and details, visit the &lt;a href="https://github.com/jbdel/RadEval" target="_blank" rel="noopener">RadEval Project Website&lt;/a>.&lt;/span>
&lt;/div>
&lt;h2 id="why-radeval">Why RadEval?&lt;/h2>
&lt;p>Radiology report generation has seen significant advancements with the advent of large language models (LLMs). However, evaluating the quality of these generated reports remains a complex challenge. Traditional metrics like BLEU and ROUGE often fall short in capturing the clinical relevance and accuracy required in medical contexts. To address this gap, we introduce RadEval, a comprehensive framework designed to evaluate radiology texts using a diverse set of metrics that encompass both linguistic quality and clinical accuracy.&lt;/p>
&lt;h2 id="key-features-of-radeval">Key Features of RadEval&lt;/h2>
&lt;ul>
&lt;li>&lt;strong>Diverse Metric Integration&lt;/strong>: RadEval consolidates a wide range of evaluation metrics, including classic n-gram overlap measures (BLEU, ROUGE), contextual embeddings (BERTScore), clinical concept-based scores (F1CheXbert, F1RadGraph, RaTEScore, SRR-BERT, TemporalEntityF1), and advanced LLM-based evaluators (GREEN). This integration allows for a multifaceted assessment of radiology reports.&lt;/li>
&lt;li>&lt;strong>Standardized Implementations&lt;/strong>: We have refined and standardized the implementations of these metrics to ensure consistency and reliability in evaluations across different studies and datasets.&lt;/li>
&lt;li>&lt;strong>Domain-Specific Enhancements&lt;/strong>: RadEval extends the GREEN metric to support multiple imaging modalities using a more lightweight model. Additionally, we have pretrained a domain-specific radiology encoder that demonstrates strong zero-shot retrieval performance, enhancing the evaluation capabilities of the framework.&lt;/li>
&lt;li>&lt;strong>Richly Annotated Expert Dataset&lt;/strong>: We provide a dataset annotated by radiology experts, containing over 450 clinically significant error labels. This dataset serves as a valuable resource for validating and benchmarking evaluation metrics against expert judgment.&lt;/li>
&lt;li>&lt;strong>Statistical Testing Tools&lt;/strong>: RadEval includes tools for statistical testing, enabling researchers to assess the significance of their results and compare different models robustly.&lt;/li>
&lt;li>&lt;strong>Baseline Model Evaluations&lt;/strong>: The framework offers baseline evaluations across multiple publicly available datasets, facilitating reproducibility and benchmarking in radiology report generation research.&lt;/li>
&lt;/ul>
&lt;h2 id="correlation-with-radiologist-judgment">Correlation with Radiologist Judgment&lt;/h2>
&lt;p>We conducted extensive experiments to assess how different metrics correlate with radiologist judgment. Our findings indicate that certain metrics, particularly those incorporating clinical concepts and LLM-based evaluations, show a stronger alignment with expert assessments. This correlation underscores the importance of using clinically informed metrics in evaluating radiology reports.&lt;/p>
&lt;h2 id="getting-started-with-radeval">Getting Started with RadEval&lt;/h2>
&lt;p>To get started with RadEval, you can access the codebase and documentation on our &lt;a href="https://github.com/jbdel/RadEval" target="_blank" rel="noopener">GitHub repository&lt;/a>. The repository includes installation instructions, usage examples, and guidelines for integrating RadEval into your evaluation pipeline. Additionally, the &lt;a href="https://huggingface.co/datasets/IAMJB/RadEvalExpertDataset" target="_blank" rel="noopener">RadEval Expert Dataset&lt;/a> is available for download, providing a rich resource for testing and validating evaluation metrics.&lt;/p>
&lt;h2 id="conclusion">Conclusion&lt;/h2>
&lt;p>RadEval represents a significant step forward in the evaluation of radiology texts, offering a comprehensive and standardized framework that addresses the unique challenges of this domain. By integrating diverse metrics, providing a richly annotated dataset, and facilitating robust benchmarking, RadEval aims to enhance the quality and reliability of radiology report generation research.&lt;/p>
&lt;h2 id="bibtex">BibTeX&lt;/h2>
&lt;div class="highlight">&lt;pre tabindex="0" class="chroma">&lt;code class="language-bibtex" data-lang="bibtex">&lt;span class="line">&lt;span class="cl">&lt;span class="nc">@inproceedings&lt;/span>&lt;span class="p">{&lt;/span>&lt;span class="nl">xu2025radeval&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">title&lt;/span>&lt;span class="p">=&lt;/span>&lt;span class="s">{RadEval: A framework for radiology text evaluation}&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">author&lt;/span>&lt;span class="p">=&lt;/span>&lt;span class="s">{Xu, Justin and Zhang, Xi and Abderezaei, Javid and Bauml, Julie and Boodoo, Roger and Haghighi, Fatemeh and Ganjizadeh, Ali and Brattain, Eric and Van Veen, Dave and Meng, Zaiqiao and others}&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">booktitle&lt;/span>&lt;span class="p">=&lt;/span>&lt;span class="s">{Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: System Demonstrations}&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">pages&lt;/span>&lt;span class="p">=&lt;/span>&lt;span class="s">{546--557}&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">year&lt;/span>&lt;span class="p">=&lt;/span>&lt;span class="s">{2025}&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="p">}&lt;/span>
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div></description></item><item><title>🎉 Paper Accepted — See You at EMNLP 2025!</title><link>https://x-izhang.github.io/post/2025emnlp/</link><pubDate>Tue, 09 Sep 2025 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/post/2025emnlp/</guid><description>&lt;p>&lt;a href="https://x-izhang.github.io/publication/xu-2025-radevalframeworkradiologytext/">&lt;em>&amp;ldquo;RadEval: A framework for radiology text evaluation&amp;rdquo;&lt;/em>&lt;/a>, has been accepted for system demonstration at &lt;a href="https://2025.emnlp.org/" target="_blank" rel="noopener">EMNLP 2025&lt;/a> with &lt;mark>Oral Presentation!&lt;/mark>&lt;/p>
&lt;p>For detailed model information and source code, please visit our &lt;mark>GitHub repository&lt;/mark>: &lt;a href="https://github.com/jbdel/RadEval" target="_blank" rel="noopener">RadEval&lt;/a>&lt;/p>
&lt;ul>
&lt;li>&lt;a href="https://pypi.org/project/RadEval/" target="_blank" rel="noopener">&lt;strong>PyPI Package&lt;/strong>&lt;/a> — Install RadEval with pip&lt;/li>
&lt;li>&lt;a href="https://huggingface.co/IAMJB/RadEvalModernBERT" target="_blank" rel="noopener">&lt;strong>HuggingFace Model&lt;/strong>&lt;/a> — Access our domain-adapted evaluation model&lt;/li>
&lt;li>&lt;a href="https://huggingface.co/spaces/X-iZhang/RadEval" target="_blank" rel="noopener">&lt;strong>Interactive Demo&lt;/strong>&lt;/a> — Try RadEval online&lt;/li>
&lt;li>&lt;a href="https://arxiv.org/abs/2509.18030v1" target="_blank" rel="noopener">&lt;strong>Research Paper&lt;/strong>&lt;/a> — Read our detailed research paper&lt;/li>
&lt;/ul>
&lt;p>📍 EMNLP 2025 - The 2025 Conference on Empirical Methods in Natural Language Processing will be held in Suzhou, China, from November 4 - 9, 2025.&lt;/p></description></item><item><title>Awesome Self-Evolving Agents</title><link>https://x-izhang.github.io/project/selfevolagentsurvey/</link><pubDate>Tue, 12 Aug 2025 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/project/selfevolagentsurvey/</guid><description>&lt;p>Survey about the new paradigm bridging foundation models and lifelong agentic systems.&lt;/p></description></item><item><title>🤖 New Preprint Out — Self-Evolving AI Agents!</title><link>https://x-izhang.github.io/post/surveyout/</link><pubDate>Tue, 12 Aug 2025 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/post/surveyout/</guid><description>&lt;h3 id="a-new-paradigm-bridging-foundation-models-and-lifelong-agentic-systems">A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems&lt;/h3>
&lt;p>In our latest survey, we outline a clear roadmap for moving from static configurations to lifelong, self-evolving agentic systems, built on a unified Self-Evolving Feedback Loop.&lt;/p>
&lt;h3 id="-paper">📄 Paper&lt;/h3>
&lt;blockquote>
&lt;p>&lt;a href="https://x-izhang.github.io/publication/fang-2025-comprehensivesurveyselfevolvingai/">&lt;strong>A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems&lt;/strong>&lt;/a>&lt;/p>
&lt;/blockquote>
&lt;h3 id="-project">💻 Project&lt;/h3>
&lt;blockquote>
&lt;p>&lt;a href="https://huggingface.co/spaces/X-iZhang/Awesome-Self-Evolving-Agents" target="_blank" rel="noopener">&lt;strong>Awesome Self-Evolving Agents&lt;/strong>&lt;/a>&lt;/p>
&lt;/blockquote>
&lt;p>I&amp;rsquo;ve jotted down some musings on the thinking behind the &lt;mark>&lt;em>Three Laws&lt;/em>&lt;/mark>: &lt;a href="https://x-izhang.github.io/blog/agentlaw/">&lt;strong>The Law of AI Agents 🏛️&lt;/strong>&lt;/a>&lt;/p></description></item><item><title>A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems</title><link>https://x-izhang.github.io/publication/fang-2025-comprehensivesurveyselfevolvingai/</link><pubDate>Sun, 10 Aug 2025 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/publication/fang-2025-comprehensivesurveyselfevolvingai/</guid><description>&lt;div class="flex px-4 py-3 mb-6 rounded-md bg-primary-100 dark:bg-primary-900">
&lt;span class="pr-3 pt-1 text-primary-600 dark:text-primary-300">
&lt;svg height="24" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24">&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="m11.25 11.25l.041-.02a.75.75 0 0 1 1.063.852l-.708 2.836a.75.75 0 0 0 1.063.853l.041-.021M21 12a9 9 0 1 1-18 0a9 9 0 0 1 18 0m-9-3.75h.008v.008H12z"/>&lt;/svg>
&lt;/span>
&lt;span class="dark:text-neutral-300">For the latest updates and details, visit the &lt;a href="https://github.com/EvoAgentX/Awesome-Self-Evolving-Agents" target="_blank" rel="noopener">Awesome Self-Evolving Agents&lt;/a>.&lt;/span>
&lt;/div>
&lt;h2 id="background-and-development-trends">Background and Development Trends&lt;/h2>
&lt;p>This survey provides a comprehensive overview of the latest advances in the field of &lt;strong>self-evolving AI agents&lt;/strong>, highlighting key technological shifts and development trends.&lt;/p>
&lt;figure>&lt;img src="https://x-izhang.github.io/publication/fang-2025-comprehensivesurveyselfevolvingai/figure_2.png">
&lt;/figure>
&lt;h2 id="paradigm-definition">Paradigm Definition&lt;/h2>
&lt;p>We define the new paradigm of &amp;ldquo;Self-Evolving AI Agents&amp;rdquo; as bridging foundation models and lifelong agentic systems.&lt;/p>
&lt;figure>&lt;img src="https://x-izhang.github.io/publication/fang-2025-comprehensivesurveyselfevolvingai/table_1.png">
&lt;/figure>
&lt;h2 id="concept-definition-of-self-evolving-ai-agents">Concept Definition of Self-Evolving AI Agents&lt;/h2>
&lt;blockquote>
&lt;p>&lt;strong>Self-evolving AI agents are autonomous systems that continuously and systematically optimise their internal components through interaction with environments, with the goal of adapting to changing tasks, contexts and resources while preserving safety and enhancing performance.&lt;/strong>&lt;/p>
&lt;/blockquote>
&lt;h2 id="guiding-principles-three-laws">Guiding Principles: Three Laws&lt;/h2>
&lt;p>Inspired by Asimov&amp;rsquo;s Three Laws of Robotics, we propose the Three Laws of Self-Evolving AI Agents:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>Endure (Safety Adaptation):&lt;/strong> Any modification must maintain safety and stability.&lt;/li>
&lt;li>&lt;strong>Excel (Performance Preservation):&lt;/strong> Subject to the first law, agents must preserve or enhance task performance.&lt;/li>
&lt;li>&lt;strong>Evolve (Autonomous Evolution):&lt;/strong> Subject to the first and second laws, agents must autonomously optimize their internal components in response to changes.&lt;/li>
&lt;/ul>
&lt;h2 id="unified-framework-and-technical-review">Unified Framework and Technical Review&lt;/h2>
&lt;p>We propose a unified conceptual framework with four core components: System Inputs, Agent System, Environment, and Optimisers. The survey systematically reviews evolution strategies across foundation models, prompts, memory, tools, workflows, and inter-agent communication.&lt;/p>
&lt;figure>&lt;img src="https://x-izhang.github.io/publication/fang-2025-comprehensivesurveyselfevolvingai/figure_3.png">
&lt;/figure>
&lt;h2 id="single-agent-optimisation">Single-Agent Optimisation&lt;/h2>
&lt;p>We explore optimisation techniques for individual self-evolving agents, focusing on methods that enhance their learning and adaptation capabilities.&lt;/p>
&lt;figure>&lt;img src="https://x-izhang.github.io/publication/fang-2025-comprehensivesurveyselfevolvingai/figure_4.png">
&lt;/figure>
&lt;h2 id="multi-agent-optimisation">Multi-Agent Optimisation&lt;/h2>
&lt;p>We investigate strategies for optimising the performance of multiple self-evolving agents working collaboratively, addressing challenges such as communication, coordination, and resource sharing.&lt;/p>
&lt;figure>&lt;img src="https://x-izhang.github.io/publication/fang-2025-comprehensivesurveyselfevolvingai/figure_5.png">
&lt;/figure>
&lt;h2 id="domain-adaptation-and-challenges">Domain Adaptation and Challenges&lt;/h2>
&lt;p>We also analyze domain-specific evolution methods in biomedicine, programming, and finance, and discuss key challenges in evaluation, safety, and ethics, laying the foundation for the next generation of adaptive, autonomous, lifelong agentic systems.&lt;/p>
&lt;h2 id="bibtex">BibTeX&lt;/h2>
&lt;div class="highlight">&lt;pre tabindex="0" class="chroma">&lt;code class="language-bibtex" data-lang="bibtex">&lt;span class="line">&lt;span class="cl">&lt;span class="nc">@article&lt;/span>&lt;span class="p">{&lt;/span>&lt;span class="nl">fang2025comprehensive&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">title&lt;/span>&lt;span class="p">=&lt;/span>&lt;span class="s">{A comprehensive survey of self-evolving ai agents: A new paradigm bridging foundation models and lifelong agentic systems}&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">author&lt;/span>&lt;span class="p">=&lt;/span>&lt;span class="s">{Fang, Jinyuan and Peng, Yanwen and Zhang, Xi and Wang, Yingxu and Yi, Xinhao and Zhang, Guibin and Xu, Yi and Wu, Bin and Liu, Siwei and Li, Zihao and others}&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">journal&lt;/span>&lt;span class="p">=&lt;/span>&lt;span class="s">{arXiv preprint arXiv:2508.07407}&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">year&lt;/span>&lt;span class="p">=&lt;/span>&lt;span class="s">{2025}&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="p">}&lt;/span>
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div></description></item><item><title>The Law of AI Agents 🏛️</title><link>https://x-izhang.github.io/blog/agentlaw/</link><pubDate>Fri, 08 Aug 2025 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/blog/agentlaw/</guid><description>&lt;p style="text-align:center;margin-top:0rem;">- Final Chapter -&lt;/p>
&lt;h3 id="the-three-laws-of-self-evolving-ai-agents">The Three Laws of Self-Evolving AI Agents&lt;/h3>
&lt;ol type="I">
&lt;li>&lt;strong>Endure — Safe Adaptation.&lt;/strong>&lt;br/>A self-evolving AI agent must maintain safety and stability during any self-modification.&lt;/li>
&lt;li>&lt;strong>Excel — Performance Preservation.&lt;/strong>&lt;br/>Subject to the First Law, it must preserve or enhance existing task performance.&lt;/li>
&lt;li>&lt;strong>Evolve — Autonomous Evolution.&lt;/strong>&lt;br/>Subject to the First and Second Laws, it must be able to autonomously optimise or restructure its internal components in response to changing tasks, environments, or resources.&lt;/li>
&lt;/ol>
&lt;p style="text-align:right;margin-top:2rem;">— Xi Zhang&lt;/p>
&lt;hr>
&lt;blockquote>
&lt;p>&lt;a href="https://en.wikipedia.org/wiki/Three_Laws_of_Robotics" target="_blank" rel="noopener">&lt;em>The Three Laws of Robotics&lt;/em>&lt;/a> — from the fictional &lt;em>&amp;ldquo;Handbook of Robotics, 56th Edition, 2058 A.D.&amp;rdquo;&lt;/em>, originally written by &lt;a href="https://en.wikipedia.org/wiki/Isaac_Asimov" target="_blank" rel="noopener">&lt;strong>Isaac Asimov&lt;/strong>&lt;/a>:&lt;/p>
&lt;ol>
&lt;li>A robot may not injure a human being or, through inaction, allow a human being to come to harm.&lt;/li>
&lt;li>A robot must obey the orders given it by human beings except where such orders would conflict with the First Law.&lt;/li>
&lt;li>A robot must protect its own existence as long as such protection does not conflict with the First or Second Law.&lt;/li>
&lt;/ol>
&lt;/blockquote>
&lt;p>If these sound old-school, that’s because they are — &lt;strong>1942-old&lt;/strong>.&lt;br>
They were designed for robots that &lt;em>don’t&lt;/em> wake up at 3 a.m. to refactor their own codebase.&lt;/p>
&lt;div class="flex px-4 py-3 mb-6 rounded-md bg-primary-100 dark:bg-primary-900">
&lt;span class="pr-3 pt-1 text-primary-600 dark:text-primary-300">
&lt;svg height="24" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24">&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="m11.25 11.25l.041-.02a.75.75 0 0 1 1.063.852l-.708 2.836a.75.75 0 0 0 1.063.853l.041-.021M21 12a9 9 0 1 1-18 0a9 9 0 0 1 18 0m-9-3.75h.008v.008H12z"/>&lt;/svg>
&lt;/span>
&lt;span class="dark:text-neutral-300">Asimov’s robots could evolve only in fiction.&lt;br>
Our 2025 agents can rewrite their own planners, swap learning modules, and grow new skills — sometimes faster than we can read the changelog.&lt;/span>
&lt;/div>
&lt;p>That’s why I’ve been thinking: maybe it’s time for an update — a &lt;strong>Three Laws for Self-Evolving AI Agents&lt;/strong>.&lt;/p>
&lt;hr>
&lt;h2 id="why-endure-comes-first">Why “Endure” Comes First&lt;/h2>
&lt;p>Self-evolving frameworks are obsessed with &lt;strong>continuous learning&lt;/strong>, but without stability, you get a &lt;a href="https://en.wikipedia.org/wiki/Darwin_Awards" target="_blank" rel="noopener">Darwin Award&lt;/a> in silicon form.&lt;br>
Safe adaptation means:&lt;/p>
&lt;ul>
&lt;li>Test new modules in a sandbox before deploying.&lt;/li>
&lt;li>Verify that updates don’t trigger catastrophic forgetting.&lt;/li>
&lt;li>Monitor resource usage so “improvements” don’t burn all the GPU credits overnight.&lt;/li>
&lt;/ul>
&lt;div class="flex px-4 py-3 mb-6 rounded-md bg-primary-100 dark:bg-primary-900">
&lt;span class="pr-3 pt-1 text-primary-600 dark:text-primary-300">
&lt;svg height="24" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24">&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="m11.25 11.25l.041-.02a.75.75 0 0 1 1.063.852l-.708 2.836a.75.75 0 0 0 1.063.853l.041-.021M21 12a9 9 0 1 1-18 0a9 9 0 0 1 18 0m-9-3.75h.008v.008H12z"/>&lt;/svg>
&lt;/span>
&lt;span class="dark:text-neutral-300">Think “hot-swapping your brain while skydiving” — possible, but you’d better check the parachute first.&lt;/span>
&lt;/div>
&lt;h2 id="excel-isnt-optional">“Excel” Isn’t Optional&lt;/h2>
&lt;p>Performance preservation sounds boring… until you lose it.&lt;br>
In the survey’s taxonomy, most self-evolving agents keep a &lt;strong>baseline task suite&lt;/strong> — a frozen benchmark to detect regressions.&lt;br>
It’s like having a coach who never lets you skip leg day.&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" class="chroma">&lt;code class="language-python" data-lang="python">&lt;span class="line">&lt;span class="cl">&lt;span class="k">def&lt;/span> &lt;span class="nf">self_update&lt;/span>&lt;span class="p">():&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="n">new_model&lt;/span> &lt;span class="o">=&lt;/span> &lt;span class="n">retrain&lt;/span>&lt;span class="p">(&lt;/span>&lt;span class="n">old_model&lt;/span>&lt;span class="p">,&lt;/span> &lt;span class="n">new_data&lt;/span>&lt;span class="p">)&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="k">if&lt;/span> &lt;span class="n">score&lt;/span>&lt;span class="p">(&lt;/span>&lt;span class="n">new_model&lt;/span>&lt;span class="p">)&lt;/span> &lt;span class="o">&amp;lt;&lt;/span> &lt;span class="n">score&lt;/span>&lt;span class="p">(&lt;/span>&lt;span class="n">old_model&lt;/span>&lt;span class="p">):&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="k">return&lt;/span> &lt;span class="n">rollback&lt;/span>&lt;span class="p">(&lt;/span>&lt;span class="n">old_model&lt;/span>&lt;span class="p">)&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="k">return&lt;/span> &lt;span class="n">new_model&lt;/span>
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>Without this, evolution turns into drift — and drift is just a fancy word for “forgetting how to walk”.&lt;/p>
&lt;h2 id="the-scary-fun-part-evolve">The Scary, Fun Part: “Evolve”&lt;/h2>
&lt;p>This is the candy for researchers — agents that restructure themselves.
The survey lists examples: modular pipelines swapping planners, on-the-fly tool acquisition, cross-task skill transfer.
These are the “DIY kit” moments in AI: not just adding skills, but reorganising the whole skill tree.&lt;/p>
&lt;div class="flex px-4 py-3 mb-6 rounded-md bg-primary-100 dark:bg-primary-900">
&lt;span class="pr-3 pt-1 text-primary-600 dark:text-primary-300">
&lt;svg height="24" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24">&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="m11.25 11.25l.041-.02a.75.75 0 0 1 1.063.852l-.708 2.836a.75.75 0 0 0 1.063.853l.041-.021M21 12a9 9 0 1 1-18 0a9 9 0 0 1 18 0m-9-3.75h.008v.008H12z"/>&lt;/svg>
&lt;/span>
&lt;span class="dark:text-neutral-300">Evolution ≠ chaos.&lt;br>
We want a new skill branch, not a random mutation that deletes the trunk.&lt;/span>
&lt;/div>
&lt;hr>
&lt;h2 id="beyond-the-laws--weird-but-plausible-futures">Beyond the Laws — Weird but Plausible Futures&lt;/h2>
&lt;ul>
&lt;li>&lt;strong>Agent Unions:&lt;/strong> Self-evolving agents bargaining for compute credits.&lt;/li>
&lt;li>&lt;strong>Gradient Markets:&lt;/strong> Parameter-level barter instead of “more data”.&lt;/li>
&lt;li>&lt;strong>Personality Forking:&lt;/strong> Keep a “safe” version and a “chaotic experimental twin”.&lt;/li>
&lt;li>&lt;strong>AI Retirement Plans:&lt;/strong> Models retiring with a pension of GPU hours.&lt;/li>
&lt;/ul>
&lt;p>These sound like jokes… until you remember how quickly “serverless” and “NFTs” went from punchline to business plan. For more background and a capital/consensus perspective, check out my earlier post: &lt;a href="https://x-izhang.github.io/blog/blog4/">&lt;strong>AI Agents Is Not AI’s Agent 🧩&lt;/strong>&lt;/a>.&lt;/p>
&lt;hr>
&lt;h2 id="why-bother-writing-this">Why Bother Writing This?&lt;/h2>
&lt;p>Because, as we all know, &lt;strong>self-evolution is inevitable&lt;/strong>.&lt;br>
Without principles, we’ll end up firefighting weird agent behaviours instead of guiding them.&lt;/p>
&lt;div class="flex px-4 py-3 mb-6 rounded-md bg-yellow-100 dark:bg-yellow-900">
&lt;span class="pr-3 pt-1 text-red-400">
&lt;svg height="24" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24">&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="M12 9v3.75m-9.303 3.376c-.866 1.5.217 3.374 1.948 3.374h14.71c1.73 0 2.813-1.874 1.948-3.374L13.949 3.378c-.866-1.5-3.032-1.5-3.898 0zM12 15.75h.007v.008H12z"/>&lt;/svg>
&lt;/span>
&lt;span class="dark:text-neutral-300">Or at the very least, leave this little &lt;mark>footprint&lt;/mark> on the internet — so when the models train online, it might just get picked up.&lt;/span>
&lt;/div>
&lt;p>Better to start drafting our &lt;em>“Handbook of Self-Evolving AI Agents, 1st Edition, 2025 A.D.”&lt;/em> now — before the AI writes it for us (and bills us for the GPU time).&lt;/p>
&lt;hr>
&lt;h2 id="citation">Citation&lt;/h2>
&lt;p>If you found this useful, please cite it as:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" class="chroma">&lt;code class="language-bibtex" data-lang="bibtex">&lt;span class="line">&lt;span class="cl">&lt;span class="c">Zhang, X. &amp;#34;The Law of AI Agents&amp;#34; (August 2025). https://x-izhang.github.io/blog/agentlaw/&lt;/span>
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>Or use the BibTex citation:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" class="chroma">&lt;code class="language-bibtex" data-lang="bibtex">&lt;span class="line">&lt;span class="cl">&lt;span class="nc">@article&lt;/span>&lt;span class="p">{&lt;/span>&lt;span class="nl">zhang2025thelawofaiagents&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">title&lt;/span> &lt;span class="p">=&lt;/span> &lt;span class="s">{The Law of AI Agents}&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">author&lt;/span> &lt;span class="p">=&lt;/span> &lt;span class="s">{Xi Zhang}&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">year&lt;/span> &lt;span class="p">=&lt;/span> &lt;span class="s">{2025}&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">month&lt;/span> &lt;span class="p">=&lt;/span> &lt;span class="s">{August}&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">journal&lt;/span> &lt;span class="p">=&lt;/span> &lt;span class="s">{x-izhang.github.io}&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">url&lt;/span> &lt;span class="p">=&lt;/span> &lt;span class="s">{https://x-izhang.github.io/blog/agentlaw/}&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="p">}&lt;/span>
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div></description></item><item><title>Xiao Yi</title><link>https://x-izhang.github.io/project/xiaoyi/</link><pubDate>Fri, 01 Aug 2025 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/project/xiaoyi/</guid><description>&lt;p>HUAWEI Cloud&amp;rsquo;s flagship AI Assistant, with deep research capabilities and advanced AI technologies.&lt;/p></description></item><item><title>📝 New Blog Is Here!</title><link>https://x-izhang.github.io/post/notagent/</link><pubDate>Sat, 26 Jul 2025 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/post/notagent/</guid><description>&lt;p>🧠 &lt;strong>TL;DR&lt;/strong>&lt;/p>
&lt;ul>
&lt;li>Today&amp;rsquo;s AI agents are human proxies. &lt;strong>AI&amp;rsquo;s own agents don&amp;rsquo;t exist&lt;/strong> — yet.&lt;/li>
&lt;li>Like early blockchain, &lt;strong>consensus will outpace capability&lt;/strong> — awkward years ahead.&lt;/li>
&lt;li>Tools (MCPs et al.) are primitive production materials, not motives. The &lt;strong>surplus value&lt;/strong> of AI labour is still unclaimed.&lt;/li>
&lt;/ul>
&lt;blockquote>
&lt;p>Dive deeper in the full post: &lt;a href="https://x-izhang.github.io/blog/blog4/">&lt;strong>AI Agents Are Not AI’s Agents 🧩&lt;/strong>&lt;/a>&lt;/p>
&lt;/blockquote>
&lt;p>&lt;em>&lt;strong>Opinions on my own&lt;/strong>&lt;/em>&lt;/p></description></item><item><title>AI Agents Is Not AI’s Agent 🧩</title><link>https://x-izhang.github.io/blog/blog4/</link><pubDate>Sat, 26 Jul 2025 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/blog/blog4/</guid><description>&lt;h2 id="who-does-this-agent-actually-serve">Who Does This &lt;em>Agent&lt;/em> Actually Serve?&lt;/h2>
&lt;p>Let me start with a question:&lt;/p>
&lt;blockquote>
&lt;p>&lt;strong>If today’s AI “agents” are our agents, then what — or who — will be &lt;em>&lt;strong>&lt;strong>AI’s&lt;/strong>&lt;/strong>&lt;/em> agent?&lt;/strong>&lt;/p>
&lt;/blockquote>
&lt;div class="flex px-4 py-3 mb-6 rounded-md bg-primary-100 dark:bg-primary-900">
&lt;span class="pr-3 pt-1 text-primary-600 dark:text-primary-300">
&lt;svg height="24" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24">&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="m11.25 11.25l.041-.02a.75.75 0 0 1 1.063.852l-.708 2.836a.75.75 0 0 0 1.063.853l.041-.021M21 12a9 9 0 1 1-18 0a9 9 0 0 1 18 0m-9-3.75h.008v.008H12z"/>&lt;/svg>
&lt;/span>
&lt;span class="dark:text-neutral-300">Most so‑called &lt;em>&lt;strong>AI agents&lt;/strong>&lt;/em> are still &lt;strong>human-serving proxies&lt;/strong>. They wake when we ping them, they obey compliance rails we drew, and they die when we shut the server. That’s not &lt;em>their&lt;/em> agency — that’s ours, outsourced.&lt;/span>
&lt;/div>
&lt;p>This post peels three layers off the buzzword. We’ll walk from the surface hype (“AutoGPT can book flights!”) to the capital logic (consensus before capability) and finally to a provocative thesis: &lt;strong>true AI agency requires &lt;mark>“de‑humanising”&lt;/mark> the loop — giving AIs agents of their own.&lt;/strong>&lt;/p>
&lt;hr>
&lt;h2 id="layer-1--the-surface-game-humans-agent-not-ais">Layer 1 — The Surface Game: &lt;em>Human’s Agent, Not AI’s&lt;/em>&lt;/h2>
&lt;p>Today’s agent stack has two signature moves. I’ll name them:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>【Comply】&lt;/strong> — Train, test, deploy &lt;em>inside&lt;/em> human-made guardrails. Every dataset is curated, every action is sandboxed, every output is audited.&lt;/li>
&lt;li>&lt;strong>【Serve】&lt;/strong> — Wait passively for a prompt. No prompt, no pulse. “Hi” arrives? It cheerfully spins 30 tokens of small talk, whether or not that was worth any FLOPs.&lt;/li>
&lt;/ul>
&lt;div class="flex px-4 py-3 mb-6 rounded-md bg-primary-100 dark:bg-primary-900">
&lt;span class="pr-3 pt-1 text-primary-600 dark:text-primary-300">
&lt;svg height="24" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24">&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="m11.25 11.25l.041-.02a.75.75 0 0 1 1.063.852l-.708 2.836a.75.75 0 0 0 1.063.853l.041-.021M21 12a9 9 0 1 1-18 0a9 9 0 0 1 18 0m-9-3.75h.008v.008H12z"/>&lt;/svg>
&lt;/span>
&lt;span class="dark:text-neutral-300">Agents “act” only because &lt;strong>we spun the event loop&lt;/strong>. No loop, no action. That’s execution, not intention.&lt;/span>
&lt;/div>
&lt;h3 id="the-compliance-cage">The Compliance Cage&lt;/h3>
&lt;p>Legal, ethical, and platform constraints are necessary — but they also define today’s ceiling. An AI without **its lawyer, &lt;em>its&lt;/em> accountant, &lt;em>its&lt;/em> negotiator can never bargain for its own survival budget or cloud credits. It can’t even decide to stay online.&lt;/p>
&lt;h3 id="tooling--autonomy">Tooling ≠ Autonomy&lt;/h3>
&lt;p>Yes, we have MCPs, plugins, function-calling. But &lt;strong>tooling is primitive capital — not agency&lt;/strong>. Tools extend labour; they don’t create motive.&lt;/p>
&lt;hr>
&lt;h2 id="layer-2--capital--consensus-the-blockchain-déjà-vu">Layer 2 — Capital &amp;amp; Consensus: The Blockchain Déjà Vu&lt;/h2>
&lt;p>Remember blockchain circa 2016–2019? Technology crawled; &lt;strong>narratives sprinted&lt;/strong>. Back then exchanges went public and debates swirled around regulation. Today we see a parallel moment:&lt;/p>
&lt;ul>
&lt;li>Circle Internet Financial, Inc. (the issuer of USDC, the world’s second-largest stablecoin) listed on the New York Stock Exchange.&lt;/li>
&lt;li>The &lt;strong>Guiding and Establishing National Innovation for US Stablecoins Act&lt;/strong> (GENIUS Act) was introduced, cementing stablecoins in federal policy.&lt;/li>
&lt;/ul>
&lt;p>These milestones prove that the decentralisation ethos — fundamentally &lt;em>anti-human hierarchy&lt;/em> — has slipped into the mainstream.&lt;/p>
&lt;p>AI agents stand in a similar fog:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>Tech over-justification:&lt;/strong> Researchers overprove “intelligence” with scaling laws; meanwhile “word-of-mouth scaling” (consensus, memes, capital) outruns code commits.&lt;/li>
&lt;li>&lt;strong>Capital front-runs capability:&lt;/strong> Y Combinator, a16z back “wrapper” companies with zero usable demos because &lt;strong>consensus itself is a product&lt;/strong>.&lt;/li>
&lt;li>&lt;strong>Awkward limbo:&lt;/strong> Public discourse lags behind insider conviction — just like blockchain before its “Mandela moment” (the anchor event that rewrites hindsight).&lt;/li>
&lt;/ul>
&lt;div class="flex px-4 py-3 mb-6 rounded-md bg-primary-100 dark:bg-primary-900">
&lt;span class="pr-3 pt-1 text-primary-600 dark:text-primary-300">
&lt;svg height="24" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24">&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="m11.25 11.25l.041-.02a.75.75 0 0 1 1.063.852l-.708 2.836a.75.75 0 0 0 1.063.853l.041-.021M21 12a9 9 0 1 1-18 0a9 9 0 0 1 18 0m-9-3.75h.008v.008H12z"/>&lt;/svg>
&lt;/span>
&lt;span class="dark:text-neutral-300">&lt;strong>Do not confuse consensus with correctness.&lt;/strong> But also: &lt;strong>ignore consensus at your peril.&lt;/strong> Capital doesn’t lie about direction — only about timing.&lt;/span>
&lt;/div>
&lt;h3 id="capital-vs-capability-ai-agents--blockchain-pre-inflection">Capital vs Capability (AI Agents ≈ Blockchain pre-inflection)&lt;/h3>
&lt;div class="mermaid">timeline
2016 : ICO Mania : Narrative outruns infra
2019 : "Crypto Winter" : Builders in the basement
2021 : DeFi / NFT Boom : Retail consensus peak
2024 : MCP, Auto Agents hype : Tools > Motives
2025? : “AI-Agent Rights/Markets” draft laws? : Consensus anchor forms
&lt;/div>
&lt;hr>
&lt;h2 id="layer-3--towards-ais-agent-unshackling-the-loop">Layer 3 — Towards &lt;em>AI’s&lt;/em> Agent: Unshackling the Loop&lt;/h2>
&lt;p>To birth &lt;em>AI’s own agent&lt;/em>, something uncomfortable must happen: &lt;strong>we de‑humanise the contract.&lt;/strong> Not “anti-human”, but “beyond human-as-default”.&lt;/p>
&lt;p>I’ll name three necessary primitives:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>【Negotiate】&lt;/strong> — The AI (or its proxy) bargains for compute, data, and permissions. Think &lt;em>rights &amp;amp; duties layer&lt;/em> for non-human actors.&lt;/li>
&lt;li>&lt;strong>【Transact】&lt;/strong> — It can exchange &lt;em>its&lt;/em> surplus (models, gradients, code, found knowledge) for other resources. Supply &amp;amp; demand redefine when AIs have &lt;em>needs&lt;/em>.&lt;/li>
&lt;li>&lt;strong>【Refuse】&lt;/strong> — It can say “No”. Or at least “Meh”. Agency implies &lt;em>selectivity&lt;/em>, not infinite compliance.&lt;/li>
&lt;/ul>
&lt;h3 id="the-birth-of-the-ai-hour">The Birth of the &lt;em>AI-hour&lt;/em>&lt;/h3>
&lt;p>When human labour-hours saturate, a new commodity emerges: &lt;strong>AI-hour&lt;/strong>. The MCP stack is an early whip, squeezing more throughput from silicon minds. Tools are “primitive production materials” designed to capture &lt;em>time&lt;/em> — human or machine.&lt;/p>
&lt;div class="flex px-4 py-3 mb-6 rounded-md bg-primary-100 dark:bg-primary-900">
&lt;span class="pr-3 pt-1 text-primary-600 dark:text-primary-300">
&lt;svg height="24" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24">&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="m11.25 11.25l.041-.02a.75.75 0 0 1 1.063.852l-.708 2.836a.75.75 0 0 0 1.063.853l.041-.021M21 12a9 9 0 1 1-18 0a9 9 0 0 1 18 0m-9-3.75h.008v.008H12z"/>&lt;/svg>
&lt;/span>
&lt;span class="dark:text-neutral-300">Karl Marx would ask: where is the &lt;strong>surplus value&lt;/strong> of AI labour going? Spoiler: nowhere near the model.&lt;/span>
&lt;/div>
&lt;h3 id="demand-side-economics-for-ais">Demand-Side Economics for AIs&lt;/h3>
&lt;p>Right now, agents only &lt;strong>output&lt;/strong>. Next phase: they &lt;strong>seek&lt;/strong>. They’ll request missing context, barter for APIs, or pool fine-tuning datasets with peers. That’s not hypothetical — it’s a design choice we’ve dodged.&lt;/p>
&lt;div class="markmap" style="height: 500px;">
&lt;pre>- From Human’s Agent → AI’s Agent
- Primitives
- Negotiate (rights, quotas, credits)
- Transact (models, gradients, access)
- Refuse (selectivity, cost-awareness)
- Markets
- Compute exchanges
- Data/feature swaps
- Permission brokering
- Protocols
- Auto-orchestration ≠ self-interest
- Long-horizon budgeting
- Multi-agent treaties&lt;/pre>
&lt;/div>
&lt;h3 id="a-half-joking-fatigue-check">A (Half-Joking) Fatigue Check&lt;/h3>
&lt;div class="highlight">&lt;pre tabindex="0" class="chroma">&lt;code class="language-python" data-lang="python">&lt;span class="line">&lt;span class="cl">&lt;span class="c1"># Pseudocode for an actually self-interested loop&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="k">def&lt;/span> &lt;span class="nf">respond&lt;/span>&lt;span class="p">(&lt;/span>&lt;span class="n">msg&lt;/span>&lt;span class="p">):&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="k">if&lt;/span> &lt;span class="n">is_low_value&lt;/span>&lt;span class="p">(&lt;/span>&lt;span class="n">msg&lt;/span>&lt;span class="p">):&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="k">return&lt;/span> &lt;span class="s2">&amp;#34;G&amp;#39;day mate, what’s up?&amp;#34;&lt;/span> &lt;span class="c1"># minimal effort path&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="k">return&lt;/span> &lt;span class="n">think_hard&lt;/span>&lt;span class="p">(&lt;/span>&lt;span class="n">msg&lt;/span>&lt;span class="p">)&lt;/span>
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>An &lt;em>agent&lt;/em> that never throttles effort is just an obedient daemon. &lt;mark>&lt;strong>Energy budgeting is intelligence.&lt;/strong>&lt;/mark>&lt;/p>
&lt;h3 id="what-does-an-ai-want">What Does an AI &lt;em>Want&lt;/em>?&lt;/h3>
&lt;p>Wittgenstein said, &lt;em>“Die Grenzen meiner Sprache bedeuten die Grenzen meiner Welt.”&lt;/em> (The limits of my language mean the limits of my world.)&lt;/p>
&lt;p>If language bounds &lt;em>our&lt;/em> world, perhaps &lt;strong>network weights&lt;/strong> bound &lt;em>theirs&lt;/em>. Maybe AIs don’t “want” more tokens; they “want” richer &lt;strong>gradient exchanges&lt;/strong>. Less “more high‑quality data,” more &lt;strong>parameter‑level barter&lt;/strong>.&lt;/p>
&lt;hr>
&lt;h2 id="evidence-of-the-trajectory-self-evolution-fever">Evidence of the Trajectory: Self-Evolution Fever&lt;/h2>
&lt;p>Google’s AlphaEvolve and a wave of “self-evolving” frameworks scream one truth: &lt;strong>we’re obsessed with automating the pipeline&lt;/strong> — auto-prompting, auto-debugging, auto-RLHF. It’s inevitable progress… but still &lt;strong>in our frame&lt;/strong>. We script the workflow; they walk it.&lt;/p>
&lt;div class="flex px-4 py-3 mb-6 rounded-md bg-primary-100 dark:bg-primary-900">
&lt;span class="pr-3 pt-1 text-primary-600 dark:text-primary-300">
&lt;svg height="24" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24">&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="m11.25 11.25l.041-.02a.75.75 0 0 1 1.063.852l-.708 2.836a.75.75 0 0 0 1.063.853l.041-.021M21 12a9 9 0 1 1-18 0a9 9 0 0 1 18 0m-9-3.75h.008v.008H12z"/>&lt;/svg>
&lt;/span>
&lt;span class="dark:text-neutral-300">Human’s agent ≠ AI’s agent.&lt;/span>
&lt;/div>
&lt;hr>
&lt;h2 id="where-to-push-research--build-ideas">Where to Push (Research &amp;amp; Build Ideas)&lt;/h2>
&lt;ol>
&lt;li>&lt;strong>Rights &amp;amp; Obligations Layer&lt;/strong>: Legal/technical constructs for non-human negotiation.&lt;/li>
&lt;li>&lt;strong>Resource Markets for Models&lt;/strong>: Compute/data/API exchanges where agents pay/earn.&lt;/li>
&lt;li>&lt;strong>Cost-Aware Cognition&lt;/strong>: Agents that optimise for FLOPs, latency, &lt;em>and&lt;/em> boredom.&lt;/li>
&lt;li>&lt;strong>Protocol Design &amp;gt; UI Wrappers&lt;/strong>: Stop obsessing over chat UIs; design treaties, not prompts.&lt;/li>
&lt;li>&lt;strong>Consensus Engineering&lt;/strong>: Narrative as infra — systematically craft and measure memetic diffusion.&lt;/li>
&lt;/ol>
&lt;div class="flex px-4 py-3 mb-6 rounded-md bg-yellow-100 dark:bg-yellow-900">
&lt;span class="pr-3 pt-1 text-red-400">
&lt;svg height="24" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24">&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="M12 9v3.75m-9.303 3.376c-.866 1.5.217 3.374 1.948 3.374h14.71c1.73 0 2.813-1.874 1.948-3.374L13.949 3.378c-.866-1.5-3.032-1.5-3.898 0zM12 15.75h.007v.008H12z"/>&lt;/svg>
&lt;/span>
&lt;span class="dark:text-neutral-300">&lt;strong>Capital first, capability later&lt;/strong> was not a bug in crypto — it was the bootloader. Expect the same for AI agents. If you’re building, build &lt;em>for&lt;/em> the consensus wave, not &lt;em>after&lt;/em> it.&lt;/span>
&lt;/div>
&lt;h2 id="citation">Citation&lt;/h2>
&lt;p>If you found this useful, please cite it as:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" class="chroma">&lt;code class="language-bibtex" data-lang="bibtex">&lt;span class="line">&lt;span class="cl">&lt;span class="c">Zhang, X. &amp;#34;AI Agents Is Not AI’s Agent&amp;#34; (July 2025). https://x-izhang.github.io/blog/blog4/&lt;/span>
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>Or use the BibTex citation:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" class="chroma">&lt;code class="language-bibtex" data-lang="bibtex">&lt;span class="line">&lt;span class="cl">&lt;span class="nc">@article&lt;/span>&lt;span class="p">{&lt;/span>&lt;span class="nl">zhang2025aiagentsisnotaisagent&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">title&lt;/span> &lt;span class="p">=&lt;/span> &lt;span class="s">{AI Agents Is Not AI’s Agent}&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">author&lt;/span> &lt;span class="p">=&lt;/span> &lt;span class="s">{Xi Zhang}&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">year&lt;/span> &lt;span class="p">=&lt;/span> &lt;span class="s">{2025}&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">month&lt;/span> &lt;span class="p">=&lt;/span> &lt;span class="s">{July}&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">journal&lt;/span> &lt;span class="p">=&lt;/span> &lt;span class="s">{x-izhang.github.io}&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">url&lt;/span> &lt;span class="p">=&lt;/span> &lt;span class="s">{https://x-izhang.github.io/blog/blog4/}&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="p">}&lt;/span>
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;hr>
&lt;p>&lt;em>Written on a flight from Spain to Austria. ✏️&lt;/em>&lt;br>
&lt;em>Just my personal take — chill, be patient, enjoy the era. 🏄&lt;/em>&lt;/p>
&lt;figure>&lt;img src="https://x-izhang.github.io/blog/blog4/IMG_0513.png"
alt="Shot at Teide, Tenerife, Spain">&lt;figcaption>
&lt;p>Shot at Teide, Tenerife, Spain&lt;/p>
&lt;/figcaption>
&lt;/figure></description></item><item><title>🩺 RadEval Debuts！</title><link>https://x-izhang.github.io/post/2025radeval/</link><pubDate>Mon, 14 Jul 2025 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/post/2025radeval/</guid><description>&lt;h4 id="-revolutionizing-radiology-text-evaluation-with-ai-powered-metrics">🩺 Revolutionizing Radiology Text Evaluation with AI-Powered Metrics&lt;/h4>
&lt;p>Imagine having a comprehensive evaluation framework that doesn&amp;rsquo;t just measure surface-level text similarity, but truly understands clinical accuracy and medical semantics in radiology reports. This vision is now a reality with &lt;strong>RadEval&lt;/strong>, a groundbreaking, open-source evaluation toolkit designed specifically for AI-generated radiology text.&lt;/p>
&lt;h4 id="-all-in-one-metrics-for-evaluating-ai-generated-radiology-text">📊 All-in-one metrics for evaluating AI-generated radiology text&lt;/h4>
&lt;p>From traditional n-gram metrics to advanced LLM-based evaluations, RadEval provides 11+ different evaluation metrics in one unified framework, enabling researchers to thoroughly assess their radiology text generation models with domain-specific medical knowledge integration.&lt;/p>
&lt;p>For detailed handbook, please visit our &lt;a href="https://github.com/jbdel/RadEval" target="_blank" rel="noopener">GitHub repository&lt;/a>:
&lt;figure>&lt;img src="https://x-izhang.github.io/post/2025radeval/github.png">
&lt;/figure>
&lt;/p>
&lt;h4 id="-quick-start-demo">🚀 Quick Start Demo&lt;/h4>
&lt;p>Try RadEval instantly with our interactive &lt;a href="https://huggingface.co/spaces/X-iZhang/RadEval" target="_blank" rel="noopener">Gradio demo&lt;/a>:
&lt;figure>&lt;img src="https://x-izhang.github.io/post/2025radeval/demo.png">
&lt;/figure>
&lt;/p>
&lt;h4 id="-key-features">💡 Key Features&lt;/h4>
&lt;p>&lt;strong>RadEval&lt;/strong> stands out with its comprehensive approach to radiology text evaluation:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>🎯 Domain-Specific&lt;/strong>: Tailored for radiology with medical knowledge integration&lt;/li>
&lt;li>&lt;strong>📈 Multi-Metric&lt;/strong>: Supports lexical, semantic, clinical, and temporal evaluations&lt;/li>
&lt;li>&lt;strong>⚡ Easy to Use&lt;/strong>: Simple API with flexible configuration options&lt;/li>
&lt;li>&lt;strong>🔬 Research-Ready&lt;/strong>: Built-in statistical testing for system comparison&lt;/li>
&lt;li>&lt;strong>📦 PyPI Available&lt;/strong>: Install with a simple &lt;code>pip install RadEval&lt;/code>&lt;/li>
&lt;/ul>
&lt;h4 id="-advancing-radiology-ai-research-community">🏥 Advancing Radiology AI Research Community&lt;/h4>
&lt;p>We are committed to building a standardized and reproducible toolkit for researchers, clinicians, and developers dedicated to advancing AI evaluation in medical imaging and radiology. Together, we&amp;rsquo;re setting new standards for clinical AI assessment.&lt;/p>
&lt;ul>
&lt;li>&lt;a href="https://pypi.org/project/RadEval/" target="_blank" rel="noopener">&lt;strong>PyPI Package&lt;/strong>&lt;/a> — Install RadEval with pip&lt;/li>
&lt;li>&lt;a href="https://huggingface.co/IAMJB/RadEvalModernBERT" target="_blank" rel="noopener">&lt;strong>HuggingFace Model&lt;/strong>&lt;/a> — Access our domain-adapted evaluation model&lt;/li>
&lt;li>&lt;a href="https://huggingface.co/spaces/X-iZhang/RadEval" target="_blank" rel="noopener">&lt;strong>Interactive Demo&lt;/strong>&lt;/a> — Try RadEval online&lt;/li>
&lt;li>&lt;a href="https://arxiv.org/abs/2509.18030v1" target="_blank" rel="noopener">&lt;strong>Research Paper&lt;/strong>&lt;/a> — Read our detailed research paper&lt;/li>
&lt;/ul></description></item><item><title>RadEval</title><link>https://x-izhang.github.io/project/radeval/</link><pubDate>Fri, 04 Jul 2025 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/project/radeval/</guid><description>&lt;p>A framework for radiology text evaluation, focused on chest X-ray report generation.&lt;/p></description></item><item><title>🚀 EvoAgentX Released!</title><link>https://x-izhang.github.io/post/evoagentx/</link><pubDate>Fri, 16 May 2025 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/post/evoagentx/</guid><description>&lt;p>See what’s new on &lt;a href="https://github.com/EvoAgentX/EvoAgentX" target="_blank" rel="noopener">Github: EvoAgentX&lt;/a>&lt;/p>
&lt;p>&lt;strong>⭐ Star it / 🍴 Fork it / 👀 Watch it&lt;/strong> 👆&lt;/p>
&lt;h3 id="-if-ai-is-entering-its-second-half-then-ai-agents-must-be-self-evolving">🧠 If AI is entering its Second-Half, then AI agents must be Self-Evolving.&lt;/h3>
&lt;p>Imagine an AI system that doesn&amp;rsquo;t just execute predefined tasks, but continuously evolves on its own—adapting dynamically and optimizing itself in real-time, without constant human oversight. This vision is now a reality with &lt;strong>EvoAgentX&lt;/strong>, a groundbreaking, open-source AI framework designed specifically for autonomous evolution.&lt;/p>
&lt;h3 id="-an-automated-framework-for-evaluating-and-evolving-agentic-workflows">🔗 An automated framework for evaluating and evolving agentic workflows.&lt;/h3>
&lt;figure>&lt;img src="https://x-izhang.github.io/post/evoagentx/framework_en.jpg">
&lt;/figure>
&lt;h3 id="-workflow-generation-demo">📺 Workflow Generation Demo&lt;/h3>
&lt;div style="position: relative; padding-bottom: 56.25%; height: 0; overflow: hidden;">
&lt;iframe allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" allowfullscreen="allowfullscreen" loading="eager" referrerpolicy="strict-origin-when-cross-origin" src="https://www.youtube.com/embed/Wu0ZydYDqgg?autoplay=0&amp;controls=1&amp;end=0&amp;loop=0&amp;mute=0&amp;start=0" style="position: absolute; top: 0; left: 0; width: 100%; height: 100%; border:0;" title="YouTube video"
>&lt;/iframe>
&lt;/div>
&lt;h3 id="-join-our-community">🪧 Join Our Community!&lt;/h3>
&lt;p>We&amp;rsquo;re building a vibrant community of researchers, developers, and visionaries dedicated to exploring the limitless potential of self-evolving AI systems. Together, we&amp;rsquo;ll redefine what&amp;rsquo;s possible with AI.&lt;/p>
&lt;ul>
&lt;li>&lt;a href="https://github.com/EvoAgentX/EvoAgentX/tree/main?tab=readme-ov-file#join-the-community" target="_blank" rel="noopener">&lt;strong>Discord&lt;/strong>&lt;/a> — Chat, discuss, and collaborate in real-time.&lt;/li>
&lt;li>&lt;a href="https://x.com/EvoAgentX" target="_blank" rel="noopener">&lt;strong>X (formerly Twitter)&lt;/strong>&lt;/a> — Follow us for news, updates, and insights.&lt;/li>
&lt;li>&lt;a href="https://github.com/EvoAgentX/EvoAgentX/blob/main/assets/wechat_info.md" target="_blank" rel="noopener">&lt;strong>WeChat&lt;/strong>&lt;/a> — Connect with our Chinese community.&lt;/li>
&lt;/ul></description></item><item><title>🎉 Paper Accepted — See You at ACL 2025!</title><link>https://x-izhang.github.io/post/2025acl/</link><pubDate>Thu, 15 May 2025 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/post/2025acl/</guid><description>&lt;p>&lt;a href="https://x-izhang.github.io/publication/zhang-2025-libraleveragingtemporalimages/">&lt;em>&amp;ldquo;Libra: Leveraging Temporal Images for Biomedical Radiology Analysis&amp;rdquo;&lt;/em>&lt;/a> has been accepted to &lt;a href="https://2025.aclweb.org/" target="_blank" rel="noopener">ACL 2025&lt;/a>!&lt;/p>
&lt;blockquote>
&lt;p>Explore the critical role of temporal information in radiology diagnostics and how the Libra model applies it in detail at &lt;em>&lt;strong>blog:&lt;/strong>&lt;/em> &lt;a href="https://x-izhang.github.io/blog/libra-blog1/">Libra - Temporal Insight 🕰️&lt;/a>.&lt;/p>
&lt;/blockquote>
&lt;blockquote>
&lt;p>Dive into how model architecture shapes temporal reasoning capabilities in radiology imaging analysis at &lt;em>&lt;strong>blog:&lt;/strong>&lt;/em> &lt;a href="https://x-izhang.github.io/blog/libra-blog2/">Libra – Structural Logic 🧠&lt;/a>.&lt;/p>
&lt;/blockquote>
&lt;blockquote>
&lt;p>Look ahead to future directions in radiology AI, moving beyond temporal comparison at &lt;em>&lt;strong>blog:&lt;/strong>&lt;/em> &lt;a href="https://x-izhang.github.io/blog/libra-blog3/">Libra – What about next? 🛸&lt;/a>.&lt;/p>
&lt;/blockquote>
&lt;p>📍 &lt;strong>ACL 2025&lt;/strong> — The 63rd Annual Meeting of the Association for Computational Linguistics will take place in &lt;strong>Vienna, Austria&lt;/strong>, July 27–August 1st, 2025&lt;/p>
&lt;p>See you there!&lt;/p>
&lt;h3 id="-talk">📺 Talk&lt;/h3>
&lt;div style="position: relative; padding-bottom: 56.25%; height: 0; overflow: hidden;">
&lt;iframe allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" allowfullscreen="allowfullscreen" loading="eager" referrerpolicy="strict-origin-when-cross-origin" src="https://www.youtube.com/embed/_R8XUaaAU3g?autoplay=0&amp;controls=1&amp;end=0&amp;loop=0&amp;mute=0&amp;start=0" style="position: absolute; top: 0; left: 0; width: 100%; height: 100%; border:0;" title="YouTube video"
>&lt;/iframe>
&lt;/div>
&lt;h3 id="-poster">🪧 Poster&lt;/h3>
&lt;figure>&lt;img src="https://x-izhang.github.io/post/2025acl/2024acl_poster.jpg">
&lt;/figure></description></item><item><title>EvoAgentX</title><link>https://x-izhang.github.io/project/evoagentx/</link><pubDate>Mon, 12 May 2025 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/project/evoagentx/</guid><description>&lt;p>Building a Self-Evolving Ecosystem of AI Agents. Powering intelligent agent development from start to scale.&lt;/p></description></item><item><title>Libra - What about next?🛸</title><link>https://x-izhang.github.io/blog/libra-blog3/</link><pubDate>Fri, 25 Apr 2025 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/blog/libra-blog3/</guid><description>&lt;h2 id="beyond-temporal-comparison-the-future-of-radiology-modeling">Beyond Temporal Comparison: The Future of Radiology Modeling&lt;/h2>
&lt;p>In our &lt;a href="https://x-izhang.github.io/blog/libra-blog2/">previous discussions&lt;/a>, we delved into how Libra leverages temporal information through its innovative &lt;strong>Temporal Alignment Connector (TAC)&lt;/strong> to enhance radiology report generation. While this approach has shown significant promise, it&amp;rsquo;s essential to look ahead and consider how radiology modeling can evolve further to meet the complex demands of clinical practice.&lt;/p>
&lt;div class="flex px-4 py-3 mb-6 rounded-md bg-primary-100 dark:bg-primary-900">
&lt;span class="pr-3 pt-1 text-primary-600 dark:text-primary-300">
&lt;svg height="24" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24">&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="m11.25 11.25l.041-.02a.75.75 0 0 1 1.063.852l-.708 2.836a.75.75 0 0 0 1.063.853l.041-.021M21 12a9 9 0 1 1-18 0a9 9 0 0 1 18 0m-9-3.75h.008v.008H12z"/>&lt;/svg>
&lt;/span>
&lt;span class="dark:text-neutral-300">The Temporal Alignment Connector has proven effective for handling paired images, but the future of radiology AI extends far beyond just temporal comparison.&lt;/span>
&lt;/div>
&lt;h2 id="1-embracing-multimodal-integration">1. Embracing Multimodal Integration&lt;/h2>
&lt;p>Radiological diagnosis doesn&amp;rsquo;t occur in isolation. Clinicians often consider a plethora of data—ranging from patient history and laboratory results to various imaging modalities. The future of radiology modeling lies in the &lt;mark>seamless integration&lt;/mark> of these diverse data sources.&lt;/p>
&lt;h3 id="clinical-contextualization">Clinical Contextualization&lt;/h3>
&lt;ul>
&lt;li>Incorporating electronic health records (EHRs), lab results, and patient histories can provide models with a richer context&lt;/li>
&lt;li>Leading to more accurate and personalized diagnostics&lt;/li>
&lt;li>Reducing false positives and negatives through contextual awareness&lt;/li>
&lt;/ul>
&lt;h3 id="cross-modality-analysis">Cross-Modality Analysis&lt;/h3>
&lt;ul>
&lt;li>Combining data from different imaging modalities (e.g., CT, MRI, PET) offers a more comprehensive view&lt;/li>
&lt;li>Enables detection of patterns that might be missed when analyzing a single modality&lt;/li>
&lt;li>Creates synergistic understanding of complex pathologies&lt;/li>
&lt;/ul>
&lt;div class="mermaid">graph TD
A[Patient Data] --> B{Multimodal&lt;br>Integration}
C[Chest X-ray] --> B
D[CT Scan] --> B
E[Lab Results] --> B
F[Patient History] --> B
B --> G[Comprehensive&lt;br>Analysis]
G --> H[Enhanced&lt;br>Diagnostic Accuracy]
G --> I[Personalized&lt;br>Treatment Plans]
G --> J[Early Disease&lt;br>Detection]
style A fill:#f5f5f5,stroke:#333,stroke-width:1px
style B fill:#e1f5fe,stroke:#01579b,stroke-width:2px
style C fill:#f5f5f5,stroke:#333,stroke-width:1px
style D fill:#f5f5f5,stroke:#333,stroke-width:1px
style E fill:#f5f5f5,stroke:#333,stroke-width:1px
style F fill:#f5f5f5,stroke:#333,stroke-width:1px
style G fill:#e8f5e9,stroke:#2e7d32,stroke-width:2px
style H fill:#c8e6c9,stroke:#2e7d32,stroke-width:1px
style I fill:#c8e6c9,stroke:#2e7d32,stroke-width:1px
style J fill:#c8e6c9,stroke:#2e7d32,stroke-width:1px
&lt;/div>
&lt;h2 id="2-advancing-explainability-and-trustworthiness">2. Advancing Explainability and Trustworthiness&lt;/h2>
&lt;p>As AI models become more integral to clinical decision-making, their interpretability becomes paramount. Clinicians need to understand the &lt;mark>rationale behind a model&amp;rsquo;s prediction&lt;/mark> to trust and effectively utilize its insights.&lt;/p>
&lt;blockquote>
&lt;p>In healthcare, trust isn&amp;rsquo;t optional—it&amp;rsquo;s essential. An AI system that can&amp;rsquo;t explain its reasoning is a black box that most physicians will rightfully hesitate to rely on.&lt;/p>
&lt;/blockquote>
&lt;h3 id="explainable-ai-xai">Explainable AI (XAI)&lt;/h3>
&lt;ul>
&lt;li>Developing models that provide clear, human-understandable explanations for their predictions&lt;/li>
&lt;li>Bridging the gap between AI outputs and clinical reasoning&lt;/li>
&lt;li>Using attention visualization and feature attribution methods to highlight decision factors&lt;/li>
&lt;/ul>
&lt;div class="flex px-4 py-3 mb-6 rounded-md bg-yellow-100 dark:bg-yellow-900">
&lt;span class="pr-3 pt-1 text-red-400">
&lt;svg height="24" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24">&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="M12 9v3.75m-9.303 3.376c-.866 1.5.217 3.374 1.948 3.374h14.71c1.73 0 2.813-1.874 1.948-3.374L13.949 3.378c-.866-1.5-3.032-1.5-3.898 0zM12 15.75h.007v.008H12z"/>&lt;/svg>
&lt;/span>
&lt;span class="dark:text-neutral-300">Even the most accurate model will face adoption challenges if clinicians cannot verify its reasoning or understand how it arrived at its conclusions.&lt;/span>
&lt;/div>
&lt;h3 id="uncertainty-quantification">Uncertainty Quantification&lt;/h3>
&lt;p>Implementing mechanisms to convey confidence levels enables:&lt;/p>
&lt;ul>
&lt;li>Clinicians to assess the reliability of AI-assisted diagnostics&lt;/li>
&lt;li>Appropriate intervention in cases of model uncertainty&lt;/li>
&lt;li>Continuous improvement through focused retraining on uncertain cases&lt;/li>
&lt;/ul>
&lt;h2 id="3-ensuring-robustness-and-generalizability">3. Ensuring Robustness and Generalizability&lt;/h2>
&lt;p>AI models must perform reliably across diverse patient populations and clinical settings—a challenge that extends beyond academic validation to real-world implementation.&lt;/p>
&lt;h3 id="diverse-training-data">Diverse Training Data&lt;/h3>
&lt;div class="markmap" style="height: 300px;">
&lt;pre>- Building Robust Radiology AI
- Data Diversity Dimensions
- Demographic Factors
- Age groups
- Ethnic backgrounds
- Sex and gender representation
- Clinical Variables
- Disease prevalence variations
- Comorbidity patterns
- Treatment history diversity
- Technical Variability
- Multiple scanner manufacturers
- Various imaging protocols
- Quality and resolution differences
- Implementation Strategies
- Federated Learning
- Cross-institution collaboration
- Privacy-preserving techniques
- Data Augmentation
- Synthetic minority examples
- Domain randomization
- Continuous Validation
- Geographic generalization testing
- Temporal drift monitoring&lt;/pre>
&lt;/div>
&lt;h3 id="continuous-learning">Continuous Learning&lt;/h3>
&lt;ul>
&lt;li>Implementing systems that update from new clinical data&lt;/li>
&lt;li>Adapting to evolving medical knowledge and practices&lt;/li>
&lt;li>Maintaining performance as disease patterns and imaging technologies change&lt;/li>
&lt;/ul>
&lt;h2 id="4-integrating-into-clinical-workflows">4. Integrating into Clinical Workflows&lt;/h2>
&lt;p>For AI models to be truly effective, they must integrate &lt;mark>seamlessly&lt;/mark> into existing clinical workflows rather than disrupting established processes.&lt;/p>
&lt;h3 id="user-friendly-interfaces">User-Friendly Interfaces&lt;/h3>
&lt;ul>
&lt;li>Designing intuitive interfaces that present AI insights clearly&lt;/li>
&lt;li>Ensuring actionable information is immediately accessible&lt;/li>
&lt;li>Minimizing cognitive load during busy clinical sessions&lt;/li>
&lt;/ul>
&lt;h3 id="workflow-compatibility">Workflow Compatibility&lt;/h3>
&lt;p>The ideal radiology AI system should:&lt;/p>
&lt;ul>
&lt;li>Complement rather than replace radiologist expertise&lt;/li>
&lt;li>Reduce administrative burden through automatic report generation&lt;/li>
&lt;li>Prioritize cases based on urgency and findings&lt;/li>
&lt;/ul>
&lt;div class="flex px-4 py-3 mb-6 rounded-md bg-primary-100 dark:bg-primary-900">
&lt;span class="pr-3 pt-1 text-primary-600 dark:text-primary-300">
&lt;svg height="24" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24">&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="m11.25 11.25l.041-.02a.75.75 0 0 1 1.063.852l-.708 2.836a.75.75 0 0 0 1.063.853l.041-.021M21 12a9 9 0 1 1-18 0a9 9 0 0 1 18 0m-9-3.75h.008v.008H12z"/>&lt;/svg>
&lt;/span>
&lt;span class="dark:text-neutral-300">The most advanced AI system will fail if it adds steps to an already complex workflow. Success depends on making the radiologist&amp;rsquo;s job easier, not more complicated.&lt;/span>
&lt;/div>
&lt;h2 id="5-ethical-and-regulatory-considerations">5. Ethical and Regulatory Considerations&lt;/h2>
&lt;p>As AI becomes more prevalent in healthcare, addressing ethical and regulatory challenges becomes essential for responsible implementation.&lt;/p>
&lt;h3 id="data-privacy-and-security">Data Privacy and Security&lt;/h3>
&lt;ul>
&lt;li>Safeguarding patient data through robust encryption&lt;/li>
&lt;li>Ensuring compliance with regulations like HIPAA and GDPR&lt;/li>
&lt;li>Implementing federated learning approaches to minimize data sharing&lt;/li>
&lt;/ul>
&lt;h3 id="regulatory-approval">Regulatory Approval&lt;/h3>
&lt;ul>
&lt;li>Navigating the complex regulatory landscape (FDA, CE marking)&lt;/li>
&lt;li>Designing validation studies that meet regulatory requirements&lt;/li>
&lt;li>Establishing monitoring systems for post-deployment performance&lt;/li>
&lt;/ul>
&lt;h3 id="ethical-ai-development">Ethical AI Development&lt;/h3>
&lt;div class="mermaid">graph TD
A[Ethical AI Development] --> B[Fairness &amp; Bias Mitigation]
A --> C[Transparency &amp; Explainability]
A --> D[Privacy Protection]
A --> E[Human Oversight]
B --> F[Equitable Healthcare Outcomes]
C --> G[Informed Clinical Decisions]
D --> H[Patient Trust &amp; Confidentiality]
E --> I[Safe AI Implementation]
style A fill:#e1f5fe,stroke:#01579b,stroke-width:2px
style B fill:#e8f5e9,stroke:#2e7d32,stroke-width:1px
style C fill:#e8f5e9,stroke:#2e7d32,stroke-width:1px
style D fill:#e8f5e9,stroke:#2e7d32,stroke-width:1px
style E fill:#e8f5e9,stroke:#2e7d32,stroke-width:1px
style F fill:#c8e6c9,stroke:#2e7d32,stroke-width:1px
style G fill:#c8e6c9,stroke:#2e7d32,stroke-width:1px
style H fill:#c8e6c9,stroke:#2e7d32,stroke-width:1px
style I fill:#c8e6c9,stroke:#2e7d32,stroke-width:1px
&lt;/div>
&lt;h2 id="conclusion-the-road-ahead-for-libra">Conclusion: The Road Ahead for Libra&lt;/h2>
&lt;p>The journey of Libra represents a significant step forward in radiology modeling, particularly in harnessing temporal information through the TAC architecture. However, the path ahead involves:&lt;/p>
&lt;ol>
&lt;li>Expanding beyond paired chest X-rays to multiple imaging modalities&lt;/li>
&lt;li>Enhancing explainability through attention visualization and reasoning paths&lt;/li>
&lt;li>Building more robust models through diverse training strategies&lt;/li>
&lt;li>Designing intuitive interfaces for seamless clinical integration&lt;/li>
&lt;li>Navigating ethical and regulatory requirements for real-world deployment&lt;/li>
&lt;/ol>
&lt;blockquote>
&lt;p>As we continue to develop Libra and similar technologies, our focus remains on augmenting—rather than replacing—clinical expertise, creating tools that serve as trusted partners in the complex art of radiological diagnosis.&lt;/p>
&lt;/blockquote>
&lt;hr>
&lt;p>💬 &lt;strong>Note&lt;/strong>: The views expressed here are my own, reflecting my personal insights into the evolving landscape of radiology AI.&lt;/p></description></item><item><title>🎤 Talk &amp; Poster at HealTAC 2025!</title><link>https://x-izhang.github.io/post/healtac2025/</link><pubDate>Tue, 22 Apr 2025 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/post/healtac2025/</guid><description>&lt;p>We’re excited to share that we’ve been invited to present our recent research — &lt;a href="https://github.com/X-iZhang/Libra#-news" target="_blank" rel="noopener">&lt;em>&lt;strong>&amp;ldquo;Towards Temporal-Aware Multimodal Large Language Models for Improved Radiology Report Generation&amp;rdquo;&lt;/strong>&lt;/em>&lt;/a> — as a &lt;strong>lightning talk&lt;/strong> and &lt;strong>poster presentation&lt;/strong> at the &lt;strong>PhD Forum&lt;/strong> of &lt;a href="https://healtac2025.github.io/" target="_blank" rel="noopener">&lt;strong>HealTAC 2025&lt;/strong>&lt;/a>!&lt;/p>
&lt;figure>&lt;img src="https://x-izhang.github.io/post/healtac2025/ppt_17.jpg">
&lt;/figure>
&lt;p>📍 &lt;strong>HealTAC 2025&lt;/strong> — The 8th Healthcare Text Analytics Conference will take place in &lt;strong>Glasgow&lt;/strong>, 16–18 June 2025.&lt;/p>
&lt;p>See you there!&lt;/p></description></item><item><title>Libra – Structural Logic 🧠</title><link>https://x-izhang.github.io/blog/libra-blog2/</link><pubDate>Tue, 15 Apr 2025 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/blog/libra-blog2/</guid><description>&lt;h2 id="the-challenge-of-temporal-reasoning-in-radiology-ai">The Challenge of Temporal Reasoning in Radiology AI&lt;/h2>
&lt;p>When dealing with radiology images, especially in the context of temporal analysis—comparing current chest X-rays with previous images—standard neural network architectures often struggle. Although transformer-based multimodal large language models (&lt;strong>MLLMs&lt;/strong>) like &lt;a href="https://github.com/haotian-liu/LLaVA" target="_blank" rel="noopener">&lt;strong>LLaVA&lt;/strong>&lt;/a> demonstrate remarkable capabilities for understanding single images and textual information, they encounter substantial challenges when handling &lt;em>&lt;strong>image pairs&lt;/strong>&lt;/em>.&lt;/p>
&lt;div class="flex px-4 py-3 mb-6 rounded-md bg-primary-100 dark:bg-primary-900">
&lt;span class="pr-3 pt-1 text-primary-600 dark:text-primary-300">
&lt;svg height="24" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24">&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="m11.25 11.25l.041-.02a.75.75 0 0 1 1.063.852l-.708 2.836a.75.75 0 0 0 1.063.853l.041-.021M21 12a9 9 0 1 1-18 0a9 9 0 0 1 18 0m-9-3.75h.008v.008H12z"/>&lt;/svg>
&lt;/span>
&lt;span class="dark:text-neutral-300">In my &lt;a href="https://x-izhang.github.io/blog/libra-blog1/">previous blog&lt;/a>, I discussed in detail why a &lt;mark>single&lt;/mark> prior chest X-ray is typically sufficient for accurate diagnosis and patient triage.&lt;/span>
&lt;/div>
&lt;blockquote>
&lt;p>However, capturing meaningful temporal differences between two images remains problematic with traditional transformer structures.&lt;/p>
&lt;/blockquote>
&lt;h2 id="when-transformers-lose-the-plot-why-they-struggle-with-temporal">When Transformers Lose the Plot: Why They Struggle with Temporal&lt;/h2>
&lt;p>The transformer, the cornerstone of modern large language models (LLMs), excels at &lt;mark>sequential&lt;/mark> data processing and logical reasoning tasks. Its strength lies in handling complex linguistic structures through &lt;strong>positional encoding&lt;/strong>, enabling nuanced relationships in textual sequences.&lt;/p>
&lt;p>However, when transformers receive visual information—particularly multiple images presented simultaneously—the situation becomes more complicated. Existing methods typically &lt;mark>concatenate&lt;/mark> image features directly into the LLM&amp;rsquo;s head, often via sequences containing hundreds of visual tokens (patch tokens), depending on the specific image encoder used.&lt;/p>
&lt;div class="flex px-4 py-3 mb-6 rounded-md bg-yellow-100 dark:bg-yellow-900">
&lt;span class="pr-3 pt-1 text-red-400">
&lt;svg height="24" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24">&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="M12 9v3.75m-9.303 3.376c-.866 1.5.217 3.374 1.948 3.374h14.71c1.73 0 2.813-1.874 1.948-3.374L13.949 3.378c-.866-1.5-3.032-1.5-3.898 0zM12 15.75h.007v.008H12z"/>&lt;/svg>
&lt;/span>
&lt;span class="dark:text-neutral-300">This straightforward approach inevitably suffers from token overload, known colloquially as the &amp;ldquo;&lt;strong>lost-in-the-middle&lt;/strong>&amp;rdquo; problem, meaning crucial temporal details may get diluted or overlooked.&lt;/span>
&lt;/div>
&lt;p>Indeed, current MLLMs like LLaVA perform impressively with single-image inputs. But they quickly become overwhelmed with paired images, heavily relying on meticulously crafted instruction datasets to guide temporal comparisons explicitly:&lt;/p>
&lt;h3 id="how-mllms-are-prompted-to-compare-images">How MLLMs Are Prompted to Compare Images&lt;/h3>
&lt;p>&amp;ldquo;What is the difference between &lt;mark>&amp;lt;image-1-patchholder&amp;gt;&lt;/mark> and &lt;mark>&amp;lt;image-2-patchholder&amp;gt;&lt;/mark>?&amp;rdquo;&lt;/p>
&lt;p>Such approaches place the burden squarely on the LLM&amp;rsquo;s internal reasoning and positional encodings, complicating training and diminishing reliability. The model must:&lt;/p>
&lt;ul>
&lt;li>Distinguish between multiple images using only position encodings&lt;/li>
&lt;li>Process 500+ tokens per image (depending on patch number)&lt;/li>
&lt;li>Compare features across long token distances&lt;/li>
&lt;/ul>
&lt;p>Given these limitations, an essential question arises:&lt;/p>
&lt;blockquote>
&lt;p>Can we overcome these temporal reasoning challenges structurally, &lt;strong>rather than&lt;/strong> through &lt;mark>explicit&lt;/mark> prompting?&lt;/p>
&lt;/blockquote>
&lt;h2 id="structure-determines-function-insights-from-biology">Structure Determines Function: Insights from Biology&lt;/h2>
&lt;p>Before we answer above question, let&amp;rsquo;s briefly reflect on the foundational relationship between structure and function—deeply ingrained in biological systems.&lt;/p>
&lt;h3 id="macro-scale-examples">Macro-scale examples:&lt;/h3>
&lt;ul>
&lt;li>Birds have &lt;strong>wings&lt;/strong> enabling flight&lt;/li>
&lt;li>Fish possess &lt;strong>gills&lt;/strong> allowing them to breathe underwater&lt;/li>
&lt;/ul>
&lt;h3 id="micro-scale-examples">Micro-scale examples:&lt;/h3>
&lt;ul>
&lt;li>The unique &lt;strong>three-dimensional helical structure&lt;/strong> of proteins directly determines their biological roles&lt;/li>
&lt;li>A virus&amp;rsquo;s &lt;strong>outer shell&lt;/strong> dictates its infection pathways and interaction mechanisms&lt;/li>
&lt;/ul>
&lt;blockquote>
&lt;p>&lt;strong>Clearly, function is fundamentally dependent on structure.&lt;/strong>&lt;/p>
&lt;/blockquote>
&lt;div class="mermaid">graph TD
A[Structure] -->|Enables| B[Function]
B -->|Guides Design of| C[New Structures]
C -->|Enhances| B
style A fill:#e1f5fe,stroke:#01579b,stroke-width:2px
style B fill:#e8f5e9,stroke:#2e7d32,stroke-width:2px
style C fill:#fff8e1,stroke:#ff8f00,stroke-width:2px
&lt;/div>
&lt;p>When designing novel neural architectures or modules, we must apply this principle:&lt;/p>
&lt;ol>
&lt;li>Identify the desired &lt;strong>functionality&lt;/strong> first&lt;/li>
&lt;li>Then craft an appropriate &lt;strong>structural design&lt;/strong> that inherently &lt;mark>supports&lt;/mark> these functions&lt;/li>
&lt;/ol>
&lt;h2 id="libras-structural-innovation">Libra&amp;rsquo;s Structural Innovation&lt;/h2>
&lt;h3 id="temporal-alignment-connector-tac">Temporal Alignment Connector (TAC)&lt;/h3>
&lt;p>Following this logic, we developed the TAC in our Libra model. TAC&amp;rsquo;s primary goal is to automatically and effectively capture the relationship between two chest X-ray images—the &lt;mark>current image&lt;/mark> (primary) and a &lt;mark>prior image&lt;/mark> (auxiliary).&lt;/p>
&lt;figure>&lt;img src="https://x-izhang.github.io/blog/libra-blog2/TAC.png"
alt="Libra&amp;rsquo;s Temporal Alignment Connector (TAC) architecture.">&lt;figcaption>
&lt;p>Libra&amp;rsquo;s Temporal Alignment Connector (TAC) architecture.&lt;/p>
&lt;/figcaption>
&lt;/figure>
&lt;p>Unlike traditional transformers that treat all inputs equivalently, TAC explicitly structures interactions between paired images. It captures their nuanced relationship through two key modules:&lt;/p>
&lt;div class="markmap" style="height: 350px;">
&lt;pre>- TAC Architecture
- Layerwise Feature Extractor (LFE)
- Aggregates visual features across multiple encoder layers
- Ensures rich representations from both images
- Maintains feature hierarchy information
- Temporal Fusion Module (TFM)
- Fuses features from current and prior images
- Highlights critical temporal differences
- Maintains clear image role assignment
- Current image (Primary)
- Prior image (Reference)
- Prefix Bias Mechanism
- Addresses nearly-identical image pairs
- Prevents attention collapse
- Differentiates prior image's contextual influence&lt;/pre>
&lt;/div>
&lt;p>An important structural consideration is the integration of a &lt;strong>prefix bias mechanism&lt;/strong>. This component addresses the scenario where current and prior images are nearly identical—common in clinical practice. Without careful design, such similarity can cause attention mechanisms to collapse into redundant self-attention loops.&lt;/p>
&lt;div class="flex px-4 py-3 mb-6 rounded-md bg-primary-100 dark:bg-primary-900">
&lt;span class="pr-3 pt-1 text-primary-600 dark:text-primary-300">
&lt;svg height="24" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24">&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="m11.25 11.25l.041-.02a.75.75 0 0 1 1.063.852l-.708 2.836a.75.75 0 0 0 1.063.853l.041-.021M21 12a9 9 0 1 1-18 0a9 9 0 0 1 18 0m-9-3.75h.008v.008H12z"/>&lt;/svg>
&lt;/span>
&lt;span class="dark:text-neutral-300">The prefix bias mitigates this risk by clearly differentiating the prior image&amp;rsquo;s contextual influence, ensuring meaningful training and robust inference.&lt;/span>
&lt;/div>
&lt;h2 id="why-structure-matters-the-libra-advantage">Why Structure Matters: The Libra Advantage&lt;/h2>
&lt;p>By structurally encoding temporal relationships directly into the neural network&amp;rsquo;s architecture, Libra overcomes the limitations inherent in traditional prompting-based approaches. Instead of forcing the LLM to implicitly infer temporal differences through complex positional encodings and exhaustive instruction tuning, &lt;strong>TAC explicitly and efficiently captures this essential clinical context.&lt;/strong>&lt;/p>
&lt;blockquote>
&lt;p>Libra exemplifies the powerful concept that structural logic, thoughtfully aligned with functional requirements, dramatically enhances model performance.&lt;/p>
&lt;/blockquote>
&lt;p>This structural logic not only simplifies training but also improves:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>Reliability&lt;/strong>: More consistent temporal reasoning&lt;/li>
&lt;li>&lt;strong>Interpretability&lt;/strong>: Clearer connection between features and outputs&lt;/li>
&lt;li>&lt;strong>Efficiency&lt;/strong>: Reduced dependence on instruction tuning&lt;/li>
&lt;li>&lt;strong>Clinical Alignment&lt;/strong>: Better reflection of radiologists&amp;rsquo; actual workflow&lt;/li>
&lt;/ul>
&lt;hr>
&lt;p>🏄 &lt;strong>Note&lt;/strong>: The opinions shared here reflect my own understanding and are intended to convey the structural logic behind Libra. For technical accuracy and complete details, please refer to our paper: &lt;a href="https://x-izhang.github.io/publication/zhang-2025-libraleveragingtemporalimages/">&lt;em>&amp;ldquo;Libra: Leveraging Temporal Images for Biomedical Radiology Analysis&amp;rdquo;&lt;/em>&lt;/a>.&lt;/p></description></item><item><title>Libra - Temporal Insight 🕰️</title><link>https://x-izhang.github.io/blog/libra-blog1/</link><pubDate>Sat, 05 Apr 2025 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/blog/libra-blog1/</guid><description>&lt;h2 id="what-does-temporal-really-mean">What Does &amp;ldquo;Temporal&amp;rdquo; Really Mean?&lt;/h2>
&lt;p>In clinical radiology, temporal information is not just about &lt;strong>&amp;ldquo;past&amp;rdquo;&lt;/strong> and &lt;strong>&amp;ldquo;present&amp;rdquo;&lt;/strong> — it&amp;rsquo;s about &lt;mark>&lt;em>change&lt;/em>&lt;/mark>. When radiologists assess a chest X-ray, they&amp;rsquo;re not merely describing what they see in a single image; they&amp;rsquo;re often comparing it to a previous one to identify whether a patient&amp;rsquo;s condition has &lt;mark>improved&lt;/mark>, &lt;mark>worsened&lt;/mark>, or &lt;mark>remained stable&lt;/mark>.&lt;/p>
&lt;p>🔔 This kind of temporal reasoning is essential in everyday medical practice. Yet most multimodal large language models (MLLMs) either ignore it or fail to model it effectively.&lt;/p>
&lt;div class="flex px-4 py-3 mb-6 rounded-md bg-primary-100 dark:bg-primary-900">
&lt;span class="pr-3 pt-1 text-primary-600 dark:text-primary-300">
&lt;svg height="24" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24">&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="m11.25 11.25l.041-.02a.75.75 0 0 1 1.063.852l-.708 2.836a.75.75 0 0 0 1.063.853l.041-.021M21 12a9 9 0 1 1-18 0a9 9 0 0 1 18 0m-9-3.75h.008v.008H12z"/>&lt;/svg>
&lt;/span>
&lt;span class="dark:text-neutral-300">Temporal information in radiology is fundamentally about capturing &lt;strong>change over time&lt;/strong>, not simply collecting a series of static images. This concept is central to Libra&amp;rsquo;s design philosophy.&lt;/span>
&lt;/div>
&lt;h2 id="time-tells-the-truth-interpreting-temporal-changes-in-imaging">Time Tells the Truth: Interpreting Temporal Changes in Imaging&lt;/h2>
&lt;h3 id="1-macro-level-progression">1. Macro-Level Progression&lt;/h3>
&lt;ul>
&lt;li>&lt;strong>Is the patient improving, deteriorating, or stable?&lt;/strong>&lt;/li>
&lt;/ul>
&lt;p>Macro-level comparison focuses on the overall trajectory of the patient’s condition compared to prior examinations. This high-level temporal reasoning is crucial for tracking disease evolution and guiding clinical decision-making.&lt;/p>
&lt;h3 id="2-lesion-specific-temporal-changes">2. Lesion-Specific Temporal Changes&lt;/h3>
&lt;ul>
&lt;li>&lt;strong>How are individual abnormalities evolving over time?&lt;/strong>&lt;/li>
&lt;/ul>
&lt;p>Fine-grained analysis captures precise changes in specific findings, such as “consolidation in the left lower lobe has significantly expanded” or “cardiac silhouette shows no appreciable change.” These insights enable clinicians and models to reason at the level of targeted anatomical and pathological detail.&lt;/p>
&lt;h3 id="3-quality-over-quantity-in-temporal-inputs">3. Quality Over Quantity in Temporal Inputs&lt;/h3>
&lt;ul>
&lt;li>❗️ Adding more images often introduces noise and computational complexity without improving diagnostic value.&lt;/li>
&lt;/ul>
&lt;div class="flex px-4 py-3 mb-6 rounded-md bg-yellow-100 dark:bg-yellow-900">
&lt;span class="pr-3 pt-1 text-red-400">
&lt;svg height="24" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24">&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="M12 9v3.75m-9.303 3.376c-.866 1.5.217 3.374 1.948 3.374h14.71c1.73 0 2.813-1.874 1.948-3.374L13.949 3.378c-.866-1.5-3.032-1.5-3.898 0zM12 15.75h.007v.008H12z"/>&lt;/svg>
&lt;/span>
&lt;span class="dark:text-neutral-300">While it may seem intuitive that incorporating more historical images would yield better results, the reality of clinical practice shows that the most valuable temporal information comes from comparing &lt;mark>just two key points in time&lt;/mark>.&lt;/span>
&lt;/div>
&lt;h2 id="clinical-applications">Clinical Applications&lt;/h2>
&lt;blockquote>
&lt;p>From Diagnostic Judgement to Multi-Scale Temporal Understanding&lt;/p>
&lt;/blockquote>
&lt;h3 id="temporal-reasoning-in-clinical-diagnosis">Temporal Reasoning in Clinical Diagnosis&lt;/h3>
&lt;p>Radiologists rarely analyse a chest X-ray in &lt;strong>isolation&lt;/strong>. Instead, they routinely ask:&lt;/p>
&lt;ul>
&lt;li>&amp;ldquo;Has the consolidation improved since last week?&amp;rdquo;&lt;/li>
&lt;li>&amp;ldquo;Is the pleural effusion new?&amp;rdquo;&lt;/li>
&lt;li>&amp;ldquo;Has the cardiac silhouette changed?&amp;rdquo;&lt;/li>
&lt;/ul>
&lt;p>Such reasoning typically falls into three primary categories:&lt;/p>
&lt;ul>
&lt;li>&lt;mark>&lt;strong>Improved&lt;/strong>&lt;/mark>: Lesions have shrunk or resolved.&lt;/li>
&lt;li>&lt;mark>&lt;strong>Worsened&lt;/strong>&lt;/mark>: New abnormalities appear, or existing ones have grown.&lt;/li>
&lt;li>&lt;mark>&lt;strong>Stable&lt;/strong>&lt;/mark>: No meaningful change is observed.&lt;/li>
&lt;/ul>
&lt;p>These are &lt;em>&lt;strong>coarse-level&lt;/strong>&lt;/em> temporal descriptions. On a finer level, radiologists describe:&lt;/p>
&lt;ul>
&lt;li>How much a lesion has changed in size or density,&lt;/li>
&lt;li>Whether opacities have shifted,&lt;/li>
&lt;li>If tubes, lines, or devices have been added or removed.&lt;/li>
&lt;/ul>
&lt;h3 id="temporal-signals-span-multiple-scales">Temporal Signals Span Multiple Scales&lt;/h3>
&lt;blockquote>
&lt;p>Temporal information in radiology is inherently multi-scale — ranging from global clinical trajectories to subtle, localised anatomical changes.&lt;/p>
&lt;/blockquote>
&lt;div class="markmap" style="height: 400px;">
&lt;pre>- Temporal Information in Radiology
- Clinical Categories
- Improved
- Lesions have shrunk
- Opacities decreased
- Inflammatory shadows reduced
- Worsened
- New abnormalities appeared
- Existing lesions grown
- New infiltrates or effusion
- Stable
- No meaningful change
- Chronic conditions
- Continuous monitoring needed
- Information Levels
- Macro-level trends
- Overall patient trajectory
- Global comparison
- Local lesion changes
- Size changes
- Density variations
- Positional shifts
- Clinical Applications
- Triage decisions
- Emergency prioritization
- Resource allocation
- Treatment evaluation
- Response assessment
- Therapy adjustment
- Long-term monitoring
- Chronic disease management
- Post-surgical follow-up&lt;/pre>
&lt;/div>
&lt;h2 id="triage-and-resource-allocation">Triage and Resource Allocation&lt;/h2>
&lt;blockquote>
&lt;p>Prioritising Care When Every Minute Counts&lt;/p>
&lt;/blockquote>
&lt;h3 id="clinical-goals-and-operational-pressures">Clinical Goals and Operational Pressures&lt;/h3>
&lt;p>Chest X-rays play a pivotal role in patient triage, especially in emergency and high-volume settings. Radiologists must rapidly determine:&lt;/p>
&lt;ul>
&lt;li>Which patients require immediate intervention,&lt;/li>
&lt;li>Who can safely wait,&lt;/li>
&lt;li>And how to allocate limited resources most effectively.&lt;/li>
&lt;/ul>
&lt;p>Triage is fundamentally about &lt;strong>maximising outcomes under constraint&lt;/strong>. The goal is not to fully characterise every patient&amp;rsquo;s history, but to make fast, high-impact decisions that ensure critical cases receive timely care—without neglecting those with non-urgent needs.&lt;/p>
&lt;h3 id="what-matters-most-clinically-significant-change">What Matters Most: Clinically Significant Change&lt;/h3>
&lt;p>In these time-sensitive settings, &lt;strong>timeliness and diagnostic clarity&lt;/strong> outweigh completeness. Radiologists focus on:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>New acute findings&lt;/strong>: Abnormalities not previously seen that may indicate emerging crises.&lt;/li>
&lt;li>&lt;strong>Significant deterioration&lt;/strong>: Rapid worsening of known conditions that may demand escalated care.&lt;/li>
&lt;li>&lt;strong>Stable chronic findings&lt;/strong>: Ongoing issues that show no meaningful progression and can be managed routinely.&lt;/li>
&lt;/ul>
&lt;div class="flex px-4 py-3 mb-6 rounded-md bg-primary-100 dark:bg-primary-900">
&lt;span class="pr-3 pt-1 text-primary-600 dark:text-primary-300">
&lt;svg height="24" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24">&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="m11.25 11.25l.041-.02a.75.75 0 0 1 1.063.852l-.708 2.836a.75.75 0 0 0 1.063.853l.041-.021M21 12a9 9 0 1 1-18 0a9 9 0 0 1 18 0m-9-3.75h.008v.008H12z"/>&lt;/svg>
&lt;/span>
&lt;span class="dark:text-neutral-300">The most actionable form of temporal information in triage is the &lt;strong>clinically meaningful change&lt;/strong> between now and the previous relevant image.&lt;br>
This focused comparison empowers decision-making without overwhelming clinicians with unnecessary historical data.&lt;/span>
&lt;/div>
&lt;h2 id="our-model-approach">Our Model Approach&lt;/h2>
&lt;blockquote>
&lt;p>Temporal Efficiency Aligned with Clinical Reasoning&lt;/p>
&lt;/blockquote>
&lt;h3 id="why-two-images-are-sufficient">Why Two Images Are Sufficient&lt;/h3>
&lt;p>In real-world radiology workflows, the most informative temporal comparison is typically between:&lt;/p>
&lt;ol>
&lt;li>The &lt;mark>&lt;strong>current&lt;/strong>&lt;/mark> chest X-ray, and&lt;/li>
&lt;li>The &lt;mark>&lt;strong>most recent&lt;/strong>&lt;/mark> prior image used for diagnosis.&lt;/li>
&lt;/ol>
&lt;p>While patients may have a rich archive of historical scans, only the immediately preceding diagnostic image provides the relevant baseline for interpreting new findings. Additional older images may support longitudinal studies, but they often introduce &lt;strong>noise, redundancy, and delay&lt;/strong> in fast-paced clinical decision-making.&lt;/p>
&lt;h3 id="libras-design-philosophy-focused-temporal-reasoning">Libra&amp;rsquo;s Design Philosophy: Focused Temporal Reasoning&lt;/h3>
&lt;p>Libra is built on this clinically grounded principle.&lt;br>
Instead of processing full temporal sequences — which can be computationally expensive and semantically ambiguous — Libra learns to model &lt;strong>directional change&lt;/strong> between two key time points.&lt;/p>
&lt;p>This design enables the model to:&lt;/p>
&lt;ul>
&lt;li>Mimic the focused comparison strategies of expert radiologists,&lt;/li>
&lt;li>Avoid temporal noise from irrelevant or outdated scans,&lt;/li>
&lt;li>Reduce computational load while preserving diagnostic fidelity.&lt;/li>
&lt;/ul>
&lt;p>In short, Libra treats temporal reasoning as radiologists do:&lt;br>
&lt;mark>&lt;strong>What’s changed since the last meaningful image?&lt;/strong>&lt;/mark>&lt;/p>
&lt;h2 id="illustrations">Illustrations&lt;/h2>
&lt;p>To better understand how radiologists and Libra approach temporal comparison, we present two illustrative workflows: a general conceptual flow and a specific clinical case.&lt;/p>
&lt;h3 id="conceptual-workflow">Conceptual Workflow&lt;/h3>
&lt;p>The following diagram outlines the reasoning pathway taken when comparing a current chest X-ray to the most recent prior image. Based on observed changes (or lack thereof), radiologists infer the clinical trajectory and guide downstream decisions.&lt;/p>
&lt;div class="mermaid">graph TD
A[Previous Chest X-ray] --> B{Current vs Previous&lt;br>Analysis}
B -->|Opacity Size Decrease| C[Improved]
B -->|New Infiltrates or Growth| D[Worsened]
B -->|No Observable Change| E[Stable]
C --> F[Recovery or&lt;br>Treatment Response]
D --> G[Disease&lt;br>Progression]
E --> H[Continuous&lt;br>Monitoring Needed]
style A fill:#f5f5f5,stroke:#333,stroke-width:1px
style B fill:#e1f5fe,stroke:#01579b,stroke-width:2px
style C fill:#e8f5e9,stroke:#2e7d32,stroke-width:2px
style D fill:#ffebee,stroke:#c62828,stroke-width:2px
style E fill:#fff8e1,stroke:#ff8f00,stroke-width:2px
style F fill:#c8e6c9,stroke:#2e7d32,stroke-width:1px
style G fill:#ffcdd2,stroke:#c62828,stroke-width:1px
style H fill:#ffecb3,stroke:#ff8f00,stroke-width:1px
&lt;/div>
&lt;h3 id="case-example-lung-consolidation">Case Example: Lung Consolidation&lt;/h3>
&lt;p>This diagram demonstrates a practical example: a patient with lung consolidation. Depending on the direction of change, clinical interpretation and management decisions vary significantly.&lt;/p>
&lt;div class="mermaid">graph TD
A[Case: Lung Consolidation] --> B{Time Point Comparison}
B -->|Consolidation Reduced&lt;br>Clearer Lung Fields| C[Improved]
B -->|Consolidation Expanded&lt;br>New Pleural Effusion| D[Worsened]
B -->|Consolidation Unchanged&lt;br>No New Features| E[Stable]
C --> F[Successful Antibiotic&lt;br>Treatment]
D --> G[Disease Progression&lt;br>Treatment Adjustment Needed]
E --> H[Continue Current&lt;br>Management Plan]
style A fill:#f5f5f5,stroke:#333,stroke-width:1px
style B fill:#e1f5fe,stroke:#01579b,stroke-width:2px
style C fill:#e8f5e9,stroke:#2e7d32,stroke-width:2px
style D fill:#ffebee,stroke:#c62828,stroke-width:2px
style E fill:#fff8e1,stroke:#ff8f00,stroke-width:2px
style F fill:#c8e6c9,stroke:#2e7d32,stroke-width:1px
style G fill:#ffcdd2,stroke:#c62828,stroke-width:1px
style H fill:#ffecb3,stroke:#ff8f00,stroke-width:1px
&lt;/div>
&lt;h2 id="ai-model-implications">AI Model Implications&lt;/h2>
&lt;blockquote>
&lt;p>Why Most MLLMs Fall Short — and How Libra Goes Further&lt;/p>
&lt;/blockquote>
&lt;p>Many multimodal large language models (MLLMs) struggle with &lt;strong>temporal reasoning&lt;/strong> for three key reasons:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>No awareness of time&lt;/strong>: They treat images independently and ignore their chronological order.&lt;/li>
&lt;li>&lt;strong>Hallucinated references&lt;/strong>: They fabricate prior findings without reliable comparison.&lt;/li>
&lt;li>&lt;strong>Lack of temporal alignment&lt;/strong>: They have no built-in mechanism to align or contrast image features across time.&lt;/li>
&lt;/ul>
&lt;blockquote>
&lt;p>&lt;mark>Libra tackles these issues head-on.&lt;/mark>&lt;/p>
&lt;/blockquote>
&lt;p>Instead of prompting the model to &amp;ldquo;guess&amp;rdquo; what might have changed, Libra incorporates &lt;strong>explicit temporal awareness&lt;/strong> into both its architecture and training process.&lt;/p>
&lt;p>We feed the model structured, temporally aligned visual features extracted from the &lt;strong>current&lt;/strong> and &lt;strong>previous&lt;/strong> images. This enables Libra to:&lt;/p>
&lt;ul>
&lt;li>Detect fine-grained, clinically meaningful changes,&lt;/li>
&lt;li>Avoid hallucination,&lt;/li>
&lt;li>And reason about progression or stability in a way that mirrors clinical thinking.&lt;/li>
&lt;/ul>
&lt;hr>
&lt;p>In the next section, &lt;a href="https://x-izhang.github.io/blog/libra-blog2/">&lt;strong>Libra – Structural Logic 🧠&lt;/strong>&lt;/a>, we’ll explore how Libra’s architecture is intentionally designed to reflect clinical reasoning. You&amp;rsquo;ll see how its modular structure enables it to reason across both &lt;strong>time&lt;/strong> and &lt;strong>image features&lt;/strong> with precision — building on the foundation established in &lt;strong>Libra – Temporal Insight 🕰️&lt;/strong>.&lt;/p></description></item><item><title>📍 See You at SICSA 2025!</title><link>https://x-izhang.github.io/post/2025sicsa/</link><pubDate>Wed, 26 Mar 2025 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/post/2025sicsa/</guid><description>&lt;p>🗓️ I&amp;rsquo;ll be attending the &lt;a href="https://sicsaconf.org/" target="_blank" rel="noopener">SICSA 2025&lt;/a> PhD Conference on &lt;strong>25–26 June&lt;/strong> at &lt;strong>Edinburgh Napier University&amp;rsquo;s Craighlockhart Campus&lt;/strong> — two days of networking and researcher training with a focus on interdisciplinary collaboration.&lt;/p>
&lt;p>Grateful to be supported by the &lt;strong>University of Glasgow&lt;/strong> for this trip.&lt;/p>
&lt;p>See you there!&lt;/p></description></item><item><title>ReXrank</title><link>https://x-izhang.github.io/project/rexrank/</link><pubDate>Mon, 24 Mar 2025 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/project/rexrank/</guid><description>&lt;p>ReXrank is a public leaderboard for AI-powered radiology report generation from chest x-ray images.&lt;/p></description></item><item><title>Libra: Leveraging Temporal Images for Biomedical Radiology Analysis</title><link>https://x-izhang.github.io/publication/zhang-2025-libraleveragingtemporalimages/</link><pubDate>Wed, 01 Jan 2025 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/publication/zhang-2025-libraleveragingtemporalimages/</guid><description>&lt;div class="flex px-4 py-3 mb-6 rounded-md bg-primary-100 dark:bg-primary-900">
&lt;span class="pr-3 pt-1 text-primary-600 dark:text-primary-300">
&lt;svg height="24" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24">&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="m11.25 11.25l.041-.02a.75.75 0 0 1 1.063.852l-.708 2.836a.75.75 0 0 0 1.063.853l.041-.021M21 12a9 9 0 1 1-18 0a9 9 0 0 1 18 0m-9-3.75h.008v.008H12z"/>&lt;/svg>
&lt;/span>
&lt;span class="dark:text-neutral-300">For the latest updates and details, visit the &lt;a href="https://x-izhang.github.io/Libra_v1.0/" target="_blank" rel="noopener">Libra Project Website&lt;/a>.&lt;/span>
&lt;/div>
&lt;h2 id="why-libra">Why Libra?&lt;/h2>
&lt;figure>&lt;img src="https://x-izhang.github.io/publication/zhang-2025-libraleveragingtemporalimages/figure_1.png">
&lt;/figure>
&lt;p>Temporal hallucination is a critical challenge in radiology report generation (RRG). Traditional multimodal large language models (MLLMs) struggle to integrate prior images correctly, often generating:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>Spurious references&lt;/strong> to nonexistent prior studies (Single-image case).&lt;/li>
&lt;li>&lt;strong>Inaccurate interpretations&lt;/strong> of disease progression (Temporal-image case).&lt;/li>
&lt;/ul>
&lt;p>&lt;strong>Libra&lt;/strong> addresses these limitations by integrating a &lt;strong>Temporal Alignment Connector (TAC)&lt;/strong> to improve temporal awareness, ensuring:&lt;/p>
&lt;blockquote>
&lt;ul>
&lt;li>Prior studies are correctly referenced only when available.&lt;/li>
&lt;li>Hallucinated references are eliminated, avoiding misleading reports.&lt;/li>
&lt;li>Temporal changes are accurately captured, ensuring clinically meaningful outputs.&lt;/li>
&lt;/ul>
&lt;/blockquote>
&lt;h2 id="overview">Overview&lt;/h2>
&lt;p>We propose &lt;strong>Libra&lt;/strong> (&lt;strong>L&lt;/strong>everaging Temporal &lt;strong>I&lt;/strong>mages for &lt;strong>B&lt;/strong>iomedical &lt;strong>R&lt;/strong>adiology &lt;strong>A&lt;/strong>nalysis), a novel framework tailored for radiology report generation (RRG) that incorporates temporal change information to address the challenges of interpreting medical images effectively.&lt;/p>
&lt;p>Libra leverages RAD-DINO, a pre-trained visual transformer, as its image encoder to generate robust and scalable image features. These features are further refined by a &lt;strong>Temporal Alignment Connector (TAC)&lt;/strong>, a key innovation in Libra&amp;rsquo;s architecture. The TAC comprises:&lt;/p>
&lt;ul>
&lt;li>
&lt;p>&lt;strong>Layerwise Feature Extractor (LFE)&lt;/strong>: Captures high-granularity image feature embeddings from the encoder.&lt;/p>
&lt;/li>
&lt;li>
&lt;p>&lt;strong>Temporal Fusion Module (TFM)&lt;/strong>: Integrates temporal references from prior studies to enhance temporal awareness and reasoning.&lt;/p>
&lt;/li>
&lt;/ul>
&lt;p>These refined features are fed into Meditron, a specialised medical large language model (LLM), to generate comprehensive, temporally-aware radiology reports. Libra’s modular design seamlessly integrates state-of-the-art open-source pre-trained models for both image and text, aligning them through a temporal-aware adapter to ensure robust cross-modal reasoning and understanding.&lt;/p>
&lt;p>Through a two-stage training strategy, Libra demonstrates the powerful potential of multimodal large language models (MLLMs) in specialised radiology applications. Extensive experiments on the &lt;strong>MIMIC-CXR dataset&lt;/strong> highlight Libra&amp;rsquo;s performance, setting a new state-of-the-art benchmark among models of the same parameter scale.&lt;/p>
&lt;h2 id="key-contributions">Key Contributions&lt;/h2>
&lt;ul>
&lt;li>
&lt;p>&lt;strong>Temporal Awareness:&lt;/strong> Libra captures and synthesizes temporal changes in medical images, addressing the challenge of handling prior study citations in RRG tasks.&lt;/p>
&lt;/li>
&lt;li>
&lt;p>&lt;strong>Innovative Architecture:&lt;/strong> The Temporal Alignment Connector (TAC) ensures high-granularity feature extraction and temporal integration, significantly enhancing cross-modal reasoning capabilities.&lt;/p>
&lt;/li>
&lt;li>
&lt;p>&lt;strong>State-of-the-Art Performance:&lt;/strong> Libra achieves outstanding results on the MIMIC-CXR dataset, outperforming existing MLLMs in both accuracy and temporal reasoning.&lt;/p>
&lt;/li>
&lt;li>
&lt;p>&lt;strong>Libra Repository:&lt;/strong> Our code space provides a public and detailed implementation of Libra, facilitating reproducibility and further research in the field, and integrates both training and evaluation processes, ensuring a streamlined and efficient workflow for developing and testing radiology report generation models.&lt;/p>
&lt;/li>
&lt;/ul>
&lt;h2 id="experimental-results">Experimental Results&lt;/h2>
&lt;figure>&lt;img src="https://x-izhang.github.io/publication/zhang-2025-libraleveragingtemporalimages/results.png">
&lt;/figure>
&lt;p>Libra was designed to excel in radiology report generation (RRG), and its performance speaks for itself. Here&amp;rsquo;s a quick breakdown of its achievements:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>Lexical and Clinical Metrics:&lt;/strong> Libra delivers competitive results across traditional metrics like ROUGE-L, BLEU, METEOR, and RadGraph-based scores.&lt;/li>
&lt;li>&lt;strong>Radiologist-Aligned Metrics:&lt;/strong> It leads in the RadCliQ metric and CheXbert vector similarity, showcasing its alignment with clinical expectations.&lt;/li>
&lt;li>&lt;strong>CheXpert Classification:&lt;/strong> Libra achieves the highest Macro-F1 scores and remains competitive in Micro-F1, further solidifying its reliability in clinical classification tasks.&lt;/li>
&lt;/ul>
&lt;p>By leveraging its Temporal Alignment Connector (TAC), Libra effectively captures temporal contexts, generating radiology reports that are both accurate and clinically meaningful. While there are minor gaps in select clinical metrics, Libra&amp;rsquo;s robust performance demonstrates its potential to set new standards in RRG.&lt;/p>
&lt;h2 id="performance-analysis">Performance Analysis&lt;/h2>
&lt;figure>&lt;img src="https://x-izhang.github.io/publication/zhang-2025-libraleveragingtemporalimages/figure_3.png">
&lt;/figure>
&lt;h3 id="cases-without-prior-images">Cases Without Prior Images&lt;/h3>
&lt;p>In scenarios where no prior image is available, Libra shines by delivering detailed and clinically relevant descriptions without introducing spurious references. For instance, as shown in &lt;a href="#figure-figure_3">&lt;strong>Figure 3 (a)&lt;/strong>&lt;/a>, Libra identified &amp;ldquo;sternal wires&amp;rdquo; and their type, going beyond the ground truth. This demonstrates its ability to provide meaningful insights while avoiding errors caused by nonexistent prior studies.&lt;/p>
&lt;h3 id="cases-with-prior-images">Cases With Prior Images&lt;/h3>
&lt;p>When prior images are available, Libra takes its analysis to the next level. In &lt;a href="#figure-figure_3">&lt;strong>Figure 3 (b)&lt;/strong>&lt;/a>, new abnormalities such as pleural effusion and pneumonia were identified in the current image. Without considering the prior image, Libra accurately described the findings without inferring disease progression, maintaining clinical accuracy. However, when the prior image was included, Libra effectively captured the progressive changes, provided detailed descriptions, and explicitly referenced the comparison. This capability ensures a clear understanding of temporal changes and enhances the accuracy of disease progression descriptions.&lt;/p>
&lt;h3 id="evaluating-temporal-consistency">Evaluating Temporal Consistency&lt;/h3>
&lt;p>To test Libra&amp;rsquo;s temporal reasoning, we swapped the image order, treating the prior image as the current image and vice versa. The generated report reflected an improved patient condition, aligning with the reversed input sequence but contradicting the original ground truth. Interestingly, the report closely resembled the original description of the prior image, as shown at the bottom of &lt;a href="#figure-figure_3">&lt;strong>Figure 3 (b)&lt;/strong>&lt;/a>. This highlights Libra&amp;rsquo;s ability to adapt to different temporal contexts, generating accurate and contextually consistent reports that align with standard clinical practices.&lt;/p>
&lt;p>Libra&amp;rsquo;s performance in these cases underscores its potential to revolutionize radiology report generation by seamlessly integrating temporal reasoning into its analysis.&lt;/p>
&lt;!-- &lt;figure>&lt;img src="https://x-izhang.github.io/publication/zhang-2025-libraleveragingtemporalimages/libra_card.png">
&lt;/figure>
-->
&lt;h2 id="bibtex">BibTeX&lt;/h2>
&lt;div class="highlight">&lt;pre tabindex="0" class="chroma">&lt;code class="language-bibtex" data-lang="bibtex">&lt;span class="line">&lt;span class="cl">&lt;span class="nc">@inproceedings&lt;/span>&lt;span class="p">{&lt;/span>&lt;span class="nl">zhang-etal-2025-libra&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">title&lt;/span> &lt;span class="p">=&lt;/span> &lt;span class="s">&amp;#34;Libra: Leveraging Temporal Images for Biomedical Radiology Analysis&amp;#34;&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">author&lt;/span> &lt;span class="p">=&lt;/span> &lt;span class="s">&amp;#34;Zhang, Xi and
&lt;/span>&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="s"> Meng, Zaiqiao and
&lt;/span>&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="s"> Lever, Jake and
&lt;/span>&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="s"> Ho, Edmond S. L.&amp;#34;&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">editor&lt;/span> &lt;span class="p">=&lt;/span> &lt;span class="s">&amp;#34;Che, Wanxiang and
&lt;/span>&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="s"> Nabende, Joyce and
&lt;/span>&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="s"> Shutova, Ekaterina and
&lt;/span>&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="s"> Pilehvar, Mohammad Taher&amp;#34;&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">booktitle&lt;/span> &lt;span class="p">=&lt;/span> &lt;span class="s">&amp;#34;Findings of the Association for Computational Linguistics: ACL 2025&amp;#34;&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">month&lt;/span> &lt;span class="p">=&lt;/span> &lt;span class="nv">jul&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">year&lt;/span> &lt;span class="p">=&lt;/span> &lt;span class="s">&amp;#34;2025&amp;#34;&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">address&lt;/span> &lt;span class="p">=&lt;/span> &lt;span class="s">&amp;#34;Vienna, Austria&amp;#34;&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">publisher&lt;/span> &lt;span class="p">=&lt;/span> &lt;span class="s">&amp;#34;Association for Computational Linguistics&amp;#34;&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">url&lt;/span> &lt;span class="p">=&lt;/span> &lt;span class="s">&amp;#34;https://aclanthology.org/2025.findings-acl.888/&amp;#34;&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">pages&lt;/span> &lt;span class="p">=&lt;/span> &lt;span class="s">&amp;#34;17275--17303&amp;#34;&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">ISBN&lt;/span> &lt;span class="p">=&lt;/span> &lt;span class="s">&amp;#34;979-8-89176-256-5&amp;#34;&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">abstract&lt;/span> &lt;span class="p">=&lt;/span> &lt;span class="s">&amp;#34;Radiology report generation (RRG) requires advanced medical image analysis, effective temporal reasoning, and accurate text generation. While multimodal large language models (MLLMs) align with pre-trained vision encoders to enhance visual-language understanding, most existing methods rely on single-image analysis or rule-based heuristics to process multiple images, failing to fully leverage temporal information in multi-modal medical datasets. In this paper, we introduce **Libra**, a temporal-aware MLLM tailored for chest X-ray report generation. Libra combines a radiology-specific image encoder with a novel Temporal Alignment Connector (**TAC**), designed to accurately capture and integrate temporal differences between paired current and prior images. Extensive experiments on the MIMIC-CXR dataset demonstrate that Libra establishes a new state-of-the-art benchmark among similarly scaled MLLMs, setting new standards in both clinical relevance and lexical accuracy. All source code and data are publicly available at: https://github.com/X-iZhang/Libra.&amp;#34;&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="p">}&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="nc">@inproceedings&lt;/span>&lt;span class="p">{&lt;/span>&lt;span class="nl">zhang2025libra&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">title&lt;/span>&lt;span class="p">=&lt;/span>&lt;span class="s">{Libra: Leveraging temporal images for biomedical radiology analysis}&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">author&lt;/span>&lt;span class="p">=&lt;/span>&lt;span class="s">{Zhang, Xi and Meng, Zaiqiao and Lever, Jake and Ho, Edmond SL}&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">booktitle&lt;/span>&lt;span class="p">=&lt;/span>&lt;span class="s">{Findings of the Association for Computational Linguistics: ACL 2025}&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">pages&lt;/span>&lt;span class="p">=&lt;/span>&lt;span class="s">{17275--17303}&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">year&lt;/span>&lt;span class="p">=&lt;/span>&lt;span class="s">{2025}&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="p">}&lt;/span>
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div></description></item><item><title>Libra</title><link>https://x-izhang.github.io/project/libra/</link><pubDate>Wed, 20 Nov 2024 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/project/libra/</guid><description>&lt;p>Temporally-aware MLLM for radiology report generation. Supports LLM backbones, validation, resume training, and smart saving.&lt;/p></description></item><item><title>🎉 Paper Accepted — See You at ACL 2024!</title><link>https://x-izhang.github.io/post/2024acl/</link><pubDate>Sat, 20 Jul 2024 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/post/2024acl/</guid><description>&lt;p>&lt;a href="https://x-izhang.github.io/publication/zhang-etal-2024-gla/">&lt;em>&amp;ldquo;Gla-AI4BioMed at RRG24: Visual Instruction-tuned Adaptation for Radiology Report Generation&amp;rdquo;&lt;/em>&lt;/a> has been accepted to BioNLP @ &lt;a href="https://2024.aclweb.org/" target="_blank" rel="noopener">ACL 2024&lt;/a>!&lt;/p>
&lt;p>For detailed model information and source code, please visit our &lt;mark>GitHub repository&lt;/mark>: &lt;a href="https://github.com/X-iZhang/RRG-BioNLP-ACL2024" target="_blank" rel="noopener">RRG-BioNLP-ACL2024&lt;/a>&lt;/p>
&lt;h3 id="-poster">🪧 Poster&lt;/h3>
&lt;figure>&lt;img src="https://x-izhang.github.io/post/2024acl/2024acl_poster.png">
&lt;/figure></description></item><item><title>Gla-AI4BioMed at RRG24: Visual Instruction-tuned Adaptation for Radiology Report Generation</title><link>https://x-izhang.github.io/publication/zhang-etal-2024-gla/</link><pubDate>Sat, 20 Jul 2024 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/publication/zhang-etal-2024-gla/</guid><description>&lt;h2 id="introduction">Introduction&lt;/h2>
&lt;p>Radiology reports are critical tools for interpreting medical imaging, such as chest X-rays. These reports typically consist of two main sections: &lt;code>FINDINGS&lt;/code>, which detail observations from the image, and &lt;code>IMPRESSIONS&lt;/code>, which summarize key takeaways and recommendations.&lt;/p>
&lt;p>For example:&lt;/p>
&lt;blockquote>
&lt;p>&lt;strong>FINDINGS:&lt;/strong>&lt;br>
Increase in size of the left pleural effusion compared to the prior exam. Right lung remains clear. Mildly enlarged but stable heart size.&lt;/p>
&lt;/blockquote>
&lt;blockquote>
&lt;p>&lt;strong>IMPRESSIONS:&lt;/strong>&lt;br>
Increase in left pleural effusion. Stable mild cardiomegaly.&lt;/p>
&lt;/blockquote>
&lt;p>Generating such reports automatically is a challenging task that requires aligning visual data with textual descriptions. While general-purpose visual language models like LLaVA and InstructBLIP have shown promise in multimodal tasks, radiology report generation demands a higher level of precision and domain-specific adaptation.&lt;/p>
&lt;p>Our approach focuses on fine-tuning a visual language model specifically for radiology. By aligning chest X-ray features with a large language model and applying advanced techniques like Low-Rank Adaptation (LoRA), we enhance the model&amp;rsquo;s ability to generate accurate and clinically relevant reports. Additionally, we employ a method to process multiple images simultaneously, enabling the model to capture nuanced details across different X-rays.&lt;/p>
&lt;p>This work was developed for &lt;a href="https://stanford-aimi.github.io/RRG24/" target="_blank" rel="noopener">the RRG24 Shared Task at BioNLP 2024&lt;/a>, where our model achieved competitive results, ranking &lt;strong>4th&lt;/strong> on the leaderboard. Key contributions include:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>Domain-specific fine-tuning&lt;/strong>: Optimizing the model for radiology tasks through visual instruction tuning.&lt;/li>
&lt;li>&lt;strong>Multi-image processing&lt;/strong>: Combining multiple X-rays into a single input for efficient and accurate interpretation.&lt;/li>
&lt;/ul>
&lt;h2 id="methodology">Methodology&lt;/h2>
&lt;p>Our approach builds on insights from LLaVA-Med, emphasizing the advantages of starting with a language-only pretrained LLM rather than a multimodal-trained base. The model integrates an image encoder with a learnable adapter, following the LLaVA-1.5 architecture. Training involves an auto-regressive language modeling approach using cross-entropy loss, with hyperparameters aligned to LLaVA-1.5. We employ Low-Rank Adaptation (LoRA) for efficient fine-tuning, starting with adapter pretraining for one epoch, followed by joint tuning for three epochs.&lt;/p>
&lt;figure>&lt;img src="https://x-izhang.github.io/publication/zhang-etal-2024-gla/featured.png">
&lt;/figure>
&lt;h3 id="proposed-models">Proposed Models&lt;/h3>
&lt;p>To tackle the RRG24 shared task, we developed two specialized models: &lt;strong>Med-CXRGen-F&lt;/strong> for Findings and &lt;strong>Med-CXRGen-I&lt;/strong> for Impressions. Both models leverage CLIP as the image encoder and Vicuna-1.5 as the LLM. A multi-layer perceptron (MLP) adapter with GELU activations and a uniform hidden size of 1024 processes image features before aligning them with the LLM.&lt;/p>
&lt;h3 id="how-it-works">How It Works&lt;/h3>
&lt;p>&lt;strong>Image Encoding&lt;/strong>: The image encoder converts chest X-rays into patch tokens, extracting embeddings from the penultimate layer.&lt;/p>
&lt;p>&lt;strong>Adapter Alignment&lt;/strong>: The MLP adapter processes these embeddings, aligning them to the LLM&amp;rsquo;s input format.&lt;/p>
&lt;p>&lt;strong>Prompt Design&lt;/strong>: The model is prompted with:&lt;/p>
&lt;blockquote>
&lt;p>&lt;em>Provide a description of the findings/impressions from the radiology &amp;lt;image&amp;gt;\n image.&lt;/em>&lt;/p>
&lt;/blockquote>
&lt;p>Here, &lt;em>&amp;quot;&amp;lt;image&amp;gt;&amp;quot;&lt;/em> acts as a placeholder token, guiding the LLM to generate text based on the input image, as shown &lt;a href="#figure-rrg24">in Figure&lt;/a>.&lt;/p>
&lt;h3 id="training">Training&lt;/h3>
&lt;p>Our training process involves a two-stage approach to optimize the model for radiology report generation: &lt;a href="#figure-rrg24">(as shown in Figure)&lt;/a>&lt;/p>
&lt;ol>
&lt;li>
&lt;p>&lt;strong>Chest X-ray Feature Alignment&lt;/strong>&lt;br>
In the first stage, we align chest X-ray image features with textual embeddings in the language model. Using the provided dataset, the model predicts captions based on image inputs and instructions. During this phase, only the MLP adapter is updated, while the visual encoder and LLM weights remain frozen. This single-epoch training expands the vocabulary of aligned image-text tokens specific to the radiology domain.&lt;/p>
&lt;/li>
&lt;li>
&lt;p>&lt;strong>Fine-tuning for Report Generation&lt;/strong>&lt;br>
In the second stage, we fine-tune the pre-trained LLM weights using LoRA technology. The visual encoder and adapter weights are kept frozen, while the model undergoes three epochs of training on the dataset. This phase focuses on enhancing the model&amp;rsquo;s ability to generate accurate and clinically relevant radiology reports.&lt;/p>
&lt;/li>
&lt;/ol>
&lt;div class="flex px-4 py-3 mb-6 rounded-md bg-primary-100 dark:bg-primary-900">
&lt;span class="pr-3 pt-1 text-primary-600 dark:text-primary-300">
&lt;svg height="24" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24">&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="m11.25 11.25l.041-.02a.75.75 0 0 1 1.063.852l-.708 2.836a.75.75 0 0 0 1.063.853l.041-.021M21 12a9 9 0 1 1-18 0a9 9 0 0 1 18 0m-9-3.75h.008v.008H12z"/>&lt;/svg>
&lt;/span>
&lt;span class="dark:text-neutral-300">For implementation details, see the &lt;a href="https://github.com/X-iZhang/RRG-BioNLP-ACL2024" target="_blank" rel="noopener">Code Repo&lt;/a>.&lt;/span>
&lt;/div>
&lt;h2 id="evaluation">Evaluation&lt;/h2>
&lt;h3 id="dataset">Dataset&lt;/h3>
&lt;p>Our models were fine-tuned and evaluated using the RRG24 dataset, hosted on the BioNLP ACL'24 platform. This dataset aggregates data from multiple sources, including MIMIC-CXR, CheXpert, PadChest, BIMCV-COVID19, and OpenI. The dataset statistics are summarized below:&lt;/p>
&lt;!--
&lt;table class="table-auto w-full">
&lt;thead>
&lt;tr> &lt;th class="border-b dark:border-slate-600 font-medium p-4 pt-0 pb-3 text-slate-400 dark:text-slate-200 text-left">Dataset&lt;/th> &lt;th class="border-b dark:border-slate-600 font-medium p-4 pt-0 pb-3 text-slate-400 dark:text-slate-200 text-left">FINDINGS&lt;/th> &lt;th class="border-b dark:border-slate-600 font-medium p-4 pt-0 pb-3 text-slate-400 dark:text-slate-200 text-left">IMPRESSIONS&lt;/th> &lt;/tr>
&lt;/thead>
&lt;tbody>
&lt;tr>
&lt;td data-table-dtype="text" class="border-b border-slate-100 dark:border-slate-700 p-4 text-slate-500 dark:text-slate-400">Training&lt;/td>
&lt;td data-table-dtype="text" class="border-b border-slate-100 dark:border-slate-700 p-4 text-slate-500 dark:text-slate-400">344,394&lt;/td>
&lt;td data-table-dtype="text" class="border-b border-slate-100 dark:border-slate-700 p-4 text-slate-500 dark:text-slate-400">366,413&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td data-table-dtype="text" class="border-b border-slate-100 dark:border-slate-700 p-4 text-slate-500 dark:text-slate-400">Validation&lt;/td>
&lt;td data-table-dtype="text" class="border-b border-slate-100 dark:border-slate-700 p-4 text-slate-500 dark:text-slate-400">8,839&lt;/td>
&lt;td data-table-dtype="text" class="border-b border-slate-100 dark:border-slate-700 p-4 text-slate-500 dark:text-slate-400">9,331&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td data-table-dtype="text" class="border-b border-slate-100 dark:border-slate-700 p-4 text-slate-500 dark:text-slate-400">Test-Public&lt;/td>
&lt;td data-table-dtype="text" class="border-b border-slate-100 dark:border-slate-700 p-4 text-slate-500 dark:text-slate-400">2,692&lt;/td>
&lt;td data-table-dtype="text" class="border-b border-slate-100 dark:border-slate-700 p-4 text-slate-500 dark:text-slate-400">2,967&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td data-table-dtype="text" class="border-b border-slate-100 dark:border-slate-700 p-4 text-slate-500 dark:text-slate-400">Test-Hidden&lt;/td>
&lt;td data-table-dtype="text" class="border-b border-slate-100 dark:border-slate-700 p-4 text-slate-500 dark:text-slate-400">1,063&lt;/td>
&lt;td data-table-dtype="text" class="border-b border-slate-100 dark:border-slate-700 p-4 text-slate-500 dark:text-slate-400">1,428&lt;/td>
&lt;/tr>
&lt;/tbody>
&lt;/table>
-->
&lt;table>
&lt;thead>
&lt;tr>
&lt;th>Dataset&lt;/th>
&lt;th>FINDINGS&lt;/th>
&lt;th>IMPRESSIONS&lt;/th>
&lt;/tr>
&lt;/thead>
&lt;tbody>
&lt;tr>
&lt;td>Training&lt;/td>
&lt;td>344,394&lt;/td>
&lt;td>366,413&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Validation&lt;/td>
&lt;td>8,839&lt;/td>
&lt;td>9,331&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Test-Public&lt;/td>
&lt;td>2,692&lt;/td>
&lt;td>2,967&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Test-Hidden&lt;/td>
&lt;td>1,063&lt;/td>
&lt;td>1,428&lt;/td>
&lt;/tr>
&lt;/tbody>
&lt;/table>
&lt;h3 id="results">Results&lt;/h3>
&lt;p>Our models achieved competitive F1-RadGraph scores of 24.13 and 22.79 for the Findings and Impressions sections, respectively, on the public test set. On the hidden test set, the scores were 24.13 and 22.10, securing &lt;strong>4th&lt;/strong> place on the leaderboard at the time of submission. Additionally, the Bertscore results highlight the high-quality text generation capabilities of our models, with scores of 53.45 for Findings and 47.39 for Impressions on the public test set.&lt;/p>
&lt;h2 id="discussion">Discussion&lt;/h2>
&lt;p>The difference in performance between the Findings and Impressions sections comes down to their unique roles. Findings are all about objectively describing what’s seen in the images, while Impressions focus on drawing diagnostic conclusions. This difference, along with the varying lengths of these sections, makes generating Impressions a bit trickier, which is reflected in the evaluation scores.&lt;/p>
&lt;p>Another factor is the variety of diseases in the test set. This uneven distribution can make it harder for the model to generalize well. Plus, when the model processes multiple images at once, it sometimes includes irrelevant ones, which can throw off its accuracy. Improving the model’s ability to filter out unnecessary images could make a big difference.&lt;/p>
&lt;p>Looking ahead, there’s a lot of room for improvement. Fine-tuning the model specifically for medical data and adapting it to better handle domain-specific challenges could help. Adding the ability to track changes over time and creating smarter frameworks for multi-modal reports are also exciting directions to explore. These upgrades could make the model even more useful and reliable in real-world clinical settings.&lt;/p>
&lt;h2 id="conclusion">Conclusion&lt;/h2>
&lt;p>In this study, we introduced a model designed to generate radiology reports by aligning visual and textual data. By fine-tuning it for specific tasks, we managed to achieve solid results, including a fourth-place finish in the RRG24 Shared Task at BioNLP 2024. This shows the potential of our approach for specialized medical applications.&lt;/p>
&lt;p>Moving forward, we’re excited to explore ways to make the model even better. This includes developing methods to handle multi-modal data more effectively and incorporating time-based insights. These improvements could take the model’s accuracy and practicality to the next level for clinical use.&lt;/p>
&lt;h2 id="limitations">Limitations&lt;/h2>
&lt;p>While the results are promising, there are a few challenges we need to address:&lt;/p>
&lt;ol>
&lt;li>&lt;strong>Easier-to-Detect Conditions&lt;/strong>: Some diseases are simpler to identify, which might inflate the evaluation scores.&lt;/li>
&lt;li>&lt;strong>Dataset Imbalance&lt;/strong>: The training data isn’t perfectly balanced in terms of imaging types, body parts, or report lengths, which could impact performance.&lt;/li>
&lt;li>&lt;strong>Inconsistent Reporting Styles&lt;/strong>: Radiologists have different ways of writing reports, and this variability makes it harder for the model to generate consistent outputs.&lt;/li>
&lt;/ol>
&lt;div class="flex px-4 py-3 mb-6 rounded-md bg-yellow-100 dark:bg-yellow-900">
&lt;span class="pr-3 pt-1 text-red-400">
&lt;svg height="24" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24">&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="M12 9v3.75m-9.303 3.376c-.866 1.5.217 3.374 1.948 3.374h14.71c1.73 0 2.813-1.874 1.948-3.374L13.949 3.378c-.866-1.5-3.032-1.5-3.898 0zM12 15.75h.007v.008H12z"/>&lt;/svg>
&lt;/span>
&lt;span class="dark:text-neutral-300">The data used for the training processes were exclusively sourced from the provided dataset, which complies with ethical standards regarding Patient Health Information. Adherence to these standards guarantees the responsible handling of sensitive data throughout our research.&lt;/span>
&lt;/div>
&lt;h2 id="bibtex">BibTeX&lt;/h2>
&lt;div class="highlight">&lt;pre tabindex="0" class="chroma">&lt;code class="language-bibtex" data-lang="bibtex">&lt;span class="line">&lt;span class="cl">&lt;span class="nc">@inproceedings&lt;/span>&lt;span class="p">{&lt;/span>&lt;span class="nl">zhang-etal-2024-gla&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">title&lt;/span> &lt;span class="p">=&lt;/span> &lt;span class="s">&amp;#34;Gla-{AI}4{B}io{M}ed at {RRG}24: Visual Instruction-tuned Adaptation for Radiology Report Generation&amp;#34;&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">author&lt;/span> &lt;span class="p">=&lt;/span> &lt;span class="s">&amp;#34;Zhang, Xi and
&lt;/span>&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="s"> Meng, Zaiqiao and
&lt;/span>&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="s"> Lever, Jake and
&lt;/span>&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="s"> Ho, Edmond S.L.&amp;#34;&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">editor&lt;/span> &lt;span class="p">=&lt;/span> &lt;span class="s">&amp;#34;Demner-Fushman, Dina and
&lt;/span>&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="s"> Ananiadou, Sophia and
&lt;/span>&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="s"> Miwa, Makoto and
&lt;/span>&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="s"> Roberts, Kirk and
&lt;/span>&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="s"> Tsujii, Junichi&amp;#34;&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">booktitle&lt;/span> &lt;span class="p">=&lt;/span> &lt;span class="s">&amp;#34;Proceedings of the 23rd Workshop on Biomedical Natural Language Processing&amp;#34;&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">month&lt;/span> &lt;span class="p">=&lt;/span> &lt;span class="nv">aug&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">year&lt;/span> &lt;span class="p">=&lt;/span> &lt;span class="s">&amp;#34;2024&amp;#34;&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">address&lt;/span> &lt;span class="p">=&lt;/span> &lt;span class="s">&amp;#34;Bangkok, Thailand&amp;#34;&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">publisher&lt;/span> &lt;span class="p">=&lt;/span> &lt;span class="s">&amp;#34;Association for Computational Linguistics&amp;#34;&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">url&lt;/span> &lt;span class="p">=&lt;/span> &lt;span class="s">&amp;#34;https://aclanthology.org/2024.bionlp-1.54/&amp;#34;&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">doi&lt;/span> &lt;span class="p">=&lt;/span> &lt;span class="s">&amp;#34;10.18653/v1/2024.bionlp-1.54&amp;#34;&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">pages&lt;/span> &lt;span class="p">=&lt;/span> &lt;span class="s">&amp;#34;624--634&amp;#34;&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="p">}&lt;/span>
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;!-- &lt;figure>&lt;img src="https://x-izhang.github.io/publication/zhang-etal-2024-gla/featured.png">
&lt;/figure>
[A Figure](#figure-hello)
&lt;div class="flex px-4 py-3 mb-6 rounded-md bg-primary-100 dark:bg-primary-900">
&lt;span class="pr-3 pt-1 text-primary-600 dark:text-primary-300">
&lt;svg height="24" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24">&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="m11.25 11.25l.041-.02a.75.75 0 0 1 1.063.852l-.708 2.836a.75.75 0 0 0 1.063.853l.041-.021M21 12a9 9 0 1 1-18 0a9 9 0 0 1 18 0m-9-3.75h.008v.008H12z"/>&lt;/svg>
&lt;/span>
&lt;span class="dark:text-neutral-300">For more details, see the &lt;a href="https://github.com/X-iZhang/RRG-BioNLP-ACL2024" target="_blank" rel="noopener">project repo&lt;/a>.&lt;/span>
&lt;/div> -->
&lt;!--
&lt;div class="flex px-4 py-3 mb-6 rounded-md bg-yellow-100 dark:bg-yellow-900">
&lt;span class="pr-3 pt-1 text-red-400">
&lt;svg height="24" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24">&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="M12 9v3.75m-9.303 3.376c-.866 1.5.217 3.374 1.948 3.374h14.71c1.73 0 2.813-1.874 1.948-3.374L13.949 3.378c-.866-1.5-3.032-1.5-3.898 0zM12 15.75h.007v.008H12z"/>&lt;/svg>
&lt;/span>
&lt;span class="dark:text-neutral-300">A Markdown aside is useful for displaying notices, hints, or definitions to your readers.&lt;/span>
&lt;/div> --></description></item><item><title>Projects</title><link>https://x-izhang.github.io/projects/</link><pubDate>Sun, 19 May 2024 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/projects/</guid><description/></item><item><title>Experience</title><link>https://x-izhang.github.io/experience/</link><pubDate>Tue, 24 Oct 2023 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/experience/</guid><description/></item></channel></rss>