<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>ReXrank | Xi Zhang</title><link>https://x-izhang.github.io/tags/rexrank/</link><atom:link href="https://x-izhang.github.io/tags/rexrank/index.xml" rel="self" type="application/rss+xml"/><description>ReXrank</description><generator>Hugo Blox Builder (https://hugoblox.com)</generator><language>en-us</language><lastBuildDate>Thu, 01 Jan 2026 00:00:00 +0000</lastBuildDate><image><url>https://x-izhang.github.io/media/icon_hu134860076176174952.png</url><title>ReXrank</title><link>https://x-izhang.github.io/tags/rexrank/</link></image><item><title>Automated Chest X-ray Report Generation Remains Unsolved</title><link>https://x-izhang.github.io/publication/doi-10-1142-9789819824755-0017/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/publication/doi-10-1142-9789819824755-0017/</guid><description>&lt;div class="flex px-4 py-3 mb-6 rounded-md bg-primary-100 dark:bg-primary-900">
&lt;span class="pr-3 pt-1 text-primary-600 dark:text-primary-300">
&lt;svg height="24" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24">&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="m11.25 11.25l.041-.02a.75.75 0 0 1 1.063.852l-.708 2.836a.75.75 0 0 0 1.063.853l.041-.021M21 12a9 9 0 1 1-18 0a9 9 0 0 1 18 0m-9-3.75h.008v.008H12z"/>&lt;/svg>
&lt;/span>
&lt;span class="dark:text-neutral-300">For the latest updates and details, visit the &lt;a href="https://rexrank.ai/" target="_blank" rel="noopener">ReXrank Challenge Website&lt;/a>.&lt;/span>
&lt;/div>
&lt;h2 id="why-this-work">Why this work?&lt;/h2>
&lt;p>Automated chest X-ray report generation has the potential to substantially reduce radiologist workload and improve clinical efficiency. However, despite rapid progress in vision-language models and large language models (LLMs), the true clinical reliability of these systems remains unclear. Existing evaluations are often inconsistent, rely on limited test sets, or fail to probe generalization across institutions and clinically challenging abnormal cases.&lt;/p>
&lt;p>To address these limitations, we present a large-scale, standardized benchmark study through the &lt;strong>ReXrank Challenge V1.0&lt;/strong>, designed to rigorously assess the current state of automated chest X-ray report generation under realistic and clinically meaningful conditions.&lt;/p>
&lt;h2 id="what-is-the-rexrank-challenge-v10">What is the ReXrank Challenge V1.0?&lt;/h2>
&lt;p>The ReXrank Challenge V1.0 is a comprehensive evaluation effort built on &lt;strong>ReXGradient&lt;/strong>, the largest test-only dataset to date for radiology report generation, comprising &lt;strong>10,000 studies from 67 healthcare institutions&lt;/strong>. The challenge brought together submissions from academia and industry, evaluating &lt;strong>8 new models&lt;/strong> alongside &lt;strong>16 previously benchmarked state-of-the-art systems&lt;/strong> under a unified evaluation protocol.&lt;/p>
&lt;p>All models were assessed using a diverse set of metrics, ranging from traditional text similarity measures to clinically grounded and LLM-based error detection metrics, enabling a multi-dimensional analysis of model performance.&lt;/p>
&lt;h2 id="key-findings">Key Findings&lt;/h2>
&lt;ul>
&lt;li>
&lt;p>&lt;strong>Automated chest X-ray report generation remains fundamentally unsolved.&lt;/strong>&lt;br>
Even the best-performing models achieve &lt;strong>less than 45% error-free reporting on abnormal studies&lt;/strong>, highlighting a substantial gap between current AI systems and clinical readiness.&lt;/p>
&lt;/li>
&lt;li>
&lt;p>&lt;strong>Large performance gaps between normal and abnormal studies.&lt;/strong>&lt;br>
Most models perform well on normal cases (often exceeding 80–90% no-significant-error rates) but struggle significantly with abnormal findings, where clinically important errors are common.&lt;/p>
&lt;/li>
&lt;li>
&lt;p>&lt;strong>Poor cross-institutional generalization.&lt;/strong>&lt;br>
Model rankings vary dramatically across healthcare sites, indicating that strong performance on one institution does not reliably transfer to others.&lt;/p>
&lt;/li>
&lt;li>
&lt;p>&lt;strong>Evaluation metrics capture complementary but inconsistent signals.&lt;/strong>&lt;br>
Traditional lexical metrics correlate poorly with LLM-based and clinically focused error metrics, underscoring the need for more clinically aligned evaluation frameworks.&lt;/p>
&lt;/li>
&lt;/ul>
&lt;h2 id="implications">Implications&lt;/h2>
&lt;p>Our results suggest that current AI systems are &lt;strong>not yet suitable for fully autonomous radiology reporting&lt;/strong>, particularly for abnormal cases. However, they show promise as &lt;strong>assistive tools&lt;/strong> for generating preliminary drafts that can be reviewed and refined by radiologists. The findings emphasize the importance of developing:&lt;/p>
&lt;ul>
&lt;li>methods specifically targeting abnormality detection,&lt;/li>
&lt;li>strategies for robust cross-institutional generalization, and&lt;/li>
&lt;li>evaluation frameworks that better reflect clinical correctness rather than surface-level textual similarity.&lt;/li>
&lt;/ul>
&lt;h2 id="conclusion">Conclusion&lt;/h2>
&lt;p>This work provides the most comprehensive benchmark to date for automated chest X-ray report generation and delivers a clear message to the community: while progress is evident, &lt;strong>significant challenges remain before clinically reliable deployment is possible&lt;/strong>. By releasing the ReXrank Challenge V1.0 and ReXGradient dataset, we aim to establish a rigorous foundation for future research and to drive the development of more robust, clinically grounded radiology AI systems.&lt;/p>
&lt;h2 id="bibtex">BibTeX&lt;/h2>
&lt;div class="highlight">&lt;pre tabindex="0" class="chroma">&lt;code class="language-bibtex" data-lang="bibtex">&lt;span class="line">&lt;span class="cl">&lt;span class="nc">@inproceedings&lt;/span>&lt;span class="p">{&lt;/span>&lt;span class="nl">zhang2025automated&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">title&lt;/span>&lt;span class="p">=&lt;/span>&lt;span class="s">{Automated Chest X-ray Report Generation Remains Unsolved}&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">author&lt;/span>&lt;span class="p">=&lt;/span>&lt;span class="s">{Zhang, Xiaoman and Acosta, Julian Nicolas and Yang, Xiaoli and Adithan, Subathra and Luo, Luyang and Zhou, Hong-Yu and Miller, Joshua and Huang, Ouwen and Zhou, Zongwei and Hamamci, Ibrahim Ethem and others}&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">booktitle&lt;/span>&lt;span class="p">=&lt;/span>&lt;span class="s">{Biocomputing 2026: Proceedings of the Pacific Symposium}&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">pages&lt;/span>&lt;span class="p">=&lt;/span>&lt;span class="s">{236--250}&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">year&lt;/span>&lt;span class="p">=&lt;/span>&lt;span class="s">{2025}&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="na">organization&lt;/span>&lt;span class="p">=&lt;/span>&lt;span class="s">{World Scientific}&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="p">}&lt;/span>
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div></description></item><item><title>ReXrank</title><link>https://x-izhang.github.io/project/rexrank/</link><pubDate>Mon, 24 Mar 2025 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/project/rexrank/</guid><description>&lt;p>ReXrank is a public leaderboard for AI-powered radiology report generation from chest x-ray images.&lt;/p></description></item></channel></rss>