<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Evaluation | Xi Zhang</title><link>https://x-izhang.github.io/tags/evaluation/</link><atom:link href="https://x-izhang.github.io/tags/evaluation/index.xml" rel="self" type="application/rss+xml"/><description>Evaluation</description><generator>Hugo Blox Builder (https://hugoblox.com)</generator><language>en-us</language><lastBuildDate>Mon, 14 Jul 2025 00:00:00 +0000</lastBuildDate><image><url>https://x-izhang.github.io/media/icon_hu134860076176174952.png</url><title>Evaluation</title><link>https://x-izhang.github.io/tags/evaluation/</link></image><item><title>🩺 RadEval Debuts！</title><link>https://x-izhang.github.io/post/2025radeval/</link><pubDate>Mon, 14 Jul 2025 00:00:00 +0000</pubDate><guid>https://x-izhang.github.io/post/2025radeval/</guid><description>&lt;h4 id="-revolutionizing-radiology-text-evaluation-with-ai-powered-metrics">🩺 Revolutionizing Radiology Text Evaluation with AI-Powered Metrics&lt;/h4>
&lt;p>Imagine having a comprehensive evaluation framework that doesn&amp;rsquo;t just measure surface-level text similarity, but truly understands clinical accuracy and medical semantics in radiology reports. This vision is now a reality with &lt;strong>RadEval&lt;/strong>, a groundbreaking, open-source evaluation toolkit designed specifically for AI-generated radiology text.&lt;/p>
&lt;h4 id="-all-in-one-metrics-for-evaluating-ai-generated-radiology-text">📊 All-in-one metrics for evaluating AI-generated radiology text&lt;/h4>
&lt;p>From traditional n-gram metrics to advanced LLM-based evaluations, RadEval provides 11+ different evaluation metrics in one unified framework, enabling researchers to thoroughly assess their radiology text generation models with domain-specific medical knowledge integration.&lt;/p>
&lt;p>For detailed handbook, please visit our &lt;a href="https://github.com/jbdel/RadEval" target="_blank" rel="noopener">GitHub repository&lt;/a>:
&lt;figure>&lt;img src="https://x-izhang.github.io/post/2025radeval/github.png">
&lt;/figure>
&lt;/p>
&lt;h4 id="-quick-start-demo">🚀 Quick Start Demo&lt;/h4>
&lt;p>Try RadEval instantly with our interactive &lt;a href="https://huggingface.co/spaces/X-iZhang/RadEval" target="_blank" rel="noopener">Gradio demo&lt;/a>:
&lt;figure>&lt;img src="https://x-izhang.github.io/post/2025radeval/demo.png">
&lt;/figure>
&lt;/p>
&lt;h4 id="-key-features">💡 Key Features&lt;/h4>
&lt;p>&lt;strong>RadEval&lt;/strong> stands out with its comprehensive approach to radiology text evaluation:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>🎯 Domain-Specific&lt;/strong>: Tailored for radiology with medical knowledge integration&lt;/li>
&lt;li>&lt;strong>📈 Multi-Metric&lt;/strong>: Supports lexical, semantic, clinical, and temporal evaluations&lt;/li>
&lt;li>&lt;strong>⚡ Easy to Use&lt;/strong>: Simple API with flexible configuration options&lt;/li>
&lt;li>&lt;strong>🔬 Research-Ready&lt;/strong>: Built-in statistical testing for system comparison&lt;/li>
&lt;li>&lt;strong>📦 PyPI Available&lt;/strong>: Install with a simple &lt;code>pip install RadEval&lt;/code>&lt;/li>
&lt;/ul>
&lt;h4 id="-advancing-radiology-ai-research-community">🏥 Advancing Radiology AI Research Community&lt;/h4>
&lt;p>We are committed to building a standardized and reproducible toolkit for researchers, clinicians, and developers dedicated to advancing AI evaluation in medical imaging and radiology. Together, we&amp;rsquo;re setting new standards for clinical AI assessment.&lt;/p>
&lt;ul>
&lt;li>&lt;a href="https://pypi.org/project/RadEval/" target="_blank" rel="noopener">&lt;strong>PyPI Package&lt;/strong>&lt;/a> — Install RadEval with pip&lt;/li>
&lt;li>&lt;a href="https://huggingface.co/IAMJB/RadEvalModernBERT" target="_blank" rel="noopener">&lt;strong>HuggingFace Model&lt;/strong>&lt;/a> — Access our domain-adapted evaluation model&lt;/li>
&lt;li>&lt;a href="https://huggingface.co/spaces/X-iZhang/RadEval" target="_blank" rel="noopener">&lt;strong>Interactive Demo&lt;/strong>&lt;/a> — Try RadEval online&lt;/li>
&lt;li>&lt;a href="https://arxiv.org/abs/2509.18030v1" target="_blank" rel="noopener">&lt;strong>Research Paper&lt;/strong>&lt;/a> — Read our detailed research paper&lt;/li>
&lt;/ul></description></item></channel></rss>