Cited 2 times in 
Cited 2 times in 
Read like a radiologist: Efficient vision-language model for 3D medical imaging interpretation
| DC Field | Value | Language |
|---|---|---|
| dc.contributor.author | Lee, Changsun | - |
| dc.contributor.author | Park, Sangjoon | - |
| dc.contributor.author | Shin, Cheong-Il | - |
| dc.contributor.author | Choi, Woo Hee | - |
| dc.contributor.author | Park, Hyun Jeong | - |
| dc.contributor.author | Lee, Jeong Eun | - |
| dc.contributor.author | Ye, Jong Chul | - |
| dc.date.accessioned | 2026-06-18T01:50:01Z | - |
| dc.date.available | 2026-06-18T01:50:01Z | - |
| dc.date.created | 2026-06-08 | - |
| dc.date.issued | 2026-06 | - |
| dc.identifier.issn | 1361-8415 | - |
| dc.identifier.uri | https://ir.ymlib.yonsei.ac.kr/handle/22282913/212706 | - |
| dc.description.abstract | Recent medical vision-language models (VLMs) have shown promise in 2D medical image interpretation. However extending them to 3D medical imaging has been challenging due to computational complexities and data scarcity. Although a few recent VLMs specified for 3D medical imaging have emerged, all are limited to learning volumetric representation of a 3D medical image as a set of sub-volumetric features. Such process introduces overly correlated representations along the z-axis that neglect slice-specific clinical details, particularly for 3D medical images where adjacent slices have low redundancy. To address this limitation, we introduce MS-VLM that mimic radiologists' workflow in 3D medical image interpretation. Specifically, radiologists analyze 3D medical images by examining individual slices sequentially and synthesizing information across slices and views. Likewise, MS-VLM leverages self-supervised 2D transformer encoders to learn a volumetric representation that capture inter-slice dependencies from a sequence of slice-specific features. Unbound by sub-volumetric patchification, MS-VLM is capable of obtaining useful volumetric representations from 3D medical images with any slice length and from multiple images acquired from different planes and phases. We evaluate MS-VLM on publicly available chest CT dataset CT-RATE and in-house rectal MRI dataset. In both scenarios, MS-VLM surpasses existing methods in radiology report generation, producing more coherent and clinically relevant reports. These findings highlight the potential of MS-VLM to advance 3D medical image interpretation and improve the robustness of medical VLMs. | - |
| dc.language | English | - |
| dc.publisher | Elsevier | - |
| dc.relation.isPartOf | MEDICAL IMAGE ANALYSIS | - |
| dc.relation.isPartOf | MEDICAL IMAGE ANALYSIS | - |
| dc.subject.MESH | Humans | - |
| dc.subject.MESH | Image Interpretation, Computer-Assisted* / methods | - |
| dc.subject.MESH | Imaging, Three-Dimensional* / methods | - |
| dc.subject.MESH | Radiologists | - |
| dc.title | Read like a radiologist: Efficient vision-language model for 3D medical imaging interpretation | - |
| dc.type | Article | - |
| dc.contributor.googleauthor | Lee, Changsun | - |
| dc.contributor.googleauthor | Park, Sangjoon | - |
| dc.contributor.googleauthor | Shin, Cheong-Il | - |
| dc.contributor.googleauthor | Choi, Woo Hee | - |
| dc.contributor.googleauthor | Park, Hyun Jeong | - |
| dc.contributor.googleauthor | Lee, Jeong Eun | - |
| dc.contributor.googleauthor | Ye, Jong Chul | - |
| dc.identifier.doi | 10.1016/j.media.2026.104077 | - |
| dc.relation.journalcode | J02201 | - |
| dc.identifier.eissn | 1361-8423 | - |
| dc.identifier.pmid | 41990528 | - |
| dc.identifier.url | https://www.sciencedirect.com/science/article/pii/S1361841526001465 | - |
| dc.subject.keyword | 3D medical imaging | - |
| dc.subject.keyword | Radiology report generation | - |
| dc.subject.keyword | Self-supervised learning | - |
| dc.subject.keyword | Vision transformers | - |
| dc.subject.keyword | Large language models | - |
| dc.contributor.affiliatedAuthor | Park, Sangjoon | - |
| dc.identifier.scopusid | 2-s2.0-105036199050 | - |
| dc.identifier.wosid | 001747383700001 | - |
| dc.citation.volume | 111 | - |
| dc.identifier.bibliographicCitation | MEDICAL IMAGE ANALYSIS, Vol.111, 2026-06 | - |
| dc.identifier.rimsid | 93269 | - |
| dc.type.rims | ART | - |
| dc.description.journalClass | 1 | - |
| dc.description.journalClass | 1 | - |
| dc.subject.keywordAuthor | 3D medical imaging | - |
| dc.subject.keywordAuthor | Radiology report generation | - |
| dc.subject.keywordAuthor | Self-supervised learning | - |
| dc.subject.keywordAuthor | Vision transformers | - |
| dc.subject.keywordAuthor | Large language models | - |
| dc.type.docType | Article | - |
| dc.description.isOpenAccess | N | - |
| dc.description.journalRegisteredClass | scie | - |
| dc.description.journalRegisteredClass | scopus | - |
| dc.relation.journalWebOfScienceCategory | Computer Science, Artificial Intelligence | - |
| dc.relation.journalWebOfScienceCategory | Computer Science, Interdisciplinary Applications | - |
| dc.relation.journalWebOfScienceCategory | Engineering, Biomedical | - |
| dc.relation.journalWebOfScienceCategory | Radiology, Nuclear Medicine & Medical Imaging | - |
| dc.relation.journalResearchArea | Computer Science | - |
| dc.relation.journalResearchArea | Engineering | - |
| dc.relation.journalResearchArea | Radiology, Nuclear Medicine & Medical Imaging | - |
| dc.identifier.articleno | 104077 | - |
Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.