0 32

Cited 2 times in

Cited 2 times in

Read like a radiologist: Efficient vision-language model for 3D medical imaging interpretation

DC Field Value Language
dc.contributor.authorLee, Changsun-
dc.contributor.authorPark, Sangjoon-
dc.contributor.authorShin, Cheong-Il-
dc.contributor.authorChoi, Woo Hee-
dc.contributor.authorPark, Hyun Jeong-
dc.contributor.authorLee, Jeong Eun-
dc.contributor.authorYe, Jong Chul-
dc.date.accessioned2026-06-18T01:50:01Z-
dc.date.available2026-06-18T01:50:01Z-
dc.date.created2026-06-08-
dc.date.issued2026-06-
dc.identifier.issn1361-8415-
dc.identifier.urihttps://ir.ymlib.yonsei.ac.kr/handle/22282913/212706-
dc.description.abstractRecent medical vision-language models (VLMs) have shown promise in 2D medical image interpretation. However extending them to 3D medical imaging has been challenging due to computational complexities and data scarcity. Although a few recent VLMs specified for 3D medical imaging have emerged, all are limited to learning volumetric representation of a 3D medical image as a set of sub-volumetric features. Such process introduces overly correlated representations along the z-axis that neglect slice-specific clinical details, particularly for 3D medical images where adjacent slices have low redundancy. To address this limitation, we introduce MS-VLM that mimic radiologists' workflow in 3D medical image interpretation. Specifically, radiologists analyze 3D medical images by examining individual slices sequentially and synthesizing information across slices and views. Likewise, MS-VLM leverages self-supervised 2D transformer encoders to learn a volumetric representation that capture inter-slice dependencies from a sequence of slice-specific features. Unbound by sub-volumetric patchification, MS-VLM is capable of obtaining useful volumetric representations from 3D medical images with any slice length and from multiple images acquired from different planes and phases. We evaluate MS-VLM on publicly available chest CT dataset CT-RATE and in-house rectal MRI dataset. In both scenarios, MS-VLM surpasses existing methods in radiology report generation, producing more coherent and clinically relevant reports. These findings highlight the potential of MS-VLM to advance 3D medical image interpretation and improve the robustness of medical VLMs.-
dc.languageEnglish-
dc.publisherElsevier-
dc.relation.isPartOfMEDICAL IMAGE ANALYSIS-
dc.relation.isPartOfMEDICAL IMAGE ANALYSIS-
dc.subject.MESHHumans-
dc.subject.MESHImage Interpretation, Computer-Assisted* / methods-
dc.subject.MESHImaging, Three-Dimensional* / methods-
dc.subject.MESHRadiologists-
dc.titleRead like a radiologist: Efficient vision-language model for 3D medical imaging interpretation-
dc.typeArticle-
dc.contributor.googleauthorLee, Changsun-
dc.contributor.googleauthorPark, Sangjoon-
dc.contributor.googleauthorShin, Cheong-Il-
dc.contributor.googleauthorChoi, Woo Hee-
dc.contributor.googleauthorPark, Hyun Jeong-
dc.contributor.googleauthorLee, Jeong Eun-
dc.contributor.googleauthorYe, Jong Chul-
dc.identifier.doi10.1016/j.media.2026.104077-
dc.relation.journalcodeJ02201-
dc.identifier.eissn1361-8423-
dc.identifier.pmid41990528-
dc.identifier.urlhttps://www.sciencedirect.com/science/article/pii/S1361841526001465-
dc.subject.keyword3D medical imaging-
dc.subject.keywordRadiology report generation-
dc.subject.keywordSelf-supervised learning-
dc.subject.keywordVision transformers-
dc.subject.keywordLarge language models-
dc.contributor.affiliatedAuthorPark, Sangjoon-
dc.identifier.scopusid2-s2.0-105036199050-
dc.identifier.wosid001747383700001-
dc.citation.volume111-
dc.identifier.bibliographicCitationMEDICAL IMAGE ANALYSIS, Vol.111, 2026-06-
dc.identifier.rimsid93269-
dc.type.rimsART-
dc.description.journalClass1-
dc.description.journalClass1-
dc.subject.keywordAuthor3D medical imaging-
dc.subject.keywordAuthorRadiology report generation-
dc.subject.keywordAuthorSelf-supervised learning-
dc.subject.keywordAuthorVision transformers-
dc.subject.keywordAuthorLarge language models-
dc.type.docTypeArticle-
dc.description.isOpenAccessN-
dc.description.journalRegisteredClassscie-
dc.description.journalRegisteredClassscopus-
dc.relation.journalWebOfScienceCategoryComputer Science, Artificial Intelligence-
dc.relation.journalWebOfScienceCategoryComputer Science, Interdisciplinary Applications-
dc.relation.journalWebOfScienceCategoryEngineering, Biomedical-
dc.relation.journalWebOfScienceCategoryRadiology, Nuclear Medicine & Medical Imaging-
dc.relation.journalResearchAreaComputer Science-
dc.relation.journalResearchAreaEngineering-
dc.relation.journalResearchAreaRadiology, Nuclear Medicine & Medical Imaging-
dc.identifier.articleno104077-
Appears in Collections:
1. College of Medicine (의과대학) > Dept. of Radiation Oncology (방사선종양학교실) > 1. Journal Papers

qrcode

Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.