help button home button JAMIA Bigger figures
HOME HELP FEEDBACK SUBSCRIPTIONS ARCHIVE SEARCH

First published June 28, 2007 as JAMIA PrePrint; doi:10.1197/jamia.M2215
Journal of the American Medical Informatics Association 2007;14(5):651-661
© 2007 American Medical Informatics Association


A more recent version of this article appeared on September 1, 2007
This Article
Right arrow Full Text (PDF)
Right arrow Data Supplement
Right arrow All Versions of this Article:
M2215v1
14/5/651    most recent
Right arrow Submit a response
Right arrow Alert me when this article is cited
Right arrow Alert me when eLetters are posted
Right arrow Alert me if a correction is posted
Services
Right arrow Similar articles in this journal
Right arrow Similar articles in PubMed
Right arrow Alert me to new issues of the journal
Right arrow Download to citation manager
Right arrow reprints & permissions
Citing Articles
Right arrow Citing Articles via Google Scholar
Google Scholar
Right arrow Articles by Lin, Y.
Right arrow Articles by Liu, Y.
Right arrow Search for Related Content
PubMed
Right arrow PubMed Citation
Right arrow Articles by Lin, Y.
Right arrow Articles by Liu, Y.

Submitted on July 17, 2006
Accepted on May 20, 2007

A Document Clustering and Ranking System for Exploring MEDLINE Citations

Yongjing Lin MS1, Wenyuan Li PhD1, Keke Chen PhD2, and Ying Liu PhD3*

Affiliation of the authors: 1 Laboratory for Bioinformatics and Medical Informatics, University of Texas at Dallas, Richardson, TX; Department of Computer Science, University of Texas at Dallas, Richardson, TX ; 2 Yahoo!, Santa Clara, CA; 3 Laboratory for Bioinformatics and Medical Informatics, University of Texas at Dallas, Richardson, TX; Department of Computer Science, University of Texas at Dallas, Richardson, TX; Department of Molecular and Cell Biology, University of Texas at Dallas, Richardson, TX

* To whom correspondence should be addressed.

Objective A major problem faced in biomedical informatics involves how best to present information retrieval results. When a single query retrieves many results, simply showing them as a long list often provides poor overview. With a goal of presenting users with reduced sets of relevant citations, this study developed an approach that retrieved and organized MEDLINE citations into different topical groups and prioritized important citations in each group.

Design A text mining system framework for automatic document clustering and ranking organized MEDLINE citations following simple PubMed queries. The system grouped the retrieved citations, ranked the citations in each cluster, and generated a set of keywords and MeSH terms to describe the common theme of each cluster.

Measurements Several possible ranking functions were compared, including citation count per year (CCPY), citation count (CC), and journal impact factor (JIF). We evaluated this framework by identifying as "important" those articles selected by the Surgical Oncology Society.

Results Our results showed that CCPY outperforms CC and JIF, i.e., CCPY better ranked important articles than did the others. Furthermore, our text clustering and knowledge extraction strategy grouped the retrieval results into informative clusters as revealed by the keywords and MeSH terms extracted from the documents in each cluster.

Conclusions The text mining system studied effectively integrated text clustering, text summarization, and text ranking and organized MEDLINE retrieval results into different topical groups.







HOME HELP FEEDBACK SUBSCRIPTIONS ARCHIVE SEARCH
Copyright © 1994 by the American Medical Informatics Association.