跳至主導覽 跳至搜尋 跳過主要內容

Finding event-relevant content from the web using a near-duplicate detection approach

研究成果: 書籍/報告/會議論文中的章節會議投稿同行評審

8 引文 斯高帕斯(Scopus)

摘要

In online resources, such as news and weblogs, authors often extract articles, embed content, and comment on existing articles related to a popular event. Therefore, it is useful if authors can check whether two or more articles share common parts for further analysis, such as cocitation analysis and search result improvement. If articles do have parts in common, we say the content of such articles is event-relevant. Conventional text classification methods classify a complete document into categories, but they cannot represent the semantics precisely or extract meaningful event-relevant content. To resolve these problems, we propose a near-duplicate detection approach for finding event-relevant content in Web documents. The efficiency of the approach and the proposed duplicate set generation algorithms make it suitable for identifying event-relevant content. The experiment results demonstrate the potential of the proposed approach for use in weblogs.

原文English
主出版物標題Proceedings of the IEEE/WIC/ACM International Conference on Web Intelligence, WI 2007
頁面291-294
頁數4
DOIs
出版狀態Published - 2007
事件IEEE/WIC/ACM International Conference on Web Intelligence, WI 2007 - Silicon Valley, CA, United States
持續時間: 2 11月 20075 11月 2007

出版系列

名字Proceedings of the IEEE/WIC/ACM International Conference on Web Intelligence, WI 2007

Conference

ConferenceIEEE/WIC/ACM International Conference on Web Intelligence, WI 2007
國家/地區United States
城市Silicon Valley, CA
期間2/11/075/11/07

指紋

深入研究「Finding event-relevant content from the web using a near-duplicate detection approach」主題。共同形成了獨特的指紋。

引用此