標題: Power-law relationship and self-similarity in the itemset support distribution: analysis and applications
作者: Chuang, Kun-Ta
Huang, Jiun-Long
Chen, Ming-Syan
資訊工程學系
Department of Computer Science
公開日期: 1-八月-2008
摘要: In this paper, we identify and explore that the power-law relationship and the self-similar phenomenon appear in the itemset support distribution. The itemset support distribution refers to the distribution of the count of itemsets versus their supports. Exploring the characteristics of these natural phenomena is useful to many applications such as providing the direction of tuning the performance of the frequent-itemset mining. However, due to the explosive number of itemsets, it is prohibitively expensive to retrieve lots of itemsets before we identify the characteristics of the itemset support distribution in targeted data. As such, we also propose a valid and cost-effective algorithm, called algorithm PPL, to extract characteristics of the itemset support distribution. Furthermore, to fully explore the advantages of our discovery, we also propose novel mechanisms with the help of PPL to solve two important problems: (1) determining a subtle parameter for mining approximate frequent itemsets over data streams; and (2) determining the sufficient sample size for mining frequent patterns. As validated in our experimental results, PPL can efficiently and precisely identify the characteristics of the itemset support distribution in various real data. In addition, empirical studies also demonstrate that our mechanisms for those two challenging problems are in orders of magnitude better than previous works, showing the prominent advantage of PPL to be an important pre-processing means for mining applications.
URI: http://dx.doi.org/10.1007/s00778-007-0054-1
http://hdl.handle.net/11536/8529
ISSN: 1066-8888
DOI: 10.1007/s00778-007-0054-1
期刊: VLDB JOURNAL
Volume: 17
Issue: 5
起始頁: 1121
結束頁: 1141
顯示於類別:期刊論文


文件中的檔案:

  1. 000257396700008.pdf

若為 zip 檔案,請下載檔案解壓縮後,用瀏覽器開啟資料夾中的 index.html 瀏覽全文。