期刊文献+

基于OEM模型的半结构化数据的模式抽取 被引量:8

OEM-based schema extraction of semi-structured data
原文传递
导出
摘要 Web数据是典型的半结构化数据 ,缺乏明确的、预知的、与数据分离存储的外在模式 ,导致查询、浏览和集成Web数据的效率极低。该文提出一种基于 OEM (objectexchange model)模型的半结构化数据的模式抽取算法 ,采用自顶向下的剪枝策略 ,可快速发现频繁简单路径集 ,应用于半结构化数据的集成及查询回答与优化。其特点是可降低目标模式的规模 。 Web data is typical semi-structured data without an explicit structure that characterises most data sets. The lack of data structure makes querying and integrating web data very inefficient. An approach was developed to identify structures in semi-structured and hierarchical data using the OEM (object exchange model) and a pruning strategy to quickly extract simple paths from the OEM graph for integrating and querying semi-structured data. The method can effectively reduce the scale of the target structure and enhance the efficiency of structure abstraction.
出处 《清华大学学报(自然科学版)》 EI CAS CSCD 北大核心 2004年第9期1264-1267,共4页 Journal of Tsinghua University(Science and Technology)
基金 国家"九七三"重点基础研究项目 ( G19980 3 0 414 ) 中国博士后科学基金 ( 2 0 0 3 0 3 414 7)
关键词 半结构化数据 模式抽取 对象交换模型 剪枝 semi-structured data schema extraction object exchange model (OEM) pruning
  • 相关文献

参考文献6

  • 1Nestorv S,Ullman J,Wiener J,et al.Representative Objects: Concise Representation of Semistructured,Hierarchical Data [EB/OL].http://citeseer.ist.psu.edu/313886.html,1997.
  • 2Papakonstantinou Y,Garcia-Molina H,Widom J.Object Exchange Across Heterogeneous Information source [EB/OL].http://citeseer.ist.psu.edu/ papakonstantinou95object.html,1995.
  • 3Quass D,Widom J,Goldman R,et al.Lore: A Lightweight Object Repository for Semistructured Data [EB/OL].http://www.informatik.uni-trier.de/~ley/db/conf/sigmod/ QuassWGHLMNRRAUW96.html,1996.
  • 4王宁,徐宏炳,王能斌.基于带根连通有向图的对象集成模型及代数[J].软件学报,1998,9(12):894-898. 被引量:25
  • 5刘芳,胡和平,路松峰.半结构化、层次数据的模式发现[J].小型微型计算机系统,2001,22(1):84-88. 被引量:11
  • 6刘芳,胡和平.半结构化数据的模式发现[J].微型电脑应用,2000,16(2):13-15. 被引量:9

二级参考文献12

  • 1[1]Peter Buneman. Semistructured data. [C]1n Proc. of PODS.Tucson,Arizona. 1997. 117~121
  • 2[2]Serge Abiteboul. Querying semi-structured data. [C]In proceed ings of ICDT, Delphi,Greece. Jan 1997. 1~18
  • 3[3]Svetlozar Nestorov, Jeffrey Ullman, Janet Wiener, et al. Representative objects: concise representations of semistructured, hierarchical data. [C]In Proc ICDE. 1997. 79~90
  • 4[4]Bayardo R. Efficiently mining long patterns from databases. [C]In Proc. of the 1998 ACM-SIGMOD Int' 1 Conf. on Management of Data. Washington USA In Proc. of ICDE, Birmingham UK 1998. 85~93
  • 5Wang N,Proceedings of the 1997 IEEE International Conference on Intelligent Processing Systems,1997年,1589页
  • 6王能斌,数据库系统,1995年,24页
  • 7S.Abiteboul.Querying semi-structured data.In Proc.ofICDT.Delphi,Greece,January,1997.1-18
  • 8Peter Buneman,Susan Davison,Gerd Hillebrand,et al.A query language and optimizationtechniques for unstructured data.In Proc.of ACM-SIGMOD InternationalConference,Montreal,Canada,June,1996,505-516
  • 9S.Abiteboul,D.Quass,J.McHugh,et al.The lorel language for semistructureddata.Technical report,Dept.of Computer Science,Stanford University,1996(http://www-db.stanford.edu/pub/papers/lore196.ps)
  • 10Ke Wang,Huiqing Liu.Schema discovery for semistructured data.In Proc.of the ThirdInternational Conference on Knowledge Discovery and Data Mining(KDD-97),NewportBeach,California,USA,August,1997.271-274

共引文献38

同被引文献59

引证文献8

二级引证文献22

相关作者

内容加载中请稍等...

相关机构

内容加载中请稍等...

相关主题

内容加载中请稍等...

浏览历史

内容加载中请稍等...
;
使用帮助 返回顶部