%0 Journal Article %T RoadRunner for Heterogeneous Web Pages Using Extended MinHash %A A Suresh Babu %A P. Premchand %A A. Govardhan %J International Journal of Database Management Systems %D 2012 %I Academy & Industry Research Collaboration Center (AIRCC) %X The Internet presents large amount of useful information which is usually formatted for its users, which makes it hard to extract relevant data from diverse sources. Therefore, there is a significant need of robust, flexible Information Extraction (IE) systems that transform the web pages into program friendly structures such as a relational database will become essential. IE produces structured data ready for post processing. Roadrunner will be used to extract information from template web pages. In this paper, we present novel algorithm for extracting templates from a large number of web documents which are generated from heterogeneous templates. The proposed system focuses on information extraction from heterogeneous web pages. We cluster the web documents based on the common template structures so that the template for each cluster is extracted simultaneously. The resultant clusters will be given as input to the Roadrunner system. %K Information Extraction %K Clustering %K Minimum Description Length Principle %K MinHash %U http://airccse.org/journal/ijdms/papers/4112ijdms03.pdf