Electronic Thesis/Dissertation
 

Discover Aggregate Sudden-Change over The Deep Web

Open Access

Nowadays, many web databases “hidden” behind their restrictive search interfaces (e.g., Amazon, eBay) contain rich and valuable information that is of significant interests to various third parties. To enable analytics over such databases, prior work has introduced techniques to produce unbiased estimations of SUM and COUNT aggregates. The problem with such technique is that it assumes the database does not change over time.Another type of web databases is spatial databases that provide Location Based Service (LBS), i.e., standalone like Google Maps, and embedded ones like “users near me” in WeChat. Mining this type of databases is very challenging since the only access interface they offer is a limited k-Nearest-Neighbor (kNN) search interface with query rate per user or IP address per day limited too.In this dissertation, we develop two web-based systems for the two types of web databases (dynamic hidden web databases and LBS databases) that (1) reveals and tracks (the changes of) user- specified aggregate queries over such hidden web databases, especially those that are frequently updated, by issuing a small number of search queries through the public web interfaces of these databases, and (2) enables fast analytics over an LBS by issuing a small number of queries through its restricted kNN interface.In addition, we develop a novel technique for tracking and discovering exceptions (detection of sudden changes) of aggregate queries, e.g., AVG, SUM, over dynamic hidden web databases, while still adhering to the stringent query-count limitations enforced by many hidden web databases providers. We provide both theoretical analysis and extensive experiment over both real-world and synthetic datasets to demonstrate the superiority of our solution over baseline solutions. We also develop a web-based system that demonstrates the effectiveness of our proposed algorithm over real-world databases.Lastly, we consider a novel problem of enabling density based clustering over the back-end database of an LBS using nothing but limited access to the kNN interface provided by the LBS. As we only have limited access to the backend database, we aim to mine from the LBS a cluster assignment function f (·), such that for any tuple t in the database (which may or may not have been accessed), f(·) can produce the cluster assignment of t with high accuracy. We provide comprehensive set of experiments over benchmark datasets and popular real-world LBS such as Yahoo! Flickr, Zillow, Redfin and Google Maps and demonstrate the effectiveness of our proposed techniques.

Author Language Date created Type of Work License
  • All rights reserved
Rights statement GW Unit Degree Advisor Persistent URL

Notice to Authors

If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.

Thumbnail Title Date Uploaded Visibility Actions
Preview of BinSuhaim_gwu_0075A_13685.pdf BinSuhaim_gwu_0075A_13685.pdf 2018-01-16 Open Access