Abstract
Recommender systems predict users' preferences over a large number of items by pooling similar information from other users and/or items in the presence of sparse observations. Despite successes of the existing approaches, challenges arising from recommender systems remain persistently unsolved. First, how to utilize user-item specific covariates and networks describing user-item interactions in a high-dimensional situation, for accurate recommendation. Second, most state-of-the-art recommender systems focus on ordinal or continuous ratings, which are not appropriate and effective for describing binary preferences in terms of performance and evaluation criteria, and it is more useful to provide a short list of top-preferred items as oppose to a complete list of items with estimated preference scores. Thus, fully and effectively utilizing the covariate information, in a framework of the top-list recommendation are attractive in devising recommender system.In the first part of the thesis, we propose a smooth neighborhood recommender in the framework of the latent factor models. A similarity kernel is utilized to borrow neighborhood information from continuous covariates over a user-item specific network, such as a user's social network, where the grouping information defined by discrete covariates is also integrated through the network. Consequently, user-item specific information is built into the recommender to battle the "cold-start" issue in the absence of observations in collaborative and content-based filtering. Moreover, we utilize a "divide-and-conquer" version of the alternating least squares algorithm to achieve scalable computation, and establish asymptotic results for the proposed method, demonstrating that it achieves superior prediction accuracy. Finally, the proposed method is illustrated in simulated examples and real benchmark data--"Last.fm" music.
In the second part of the thesis, a new collaborative recommender ranking system is developed to predict most-preferred items for each user given search queries based on binary response. In particular, we propose a ψ-ranker based on ranking functions incorporating information on users, items, and search queries through latent factor models. Moreover, we show that the proposed nonconvex surrogate pairwise ψ-loss performs well under four popular bipartite ranking losses, such as the sum loss, pairwise zero-one loss, discounted cumulative gain, and mean average precision. We develop a parallel computing strategy to optimize the intractable loss of two levels of nonconvex components through difference of convex programming and block successive upper-bound minimization. Theoretically, we establish a probabilistic error bound for the ψ-ranker and show that its ranking error has a sharp rate of convergence in the general framework of bipartite ranking, even when the dimension of the model parameters diverges with the sample size. This result also indicates that the ψ-ranker unifies the error rate of the pairwise ranking and scoring approaches in the bipartite ranking. Finally, we illustrate our methods using data from the Mobike big data challenge, consisting of three-million bicycle sharing records.
| Date of Award | 3 Jul 2019 |
|---|---|
| Original language | English |
| Awarding Institution |
|
| Supervisor | Junhui WANG (Supervisor) |
Cite this
- Standard