Improve Decision Tree for Probability-Based Ranking by Lazy Learners

From National Research Council Canada

Author	Search for: Liang, H.; Search for: Yan, Y.
Format	Text, Article
Conference	The 18th IEEE International Conference on Tools with Artificial Intelligence (ICTAI06), November 13-15, 2006, Washington, DC
Abstract	Existing work shows that classic decision trees have inherent deficiencies in obtaining a good probability-based ranking (e.g. AUC). This paper aims to improve the ranking performance under decision-tree paradigms by presenting two new models. The intuition behind our work is that probability-based ranking is a relative metric among samples, therefore, distinct probabilities are crucial for accurate ranking. The first model, Lazy Distance-based Tree (LDTree), uses a lazy learner at each leaf to explicitly distinguish the different contributions of leaf samples when estimating the probabilities for an unlabeled sample. The second model, Eager Distance-based Tree (EDTree), improves LDTree by changing it into an eager algorithm. In both models, each unlabeled sample is assigned a set of unique probabilities of class membership instead of a set of uniformed ones, which gives finer resolution to differentiate samples and leads to the improvement of ranking. On 34 UCI sample sets, experiments verify that our models greatly outperform C4.5, C4.4 and other standard smoothing methods designed for better ranking.
Publication date	2006
In	The 18th IEEE International Conference on Tools with Artificial Intelligence (ICTAI06) [Proceedings].
Language	English
NRC number	NRCC 48784
NPARC number	5765136
Export citation	Export as RIS
Report a correction	Report a correction (opens in a new tab)
Record identifier	7756da3d-b54f-471f-b744-b9710ffe79d7
Record created	2009-03-29
Record modified	2020-10-09

Date modified:: 2025-02-05