【问题标题】:How C4.5 algorithm handles data with same attributes but different results?C4.5算法如何处理属性相同但结果不同的数据?
【发布时间】:2017-08-18 01:36:06
【问题描述】:

我正在尝试使用 C4.5 算法为学校项目创建决策树。决策树为Haberman's Survival Data Set,属性信息如下。

Attribute Information:

1. Age of patient at time of operation (numerical)
2. Patient's year of operation (year - 1900, numerical)
3. Number of positive axillary nodes detected (numerical)
4. Survival status (class attribute)
    1 = the patient survived 5 years or longer
    2 = the patient died within 5 year

我们需要实现一个决策树,其中每个叶子都必须有一个不同的结果(意味着该叶子的熵应该为 0),但是有六个实例具有相同的属性,但结果不同。

例如:

66,58,0,2
66,58,0,1

C4.5算法在这种情况下是做什么的,我到处搜索了,但没有找到任何信息。

谢谢。

【问题讨论】:

  • Yazlab başa bela dimi :)
  • @EmreKantar 哈哈,阿南。 :)

标签: algorithm decision-tree j48 c4.5


【解决方案1】:

阅读 Quinlan, J. R. C4.5:机器学习程序。 Morgan Kaufmann Publishers,1993 年。(如果你有大学作业,学习 C4.5 会很好)

从我所学的。好像在第 137 页,源代码列表 build.c
有一行
//* if all case are the same.... or there are not enough case to divide(喜欢你的问题)
它将return Node
此节点来自
Node = Leaf(ClassFreq, BestClass, Cases, Cases-NoBestClass);

ClassFreq 存储每个类的计数
BestClass 存储是 主导类(最频繁)案例存储有多少数据
NoBestClass 存放多少个BestClass的数据

这个叶子函数来自文件Trees.c 这个叶子函数将返回一个叶子为bestClass (Best class become the leaf) 的节点。

所有这些信息参考 Quinlan, J. R. C4.5:机器学习程序。摩根考夫曼出版社,1993 年。

任何知道这方面的人,如果我做错了什么,请发表评论。谢谢

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 2013-03-15
    • 1970-01-01
    • 2012-08-19
    • 1970-01-01
    • 2020-12-15
    • 2015-10-23
    • 1970-01-01
    • 2017-02-04
    相关资源
    最近更新 更多