【问题标题】:Why CausalNex output in python is wrong?为什么 python 中的 CausalNex 输出是错误的?
【发布时间】:2023-02-11 14:44:42
【问题描述】:

我在 python 中使用 causalnex 从 python 中的数据集创建 DAG。

我得到了图表,节点是正确的,但边缘完全不对。我在具有四个随机自变量(请求者、风险、规模、开发人员)和一个依赖变量(持续时间)的数据框 df 中进行了尝试,生成的图形是这样的: DAG using CausalNex

我是否错误地使用了图书馆?为什么这个数字与真实的数据生成过程相距如此之远?贝叶斯网络模型能否胜过 causalnex?

我试过这段代码:

from causalnex.structure.notears import from_pandas
import matplotlib.pyplot as plt
import networkx as nx

sm = from_pandas(df)
sm.remove_edges_below_threshold(0.8)
nx.draw_shell(sm, with_labels=True, font_weight ="bold")
plt.show()

我期待这样的事情:Expected Output

【问题讨论】:

  • 请将数据框数据添加到问题中。
  • 要重现数据集: import dumpy as np import pandas as pd np.random.seed(42) fib_list = [0, 1, 1, 2, 3, 5, 8, 13, 21, 34, 55, 89] data = {“请求者”:np.random.randint(1,4,100),“大小”:np.random.randint(1,4,100),“风险”:np.random.randint(1,4,100)} df = pd. DataFrame(data) df['Developer'] = np.random.choice(fib_list, df.shape[0]) df["Duration"] = (0.1*df["Requestor"] + 0.2*df["Size" ] + 0.2*df["风险"] + 0.5*df["开发人员"])

标签: python


【解决方案1】:

我会说变量之间的关系不容易捕获(特别是由于 Developer 的域大小)。连续“Duration”的父类的域大小为4*4*4*12 ... duration 本身并不是真正连续的,但可以取 102 个不同的值 ...

因此,在学习算法期间,大小为 100 的数据库确实不足以使测试/分数准确。

请注意,我将 Duration 乘以 10 以保留整数值。

仅供参考,推断是最后一个 BN

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 2022-10-06
    • 2015-04-12
    • 1970-01-01
    • 2020-09-24
    • 1970-01-01
    • 2020-09-03
    • 1970-01-01
    • 2015-10-30
    相关资源
    最近更新 更多