【问题标题】:Apply Chi-Square to dataset which contains categorical variables将卡方应用于包含分类变量的数据集
【发布时间】:2022-01-23 14:17:24
【问题描述】:

我的数据集包含以下列:

Voted? Political Category
Yes            Right
No             Left
Not Answered   Center
Yes            Right
Yes            Right
No             Right

我需要计算卡方来查看哪个类别与投票的人最相关。两列都包含字符串。如何为每个值赋予数字表示以应用卡方?

【问题讨论】:

标签: python pandas numpy scipy


【解决方案1】:

您可以使用pd.factorize 对分类变量进行编码:

df['nVoted?'] = pd.factorize(df['Voted?'])[0]
df['nCategory'] = pd.factorize(df['Political Category'])[0]
print(df)

# Output
         Voted? Political Category  nVoted?  nCategory
0           Yes              Right        0          0
1            No               Left        1          1
2  Not Answered             Center        2          2
3           Yes              Right        0          0
4           Yes              Right        0          0
5            No              Right        1          0

之后你可以使用scipy.stats.chisquare

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 2018-10-07
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2018-01-01
    相关资源
    最近更新 更多