【问题标题】:Is there any way to perform faster this operation? [duplicate]有什么方法可以更快地执行此操作? [复制]
【发布时间】:2020-07-01 00:06:41
【问题描述】:

我正在处理一个大型医疗数据集,我想查看是否有一些行对应于同一位患者。与患者 ID 对应的列是 ID Patient。
我想创建一个新列,如果该患者出现在多行中,则为 "yes",如果仅出现一次,则为 "no"。
这是我做的代码:

df['Repeated'] = 'No' # New Column

for i in range(0,len(df)):
    for f in range(0,len(df)):
        if df['ID NP'].iloc[i] == df['ID NP'].iloc[f]:
            df['ID NP'].iloc[i] = 'Yes'
        else:
            df['ID NP'].iloc[i] = 'No'

但是此操作花费了太多时间。有什么方法可以更快吗?

【问题讨论】:

标签: python pandas loops for-loop replace


【解决方案1】:

这不仅花费了太多时间,而且还将所有行中的所有ilocs 都设置为“是”。

您只需要计算可以找到多少次不同的 iloc。数一数:

stats = {}
for i in df['ID NP'].iloc:
    stats[i] = stats.get(i, 0) + 1

现在你只需要遍历 iloc 索引:

for i in range(0, len(df['ID NP'].iloc)):
    id_np = df['ID NP'].iloc[i]
    if stats[id_np] > 1:
        # then there should be 'Yes' in this row

【讨论】:

    猜你喜欢
    • 2016-09-02
    • 1970-01-01
    • 2022-11-24
    • 2013-02-22
    • 2020-03-06
    • 2022-11-18
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多