【问题标题】:Trying to create a predetermined answer from conditions relating to 2 other columns within a dataframe尝试根据与数据框中其他 2 列相关的条件创建预定答案
【发布时间】:2019-12-10 14:00:16
【问题描述】:

我正在尝试创建一个数据框。

df = pd.DataFrame(columns=["Year", "Fuel", "Status", "Sex", "Service", "Expected"])

其他列包含使用np.random 创建的数据。

在“预期”列中,我想根据几个条件输入通过或失败。如果里程少于 100000 并且如果服务是,则它会通过,否则它会失败。

这是我目前所拥有的

df["Expected"]  = df.loc[(df['Mileage']< 100000) | (df['Service'] == 'Yes', "Pass", "Fail")]

它正在显示错误消息

ValueError: operands could not be broadcast together with shapes (500,) (3,) 

我已经用 500 行数据填充了其他列。但我不确定这 3 与什么有关。可能是 Yes、Pass、Fail 值。

我还尝试了df['Expected'] = np.where(df ["Mileage"] &lt; 132352, ['Service'] == "Yes",'Pass','Fail') 哪种方法有效。

我是不是走错了路?

任何帮助或指点将不胜感激。

【问题讨论】:

    标签: python pandas dataframe jupyter-notebook


    【解决方案1】:

    我将创建一个函数,该函数将pd.Series 对象作为唯一参数,然后返回该单元格的值。然后使用pd.apply(lambda row: your_function(row), axis=1)。所以:

    def your_function(row):
        if row["Mileage"] <132352 and row["Service"] == "Yes" :# fill in your other conditions here
            return "Pass"
        else:
            return "Fail"
    
    df["Expected"] = df.apply(lambda row: your_function(row), axis=1)
    

    【讨论】:

    • 在最后一行 pd.apply 创建了“AttributeError: module 'pandas' has no attribute 'apply”' 所以我将其更改为 df.apply 然后创建“ValueError: ('真值的系列不明确。使用 a.empty、a.bool()、a.item()、a.any() 或 a.all()。'、'发生在索引 1')"。然后我使用了def fun(row): if row["Mileage"] &lt;100000 and df['Service'].any(): return "Pass" else: return "Fail" df["Expected"] = df.apply(lambda row: fun(row), axis=1) 谢谢。
    • 是的,对不起。它应该是df.apply。我对您使用 .any() 所做的事情感到困惑。你能再解释一下吗?
    • 这样做有意义吗df["service"] = "Yes"
    • 如果我只使用def fun(row): if row["Mileage"] &lt;100000 and df['Service']:,我会收到错误消息 ValueError: ('The truth value of a Series is ambiguous. Use a.empty, a.bool(), a.item(), a.any() 或 a.all().', '发生在索引 2')。所以我输入 .any(): 在我查找 stackoverflow.com/questions/53830081/… 并遇到你正在比较两个 pd.Series 或一个 pd.Series 与一个值之后,所以你可能有多个 True 和多个 False 值,你必须改为:if (data == ask_minute['lastUpdated']).any()
    • 完美运行。感谢您为此付出的所有时间。
    【解决方案2】:

    您可以简单地用'Fail' 填充Expected 列:

    df['Expected'] = 'Fail'
    

    然后:

    df.at[df[(df['Mileage']<100000) & (df['Service'] == 'Yes')].index,'Expected'] = 'Pass'
    

    【讨论】:

    • 每个单元格返回失败。
    • 您确定数据框中有满足要求的行吗?这应该可以正常工作。
    • 是的。我正在寻找一个有 43000 和 Yes 的服务,但它仍然失败。
    • 我已经设法在我的代码中找到错误。它现在应该可以工作了。
    • 完美运行。感谢您抽出宝贵时间完成此操作。
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2022-12-08
    • 2017-01-03
    • 1970-01-01
    • 2021-02-02
    • 1970-01-01
    相关资源
    最近更新 更多