【问题标题】:How to check that list elements exists in the dataframe?如何检查数据框中是否存在列表元素?
【发布时间】:2021-08-19 00:35:40
【问题描述】:

我的数据框中有一列包含不同长度的字符串列表,如下所示:

           names                                            venue
 
[Instagrammable, Restaurants, Vegan]                          14 Hills
[Date Night, Vibes, Drinks]                                   Upper 14
[Date Night, Drinks, After Work Drinks, Cocktail]             Hills
            .                                                   .                  
            .                                                   .
            .

现在,如果我想检查我的数据框中是否存在某些列表,该怎么做。

Example1:

Input :
        find_list=[Date Night, Vibes, Drinks]
        venue = 'Upper 14'
Output:
        Record is present in my dataframe

Example 2:

Input :
        find_list=[Date Night, Drinks]
        venue='Hills 123'
Output:
        Record is not present in my dataframe

例子

Input :
        find_list=[   Date Night, Vibes, Drinks]
        venue = 'Upper 14'
Output:
        Record is not present in my dataframe

【问题讨论】:

    标签: python pandas list dataframe


    【解决方案1】:

    您可以使用.apply() 和.any():

    find_list = ["Date Night", "Vibes", "Drinks"]
    
    if df["names"].apply(lambda x: x == find_list).any():
        print("List is present in my dataframe")
    else:
        print("List is not present in my dataframe")
    

    打印:

    List is present in my dataframe
    

    编辑:匹配记录:

    find_list = ["Date Night", "Vibes", "Drinks"]
    venue = "Upper 14"
    
    if df.apply(
        lambda x: x["names"] == find_list and x["venue"] == venue, axis=1
    ).any():
        print("Record is present in my dataframe")
    else:
        print("Record is not present in my dataframe")
    

    打印:

    Record is present in my dataframe
    

    编辑 2:从输入列表中去除空格:

    find_list = ["      Date Night", "Vibes", "Drinks"]
    venue = "Upper 14"
    
    if df.apply(
        lambda x: all(a.strip() == b.strip() for a, b in zip(x["names"], find_list))
        and x["venue"] == venue,
        axis=1,
    ).any():
        print("Record is present in my dataframe")
    else:
        print("Record is not present in my dataframe")
    

    打印:

    Record is present in my dataframe
    

    编辑 3:删除单词之间的多余空格:

    import re
    
    find_list = ["      Date     Night", "Vibes", "Drinks"]
    venue = "Upper 14"
    
    r = re.compile(r"\s{2,}")
    
    if df.apply(
        lambda x: all(
            r.sub(a.strip(), " ") == r.sub(b.strip(), " ")
            for a, b in zip(x["names"], find_list)
        )
        and x["venue"] == venue,
        axis=1,
    ).any():
        print("Record is present in my dataframe")
    else:
        print("Record is not present in my dataframe")
    

    【讨论】:

    • 我必须匹配上面问题中显示的完整记录/行,已编辑!
    • 如果我在列表中的任何值中留出一些空间,因为列表值是字符串,它不会给出正确的输出,因为我已经编辑了我的问题!
    • 嗯,这是有效的,但不是合适的答案,因为它会从单词前后删除多余的空格,如果我在单词之间给出空格,比如 Date Night ,它不会处理这个
    • @Haseeb 好吧,那么您需要在进行检查之前正确清理您的数据。
    • 你能帮我写出可以去除单词之间空格的代码吗,它应该只保留一个
    猜你喜欢
    • 2021-03-05
    • 2019-10-04
    • 1970-01-01
    • 2021-12-17
    • 2011-08-02
    • 2018-01-10
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多