【问题标题】:Getting list of address by flattening row in Dataframe pandas with a delimiter通过使用分隔符在 Dataframe pandas 中展平行来获取地址列表
【发布时间】:2021-07-28 21:09:01
【问题描述】:

我从一个 excel 文件创建了一个数据框。该表是一个联系人列表,其中一行包含位置名称、下一行街道地址、下一个城市,并为每个位置重复。现在每个位置都由一个特定的短语“站点代码”分隔。下面是我看到的一个例子。

Name and address Don't Care
Location Name x
Street addresss x
City x
State x
Site code x
Location Name x
Street addresss x
City x
State x
Site code x

我希望能够将站点代码之间的所有内容扁平化为一行,这样就可以了

[location name, street address, city, state]

我正在考虑创建一个函数来查看表格并创建一个字典,该字典的键将是位置名称,它会附加所有内容,直到到达站点代码然后跳过并移动到下一个字典条目。但也感觉我想多了。你怎么看?

【问题讨论】:

    标签: python pandas


    【解决方案1】:

    您可以使用numpy.array_split 在包含“站点代码”的行上拆分 DataFrame,然后根据需要解析生成的块:

    import numpy as np
    chunks = np.array_split(df["Name and address"], df[df["Name and address"]=="Site code"].index)
    output = [chunk[chunk!="Site code"].tolist() for chunk in chunks if chunk.tolist()!=["Site code"]]
    
    >>> output
    [['Location Name', 'Street addresss', 'City', 'State'],
     ['Location Name', 'Street addresss', 'City', 'State']]
    

    【讨论】:

    • 我无法正确拆分阵列。我有一个索引数组,其中“站点代码”为真,但由于某种原因,它并没有在元素处或根本没有完全拆分它。
    【解决方案2】:

    如果你总是有相同数量的字段,你可以使用numpy来reshape数据:

    pd.DataFrame(df['Name and address'].values.reshape((-1,5))[:, :4]).apply(list, axis=1)
    

    这是如何工作的:

    (pd.DataFrame(df['Name and address']  # get relevant column
                    .values               # access the underlying numpy array
                    .reshape((-1,5))      # reshape: new row every 5 elements
                    [:, :4]               # drop the 5th column ("Site code")
                 ).apply(list, axis=1)    # make as list again
    )
    

    输出:

    0    [Location Name , Street addresss , City , State ]
    1    [Location Name , Street addresss , City , State ]
    dtype: object
    

    【讨论】:

    • 不幸的是,我不能让我拥有相同数量的字段。谢谢!
    猜你喜欢
    • 2016-09-10
    • 2019-11-05
    • 1970-01-01
    • 2021-05-07
    • 1970-01-01
    • 2018-09-02
    • 1970-01-01
    • 2018-09-24
    • 2020-09-26
    相关资源
    最近更新 更多