【问题标题】:splitting column and extracting country, cities and organisation names拆分列并提取国家、城市和组织名称
【发布时间】:2021-02-03 12:07:43
【问题描述】:

我有一个地址列如下的数据框。我想拆分此列,以便将国家、城市和机构分成不同的列。具有挑战性的部分是每个细胞都有不同的结构。所有这些单元格的共同点是它们以城市、国家/地区结尾,但在某些情况下,例如行索引 3,有多个条目。

    id      address
------------------------------------------------------------------------------------------------
0   223     Department of GI and HPB Surgery, University Hospital Northern Norway, Breivika, Tromsø, Norway; Institute of Clinical Medicine, University of Tromsø, Tromsø, Norway
1   223     Department of Surgery, University Hospital Maastricht, Maastricht, The Netherlands; NUTRIM School for Nutrition, Toxicology and Metabolism, Maastricht University, Maastricht, The Netherlands
2   223     Department of Surgery, University Hospital Maastricht, Maastricht, The Netherlands; NUTRIM School for Nutrition, Toxicology and Metabolism, Maastricht University, Maastricht, The Netherlands
3   223     Department of Surgery, Närebro University Hospital, Närebro; Department of Molecular Medicine and Surgery, Karolinska Institutet, Stockholm, Sweden'}, {'id': '9900', 'name': 'Närebro universitet, Institutionen för läkarutbildning
4   223     Clinical Surgery, University of Edinburgh, Royal Infirmary of Edinburgh, Edinburgh, UK
5   223     Division of Gastrointestinal Surgery, Nottingham Digestive Diseases Centre, National Institute for Health Research, Biomedical Research Unit, Nottingham University Hospitals, Queen's Medical Centre, Nottingham, UK
6   223     Hospital of Lausanne (CHUV), Lausanne, Switzerland
7   223     Department of GI and HPB Surgery, University Hospital Northern Norway, Breivika, Tromsø, Norway; Institute of Clinical Medicine, University of Tromsø, Tromsø, Norway
8   223     Clinical Surgery, University of Edinburgh, Royal Infirmary of Edinburgh, Edinburgh, UK
9   223     Department of GI and HPB Surgery, University Hospital Northern Norway, Breivika, Tromsø, Norway; Institute of Clinical Medicine, University of Tromsø, Tromsø, Norway

有人可以帮忙吗?

注意上面的数据框是我的数据框的一个子集,这就是为什么 id 列对于所有行都具有相同的值。原始数据框有大约 10k 行,这就是为什么不能在这里分享。

【问题讨论】:

  • 您可以创建一个包含所有国家/地区的列表,另一个包含城市的列表,然后,您可以使用正则表达式来提取正确的字符串。 String 的其余部分将是机构。
  • 你的逻辑似乎是合理的。你能提供你的代码sn-p吗?
  • 这里可以使用named entity recognition

标签: pandas dataframe split


【解决方案1】:

这对于您的 10K 行数据库来说可能过于简单,但希望您能找到正确的方向。

请注意,行索引 3 格式不正确,因为它有大括号等 - 看起来像是创建/抓取数据时的解析问题。在下面这被忽略了,实际上你想清理你的输入或修复上游的问题。

首先,我根据您的数据创建一个玩具数据集:

import pandas as pd
from io import StringIO
raw_data = StringIO(
"""
!Id!address
0!223!Department of GI and HPB Surgery, University Hospital Northern Norway, Breivika, Tromsø, Norway; Institute of Clinical Medicine, University of Tromsø, Tromsø, Norway
1!223!Department of Surgery, University Hospital Maastricht, Maastricht, The Netherlands; NUTRIM School for Nutrition, Toxicology and Metabolism, Maastricht University, Maastricht, The Netherlands
2!223!Department of Surgery, University Hospital Maastricht, Maastricht, The Netherlands; NUTRIM School for Nutrition, Toxicology and Metabolism, Maastricht University, Maastricht, The Netherlands
3!223!Department of Surgery, Närebro University Hospital, Närebro; Department of Molecular Medicine and Surgery, Karolinska Institutet, Stockholm, Sweden'}, {'id': '9900', 'name': 'Närebro universitet, Institutionen för läkarutbildning
4!223!Clinical Surgery, University of Edinburgh, Royal Infirmary of Edinburgh, Edinburgh, UK
5!223!Division of Gastrointestinal Surgery, Nottingham Digestive Diseases Centre, National Institute for Health Research, Biomedical Research Unit, Nottingham University Hospitals, Queen's Medical Centre, Nottingham, UK
6!223!Hospital of Lausanne (CHUV), Lausanne, Switzerland
7!223!Department of GI and HPB Surgery, University Hospital Northern Norway, Breivika, Tromsø, Norway; Institute of Clinical Medicine, University of Tromsø, Tromsø, Norway
8!223!Clinical Surgery, University of Edinburgh, Royal Infirmary of Edinburgh, Edinburgh, UK
9!223!Department of GI and HPB Surgery, University Hospital Northern Norway, Breivika, Tromsø, Norway; Institute of Clinical Medicine, University of Tromsø, Tromsø, Norway
""")
data = pd.read_csv(raw_data, index_col=0, delimiter='!')

接下来,对于具有多个地址的行,我将它们拆分并放在数据框中的单独行上。我假设它们用';'分隔总是,就像你的例子一样

data['address'] = data['address'].str.split(';')
data = data.explode('address')

接下来,我通过用“,”分割地址来标记地址。这里address_tokens 列将包含此之后的标记列表

data['address_tokens'] = data['address'].str.split(',')

现在,对于每一行,我们将标记组合成一个包含三个元素的列表,其中包含 [tokens[0:N-3] 用逗号连接在一起,token[N-2],token[N-1]],我们识别这些作为机构,城市,国家

data['address_3'] = data['address_tokens'].apply(lambda tks: [','.join(tks[:-3]), tks[-2], tks[-1]] )
data[['institution', 'city', 'country']] = data['address_3'].apply(pd.Series)

我将所有中间步骤保留在数据框中,以便您查看结果。 ['institution', 'city', 'country'] 三列包含您所要求的内容,除了来自原始行索引 3 的 {,} 等问题

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 2022-11-11
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2015-04-25
    • 1970-01-01
    相关资源
    最近更新 更多