【问题标题】:combine pandas text rows based on condition根据条件组合熊猫文本行
【发布时间】:2020-04-05 05:13:29
【问题描述】:

我有这种df:

df = pd.DataFrame({"text_column" : ['question: everybody is kongfu fighting', 'panda: of course',  'question: Why is the world so great ?', 'friend: Everybody is smart', 'and everybody is cool', 'enemy: no that is just not true', 'jordan: i want to add one thing: please', 'do not talk about this.', ' 2nd question : are you sure ?', 'yeah sure' ]})

                                text_column
0   question: everybody is kongfu fighting
1   panda: of course
2   question: Why is the world so great ?
3   friend: Everybody is smart
4   and everybody is cool
5   enemy: no that is just not true
6   jordan: i want to add one thing: please
7   do not talk about this.
8    2nd question : are you sure ?
9   messi: yeah sure
10  question: you are sure about this ?
11  donald: youre questions are stupid!

我想要以下输出

                 type_column                                     new_text_column
0  question: panda:                                        everybody is kongfu fighting of course

1  question: friend: enemy: jordan: 2nd question : messi:  Why is the world so great ? Everybody is smart and everybody is cool no that is just not true i want to add one thing: please do not talk about this. are you sure ? yeah sure
2  question: donald:                                       youre questions are stupid!

基本上每个问题和答案(主题)都必须在一个单元格中。 我可以编写一个有效但会使用 apply 的函数,这通常不是最佳解决方案。 有人知道怎么做吗?

【问题讨论】:

  • 提示:第一步是查看df.text_column.str.extract('^(.*: )?(.*)$')和groupby。

标签: python-3.x string pandas text


【解决方案1】:

定义以下函数:

  1. “专门”将源文本字段拆分为 2 部分:

    def mySplit(txt):
        tbl = re.split(': ?', txt, 1)
        if len(tbl) == 1:
            tbl.insert(0, '')
        return pd.Series(tbl, index=['Qn', 'Ans'])
    
  2. 重新格式化一组行:

    def reformat(grp):
        t1 = ': '.join(grp.Qn.tolist()) + ':'
        t2 = ' '.join(grp.Ans.tolist())
        return pd.Series([t1, t2], index=['type_column', 'new_text_column'])
    

然后,得到结果运行:

df.text_column.apply(mySplit)\
    .groupby(df2.Qn.str.startswith('question').cumsum())\
    .apply(reformat).reset_index(drop=True)

它执行:

  • 将 text_column 专门拆分为 2 列(Qn 和 Ans)。
  • 从以 Qn 开始的每一行开始,以 question 开始分成几组。
  • 对每个组应用重新格式化。
  • 重置索引(丢弃旧索引)。

【讨论】:

  • 谢谢; df2.Qn.str.startswith('question').cumsum() 这是我缺少的部分
【解决方案2】:

从示例中很难看出分离的标准是什么。

我猜它在冒号上分裂,所以你可以尝试列表理解

df["type_column"] = [x.split(":")[0] for x in df["text_column"]]
df["new_text_column"] = [x.split(":")[1] for x in df["text_column"]]

【讨论】:

    猜你喜欢
    • 2019-11-03
    • 2016-12-25
    • 2022-01-25
    • 1970-01-01
    • 2020-11-11
    • 1970-01-01
    • 2021-08-27
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多