【问题标题】:Remove '\n' in text in pandas python删除熊猫python中文本中的'\ n'
【发布时间】:2019-02-14 16:12:36
【问题描述】:

以下代码是我用来删除 ['text'] 列中的 \n 的当前代码:

df = pd.read_csv('file1.csv')

df['text'].replace('\s+', ' ', regex=True, inplace=True) # remove extra whitespace
df['text'].replace('\n',' ', regex=True) # remove \n in text

header = ["text", "word_length", "author"]

df_out = df.to_csv('sn_file1.csv', columns = header, sep=',', encoding='utf-8')

我也从建议中尝试过:

df['text'].replace('\n', '')
df['text'] = df['text'].str.replace('\n', '').str.replace('\s+', ' ').str.strip()

输出:'真是个聪明人! \n好像他也对房地产交易一无所知......'

删除空格的代码正在运行。但不是删除\n。任何人都可以在这件事上帮助我吗?谢谢。

我也尝试根据此链接的建议解决问题removing newlines from messy strings in pandas dataframe cells?,但它仍然无法正常工作。

已解决:

df['text'].replace(r'\s+|\\n', ' ', regex=True, inplace=True) 

【问题讨论】:

  • df['text'].replace('\n', '') 的工作情况如何?
  • @anky_91 我试过了,但还是一样。不过谢谢你的建议
  • \s 也匹配换行符,因此它应该可以工作,除非您的输入字符串包含实际的反斜杠,后跟文字 n 而不是换行符。
  • @Lily 您能否edit 提出您的问题,然后提供所提供的解决方案及其结果,以及它们与您的期望有何不同?此刻......“不,不工作”并没有帮助任何人看到一种可能的方法。谢谢。
  • 听起来好像根本没有换行符,而是\ + n。如果你使用df['text'].replace(r'\s+|\\n', ' ', regex=True, inplace=True),它会消失吗?

标签: python regex pandas string python-2.7


【解决方案1】:

考虑到要将更改应用于“文本”列,请选择该列

df['text']

然后,为了实现这一点,可以使用pandas.DataFrame.replace。

这样可以传递正则表达式regex=True,这会将两个列表中的两个字符串都解释为正则表达式(而不是直接匹配它们)。

接@Wiktor Stribiżew suggestion,下面会做的工作

df['text'] = df['text'].replace(r'\s+|\\n', ' ', regex=True) 

This正则表达式语法参考可能会有所帮助。

【讨论】:

    猜你喜欢
    • 2020-05-11
    • 1970-01-01
    • 2016-04-11
    • 2020-11-25
    • 2019-01-06
    • 1970-01-01
    • 2014-08-19
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多