【问题标题】:Read CSV file with Pandas: Regex delimiter使用 Pandas 读取 CSV 文件:正则表达式分隔符
【发布时间】:2022-10-14 22:00:48
【问题描述】:

我在尝试为 read_csv 分隔符找到正确的正则表达式时遇到问题。 我最初的 txt 数据看起来像这样。

t = '''
[21.01.22, 07:32:11] text1
text2
[21.01.22, 07:34:18] text3
[21.01.22, 07:32:51] text4
text5
'''

我需要用换行符和方括号表达式分隔行,以便所需的结果看起来像这样

column 1 | column2
[21.01.22, 07:32:11] | text1 text2
[21.01.22, 07:34:18] | text3
[21.01.22, 07:32:51] | text4 text5

我目前正在努力解决的问题是某些行包含没有方括号的字符串。方括号内的文本始终具有相同的格式:[dd.mm.yy, hh:mm:ss]

你能帮我找到分隔符参数的正确正则表达式吗?

data = pd.read_csv('t.txt', delimiter=r"\[(..................)\]", header=None, engine="python")

【问题讨论】:

  • 您可以更新示例以添加不带方括号的行吗?你总是只有两列吗?

标签: python pandas regex


【解决方案1】:

试试(regex101):

import re
import pandas as pd

t = """
[21.01.22, 07:32:11] text1
text2
[21.01.22, 07:34:18] text3
[21.01.22, 07:32:51] text4
text5
"""

df = pd.DataFrame(
    re.findall(r"^([[^]]+])(.*?)(?=^[|Z)", t, flags=re.S | re.M),
    columns=["Column1", "Column2"],
)
df["Column2"] = df["Column2"].str.replace("
", " ").str.strip()
print(df)

印刷:

                Column1      Column2
0  [21.01.22, 07:32:11]  text1 text2
1  [21.01.22, 07:34:18]        text3
2  [21.01.22, 07:32:51]  text4 text5

【讨论】:

  • 显然不是所有的行都有方括号,所以这不起作用(等待一个例子......)
  • @Andrej Kesely 感谢您的解决方案!事实上,它看起来已经非常接近我想要的了。唯一的问题是我需要将 txt 文件转换为 pandas 数据框,而不是我的示例中的字符串。您能否详细说明,我如何在 pd.read_csv 语句中使用相同的逻辑(我假设在分隔符参数中)?
  • @mozway 也感谢您的回复。在我的初始示例中,没有括号的行表示为 text2 & text5
  • 我明白了,那么这应该可以工作,我认为它会更复杂;)
  • 使用with open('your_file.csv') as f: df = pd.DataFrame(re.findall(..., f.read(), ...)...)
【解决方案2】:

可能不优雅,但似乎有效

# readin the file
lines=''
with open("c:csv2.txt") as fi:  
    line=fi.read()
    lines += line

#replace newline with space, so that we have a single string
lines=re.sub(r'(
)+',' ', lines)

# add few delimiters to help split up the lines at set locations
# workaround: add | delimiter before [
lines=re.sub(r'( [)+','|[', lines)

#workaround: add ; delimiter after ]
lines=re.sub(r'(] )+','];', lines)

# create a dataframe by splitting on | delimiter
df1=pd.DataFrame(lines.split('|'))

# split again on ; delimiter and create new columns
df1[['column1','columns2']]= df1[0].str.split(";", expand=True) 

# drop the originally read-in column
df1.drop(columns=[0], inplace=True)
df1

    column1                 columns2
0   [21.01.22, 07:32:11]    text1 text2
1   [21.01.22, 07:34:18]    text3
2   [21.01.22, 07:32:51]    text4 text5

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 2017-05-17
    • 2015-06-30
    • 2020-09-06
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2019-07-16
    相关资源
    最近更新 更多