【问题标题】:pandas dataframe read_csv, specify columns and keep whole line as a stringpandas dataframe read_csv,指定列并将整行保留为字符串
【发布时间】:2017-02-09 10:39:00
【问题描述】:

在 pandas read_csv 中,有没有办法指定例如。 col1、col15、整行?

我正在尝试从一个文本文件中导入大约 700000 行数据,该文本文件以“^”作为分隔符,没有文本限定符和回车作为行分隔符。

从文本文件中,我需要第 1 列、第 15 列,然后是表格/数据框的三列中的整行。

我已经搜索了如何在 pandas 中执行此操作,但不太了解它以获取逻辑。我可以很好地导入所有 26 列,但这对我的问题没有帮助。

my_df = pd.read_csv("tablefile.txt", sep="^", lineterminator="\r",  low_memory=False)

或者我可以使用标准 python 将数据放入表中,但这需要大约 4 小时才能处理 700000 行。这对我来说太长了。

count_1 = 0
for line in open('tablefile.txt'):
    if count_1 > 70:
        break
    else:
        col1id = re.findall('^(\d+)\^', line)
        col15id = re.findall('^.*\^.*\^(\d+)\^.*\^.*\^.*\^.*\^.*\^.*\^.*\^.*\^.*\^.*\^.*', line)
        line = line.strip()

        count_1 = count_1 + 1

        cur.execute('''INSERT INTO mytable (mycol1id, mycol15id, wholeline) VALUES (?, ?, ?)''', 
        (col1id[0], col15id[0], line, ) )

        conn.commit()
    print('row count_1=',count_1)

在 pandas read_csv 中,有没有办法指定例如。 col1、col15、整行?

如上,col1 和 col15 是数字,wholeline 是字符串

  • 我不想在导入后重建字符串,因为我可能会在此过程中丢失一些字符。

谢谢

编辑: 为每一行提交数据库非常耗时。

【问题讨论】:

  • 仅使用 python 时,您应该在循环外编译一次正则表达式。这必须加快速度
  • 我不明白这是如何工作的,我认为 re.findall(regex, object) 需要在调用 re.findall 之前创建对象。你有例子吗?

标签: python pandas import


【解决方案1】:

使用一些准分隔符将整行作为一个 df 读取(在下面使用 &),然后使用 usecols 再次读取并指定第 1 列和第 15 列的索引并将它们加在一起。

my_df_full = pd.read_csv("tablefile.txt", sep="&", lineterminator="\r", low_memory=False)
my_df_full.columns = ['full_line']

my_df_cols = pd.read_csv("tablefile.txt", sep="^", lineterminator="\r", low_memory=False, usecols=[1,15])

my_df_full[['col1', 'col15']] = my_df_cols

【讨论】:

  • 事实证明很难找到文本中没有的分隔符,但我会继续寻找。
  • @CArnold 如果你没有找到任何分隔符,你可以连接所有的列,有点乏味但应该可以。看到这个:stackoverflow.com/questions/19377969/…。我不确定它是否有效,但您也可以尝试删除 low_memory = False 并使用字符串作为分隔符。 sep="c_arnold_pandas"
【解决方案2】:

首先,您可以编译您的正则表达式以避免为每一行解析它们

import re

reCol1id = re.compile('^(\d+)\^')
reCol15id = re.compile('^.*\^.*\^(\d+)\^.*\^.*\^.*\^.*\^.*\^.*\^.*\^.*\^.*\^.*\^.*')

count_1 = 0
for line in open('tablefile.txt'):
    if count_1 > 70:
        break
    else:
        col1id = reCol1id.findall(line)[0]
        col15id = reCol15id.findall(line)[0]
        line = line.strip()

        count_1 += 1

        cur.execute('''INSERT INTO mytable (mycol1id, mycol15id, wholeline) VALUES (?, ?, ?)''', 
        (col1id, col15id, line, ) )

        conn.commit()
    print('row count_1=',count_1)

【讨论】:

    【解决方案3】:

    我把conn.commit() 放在for 循环的外面。它将加载时间减少到几分钟,尽管我猜它不太安全。

    无论如何感谢您的帮助。

    【讨论】:

      猜你喜欢
      • 2020-03-22
      • 2017-01-13
      • 1970-01-01
      • 2016-07-17
      • 2018-06-22
      • 2019-04-28
      • 2020-11-20
      • 2022-12-29
      • 1970-01-01
      相关资源
      最近更新 更多