【问题标题】:Compare two large files and combine matching information比较两个大文件并合并匹配信息
【发布时间】:2017-10-25 10:15:37
【问题描述】:

我有两个相当大的文件,JSON(185,000 行)和 CSV(650,000)。我需要遍历 JSON 文件中的每个 dict,然后在其中遍历 part_numbers 中的每个部分,并比较它以获取 CSV 中该部分所在位置的前三个字母。

由于某种原因,我很难正确地做到这一点。我的脚本的第一个版本太慢了,所以我正在尝试加快速度

JSON 示例:

[
    {"category": "Dryer Parts", "part_numbers": ["ABC", "DEF", "GHI", "JKL", "MNO", "PQR"], "parent_category": "Dryers"},
    {"category": "Washer Parts", "part_numbers": ["ABC", "DEF", "GHI", "JKL", "MNO", "PQR"], "parent_category": "Washers"},
    {"category": "Sink Parts", "part_numbers": ["ABC", "DEF", "GHI", "JKL", "MNO", "PQR"], "parent_category": "Sinks"},
    {"category": "Other Parts", "part_numbers": ["ABC", "DEF", "GHI", "JKL", "MNO", "PQR"], "parent_category": "Others"}
]

CSV:

WCI|ABC
WPL|DEF
BSH|GHI
WCI|JKL

结束的字典如下所示:

{"category": "Other Parts",
 "part_numbers": ["WCIABC","WPLDEF","BSHGHI","JKLWCI"...]}

这是我到目前为止所做的一个示例,它在if (part.rstrip() == row[1]): 返回IndexError: list index out of range:

import csv
import json
from multiprocessing import Pool

def find_part(item):
    data = {
        'parent_category': item['parent_category'],
        'category': item['category'],
        'part_numbers': []
    }

    for part in item['part_numbers']:
        for row in reader:
            if (part.rstrip() == row[1]):
                data['part_numbers'].append(row[0] + row[1])

    with open('output.json', 'a') as outfile:
        outfile.write('    ')
        json.dump(data, outfile)
        outfile.write(',\n')


if __name__ == '__main__':
    catparts = json.load(open('catparts.json', 'r'))
    partfile = open('partfile.csv', 'r')
    reader = csv.reader(partfile, delimiter='|')


    with open('output.json', 'w+') as outfile:
        outfile.write('[\n')

    p = Pool(50)
    p.map(find_part, catparts)

    with open('output.json', 'a') as outfile:
        outfile.write('\n]')

【问题讨论】:

  • 那时row 是什么?你之前处理了多少行?此外,您能否发布 minimal 代码(根据发布指南),或者该文件是否读取了问题的重要部分?
  • 文件读取并不是真正的问题。更多的是关于如何有效地做到这一点。你还需要什么代码?
  • 没有更多代码... 更少。但是,我想我在无法运行代码的情况下发现了它。
  • 当我使用示例输入文件运行您的代码时,我在 'parent_category': item['parent_category'], 函数中的 'parent_category': item['parent_category'], 行上得到一个 KeyError: 'parent_category' — 所以我无法重现“不正确的索引错误”问题提到。
  • @martineau 对不起!我正在使用一个忘记包含在 json 示例中的密钥。它不见了parent_category

标签: python json python-3.x csv multiprocessing


【解决方案1】:

正如我在评论中所说,您的代码(现在)给了我一个 NameError: name 'reader' 未在 find_part() 函数中定义。解决方法是将csv.reader 的创建移到函数中。我还更改了文件的打开方式,以使用with 上下文管理器和newline 参数。这也解决了一堆单独的任务都试图同时读取同一个 csv 文件的问题。

您的方法效率非常低,因为它为item['part_numbers'] 中的每个部分读取整个'partfile.csv' 文件。不过,以下似乎可行:

import csv
import json
from multiprocessing import Pool

def find_part(item):
    data = {
        'parent_category': item['parent_category'],
        'category': item['category'],
        'part_numbers': []
    }

    for part in item['part_numbers']:
        with open('partfile.csv', newline='') as partfile:  # open csv in Py 3.x
            for row in csv.reader(partfile, delimiter='|'):
                if part.rstrip() == row[1]:
                    data['part_numbers'].append(row[0] + row[1])

    with open('output.json', 'a') as outfile:
        outfile.write('    ')
        json.dump(data, outfile)
        outfile.write(',\n')

if __name__ == '__main__':
    catparts = json.load(open('carparts.json', 'r'))

    with open('output.json', 'w+') as outfile:
        outfile.write('[\n')

    p = Pool(50)
    p.map(find_part, catparts)

    with open('output.json', 'a') as outfile:
        outfile.write(']')

这是一个更高效的版本,每个子进程仅读取整个 'partfile.csv' 文件一次:

import csv
import json
from multiprocessing import Pool

def find_part(item):
    data = {
        'parent_category': item['parent_category'],
        'category': item['category'],
        'part_numbers': []
    }

    with open('partfile.csv', newline='') as partfile:  # open csv for reading in Py 3.x
        partlist = [row for row in csv.reader(partfile, delimiter='|')]

    for part in item['part_numbers']:
        part = part.rstrip()
        for row in partlist:
            if row[1] == part:
                data['part_numbers'].append(row[0] + row[1])

    with open('output.json', 'a') as outfile:
        outfile.write('    ')
        json.dump(data, outfile)
        outfile.write(',\n')

if __name__ == '__main__':
    catparts = json.load(open('carparts.json', 'r'))

    with open('output.json', 'w+') as outfile:
        outfile.write('[\n')

    p = Pool(50)
    p.map(find_part, catparts)

    with open('output.json', 'a') as outfile:
        outfile.write(']')

虽然您可以在主任务中将'partfile.csv' 数据读入内存并将其作为参数传递给find_part() 子任务,但这样做只是意味着必须对每个进程的数据进行腌制和解封。您需要运行一些计时测试来确定这是否比使用csv 模块显式读取它更快,如上所示。

还要注意,在将任务提交到Pool 之前,预处理来自'carparts.json' 文件的数据加载并从每行的第一个元素中去除尾随空格也会更有效,因为这样您就不需要一遍又一遍地在find_part() 中执行part = part.rstrip()。同样,我不知道这样做是否值得努力——只有时间测试才能确定答案。

【讨论】:

  • 这个答案是否会导致多个进程尝试打开文件时出现问题?
  • @RyanScottCady:不,它不会引起问题,因为进程都只是试图读取文件。 partfile.csv 很大吗?
  • partfile.csv 是 650k 行。所以它不是超级大。我想知道将其内容读入字典是否存在问题哈哈。我尝试了你的解决方案,它正在工作!
  • Ryan:按照今天的标准,这不是很大,因此您可以将其读入内存并将其传递给每个进程。但是,这样做会引入一些开销——需要对每个进程进行腌制和取消腌制,这可能不会快得多。我正在为此寻找解决方案,并将相应地更新我的答案——如果证明可行的话……
【解决方案2】:

我想我找到了。您的 CSV 阅读器与许多其他文件访问方法一样:您按顺序读取文件,然后点击 EOF。当您尝试对第二部分执行相同操作时,该文件已经在 EOF,并且第一次 read 尝试返回 null 结果;这没有第二个元素。

如果您想再次访问所有记录,您需要重置文件书签。最简单的方法是使用

回溯到字节 0
partfile.seek(0)

另一种方法是关闭并重新打开文件。

这会让你感动吗?

【讨论】:

  • 感谢您的意见!我应该在哪里将此 sn-p 添加到我的代码中?
  • 当你想回到文件开头时插入它。这在您的控制和数据流中处于什么位置?
【解决方案3】:

只要所有部件号都存在于 csv 中,这应该可以工作。

import json

# read part codes into a dictionary
with open('partfile.csv') as fp:
    partcodes = {}
    for line in fp:
        code, number = line.strip().split('|')
        partcodes[number] = code

with open('catparts.json') as fp:
    catparts = json.load(fp)

# modify the part numbers/codes 
for cat in catparts:
    cat['part_numbers'] = [partcodes[n] + n for n in cat['part_numbers']]

# output
with open('output.json', 'w') as fp:
    json.dump(catparts, fp)

【讨论】:

  • 感谢您的回答。使用with open() 与仅使用类似:file = open('file', 'r') 之间有什么区别?
  • 另外,我收到您的代码错误:ValueError: too many values to unpack (expected 2) 我猜这是因为文件太大了。 CSV 有 650,000 行,JSON 文件有 185,000
  • 这可能是由 partfile.csv 中的某些行引起的,其中包含多个 |。这只是示例代码,所以我对一致的输入数据做了一些假设。在生产代码中,您必须处理格式错误的输入。
猜你喜欢
  • 1970-01-01
  • 1970-01-01
  • 2014-02-21
  • 1970-01-01
  • 2018-08-05
  • 1970-01-01
  • 1970-01-01
  • 2015-12-05
  • 2017-07-30
相关资源
最近更新 更多