【问题标题】:Create document index of word positions创建单词位置的文档索引
【发布时间】:2017-02-06 13:57:25
【问题描述】:

问题:

我想通过在 python 中创建一个数据结构来执行索引,该结构将存储给定文本文件中的所有单词,还将存储其行号(这些单词出现的所有行)以及单词的位置(第 # 列)在该特定行中。

到目前为止,我可以通过将所有行号附加到列表中来将单词存储在字典中,但我无法将它们的位置存储在该特定行中。

我需要这个数据结构来更快地搜索文本文件。

这是我到目前为止的代码:

from collections import defaultdict
thetextfile = open('file.txt','r')
thetextfile = thetextfile.read()
file_s = thetextfile.split("\n")
wordlist = defaultdict(list)
lineNumber = 0
for (i,line) in enumerate(file_s):

    lineNumber = i
    for word in line.split(" "):
       wordlist[word].append(lineNumber)

print(wordlist)

【问题讨论】:

  • 你的文本文件是什么格式的?
  • @Leonid ,可以是任何格式。
  • @EdwinvanMierlo ,我是 python 新手,我不能很好地进行。
  • 编辑了我的问题,我想现在应该很清楚了。

标签: python python-3.x indexing


【解决方案1】:

以下是一些代码,用于存储文本文档中单词的行号和列:

from collections import defaultdict, namedtuple

# build a named tuple for the word locations
Location = namedtuple('Location', 'line col')

# dict keyd by word in document
word_locations = defaultdict(list)

# go through each line in the document
for line_num, line in enumerate(open('my_words.txt', 'r').readlines()):
    column = -1
    prev_col = 0

    # process the line, one word at a time
    while True:   
        if prev_col < column:
            word = line[prev_col:column]
            word_locations[word].append(Location(line_num, prev_col))
        prev_col = column+1

        # find the next space
        column = line.find(' ', prev_col)

        # check for more spaces on the line
        if column == -1:

            # there are no more spaces on the line, store the last word
            word = line[prev_col:column]
            word_locations[word].append(Location(line_num, prev_col))

            # go onto the next line
            break

print(word_locations)

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2011-12-22
    • 2017-03-30
    • 2015-08-24
    相关资源
    最近更新 更多