【问题标题】:Find if value is equal to or between values in 2D array查找值是否等于或介于二维数组中的值之间
【发布时间】:2019-04-18 15:59:41
【问题描述】:

我有一个 python 脚本,它捕获日志数据并将其转换为二维数组。

脚本的下一部分旨在遍历 .csv 文件并评估每一行的第一列,并确定该值是否等于或介于 2D 数组中的值之间。如果是,则将最后一列标记为 TRUE。如果不是,则将其标记为 FALSE。

例如,如果我的二维数组如下所示:

[[1542053213, 1542053300], [1542055000, 1542060105]]

我的 csv 文件如下所示:

1542053220, Foo, Foo, Foo
1542060110, Foo, Foo, Foo

第一行的最后一列应为 TRUE(或 1),而第二行的最后一列应为 FALSE(或 0)。

我当前的代码如下所示:

from os.path import expanduser
import re
import csv
import codecs

#Setting variables
#Specifically, set the file path to the reveal log
filepath = expanduser('~/LogAutomation/programlog.txt')
csv_filepath = expanduser('~/LogAutomation/values.csv')
tempStart = ''
tempEnd = ''

print("Starting Script")

#open the log
with open(filepath) as myFile:
    #read the log
    all_logs = myFile.read()
myFile.close()

#Create regular expressions
starting_regex = re.compile(r'\[(\d+)\s+s\]\s+Starting\s+Program')
ending_regex = re.compile(r'\[(\d+)\s+s\]\s+Ending\s+Program\.\s+Stopping')

#Create arrays of start and end times
start_times = list(map(int, starting_regex.findall(all_logs)))
end_times = list(map(int, ending_regex.findall(all_logs)))

#Create 2d Array
timeArray = list(map(list, zip(start_times, end_times)))

#Print 2d Array
print(timeArray)

print("Completed timeArray construction")

#prints the csv file
with open(csv_filepath, 'rb') as csvfile:
    reader = csv.reader(codecs.iterdecode(csvfile, 'utf-8'))

    for row in reader:
        currVal = row[0]
            #if currVal is equal to or in one of the units in timeArray, mark last column as true
            #else, mark last column as false

csvfile.close()

print("Script completed")

我已经成功地遍历了我的 .csv 文件并获取了每一行的第一列的值,但我不知道如何进行比较。不幸的是,关于值之间的签入,我不熟悉二维数组数据结构。此外,我的 .csv 文件中的列数可能会波动,因此是否有人知道一种非静态方法来确定“最后一列”以便能够在文件中写入该列之后的列?

有人可以帮我吗?

【问题讨论】:

  • 我不明白预期的输出。您想将 TRUE/FALSE 作为二维数组中行的最后一个元素吗?将 TRUE/FALSE 添加到 2D 数组中的行?或添加到 csv 行?在那种情况下,您将结果保存在哪里?同一个文件?
  • 对不起,你不明白。编写的目标是,如果第一列的值等于或介于二维数组中的值之间,则将 .csv 文件的最后一列(在示例代码中的 values.csv)写入 TRUE。如果不是,则在最后一列写 FALSE。

标签: python arrays csv multidimensional-array


【解决方案1】:

您只需要遍历列表列表并检查该值是否在任何间隔内。这是一个简单的方法:

with open(csv_filepath, 'rb') as csvfile:
    reader = csv.reader(codecs.iterdecode(csvfile, 'utf-8'))
    input_rows = [row for row in reader]

with open(csv_filepath, 'w') as outputfile:
    writer = csv.writer(outputfile)

    for row in input_rows:
        currVal = int(row[0])
        ok = 'FALSE'

        for interval in timeArray:
            if interval[0] <= curVal <= interval[1]:
                ok = 'TRUE'
                break

        writer.writerow(row + [ok])

上面的代码会将结果写入同一个文件,所以要小心。我还删除了csvfile.close(),因为如果您使用with 语句,该文件将自动为您关闭。

【讨论】:

  • 嗨@Marco Bonelli,非常感谢您的回复。请给我一点时间(也许 30 分钟)来集成和测试我的代码,我会尽快回复您。你能解释一下interval in TimeArray 行吗?
  • @JerryM。当然:就像你做for row in reader 一样,你可以使用for ... in ... 语法来迭代任何列表(或一般的可迭代对象)。您的timeArray 是一个列表列表,因此您可以使用for interval in timeArray 对其进行迭代。一旦进入for,interval 变量将保存正在处理的当前元素的值,因此在您的情况下,它将首先是[1542053213, 1542053300],然后是[1542055000, 1542060105]。
  • 嗨,马可。运行此代码最终会删除我的 .csv 文件中的值。我需要写入我打开的同一个 csv 文件。希望这会有所帮助。到目前为止,我感谢您的回复。
  • @JerryM。我假设您希望将结果保存在另一个文件中。我更改了代码,现在更改直接在同一个文件上进行。看看这是否适合你。另外,“将最后一列标记为 TRUE 或 FALSE”是指要替换最后一列中的值还是要在最后一列之后添加一个新列?
  • 我想加个专栏,没说清楚很抱歉。
【解决方案2】:

我会选择更 Python 的东西。

compare = lambda x, y, t: (x <= int(t) <= y)
with open('output.csv', 'w') as outputfile:
    writer = csv.writer(outputfile)
    with open(csv_filepath, 'rb') as csvfile:
        reader = csv.reader(codecs.iterdecode(csvfile, 'utf-8'))

    for row in reader:
        currVal = row[0]
        #if currVal is equal to or in one of the units in timeArray, mark last column as true
        #else, mark last column as false
        match = any(compare(x, y, currVal) for x, y in timeArray)
        write.writerow(row + ['TRUE' if match else 'FALSE'])

    csvfile.close()
outputfile.close()

【讨论】:

  • 用您的解决方案修复了一些缩进后,我现在在line 79 上收到SyntaxError: unexpected EOF while parsing
  • missing )m sorry 在 python 上很难用缩进复制
  • 没问题@Tzomas,我知道这很奇怪。我有个问题。是否可以同时打开、读取和编辑同一个文件?理想情况下,我不想创建单独的输出文件。我想我错误地编辑了你的代码。你愿意聊天吗?
  • ` 用于阅读器中的行:文件“/usr/lib64/python3.4/codecs.py”,第 1034 行,在迭代器中用于迭代器中的输入:ValueError:对已关闭文件的 I/O 操作。 ` 是错误。
  • 我从不使用编解码器不得不说,我认为您可以正常读取文件。你不能在打开的文件上写。只需修改代码先读取(如果文件不是那么大,你应该没有问题)然后关闭文件,一旦它关闭,再次打开它但这次写在上面,你会覆盖内容所以在意。先用副本测试
猜你喜欢
  • 2018-07-27
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2015-05-25
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2022-11-29
相关资源
最近更新 更多