【问题标题】:Python merge two csv files on multiple columns and nearest datetimePython在多列和最近的日期时间上合并两个csv文件
【发布时间】:2016-02-12 15:44:23
【问题描述】:

我有两个要合并的 csv 文件。

文件1:

rel_id, acc_id, value, timestamp
1, 2, True, 2016-01-04 19:20:22
2, 3, True, 2016-01-04 18:35:56
1, 2, True, 2016-01-04 20:43:12
1, 5, False, 2016-01-04 18:15:20
2, 3, True, 2016-01-04 20:43:11

文件2:

rel_id, acc_id, value, timestamp
1, 2, 250, 2016-01-04 20:43:13
1, 5, 610, 2016-01-04 18:15:23
2, 3, 400, 2016-01-04 18:35:58
2, 3, 300, 2016-01-04 20:43:13
1, 2, 500, 2016-01-04 19:20:23

我想根据 rel_id、acc_id 和时间戳合并这两个文件。

合并(文件 1 和文件 2):

rel_id, acc_id, value_file1, timestamp, value_file2
1, 2, True, 2016-01-04 19:20:22, 500
2, 3, True, 2016-01-04 18:35:56, 400
1, 2, True, 2016-01-04 20:43:12, 250
1, 5, False, 2016-01-04 18:15:20, 610
2, 3, True, 2016-01-04 20:43:11, 300

不过 file2 的时间戳稍晚一些。

在 stackoverflow 上搜索将我带到这篇文章:pandas merge dataframes by closest time

但我不知道如何在 rel_id、acc_id 和最近的时间戳上进行匹配。

import pandas as pd


file1 = pd.read_csv('file1.csv')
file2 = pd.read_csv('file2.csv')


file1.columns = ['rel_id', 'acc_id', 'value', 'timestamp']
file2.columns = ['rel_id', 'acc_id', 'value', 'timestamp']


file1['timestamp'] = pd.to_datetime(file1['timestamp'])
file2['timestamp'] = pd.to_datetime(file2['timestamp'])


file1_dt = pd.Series(file1["timestamp"].values, file1["timestamp"])
file1_dt.reindex(file2["timestamp"], method="nearest")
file2["nearest"] = file1_dt.reindex(file2["timestamp"],    method="nearest").values

print file2

我根据另一篇文章尝试了上面的代码,但这与 rel_id 和 acc_id 不匹配。加上上面的代码已经引发了一个错误:

ValueError: index 必须单调递增或递减

非常感谢任何帮助。谢谢。

【问题讨论】:

  • 这不是总是选择提前的文件吗?没有多大意义,还是我误解了什么?

标签: python csv pandas merge


【解决方案1】:

您正在尝试根据未排序的索引重新索引。 假设您的 CSV 没有标题:

column_names = ['rel_id', 'acc_id', 'value', 'timestamp']
file1 = pd.read_csv('file1.csv',
                    index_col=['timestamp'],
                    parse_dates='timestamp',
                    header=None,
                    names=column_names).sort_index()
file2 = pd.read_csv('file2.csv',
                    index_col=['timestamp'],
                    parse_dates='timestamp',
                    header=None,
                    names=column_names).sort_index()
file1.set_index(file1.reindex(file2.index, method='nearest').index, inplace=True)



                     rel_id  acc_id  value
timestamp
2016-01-04 18:15:23       1       5  False
2016-01-04 18:35:58       2       3   True
2016-01-04 19:20:23       1       2   True
2016-01-04 20:43:13       2       3   True
2016-01-04 20:43:13       1       2   True

并合并file1和file2:

file1.reset_index().merge(file2.reset_index(), on=['acc_id', 'rel_id', 'timestamp']).set_index('timestamp')

【讨论】:

  • 感谢您的回答和示例。我觉得很愚蠢,但是如何获得合并的输出(因此显示两个值列并在 file1 或 file2 的时间戳上合并)?使用pd.merged(file1, file2, on='timestamp')时会报错。
  • 可能是更方便的方法,但类似于:file1.reset_index().merge(file2.reset_index(), on=['acc_id', 'rel_id', 'timestamp']).set_index('timestamp')。特别是因为现在您将timestamp 设置为索引,然后重新设置并再次设置...
  • Valtuarte:感谢您的帮助。我似乎还不能让它以正确的方式工作。我通过添加所需的输出和更改 file2 数据来编辑我的问题,以表明它并不总是按时间顺序排列。在您的评论中使用建议的代码时,文件 2 的时间戳在 2016-01-04 20:43:13 时出错。 rel_id = 1 和 acc_id = 2 时显示此时间戳 2 次。
  • 感谢 valtuarte。它确实适用于示例数据,因此我检查了您的答案。但不幸的是,我无法让它与我的实际数据一起使用。
  • 将索引设置为时间戳将不起作用,除非时间戳是唯一的。由于目标是合并多个列,因此每个 rel_id 的时间戳可能是唯一的,但在数据帧中可能不是唯一的。
猜你喜欢
  • 2016-07-28
  • 2020-06-30
  • 2019-05-27
  • 1970-01-01
  • 1970-01-01
  • 2016-09-04
  • 1970-01-01
  • 2017-06-27
  • 1970-01-01
相关资源
最近更新 更多