【问题标题】:Trying to find the nearest date before and after a specified date from a list of dates in a comma separated string within a pandas Dataframe试图从熊猫数据框中逗号分隔字符串中的日期列表中查找指定日期之前和之后的最近日期
【发布时间】:2021-02-07 01:21:09
【问题描述】:

tldr;我在dtype: datetime64[ns] <class 'pandas.core.series.Series'> 中有一个index_date 和一个list_of_dates 类型的<class 'list'>,其中单个元素采用str 格式。将这些转换为相同数据类型的最佳方法是什么,以便我可以将日期排序为index_date 之前和之后最接近的日期?

我有一个带有列的 pandas 数据框 (df):

ID_string                   object
indexdate           datetime64[ns]
XR_count                     int64
CT_count                     int64
studyid_concat              object
studydate_concat            object
modality_concat             object

它看起来像:

    ID_string   indexdate   XR_count    CT_count    studyid_concat      studydate_concat
0   55555555    2020-09-07  10          1           ['St1', 'St5'...]       ['06/22/2019', '09/20/2020'...]
1   66666666    2020-06-07  5           0           ['St11', 'St17'...]     ['05/22/2020', '06/24/2020'...]

studyid_concat ("St1") 中的 0 元素对应于 studydate_concat 和 modality_concat 等中的 0 元素。由于篇幅原因,我没有显示modality_concat,但它类似于['XR', 'CT', ...]

我目前的目标是找到在我的索引日期之前和之后进行的最接近的 X 射线研究,并且能够从最接近到最远对研究进行排名。我对熊猫有点陌生,但这是我目前的尝试:

df = pd.read_excel(path_to_excel, sheet_name='Sheet1')

# Convert comma separated string from Excel to lists of strings
df.studyid_concat = df.studyid_concat.str.split(',')
df.studydate_concat = df.studydate_concat.str.split(',')
df.modality_concat = df.modality_concat.str.split(',')

for x in in df['ID_string'].values:
    index_date = df.loc[df['ID_string'] == x, 'indexdate']

    # Had to use subscript [0] below because result of above was a list in an array
    studyid_list = df.loc[df['ID_string'] == x, 'studyid_concat'].values[0]
    date_list = df.loc[df['ID_string'] == x, 'studydate_concat'].values[0]
    modality_list = df.loc[df['ID_string'] == x, 'modality_concat'].values[0]

    xr_date_list = [date_list[x] for x in range(len(date_list)) if modality_list[x]=="XR"]
    xr_studyid_list = [studyid_list[x] for x in range(len(studyid_list)) if modality_list[x]=="XR"]

就我所知,因为我对这里的数据类型有些困惑。我的 indexdate 目前在 dtype: datetime64[ns] <class 'pandas.core.series.Series'> 中,我正在考虑使用 datetime 模块进行转换,但很难弄清楚如何进行转换。我也不确定是否需要。我的xr_study_list 是一个字符串列表,其中包含格式为“mm/dd/yyyy”的日期。我想如果我能以正确的格式获得数据类型,我是否能弄清楚其余的。我只是比较日期是否 >= 或 indexdate 以排序到之前/之后,然后将每个日期减去 indexdate 并排序。我认为无论我对xr_date_list 做什么,我都必须确保对xr_studyid_list 做同样的事情来跟踪唯一的学习ID

编辑:所需的输出数据框看起来像

    ID_string   indexdate   StudyIDBefore           StudyDateBefore     
0   55555555    2020-09-07  ['St33', 'St1', ...]    [2020-09-06, 2019-06-22, ...]
1   66666666    2020-06-07  ['St11', 'St2', ...]    [2020-05-22, 2020-05-01, ...]

“之前”变量将从最近到最远排序,并且存在类似的“之后”列。我目前的目标是检查在此索引日期之前和之后的 3 天内是否存在研究,但具有上述数据框如果我需要开始寻找最近的研究之外的东西,这会给我灵活性。

【问题讨论】:

  • 你想要的输出是什么?你能创建一个数据框来显示吗?
  • 刚刚更新了上面的例子。基本上想要之前和之后的列包含从索引日期最近到最远排序的列表,一列 ID 和一列日期。

标签: python pandas dataframe datetime datetime64


【解决方案1】:

在花了一些时间思考它并参考更多 pandas to_datetime 文档后,我想我找到了自己的答案。基本上意识到我可以使用 pd.to_datetime 转换我的字符串日期列表

date_list = pd.to_datetime(df.loc[df['ID_string'] == x, 'studydate_concat'].values[0]).values

然后可以从这个列表中减去我的索引日期。选择在临时数据框中执行此操作,以便我可以跟踪其他列值(如研究 ID、模态等)。

完整代码如下:

for x in df['ID_string'].values:
    index_date = df.loc[df['ID_string'] == x, 'indexdate'].values[0]
    date_list = pd.to_datetime(df.loc[df['ID_string'] == x, 'studydate_concat'].values[0]).values
    modality_list = df.loc[df['ID_string'] == x, 'modality_concat'].values[0]
    studyid_list = df.loc[df['ID_string'] == x, '_concat'].values[0]

    tempdata = list(zip(studyid_list, date_list, modality_list))
    tempdf = pd.DataFrame(tempdata, columns=['studyid', 'studydate', 'modality'])

    tempdf['indexdate'] = index_date
    tempdf['timedelta'] = tempdf['studydate']-tempdf['index_date']

    tempdf['study_done_wi_3daysbefore'] = np.where((tempdf['timedelta']>=np.timedelta64(-3,'D')) & (tempdf['timedelta']<np.timedelta64(0,'D')), True, False)
    tempdf['study_done_wi_3daysafter'] = np.where((tempdf['timedelta']<=np.timedelta64(3,'D')) & (tempdf['timedelta']>=np.timedelta64(0,'D')), True, False)
    tempdf['study_done_onindex'] = np.where(tempdf['timedelta']==np.timedelta64(0,'D'), True, False)

    XRonindex[x] = True if len(tempdf.loc[(tempdf['study_done_onindex']==True) & (tempdf['modality']=='XR'), 'studyid'])>0 else False
    XRwi3days[x] = True if len(tempdf.loc[(tempdf['study_done_wi_3daysbefore']==True) & (tempdf['modality']=='XR'), 'studyid'])>0 else False
    # can later map these values back to my original dataframe as a new column

【讨论】:

    猜你喜欢
    • 2019-04-28
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2016-01-13
    • 1970-01-01
    相关资源
    最近更新 更多