【问题标题】:pandas dataframe index remove date from datetime熊猫数据框索引从日期时间中删除日期
【发布时间】:2021-03-19 21:12:00
【问题描述】:

文件dataexample_df.txt:

2020-12-04_163024 26.15 26.37 19.40 24.57
2020-12-04_163026 26.15 26.37 19.20 24.57
2020-12-04_163028 26.05 26.37 18.78 24.57

我想将其作为 pandas 数据框读取,其中索引列只有'%H:%M:%S' 格式的时间部分,没有日期。

import pandas as pd
df = pd.read_csv("dataexample_df.txt", sep=' ', header=None, index_col=0)
print(df)

输出:

                       1      2      3      4
0
2020-12-04_163024  26.15  26.37  19.40  24.57
2020-12-04_163026  26.15  26.37  19.20  24.57
2020-12-04_163028  26.05  26.37  18.78  24.57

但是,想要的输出:

              1      2      3      4
0
16:30:24  26.15  26.37  19.40  24.57
16:30:26  26.15  26.37  19.20  24.57
16:30:28  26.05  26.37  18.78  24.57

我尝试了不同的date_parser= -functions(参见Parse_dates in Pandas 中的答案) 但只收到错误消息。另外,Python/Pandas convert string to time only 有点相关,但运气不好,我被卡住了。我正在使用 Python 3.7。

【问题讨论】:

    标签: python pandas dataframe


    【解决方案1】:

    考虑到您的df 是这样的:

    In [121]: df
    Out[121]: 
                           1      2      3      4
    0                                            
    2020-12-04_163024  26.15  26.37  19.40  24.57
    2020-12-04_163026  26.15  26.37  19.20  24.57
    2020-12-04_163028  26.05  26.37  18.78  24.57
    

    您可以将Series.replace 与Series.dt.time 一起使用:

    In [122]: df.reset_index(inplace=True)
    In [127]: df[0] = pd.to_datetime(df[0].str.replace('_', ' ')).dt.time
    
    In [130]: df.set_index(0, inplace=True)
    
    In [131]: df
    Out[131]: 
                  1      2      3      4
    0                                   
    16:30:24  26.15  26.37  19.40  24.57
    16:30:26  26.15  26.37  19.20  24.57
    16:30:28  26.05  26.37  18.78  24.57
    

    【讨论】:

    • 谢谢,但我收到了来自df['0'] 的错误消息。如果我使用df[0]和df.set_index(0, inplace=True),它确实有效
    • 好的。我从这篇文章中复制粘贴了你的数据框,所以对我来说['0'] 有效。但是解决方案的整个概念保持不变。如果答案有帮助,请不要忘记点赞并接受答案。
    【解决方案2】:

    在这里,我创建了一个简单的函数来格式化你的日期时间列,请试试这个。

    import pandas as pd
    
    df = pd.read_csv('data.txt', sep=" ", header=None)
    
    def format_time(date_str):
        # split date and time
        time =  iter(date_str.split('_')[1])
        # retun the time value adding
        return ':'.join(a+b for a,b in zip(time, time))
    
    df[0] = df[0].apply(format_time)
    
    print(df)
    

    【讨论】:

    • 谢谢,这是如何使用apply 方法的好例子,我对此一无所知。但这仍然需要最后再进行一次操作:df.set_index(0, inplace=True) 就像上面的@Mayank 回答一样。
    • @user9393931 还要补充一点,apply 基本上是引擎盖下的循环。因此,与我在答案中使用的功能相比,它要慢一些。
    • @user9393931 同意你的看法。
    【解决方案3】:

    您需要使用format 参数告诉它您的日期格式是什么(否则您会收到错误消息):

    # gives an error:
    pd.to_datetime('2020-12-04_163024')
    
    # works:
    pd.to_datetime('2020-12-04_163024', format=r'%Y-%m-%d_%H%M%S')
    

    因此,您可以将其应用于您的数据框,然后使用 dt.time 访问时间:

    df['time'] = pd.to_datetime(df.index, format=r'%Y-%m-%d_%H%M%S').dt.time
    

    这会给你时间作为一个对象,但如果你想格式化它,只需使用这样的东西:

    df['time'] = df['time'].strftime('%H:%M:%S')
    

    【讨论】:

    • 我从df['time'] = pd.to_datetime(df.index, format=r'%Y-%m-%d_%H%M%S').dt.time 行收到错误AttributeError: 'DatetimeIndex' object has no attribute 'dt'?
    猜你喜欢
    • 2023-01-12
    • 1970-01-01
    • 2017-02-09
    • 2018-02-02
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2017-02-08
    • 2021-05-03
    相关资源
    最近更新 更多