【问题标题】:ValueError: cannot convert float NaN to integerValueError:无法将浮点 NaN 转换为整数
【发布时间】:2019-12-04 16:08:09
【问题描述】:

我正在编写一个函数,它返回一个字典,其中文档的年份作为键,作为值,它指定了一个由 def do_get_citations_per_year 函数返回的元组。

这个函数处理df:

def do_process_citation_data(f_path):
    global my_ocan

    my_ocan = pd.read_csv(f_path, names=['oci', 'citing', 'cited', 'creation', 'timespan', 'journal_sc', 'author_sc'],
                          parse_dates=['creation', 'timespan'])
    my_ocan = my_ocan.iloc[1:]  # to remove the first row
    my_ocan['creation'] = pd.to_datetime(my_ocan['creation'], format="%Y-%m-%d", yearfirst=True)
    my_ocan['timespan'] = my_ocan['timespan'].map(parse_timespan)
    #print(my_ocan.info())
    print(my_ocan['timespan'])
    return my_ocan

然后我就有了这个功能,运行时不会触发任何错误:

    result = tuple()
    my_ocan['creation'] = pd.DatetimeIndex(my_ocan['creation']).year

    len_citations = len(my_ocan.loc[my_ocan["creation"] == year, "creation"])
    timespan = round(my_ocan.loc[my_ocan["creation"] == year, "timespan"].mean())
    result = (len_citations, timespan)
    print(result)


    return result

当我在另一个函数中运行该函数时:

def do_get_citations_all_years(data):
    mydict = {}
    s = set(my_ocan.creation)
    for year in s:
        mydict[year] = do_get_citations_per_year(data, year)

    return mydict

我得到错误:

  File "/Users/lisa/Desktop/yopy/execution_example.py", line 28, in <module>
    print(my_ocan.get_citations_all_years())
  File "/Users/lisa/Desktop/yopy/ocan.py", line 35, in get_citations_all_years
    return do_get_citations_all_years(self.data)
  File "/Users/lisa/Desktop/yopy/lisa.py", line 112, in do_get_citations_all_years
    mydict[year] = do_get_citations_per_year(data, year)
  File "/Users/lisa/Desktop/yopy/lisa.py", line 99, in do_get_citations_per_year
    timespan = round(my_ocan.loc[my_ocan["creation"] == year, "timespan"].mean())
ValueError: cannot convert float NaN to integer

我可以做些什么来解决这个问题?

提前谢谢你

【问题讨论】:

    标签: python-3.x pandas


    【解决方案1】:

    这个错误意味着my_ocan.loc[my_ocan["creation"] == year, "timespan"].mean()是NaN。

    您应该在计算平均值之前用0 填充NaN 值,因为它不会改变平均值。这是一个例子:

    timespan = my_ocan.loc[my_ocan["creation"] == year, "timespan"].fillna(0).mean()
    

    【讨论】:

    • * 是 nan 不是 None。
    • @Guimute 谢谢。我的错误:)
    • 您也可以通过调用dropna() apparently 完全删除NaN 值。这看起来更明确。 :)
    • 嘿!感谢您的回答,我运行了您的解决方案,但错误仍然存​​在 =( timespan = round(my_ocan.loc[my_ocan["creation"] == year, "timespan"].fillna(0).mean()) ValueError:无法将浮点 NaN 转换为整数
    • 我也在上面的问题中发布了它更易读的地方=)
    【解决方案2】:

    @Ha Bom,用零填充会改变平均值,我想解决方案是用 NaN 删除行:

    timespan = my_ocan.loc[my_ocan["creation"] == year, "timespan"].dropna().mean()
    

    如果您不想删除任何行而不是用平均值填充,例如请参阅Stackoverflow question for an example

    编辑 @Ha Bom 解决方案很好,因为关键是将平均值替换为零

    【讨论】:

    • 嗨,路易斯!我超越了运行您的解决方案的错误,但结果不是预期的:例如,通过运行 print(my_ocan.get_citations_per_year(2018)) 我得到 (32, 451.8235294117647) 但是当在另一个返回字典的函数中调用此函数时, 我得到 {2016: (0, nan), 2017: (0, nan), 2018: (0, nan), 2013: (0, nan), 2015: (0, nan)} 为什么会这样?跨度>
    • 嗨丽莎,你构建字典的功能是什么样的?
    • def do_get_citations_all_years(data): mydict = {} s = set(my_ocan.creation) for year in s: mydict[year] = do_get_citations_per_year(data, year) return mydict
    • 啊,我明白了,我认为原因是我使用 dropna() 删除了行,所以我不想这样做。我需要这些值 0.0,因为它们在稍后计算平均值时很重要。有没有其他方法可以填充这些空值?
    • 好吧,如果您的缺失值可以用零替换,那么@Ha Bom 解决方案就是正确的解决方案。但是你必须确保你有每年的数据,否则my_ocan.loc[my_ocan["creation"] == year, "timespan"] 可能只是空的
    猜你喜欢
    • 2020-08-14
    • 2020-04-10
    • 2021-11-12
    • 1970-01-01
    • 1970-01-01
    • 2015-09-10
    • 1970-01-01
    • 2018-04-30
    • 1970-01-01
    相关资源
    最近更新 更多