【问题标题】:Assignment to DataFrame not working but dtypes changed分配给 DataFrame 不起作用,但 dtypes 已更改
【发布时间】:2019-03-28 14:08:21
【问题描述】:

对 DataFrame 的分配不起作用,但 dtypes 已更改。

数据科学的新手,我想将target_frame 分配给empty_frame,但在再次分配之前它不起作用。而在分配过程中,empty_frame的dtypes从int32变为float64,最后设置为int64。

我尝试将我的模型简化为下面的代码,他们有同样的问题。

import pandas as pd
import numpy as np

dataset = [[[i for i in range(5)], ] for i in range(5)]
dataset = pd.DataFrame(dataset, columns=['test'])  

empty_numpy = np.arange(25).reshape(5, 5)
empty_numpy.fill(np.nan)

# Solution 1: change the below code into 'empty_frame = pd.DataFrame(empty_numpy)' then everything will be fine
empty_frame = pd.DataFrame(empty_numpy, columns=[str(i) for i in range(5)])

series = dataset['test']
target_frame = pd.DataFrame(list(series))

# Solution 2: run `empty_frame[:] = target_frame` twice, work fine to me.
# ==================================================================
# First try.
empty_frame[:] = target_frame
print("="*40)
print(f"Data types of empty_frame: {empty_frame.dtypes}")
print("="*40)

print("Result of first try: ")
print(empty_frame)
print("="*40)


# Second try.
empty_frame[:] = target_frame

print(f"Data types of empty_frame: {empty_frame.dtypes}")
print("="*40)

print("Result of second try: ")
print(empty_frame)
print("="*40)
# ====================================================================

我希望上面代码的输出应该是:

========================================
Data types of empty_frame: 0    int64
1    int64
2    int64
3    int64
4    int64
dtype: object
========================================
Result of first try: 
   0  1  2  3  4
0  0  1  2  3  4
1  0  1  2  3  4
2  0  1  2  3  4
3  0  1  2  3  4
4  0  1  2  3  4
========================================

但我第一次尝试时它不起作用。

这个问题有两种解决方案,但我不知道为什么:

  • 正如我在代码中显示的那样,一次运行两次尝试分配。
  • 创建empty_frame时删除列名。

我想弄清楚两件事:

  1. 为什么empty_frame 的数据类型发生了变化。
  2. 为什么我的代码中显示的解决方案可以解决这个分配问题。

谢谢。

【问题讨论】:

    标签: python pandas


    【解决方案1】:

    如果我正确理解了您的问题,那么当您创建 empty_numpy 矩阵时,您的问题就开始了。 我最喜欢的解决方案是改用 empty_numpy = np.empty([5,5]) (这里的默认 dtypes 是 float64)。那么“第一次尝试的结果:”是正确的。意思是:

    import pandas as pd
    import numpy as np
    
    dataset = [[[i for i in range(5)],] for i in range(5)]
    dataset = pd.DataFrame(dataset, columns=['test'])  
    
    empty_numpy = np.empty([5,5])
    # here you may add empty_numpy.fill(np.nan) but it's not necessary,result is the same
    
    empty_frame = pd.DataFrame(empty_numpy, columns=[str(i) for i in range(5)])
    
    series = dataset['test']
    target_frame = pd.DataFrame(list(series))
    
    # following assignment is correct then
    empty_frame[:] = target_frame
    print('='*40)
    print(f'Data types of empty_frame: {empty_frame.dtypes}')
    print('='*40)
    
    print("Result of first try: ")
    print(empty_frame)
    print("="*40)
    

    或者只是在你的 np.arrange 调用中添加 dtype 属性,就像这样:

    empty_numpy = np.arange(25, dtype=float).reshape(5, 5)
    

    然后它也可以工作(但它有点无聊;o)。

    【讨论】:

    • 感谢您的解决方案。据我了解,这是因为 np.nan 的类型是 float64 导致分配失败。但是仍然有一些让我感到困惑的地方,为什么相同的代码输出不同。如您所见,我在示例中运行了这行empty_frame[:] = target_frame 两次,但得到了不同的结果,这对我来说是不合理的。
    • @KayCrazy 您的 target_frame 的 dtype 是 int64,但您的解决方案中 empty_frame 的 dtype 是 int32 。第一个赋值 empty_frame[:] = target_frame 只是将 empty_frame 转换为 float64 - 否则你可能会失去精度(当 target_frame 将包含值 > int32 范围)。第二次运行已经分配了值(因为 float64 > int64 )。当您在 np.arrange 函数中显式设置 dtype 或使用 np.empty 时,您可以避免此问题(float 是默认设置)。解释充分吗?
    猜你喜欢
    • 2020-10-05
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2018-09-14
    • 2018-04-16
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多