【问题标题】:psycopg2: can't adapt type 'numpy.int64'psycopg2:无法适应类型“numpy.int64”
【发布时间】:2018-11-10 13:41:44
【问题描述】:

我有一个具有如下所示 dtypes 的数据框,我想将数据框插入 postgres 数据库,但由于错误 can't adapt type 'numpy.int64'

id_code               int64
sector              object
created_date         float64
updated_date    float64

如何将这些类型转换为原生 Python 类型,例如从 int64(本质上是“numpy.int64”)转换为经典的 int,然后通过 psycopg2 客户端为 postgres 所接受。

data['id_code'].astype(np.int)  defaults to int64

仍然可以从一种 numpy 类型转换为另一种类型(例如从 int 转换为 float)

data['id_code'].astype(float)

更改为

dtype: float64

如果有人知道如何将它们转换为有用的经典类型,那么 psycopg2 似乎并不理解 numpy 数据类型。

更新:插入数据库

def insert_many():
    """Add data to the table."""
    sql_query = """INSERT INTO classification(
                id_code, sector, created_date, updated_date)
                VALUES (%s, %s, %s, %s);"""
    data = pd.read_excel(fh, sheet_name=sheetname)
    data_list = list(data.to_records())

    conn = None
    try:
        conn = psycopg2.connect(db)
        cur = conn.cursor()
        cur.executemany(sql_query, data_list)
        conn.commit()
        cur.close()
    except(Exception, psycopg2.DatabaseError) as error:
        print(error)
    finally:
        if conn is not None:
            conn.close()

【问题讨论】:

  • 你能显示你插入的代码吗?
  • @Michael 当然,更新了问题描述。
  • 您是否尝试过将 id_code 定义为 BIGINT?
  • 如果您的意思是在数据库端将其定义为 BIGINT,不,我没有尝试这种方法,因为我理解问题是 psycopg2 客户端没有将 numpy 类型映射到 python 本机。
  • Ints 从 DataFrame 中返回为 numpy.int64。 data = [[1],[2]] df = pd.DataFrame(data) print(type(df.iloc[0][0])) 导致

标签: numpy psycopg2


【解决方案1】:

在代码中的某处添加以下内容:

import numpy
from psycopg2.extensions import register_adapter, AsIs
def addapt_numpy_float64(numpy_float64):
    return AsIs(numpy_float64)
def addapt_numpy_int64(numpy_int64):
    return AsIs(numpy_int64)
register_adapter(numpy.float64, addapt_numpy_float64)
register_adapter(numpy.int64, addapt_numpy_int64)

【讨论】:

  • 如何在postgresql数据库中插入numpy?
【解决方案2】:

同样的问题,我将series转换为nd.array和int后成功解决了这个问题。

你可以尝试如下:

data['id_code'].values.astype(int)

--

更新:

如果值包括NaN,它仍然是错误的。 似乎 psycopg2 无法解释 np.int64 格式,因此以下方法适用于我。

import numpy as np
from psycopg2.extensions import register_adapter, AsIs
psycopg2.extensions.register_adapter(np.int64, psycopg2._psycopg.AsIs)

【讨论】:

  • 谢谢,register_adapter 帮助了。请注意,对于这些导入,完全限定的引用不起作用,您只需要:register_adapter(np.int64, AsIs)
【解决方案3】:

我不确定为什么你的 data_list 包含 NumPy 数据类型,但是当我运行你的代码时,同样的事情发生在我身上。这是构造 data_list 的另一种方法,以便整数和浮点数最终成为它们的原生 python 类型:

data_list = [list(row) for row in data.itertuples(index=False)] 

替代方法

我认为您可以使用 pandas to_sql 以更少的代码行完成同样的事情:

import sqlalchemy
import pandas as pd
data = pd.read_excel(fh, sheet_name=sheetname)
engine = sqlalchemy.create_engine("postgresql://username@hostname/dbname")
data.to_sql(engine, 'classification', if_exists='append', index=False)

【讨论】:

  • 在我的环境中,使用推导生成列表仍然会导致 numpy 数据类型:In [341]: type([list(row) for row in data.itertuples(index=False)][0][0]) Out[341]: numpy.int64
  • 实际上注意到,尽管数据类型为 numpy,但 pyscopg2 根据建议的理解生成的 data_list 解析没有任何问题。
【解决方案4】:

我遇到了同样的问题并使用以下方法修复了它: df = df.convert_dtypes()

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 2021-08-08
    • 2019-11-10
    • 2020-08-21
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2016-10-19
    相关资源
    最近更新 更多