【问题标题】:Pandas data frame method `to_gbq` handling nested data in data frame熊猫数据框方法`to_gbq`处理数据框中的嵌套数据
【发布时间】:2018-05-27 08:44:06
【问题描述】:

我已使用 pandas read_json 将带有嵌套对象的 json 数据读入数据框中。

我想使用 pandas to_gbq 将这些数据推送到 Google Big Query,但嵌套元素会出现如下错误:

StreamingInsertError: Error at Row: 0, Reason: invalid, Location: payment_details, Message: This field is not a record.

数据框如下所示:

df['payment_details'][1]
{u'credit_card_bin': u'xxxx', u'avs_result_code': u'Y', u'credit_card_company': u'Visa', u'cvv_result_code': u'M', u'credit_card_number': u'yyy'}

应该如何处理以便 GBQ 将其作为记录使用?

将对象映射到字符串时,问题似乎出在哪里 pandas_gbq/gbq.py

type_mapping = {
    'i': 'INTEGER',
    'b': 'BOOLEAN',
    'f': 'FLOAT',
    'O': 'STRING',
    'S': 'STRING',
    'U': 'STRING',
    'M': 'TIMESTAMP'
}

【问题讨论】:

    标签: python pandas google-bigquery


    【解决方案1】:

    pandas_gbq 不支持结构或数组,在 pandas DataFrames 中存储 dicts 或列表也不是 'pandantic'

    如果您使用的是结构,您可以在 DataFrame 中使用多个列吗?

    【讨论】:

      猜你喜欢
      • 2021-02-15
      • 2016-07-07
      • 2017-11-28
      • 2019-07-08
      • 2016-11-29
      • 1970-01-01
      • 1970-01-01
      • 2023-03-23
      相关资源
      最近更新 更多