【问题标题】:Pandas df read every row, return SQL query with a new column in dfPandas df 读取每一行,在 df 中返回带有新列的 SQL 查询
【发布时间】:2023-03-30 09:22:01
【问题描述】:

我有以下 pandas 数据框作为 df,我想查询 df['item'] 的每一行,它将从 SQL Server 数据库返回相应的 item_description,并用列“id”、“qty”、“item”填充 df , 'item_description'

|  id | qty  | item |
+-----+------+------+
| 001 |  700 | CB04 |
| 002 |  500 |      |
| 003 | 1500 | AB01 |

我正在做以下事情:

query = "select item_description from item_book WHERE item in {}".format(tuple(df['item']))

并使用

将其作为 df 返回
pd.read_sql_query(query, cnxn)

结果:

| item_description |
+------------------+
| apple            |
| orange           |

我打算将两个 dfs 连接在一起,这样做可能行不通,因为我在 df 的第二行有一个空值,而我的查询只返回了两行。

有没有更有效的方法来做到这一点。

【问题讨论】:

    标签: python sql pandas dataframe


    【解决方案1】:

    更改 SQL 查询以返回 item 和 item_description 列,从而为您提供如下数据框:

       item item_description
    0  CB04            apple
    1  AB01           orange
    

    然后你就有了一个公共列,可用于使用merge 函数连接两个数据框:

    pd.merge(original_df, desc_df, on="item", how="left")
    

    我们可以省略 on 参数,因为每个数据帧中只有一列具有相同的名称,pandas 会找出它们应该加入的内容。但是how 参数对于保留第一个数据帧(left-merge 的大多数参数)中没有第二个数据帧中的相应行的任何行是必要的。结果是:

       id   qty  item item_description
    0   1   700  CB04            apple
    1   2   500                    NaN
    2   3  1500  AB01           orange
    

    您可以在pandas documentation 中阅读有关合并的更多信息。

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 1970-01-01
      • 2022-08-17
      • 1970-01-01
      • 2021-12-26
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2018-04-04
      相关资源
      最近更新 更多