【问题标题】:Pandas Fill column value with next available value in the same column熊猫用同一列中的下一个可用值填充列值
【发布时间】:2018-08-20 01:12:04
【问题描述】:

我正在处理一个数据集,其中 PLU 列中的值分散在各处,例如: 在 500 多列中,我有 4 列:

Inventory_No | Description | Group | PLU
----------------------------------------------
93120007     | Coke        |Drinks | 1000
93120008     | Diet Coke   |Drinks | 1003
93120009     | Coke Zero   |Drinks | 1104
93120010     | Fanta       |Drinks | 1105

93120011     | White Bread |Bread  | 93120011     
93120012     | whole Meal  |Bread  | 93120012     
93120013     | Whole Grains|Bread  | 110011
93120014     | Flat white  |Breads | 1115092

我希望我的输出如下所示,如果 PLU 列中有任何长度超过 6 位的值,系统会检查 PLU 序列中长度小于 4 位的下一个可用数字并添加增量1 并将 PLU 值分配给该行,并且不会更改任何现有的小于 6 位数的 PLU:

Inventory_No | Description | Group | PLU
----------------------------------------------
93120007     | Coke        |Drinks | 1000
93120011     | White Bread |Bread  | 1001
93120012     | whole Meal  |Bread  | 1002
93120008     | Diet Coke   |Drinks | 1003
93120014     | Flat white  |Breads | 1004
   .         |     .       |  .    |   .
   .         |     .       |  .    |   .
   .         |     .       |  .    |   .
93120009     | Coke Zero   |Drinks | 1104
93120010     | Fanta       |Drinks | 1105
93120013     | Whole Grains|Bread  | 110011

我想要序列中小于 6 位的下一个可用值并将其递增 1,如果它找到任意数量的递增值的序列,则跳过该序列并从序列之后的下一个可用值开始序列长度小于 6 位:
我已经检查了以下链接,它们正朝着用 0 或 Nan 值填充序列
fill-in-a-missing-values-in-range-with-pandas
missing-data-insert-rows-in-pandas-and-fill-with-nan

提前感谢您的回答。 问候,

【问题讨论】:

  • 遍历PLU中的所有值并从之前的行中分配一个递增的值还不够吗?而且我猜您隐含地假设您分配的数字不会超过 6 位数,因为它们可能与从外部分配的数字发生冲突?
  • 既然你会改变一些PLU......你能不能不重新写很多?如果它需要保持一致以引用其他很好的东西,但在你更改它们的地方不会出现这种情况......所以......看起来如果你很乐意这样做,你应该给他们所有的品牌新代码从 1000 开始

标签: python python-3.x pandas


【解决方案1】:

设置

print(df)

   Inventory_No   Description   Group       PLU
0      93120007          Coke  Drinks      1000
1      93120008     Diet Coke  Drinks      1003
2      93120009     Coke Zero  Drinks      1104
3      93120010         Fanta  Drinks      1105
4      93120011   White Bread   Bread  93120011
5      93120012    whole Meal   Bread  93120012
6      93120013  Whole Grains   Bread    110011
7      93120014    Flat white  Breads   1115092

首先,让我们创建一个值列表,我们可以使用这些值来填充df.PLU 中不包含的值:

fillers = [
    i for i in np.arange(df.PLU.min(), df.PLU.min() + len(df)) if i not in set(df.PLU)
]
# [1001, 1002, 1004, 1005, 1006, 1007]

现在我们可以用我们的新值制作一个系列并填充:

condition = df.PLU.ge(1e6)
s = df.loc[condition]
fill = pd.Series(fillers[len(s):], index=s.index)
df.assign(PLU=df.PLU.mask(condition).fillna(fill).astype(int)).sort_values('PLU')

输出:

   Inventory_No   Description   Group     PLU
0      93120007          Coke  Drinks    1000
4      93120011   White Bread   Bread    1001
5      93120012    whole Meal   Bread    1002
1      93120008     Diet Coke  Drinks    1003
7      93120014    Flat white  Breads    1004
2      93120009     Coke Zero  Drinks    1104
3      93120010         Fanta  Drinks    1105
6      93120013  Whole Grains   Bread  110011

【讨论】:

  • #user3483203 我已经尝试过你的代码填充 = [ i for i in np.arange(df.PLU.min(), df.PLU.min() + len(df)) 如果我没有in set(df.PLU) ] 并收到以下错误:TypeError: must be str, not int on df.PLU.min() + len(df))
  • PLU 列的数据类型是什么?
  • 当我使用 type(df.PLU) 检查时它是一个系列
  • df.PLU.dtype 并显示输出,我猜PLU 是一个字符串列
  • 当我完成 df.PLU.dtype 后它给了我 dtype('O')
【解决方案2】:

示例数据框:

df = pd.DataFrame({'PLU': ['1001', '1002', '1110679', '1003', '1005', '12345', '1234567', '1231231231312', '1003', '1110679']}

获取下一个未使用的 4 位数字:

start_at = int(df['PLU'][df.PLU.str.len() == 4].max()) + 1

构建一个从起始数字到 10000 的可迭代对象(因此范围最多为 9999 - 例如:只有 4 位数字):

spare_code = iter(range(start_at, 10000))

PLU长度超过6个字符时,替换为下一个备用码...

to_replace = df['PLU'].str.len() > 6
df.loc[to_replace, 'PLU'] = df.PLU[to_replace].map(lambda v: str(next(spare_code)))

给你一个修改后的df

     PLU
0   1001
1   1002
2   1006
3   1003
4   1005
5  12345
6   1007
7   1008
8   1003
9   1009

【讨论】:

    猜你喜欢
    • 2021-03-18
    • 2020-11-26
    • 1970-01-01
    • 1970-01-01
    • 2017-02-08
    • 1970-01-01
    • 2021-09-11
    • 1970-01-01
    • 2012-05-29
    相关资源
    最近更新 更多