【问题标题】:How to construct a list comprehension with nested for loops and conditionals for pandas?如何使用嵌套的 for 循环和 pandas 的条件构建列表理解?
【发布时间】:2019-03-12 00:45:23
【问题描述】:

我很难让以下复杂的列表理解按预期工作。这是一个带有条件的双重嵌套 for 循环。

让我先解释一下我在做什么:

import pandas as pd

dict1 = {'stringA':['ABCDBAABDCBD','BBXB'], 'stringB':['ABDCXXXBDDDD', 'AAAB'], 'num':[42, 13]}

df = pd.DataFrame(dict1)
print(df)
        stringA       stringB  num
0  ABCDBAABDCBD  ABDCXXXBDDDD   42
1          BBXB          AAAB   13

此 DataFrame 有两列 stringA 和 stringB,其中的字符串包含字符 A、B、C、D、X。根据定义,这两个字符串的长度相同。

基于这两列,我创建了字典,使得stringA 从索引 0 开始,stringB 从num 开始的索引开始。

这是我使用的函数:

def create_translation(x):
    x['translated_dictionary'] = {i: i +x['num'] for i, e in enumerate(x['stringA'])}
    return x

df2 = df.apply(create_translation, axis=1).groupby('stringA')['translated_dictionary']


df2.head()
0    {0: 42, 1: 43, 2: 44, 3: 45, 4: 46, 5: 47, 6: ...
1                         {0: 13, 1: 14, 2: 15, 3: 16}
Name: translated_dictionary, dtype: object

print(df2.head()[0])
{0: 42, 1: 43, 2: 44, 3: 45, 4: 46, 5: 47, 6: 48, 7: 49, 8: 50, 9: 51, 10: 52, 11: 53}

print(df2.head()[1])
{0: 13, 1: 14, 2: 15, 3: 16}

没错。

但是,这些字符串中有“X”字符。这需要一个特殊规则:如果X 在stringA 中,则不要在字典中创建键值对。如果X 在stringB 中,则该值不应为i + x['num'] 而是-500。

我尝试了以下列表理解:

def try1(x):
    for count, element in enumerate(x['stringB']):
        x['translated_dictionary'] = {i: -500 if element == 'X' else  i + x['num'] for i, e in enumerate(x['stringA']) if e != 'X'}
    return x

这给出了错误的答案。

df3 = df.apply(try1, axis=1).groupby('stringA')['translated_dictionary']

print(df3.head()[0]) ## this is wrong!
{0: 42, 1: 43, 2: 44, 3: 45, 4: 46, 5: 47, 6: 48, 7: 49, 8: 50, 9: 51, 10: 52, 11: 53}

print(df3.head()[1])   ## this is correct! There is no key for 2:15!
{0: 13, 1: 14, 3: 16}

没有 -500 值!

正确答案是:

print(df3.head()[0])
{0: 42, 1: 43, 2: 44, 3: 45, 4:-500, 5:-500, 6:-500, 7: 49, 8: 50, 9: 51, 10: 52, 11: 53}

print(df3.head()[1])
{0: 13, 1: 14, 3: 16}

【问题讨论】:

  • 为什么你的最后一个例子有 13、14、16 而不是 13、14、15?
  • @JohnZwinck 这是基于第一条规则的。 “如果 X 在 stringA 中,不要在字典中创建键值对。”在这种情况下,BBXB 在 2:15 处有一个 X。这有意义吗?

标签: python python-3.x pandas list list-comprehension


【解决方案1】:

为什么不扁平化你的 df ,你可以用这个 post 检查并重新创建 dict

n=df.stringA.str.len()
newdf=pd.DataFrame({'num':df.num.repeat(n),'stringA':sum(list(map(list,df.stringA)),[]),'stringB':sum(list(map(list,df.stringB)),[])})


newdf=newdf.loc[newdf.stringA!='X'].copy()# remove stringA value X
newdf['value']=newdf.groupby('num').cumcount()+newdf.num # using groupby create the cumcount 
newdf.loc[newdf.stringB=='X','value']=-500# assign -500 when stringB is X
[dict(zip(x.groupby('num').cumcount(),x['value']))for _,x in newdf.groupby('num')] # create the dict for different num by group
Out[390]: 
[{0: 13, 1: 14, 2: 15},
 {0: 42,
  1: 43,
  2: 44,
  3: 45,
  4: -500,
  5: -500,
  6: -500,
  7: 49,
  8: 50,
  9: 51,
  10: 52,
  11: 53}]

【讨论】:

  • 你能解释一下你在上面做什么吗?
  • @ShanZhengYang 检查一下,更新了。同样对于扁平化的,我这里没有解释,你可以查看链接的帖子。
【解决方案2】:

这是一个简单的方法,没有任何理解(因为它们无助于澄清代码):

def create_translation(x):
    out = {}
    num = x['num']
    for i, (a, b) in enumerate(zip(x['stringA'], x['stringB'])):
        if a == 'X':
            pass
        elif b == 'X':
            out[i] = -500
        else:
            out[i] = num
        num += 1
    x['translated_dictionary'] = out
    return x

【讨论】:

  • 这行得通!我没有想过将num 作为局部变量进行跟踪并进行迭代。谢谢!
猜你喜欢
  • 1970-01-01
  • 2011-04-07
  • 1970-01-01
  • 1970-01-01
  • 2022-11-07
  • 2020-03-06
  • 1970-01-01
  • 2019-11-27
  • 1970-01-01
相关资源
最近更新 更多