【问题标题】:How to check if a combination of values in 2 columns is present in a pandas data frame如何检查熊猫数据框中是否存在2列中的值组合
【发布时间】:2014-06-03 08:57:09
【问题描述】:

我有 2 个数据框。标题和部分是其中的两列。 我需要检查一个数据框中的特定标题和部分的组合是否存在于第二个数据框中。

例如数据框torange有列title、s_low、s_high等; usc 具有列标题和部分。

如果torange中有下面一行

title   s_low   s_high
  1        1      17

如果一行代码需要签入usc

title   section
 1        1

还有一行

title   section
 1          17

存在;并创建一个新表,通过扩展usc 中的 s_low 和 s_high 之间的范围,在 torange 中写入标题、部分和其余列。

我编写了以下代码,但不知何故,它在几次迭代后就无法工作/停止。我怀疑“i”的计数器有问题,可能是语法错误。另外,可能与any()的语法有关

import MySQLdb as db
from pandas import DataFrame
from pandas.io.sql import frame_query
cnxn = db.connect('127.0.0.1','xxxxx','xxxxx','xxxxxxx', charset='utf8', use_unicode=True )
torange = frame_query("SELECT title, s_low, s_high, post, pre, noy, rnum from torange", cnxn)
usc = frame_query("SELECT title, section from usc", cnxn)

i=0
for row in torange:
    t =  torange.title[i]
    s_low = torange.s_low[i]
    s_high = torange.s_high[i]
    for row in usc:
        if (any(usc.title == t) & any(usc.section == s_low)):
            print 't', t, 's_low' , s_low,'i', i
            if (any(usc.title == t) & any(usc.section == s_high)):
                print 't', t, 's_high', s_high, 'i',i
                print i, '*******************match************************'
    i=i+1

(请忽略打印语句。这是我正在做的一项更大任务的一部分,打印只是用来检查正在发生的事情。)

在这方面的任何帮助将不胜感激。

【问题讨论】:

    标签: python sql pandas dataframe


    【解决方案1】:

    您的整个检查和迭代都搞砸了。您在usc 中迭代row,但您的any() 条件检查usc,而不是row。此外,row 是两个循环的迭代器。这是一个更清晰的起点:

    for index, row in torange.iterrows(): 
        t = row['title']
        s_low = row['s_low']
        s_high = row['s_high']
        uscrow = usc[(usc.title == t) & (usc.section == slow)]
        # uscrow now contains all the rows in usc that fulfill your condition.
    

    【讨论】:

    • 此外,在遍历数据帧时,我总是使用iterrows() 和iteritems(),以确保它确实符合我的预期。
    • 非常很少如果你真的需要迭代。这是isin的简单应用
    • @FooBar 感谢您的回复。我是熊猫新手,因此迭代混乱。如何在同一语句中检查 s_high 以及/否则?我的结果集将包含满足以下条件的行 [(usc.title == t) 和 (usc.section == s_low)] 和 [(usc.title == t) 和 (usc.section == s_high )]
    • 简单地说,特定标题的 s_low 和 s_high 都必须存在于 usc 中。
    • @Jeff 如何在使用 isin 时进行迭代?我已经阅读了文档:pandas.pydata.org/pandas-docs/version/0.13.1/generated/… 如何将许多值作为参数传递给它?
    猜你喜欢
    • 1970-01-01
    • 2022-01-01
    • 2014-06-26
    • 2021-10-14
    • 1970-01-01
    • 2020-05-08
    • 2019-11-22
    • 2020-03-18
    • 1970-01-01
    相关资源
    最近更新 更多