【问题标题】:Pandas: How do I search series cells for partial string match?Pandas:如何在系列单元格中搜索部分字符串匹配?
【发布时间】:2021-03-15 01:56:24
【问题描述】:

我通过导入一个 Excel 工作表创建了一个 DataFrame,该工作表具有多列但行中的文本数据不一致。例如,标记为“Airplane 1”的一列在第 25 行具有“Gross weight: 2500”,而“Airplane 2”列在另一行具有“Gross weight: 3000”。所有列都有一个“总重量:”条目,但它们的行号相差 1 或更多。我可以遍历列和行,但似乎无法在一行中的单元格中查询特定字符串。我尝试了几种方法,下面发现一种失败。很明显,我错误地尝试从系列中生成单个布尔值,从而产生错误,但我似乎无法进入系列中的单独单元格。最终,我想确定特定参数,例如“毛重:”,提取与该参数关联的数字并将其与其特定列联系起来。是的,我是新手,在此先感谢...

只是为了表明数据在那里......

#print(df.at[2,'Aviat_A-1B'])
x = df.loc[11,"Aviat_A-1B"]
#x.partition(':')
#print(type(x))
#print(x.split(':'))
print(x)

Gross weight (lbs.): 2000

这行不通……

sub = 'Gross weight (lbs.):'
for index, row in df.iterrows():
    print(type(index))
    print(index)
    print('~~~~~~')
    print(type(row))
    print(row)
    print('------')
    if row.str.extract(sub):
        print(type(row))
        print(row)
        print('------')


<class 'int'>
3
<class 'pandas.core.series.Series'>
1997_7GCAA_American_Champion_Adventure                              Price as tested: $76,000
Aviat_A-1B                                                       Engine make/model: Lycoming
1960_Beech_Travel_Air_B95                                     Engine make/model: Lycoming...
1979_Beechcraft_Bonanza_A-36                                                        IO-520BB
1977_Bellanca_8KCAB-180_Super_Decathlon                       Engine make/model: Lycoming...
                                                                 ...                        
1974_Piper_Arrow_II_with_LoPresti_Speed       Price: $48,500 (plus mod. cost)               
1999_Piper_Archer_III                              Engine make/model: Lycoming              
Ryan_Navion                                       Engine make/model: Cont. E-185            
1997_Mooney_Ovation                              Engine make/model: Continental IO-550G     
1997_Mooney_Encore_Prototype                   Engine make/model: Cont TSIO-360-SB          
Name: 3, Length: 61, dtype: object
------
---------------------------------------------------------------------------
ValueError                                Traceback (most recent call last)
<ipython-input-52-8590ea2a7401> in <module>
     11     print(row)
     12     print('------')
---> 13     if row.str.extract(sub):
     14         print(type(row))
     15         print(row)

C:\ProgramData\Anaconda3\lib\site-packages\pandas\core\generic.py in __nonzero__(self)
   1327 
   1328     def __nonzero__(self):
-> 1329         raise ValueError(
   1330             f"The truth value of a {type(self).__name__} is ambiguous. "
   1331             "Use a.empty, a.bool(), a.item(), a.any() or a.all()."

ValueError: The truth value of a DataFrame is ambiguous. Use a.empty, a.bool(), a.item(), a.any() or a.all().

!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!
UPDATE
!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!

There has to be a smarter way than this to find and extract a number, 199 in the case below...

`wing_area = []
#print(df)
for col, item in df.iteritems():
#    print(col)
    wing_area_bool = item.str.contains("Wing area", na=False)
#    print(df.index[wing_area_bool])
#    print(item[wing_area_bool])

#    wing_area.append(item[wing_area_bool])
#    wing_area.append(item[wing_area_bool].str.split(":"))
    wing_area.append(item[wing_area_bool].str.split())
    
print(wing_area[-1])
#print(len(wing_area[-1]))
#print(str(wing_area[-1]))
#x = (str(wing_area[-1]).split(","))
#y = x[-2].split("]")
#int(y[0].strip())
int(str(wing_area[-1]).split(",")[-2].split("]")[0].strip())

22    [Wing, area, (sq., ft.):, 199]
Name: 1960_Beech_Travel_Air_B95, dtype: object
199`


【问题讨论】:

    标签: python pandas string series


    【解决方案1】:
    df.loc[df.column_name.str.contains("Gross weight", na=False)]
    

    reference

    要迭代地执行此操作,您可以在执行此操作时遍历列名

    for name in [column_name_1, column_name_2, column_name_x]:
         df.loc[df.name.str.contains("Gross weight", na=False)]
    

    根据你想对结果做什么,你会在循环中执行它

    【讨论】:

    • 不完全 - 我使用 ``` import pandas as pd import numpy as np file = 'C:/Light AC data.xls' df = pd.read_excel(file, sheet_name) 导入了 .xls 文件=0) df.dropna(inplace = True) df``` (对不起,我不知道如何在回复中输入缩进代码...)。这可能就是为什么,当我使用您的建议时(我曾以稍微不同的方式尝试过 - 我有 61 列我想迭代)我得到了这个... df[df["Aviat_A-1B"].str .contains("Gross Weight")]... ValueError: Cannot mask with non-boolean array contains NA / NaN values...
    • 试试 df.loc[df.a.str.contains("毛重", na=False)] reference
    • 成功了,谢谢!现在我只需要弄清楚如何通过所有列迭代地执行此操作。哦,谢谢你的参考指针!
    • 太棒了!请参阅我在每列上执行此操作的更新答案
    • 实际上,我将您的建议wing_area = [] 调整为col, item in df.iteritems():wing_area_bool = item.str.contains("Wing area", na=False)wing_area.append( item[wing_area_bool].str.split(":")) print(type(wing_area)) print(wing_area[0]) 22 [ Wing area (sq. ft.), 165] Name: 1997_7GCAA_American_Champion_Adventure , dtype: 对象
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2017-03-08
    • 1970-01-01
    • 2015-03-10
    • 1970-01-01
    • 2013-06-04
    • 2017-07-15
    相关资源
    最近更新 更多