【问题标题】:Given an input, return the row which has a value of a range which contains that input给定一个输入,返回具有包含该输入的范围值的行
【发布时间】:2019-10-31 21:13:08
【问题描述】:

我是新人,所以请温柔一点。我是自学成才的,所以我可能会觉得我有点白痴。

所以我有一个带有多级索引的 Pandas 数据框。

TEST     Gender     Age     Category     Level     Points     Curl Ups
IST      Female     20-24   Outstanding  High      100        105-500
                            Outstanding  Medium    95         103-104
                            Outstanding  Low       90         100-102
...       ...       ...         ...        ...      ...        ...
                    25-29   Outstanding  High      100        103-500
                            Outstanding  Medium    95         100-102
...

我希望能够输入测试、性别、年龄和个人所做的弯举次数,即;对于 IST 测试,一位 26 岁并做了 100 次弯举的女性看起来像

['IST','Female',26,100]

我正在寻找的输出将是点,类别级别。即;

[95, 'Outstanding-Medium']

所以它需要能够为 'Age' 获取 26 的输入,并且知道跳转到 25-29 块下的行,然后获取 100 的输入并知道拉100 落入,即; 100-102。这些都是包容性的(26 岁时进行 100 次弯举,包括 102 次弯举可以让您获得 95 分,这是一个出色的中等水平)。

我尝试使用 loc,但我发现的所有内容都以另一种方式使用,它会询问您在所需范围内放置的行,然后提取包含该范围内的值的行。

我需要的是,你给它一个值,它会拉出具有包含该值的范围的行。

【问题讨论】:

    标签: python pandas dataframe search multi-index


    【解决方案1】:

    这不是一个完整的答案,但我认为这个总体思路可能对您有用。这个想法是在数据框中创建代表每个范围值的最小值和最大值的新列。然后您可以非常简单地根据这些值过滤数据框:

    import numpy as np
    import pandas as pd
    
    df = pd.DataFrame()
    df['Curl Ups'] = ['100-500', '100-300', '200-300']
    df['Age'] = ['20-24', '25-29', '30-34']
    df['Points'] = [50, 60, 70]
    df['Category'] = ['Decent', 'Average', 'Good']
    
    # Create new min & max columns
    df['min_curlups'] = [int(x[:x.index('-')]) for x in df['Curl Ups']]
    df['max_curlups'] = [int(x[x.index('-')+1:]) for x in df['Curl Ups']]
    
    df['min_age'] = [int(x[:x.index('-')]) for x in df['Age']]
    df['max_age'] = [int(x[x.index('-')+1:]) for x in df['Age']]
    
    # Filter by contestant's value being within the range
    num_curlups = 200
    filtered_df = df.loc[(df['min_curlups'] < num_curlups) & (num_curlups < df['max_curlups'])]
    print(filtered_df[['Points', 'Category']])
    

    编辑

    我之前错过了多索引部分。事实上,我从未听说过 Pandas 中的这种功能。万岁学习新事物! 这是一个似乎可行的解决方案,希望您不需要进行太多调整:

    首先,设置一个测试 DataFrame 并在我的方法需要过滤的列中添加:

    import numpy as np
    import pandas as pd
    
    df = pd.DataFrame(index=[np.array(['IST', 'IST', 'IST', 'IST', 'ELSE', 'ELSE']),np.array(['Female', 'Male', 'Female', 'Male', 'Female', 'Male'])])
    df.index.names = ['TEST', 'Gender']
    df['Curl Ups'] = ['100-700', '100-300', '100-300', '100-300', '200-300', '400-500']
    df['Age'] = ['20-24', '25-29', '30-34', '20-24', '30-34', '20-24']
    df['Points'] = [90, 60, 70, 80, 70, 80]
    df['Category'] = ['Outstanding', 'Average', 'Good', 'Better-than-me', 'Good', 'Better-than-me']
    df['Level'] = ['Low', 'Medium', 'High', 'Low', 'Medium', 'High']
    
    # Create 'Category-Level' combined column
    df['Category-Level'] = df['Category'] + '-' + df['Level']
    
    # Create new min & max columns
    df['min_Curl Ups'] = [int(x[:x.index('-')]) for x in df['Curl Ups']]
    df['max_Curl Ups'] = [int(x[x.index('-')+1:]) for x in df['Curl Ups']]
    
    df['min_Age'] = [int(x[:x.index('-')]) for x in df['Age']]
    df['max_Age'] = [int(x[x.index('-')+1:]) for x in df['Age']]
    
    df
    

    这是一个函数,它应该接受与您所需类似的输入并输出所需信息:

    def filter_multiindex_df( df, inputs, input_fields=['TEST', 'Gender', 'Age', 'Curl Ups'], outputs=['Points', 'Category-Level'] ):
    
        # Pull out the inputs corresponding to multi-level indices & filter to get the DF cross-section
        input_idx = [i for i,x in enumerate(input_fields) for j,y in enumerate(df.index.names) if x == y]
        index_inputs = [inputs[i] for i in input_idx]
        filt_df = df.xs(index_inputs)
    
        # Filter based on the rest of the inputs that weren't indices
        for i, field in enumerate(input_fields):
            if i not in input_idx:
                filt_df = filt_df.loc[ (filt_df['min_'+field] <= inputs[i]) & (inputs[i] <= filt_df['max_'+field])]
    
        return filt_df[outputs].values[0]
    
    
    # Test the function    
    print(filter_multiindex_df(df, inputs=['IST', 'Female', 22, 100]))
    

    【讨论】:

    • 虽然这个想法适用于 curl-ups 列,但当您拥有多级索引时,它并不适用。你知道我们如何使用 df.xs 或类似的东西来做到这一点吗?
    • 虽然不漂亮,但请参阅上面我编辑的尝试。
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2023-04-02
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多