这不是一个完整的答案,但我认为这个总体思路可能对您有用。这个想法是在数据框中创建代表每个范围值的最小值和最大值的新列。然后您可以非常简单地根据这些值过滤数据框:
import numpy as np
import pandas as pd
df = pd.DataFrame()
df['Curl Ups'] = ['100-500', '100-300', '200-300']
df['Age'] = ['20-24', '25-29', '30-34']
df['Points'] = [50, 60, 70]
df['Category'] = ['Decent', 'Average', 'Good']
# Create new min & max columns
df['min_curlups'] = [int(x[:x.index('-')]) for x in df['Curl Ups']]
df['max_curlups'] = [int(x[x.index('-')+1:]) for x in df['Curl Ups']]
df['min_age'] = [int(x[:x.index('-')]) for x in df['Age']]
df['max_age'] = [int(x[x.index('-')+1:]) for x in df['Age']]
# Filter by contestant's value being within the range
num_curlups = 200
filtered_df = df.loc[(df['min_curlups'] < num_curlups) & (num_curlups < df['max_curlups'])]
print(filtered_df[['Points', 'Category']])
编辑
我之前错过了多索引部分。事实上,我从未听说过 Pandas 中的这种功能。万岁学习新事物!
这是一个似乎可行的解决方案,希望您不需要进行太多调整:
首先,设置一个测试 DataFrame 并在我的方法需要过滤的列中添加:
import numpy as np
import pandas as pd
df = pd.DataFrame(index=[np.array(['IST', 'IST', 'IST', 'IST', 'ELSE', 'ELSE']),np.array(['Female', 'Male', 'Female', 'Male', 'Female', 'Male'])])
df.index.names = ['TEST', 'Gender']
df['Curl Ups'] = ['100-700', '100-300', '100-300', '100-300', '200-300', '400-500']
df['Age'] = ['20-24', '25-29', '30-34', '20-24', '30-34', '20-24']
df['Points'] = [90, 60, 70, 80, 70, 80]
df['Category'] = ['Outstanding', 'Average', 'Good', 'Better-than-me', 'Good', 'Better-than-me']
df['Level'] = ['Low', 'Medium', 'High', 'Low', 'Medium', 'High']
# Create 'Category-Level' combined column
df['Category-Level'] = df['Category'] + '-' + df['Level']
# Create new min & max columns
df['min_Curl Ups'] = [int(x[:x.index('-')]) for x in df['Curl Ups']]
df['max_Curl Ups'] = [int(x[x.index('-')+1:]) for x in df['Curl Ups']]
df['min_Age'] = [int(x[:x.index('-')]) for x in df['Age']]
df['max_Age'] = [int(x[x.index('-')+1:]) for x in df['Age']]
df
这是一个函数,它应该接受与您所需类似的输入并输出所需信息:
def filter_multiindex_df( df, inputs, input_fields=['TEST', 'Gender', 'Age', 'Curl Ups'], outputs=['Points', 'Category-Level'] ):
# Pull out the inputs corresponding to multi-level indices & filter to get the DF cross-section
input_idx = [i for i,x in enumerate(input_fields) for j,y in enumerate(df.index.names) if x == y]
index_inputs = [inputs[i] for i in input_idx]
filt_df = df.xs(index_inputs)
# Filter based on the rest of the inputs that weren't indices
for i, field in enumerate(input_fields):
if i not in input_idx:
filt_df = filt_df.loc[ (filt_df['min_'+field] <= inputs[i]) & (inputs[i] <= filt_df['max_'+field])]
return filt_df[outputs].values[0]
# Test the function
print(filter_multiindex_df(df, inputs=['IST', 'Female', 22, 100]))