【问题标题】:How to do I pull out numbers from python string?如何从 python 字符串中提取数字?
【发布时间】:2019-09-13 16:46:14
【问题描述】:

我必须从 statsmodel 包生成的系数参数中提取我的节点值并将其放在它自己的列中。

下面是熊猫数据框的当前示例,下面是我正在寻找的解决方案。当使用 statsmodels 包拟合分段线性模型时,变量会以patsy 语句的形式返回。如果一个人打一个结,就会有两个系数。如果用户放两个结,三个系数。在每个变量语句的末尾,括号内都有一个数字。如果该数字 = [0],那么我需要新列中的值来表示 0。如果数字是[1],那么我需要将新列中的值改为字符串的knots= [] 部分中的第一个值。如果号码是[2],那么我需要将knots=[]语句中的第二个号码拉出来,依此类推。我已经尝试过在线帮助工具,但我没有取得任何突破。

import pandas as pd
#current

dict = {'index': ['bs(np.clip(vehicle_age_model, 0, np.inf), degree=1, knots=[10, 25])[0]'
        , 'bs(np.clip(vehicle_age_model, 0, np.inf), degree=1, knots=[10, 25])[1]'
        , 'bs(np.clip(vehicle_age_model, 0, np.inf), degree=1, knots=[10, 25])[2]'
        ,'bs(np.clip(driver_age_model, 0, np.inf), degree=1, knots=[25])[0]'
        , 'bs(np.clip(driver_age_model, 0, np.inf), degree=1, knots=[25])[1]'
        ,'bs(np.clip(length_ft_model, 0, np.inf), degree=1, knots=[32])[0]'
        ,'bs(np.clip(length_ft_model, 0, np.inf), degree=1, knots=[32])[0]']}

df1 = pd.DataFrame.from_dict(dict)

df1

# Solution

dict2 = {'index': ['bs(np.clip(vehicle_age_model, 0, np.inf), degree=1, knots=[10, 25])[0]'
        , 'bs(np.clip(vehicle_age_model, 0, np.inf), degree=1, knots=[10, 25])[1]'
        , 'bs(np.clip(vehicle_age_model, 0, np.inf), degree=1, knots=[10, 25])[2]'
        ,'bs(np.clip(driver_age_model, 0, np.inf), degree=1, knots=[10, 25])[0]'
        , 'bs(np.clip(driver_age_model, 0, np.inf), degree=1, knots=[10, 25])[1]'
        ,'bs(np.clip(length_ft_model, 0, np.inf), degree=1, knots=[32])[0]'
        ,'bs(np.clip(length_ft_model, 0, np.inf), degree=1, knots=[32])[0]'],
       'desired_1': [0,10,25,0,25,0,32]}

df2 = pd.DataFrame.from_dict(dict2)
df2

【问题讨论】:

  • 您的字符串似乎包含表达式。为什么不直接运行它们,或者评估它们?
  • 包确实提供了系数。你应该看看。还是您需要正则表达式解决方案?其中我不推荐
  • 嗨@Onyambu。我需要把它们拉出来,这样我就可以用它们加入另一个离散年龄的表。这将允许我根据年龄加入正确的系数。
  • 我提供的答案没有解决问题吗?

标签: regex python-3.x string


【解决方案1】:
import re

def pull_number_and_index(input_string):
    patt = r'.*\[(\d)\]$'
    l_idx = int(re.sub(patt, r'\g<1>', input_string))
    l_patt = r'.*knots=\[(.*)\]\).*'
    l_str = re.sub(l_patt, r'\g<1>', input_string)
    knots_list = list(l_str.split(','))
    if l_idx == 0:
        return 0
    else:
        return knots_list[l_idx-1]

df1['desired1'] = df1['index'].apply(pull_number_and_index)

正则表达式有点奇怪,patt 匹配捕获组中括号中的最后一个数字,提取该数字并将其转换为 int。

l_patt 匹配捕获组中knots= 之后的列表,使用re.sub 提取它。生成的字符串将转换为带有str.split 的列表。

那么比较就很直接了。

【讨论】:

    【解决方案2】:

    你可以这样做:

     df1.assign(desired1 = df1['index'].str.replace('.*=.','([0, ').apply(eval))
    Out: 
                                                   index  desired1
    0  bs(np.clip(vehicle_age_model, 0, np.inf), degr...         0
    1  bs(np.clip(vehicle_age_model, 0, np.inf), degr...        10
    2  bs(np.clip(vehicle_age_model, 0, np.inf), degr...        25
    3  bs(np.clip(driver_age_model, 0, np.inf), degre...         0
    4  bs(np.clip(driver_age_model, 0, np.inf), degre...        25
    5  bs(np.clip(length_ft_model, 0, np.inf), degree...         0
    6  bs(np.clip(length_ft_model, 0, np.inf), degree...         0
    

    不过,我不推荐eval,否则你应该使用ast.literal_eval

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 2020-11-13
      • 1970-01-01
      • 1970-01-01
      • 2018-05-31
      • 1970-01-01
      相关资源
      最近更新 更多