【问题标题】:string column manipulation in Data Frame in pandas熊猫数据框中的字符串列操作
【发布时间】:2015-05-18 20:16:51
【问题描述】:

我在这样的数据框中有一个字符串列(时间)。我想在数字之间加上下划线并删除月份。

Time
2- 3 months          
1- 2 months          
10-11 months          
4- 5 months
 Desired output:
2_3           
1_2           
10_11           
4_5 

这是我正在尝试但似乎不起作用的方法。

def func(string):
    a_new_string =string.replace('- ','_')
    a_new_string1 =a_new_string.replace('-','_')
    a_new_string2= a_new_string1.rstrip(' months')
    return a_new_string2

并将函数应用于数据框。

df['Time'].apply(func)

【问题讨论】:

  • 您是否将结果分配回去? df['Time'] = df['Time'].apply(func)?
  • 是的。我想在数据框的时间列中应用该函数。

标签: python regex string pandas


【解决方案1】:

一种选择是使用 3 个 str replace 调用:

In [18]:

df['Time'] = df['Time'].str.replace('- ', '_')
df['Time'] = df['Time'].str.replace('-', '_')
df['Time'] = df['Time'].str.replace(' months', '')
df
Out[18]:
    Time
0    2_3
1    1_2
2  10_11
3    4_5

我认为您的问题可能是您没有将apply 的结果分配回去:

In [21]:

def func(string):
    a_new_string =string.replace('- ','_')
    a_new_string1 =a_new_string.replace('-','_')
    a_new_string2= a_new_string1.rstrip(' months')
    return a_new_string2

df['Time'] = df['Time'].apply(func)
df
Out[21]:
    Time
0    2_3
1    1_2
2  10_11
3    4_5

您也可以将其设为单行:

In [25]:

def func(string):
    return string.replace('- ','_').replace('-','_').rstrip(' months')

df['Time'] = df['Time'].apply(func)
df
Out[25]:
    Time
0    2_3
1    1_2
2  10_11
3    4_5

【讨论】:

  • 我试过了,但我需要使用函数。我试图在函数中编写相同的内容。
猜你喜欢
  • 1970-01-01
  • 2021-12-14
  • 2017-02-16
  • 1970-01-01
  • 2018-04-26
  • 1970-01-01
  • 1970-01-01
  • 2017-02-24
  • 1970-01-01
相关资源
最近更新 更多