【问题标题】:Removing specific value from cell of dataframe and shifting the value towards left从数据框的单元格中删除特定值并将值向左移动
【发布时间】:2020-02-18 17:18:29
【问题描述】:

我正在使用 pandas 数据框。我在某些单元格中有不需要的数据。我需要清除特定单元格中的数据并将整行向左移动一个单元格。我已经尝试了几件事,但它对我不起作用。这是示例数据框

     userId             movieId  ratings  extra
0       1                 500      3.5     
1       1                 600      4.5    
2       1                www.abcd      700     2.0
3       2                1100      5.0
4       2                1200      4.0
5       3                 600      4.5
6       4                 600      5.0
7       4                1900      3.5

预期结果:

     userId             movieId  ratings   extra
0       1                 500      3.5
1       1                 600      4.5
2       1                 700      2.0
3       2                1100      5.0
4       2                1200      4.0
5       3                 600      4.5
6       4                 600      5.0
7       4                1900      3.5

我已尝试以下代码,但显示以下错误。

raw = df[f['ratings'].str.contains('www')==True] 

#Here I am trying to fix the specific cell value to empty but it shows the following error.
**AttributeError:** 'str' object has no attribute 'at'
df = df.at[raw, 'movieId'] = ' '



#code for shifting the cell value to left
df.iloc[raw,2:-1] = df.iloc[raw,2:-1].shift(-1,axis=1)


【问题讨论】:

    标签: python pandas dataframe


    【解决方案1】:

    您可以通过掩码移动值,但这是非常重要的匹配类型,这意味着如果列 movieId 由字符串填充(因为至少一个字符串)需要将其转换为 to_numeric 的数字以避免数据丢失,因为不同的类型:

    m = df['movieId'].str.contains('www')
    df['movieId'] = pd.to_numeric(df['movieId'], errors='coerce')
    
    #if want shift only missing values rows
    #m = df['movieId'].isna()   
    df[m] = df[m].shift(-1, axis=1)
    df['userId'] = df['userId'].ffill()
    df = df.drop('extra', axis=1)
    print (df)
       userId  movieId  ratings
    0     1.0    500.0      3.5
    1     1.0    600.0      4.5
    2     1.0    700.0      2.0
    3     2.0   1100.0      5.0
    4     2.0   1200.0      4.0
    5     3.0    600.0      4.5
    6     4.0    600.0      5.0
    7     4.0   1900.0      3.5
    

    如果省略转换为数字得到缺失值:

    m = df['movieId'].str.contains('www')
    df[m] = df[m].shift(-1, axis=1)
    df['userId'] = df['userId'].ffill()
    df = df.drop('extra', axis=1)
    print (df)
       userId movieId  ratings
    0     1.0     500      3.5
    1     1.0     600      4.5
    2     1.0     NaN      2.0
    3     2.0    1100      5.0
    4     2.0    1200      4.0
    5     3.0     600      4.5
    6     4.0     600      5.0
    7     4.0    1900      3.5
    

    【讨论】:

    • @jearael 当我尝试 shiftValueError: cannot index with vector containing NA / NaN values 时,我的原始数据框中出现此错误。
    • @Sanwal - 将 m = df['movieId'].str.contains('www') 更改为 m = df['movieId'].str.contains('www', na=False)
    【解决方案2】:

    你可以试试这个:-

    df['movieId'] = pd.to_numeric(df['movieId'], errors='coerce')
    df = df.sort_values(by = 'movieId', ascending = 'True')
    

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2017-08-24
      • 1970-01-01
      • 1970-01-01
      • 2014-06-07
      • 2014-12-26
      相关资源
      最近更新 更多