【问题标题】:Pandas: Replace all strings in lowercase in a column with None熊猫:用无替换列中所有小写字符串
【发布时间】:2018-11-11 00:51:16
【问题描述】:

我有一个数据集,其中包含一个名为“名称”的列,其中包括不是名称的字符串。这些都是用小写写的。

df = pd.DataFrame({'names': ['Chris Z', 'Hulk Hogan', 'notaname',
                             'whateven']})

预期输出:

     names
0    Chris Z
1    Hulk Hogan
2    NaN
3    NaN
Name: names, dtype: object

我想用 NaN 替换它们,我已经尝试过:

df['names'] = df['names'].replace(r'[a-z]{2}', None, inplace=True, regex=True)

但这会替换列中的所有条目,包括以大写字母开头的条目。能否请教一个解决方案?

【问题讨论】:

    标签: python string pandas replace


    【解决方案1】:

    使用 mask 和 ^[a-z]+$ 作为您的正则表达式:

    df = pd.DataFrame({'names': ['Chris Z', 'Hulk Hogan', 'notaname', 'whateven']})
    
    df.names.mask(df.names.str.match(r'^[a-z]+$'))
    
    0       Chris Z
    1    Hulk Hogan
    2           NaN
    3           NaN
    Name: names, dtype: object
    

    如果某些小写字符串中有空格,请改用^[a-z\s]+$。

    ^            # Asserts position at beginning of string
    [  
      a-z        # Matches any lowercase character 1 or more times
    ]+           
    $            # Asserts position at end of string
    

    【讨论】:

      【解决方案2】:

      没有正则表达式,您可以将一个系列与其自身的小写版本进行比较:

      df.loc[df['names'] == df['names'].str.lower(), 'names'] = np.nan
      
      print(df['names'])
      
      0       Chris Z
      1    Hulk Hogan
      2           NaN
      3           NaN
      Name: names, dtype: object
      

      【讨论】:

        猜你喜欢
        • 1970-01-01
        • 2019-11-23
        • 2017-03-12
        • 2019-03-04
        • 2019-06-03
        • 2020-03-07
        • 2021-07-11
        • 2023-03-19
        • 2022-10-13
        相关资源
        最近更新 更多