【问题标题】:Pandas Data Frame Partial String Replace熊猫数据框部分字符串替换
【发布时间】:2017-09-09 00:42:13
【问题描述】:

给定这个数据框:

import pandas as pd
d=pd.DataFrame({'A':['a','b',99],'B':[1,2,'99'],'C':['abcd99',4,5]})
d

    A   B   C
0   a   1   abcd*
1   b   2   4
2   99  99  5

我想用星号替换整个数据框中的所有 99。 我试过这个:

d.replace('99','*')

...但它仅适用于 B 列中的字符串 99。

提前致谢!

【问题讨论】:

    标签: python string pandas replace partial


    【解决方案1】:

    这样就可以了:

    import pandas as pd
    d=pd.DataFrame({'A':['a','b',99],'B':[1,2,'99'],'C':['abcd99',4,5]})
    d=d.astype(str)
    d.replace('99','*',regex=True)
    

    给了

        A   B   C
    0   a   1   abcd*
    1   b   2   4
    2   *   *   5
    

    请注意,这会创建一个新的数据框。您也可以改为就地执行此操作:

    d.replace('99','*',regex=True,inplace=True)
    

    【讨论】:

      【解决方案2】:

      如果要替换所有 99s ,请尝试使用正则表达式

      >>> d.astype(str).replace('99','*',regex=True)

          A   B   C
      0   a   1   abcd*
      1   b   2   4
      2   *   *   5
      

      【讨论】:

        【解决方案3】:

        使用numpys 字符函数

        d.values[:] = np.core.defchararray.replace(d.values.astype(str), '99', '*')
        d
        
           A  B      C
        0  a  1  abcd*
        1  b  2      4
        2  *  *      5
        

        幼稚时间测试

        【讨论】:

          【解决方案4】:

          问题是 A 列和 B 列中的值 99 属于不同类型:

          >>> type(d.loc[2,"A"])
          <class 'int'>
          >>> type(d.loc[2,"B"])
          <class 'str'>
          

          您可以通过df.astype() 将数据框转换为字符串类型,然后替换,结果是:

          >>> d.astype(str).replace("99","*")
             A  B       C
          0  a  1  abcd99
          1  b  2       4
          2  *  *       5
          

          编辑:使用正则表达式是其他答案给出的正确解决方案。由于某种原因,我错过了您的 DataFrame 中的 abcd*。

          将其留在这里,以防万一它对其他人有帮助。

          【讨论】:

            猜你喜欢
            • 2017-07-08
            • 2021-09-07
            • 2017-11-17
            • 2018-08-01
            • 1970-01-01
            • 2020-05-13
            • 2022-10-13
            • 2018-02-11
            相关资源
            最近更新 更多