【问题标题】:Is there a way to split row values into separate columns with Pandas?有没有办法使用 Pandas 将行值拆分为单独的列?
【发布时间】:2019-07-27 03:30:42
【问题描述】:

我目前有一个 pandas 数据框,内容如下:

0   (dev_id='A', accon_time='B', start_time='C',end_time='D')
1   (dev_id='E', accon_time='F', start_time='G',end_time='H')
2   (dev_id='I', accon_time='J', start_time='K',end_time='L')

这个数据框的当前形状是 (574,1),而我实际上希望它是 (574,4),其中每行中的 4 个逗号分隔值中的每一个实际上都拆分为 4 个单独的列。

有什么办法吗?

  • 此数据来自 SQL Alchemy 查询

我尝试先将查询转换为 pandas 系列,然后使用 Series.str.split,但结果与原始数据框相同。

ser = pd.Series(qry)
ser.str.rsplit(pat=",", n=4, expand=True)
print(ser)
df = pd.DataFrame(data=ser)
print(df)

这是我用来查询数据的:

class Trip(Base):
    __tablename__ = 'trip'
    dev_id = Column(String(50), primary_key=True)
    accon_time = Column(Integer)
    start_time = Column(Integer)
    end_time = Column(Integer)

    def __repr__(self):
        return "(dev_id='%s', accon_time='%s', start_time='%s',end_time='%s')" 
          % (self.dev_id, self.accon_time, self.start_time, self.end_time)

qry = session.query(Trip).\
        filter(Trip.accon_time.between(20190620000000, 20190621000000)).\
        filter(Trip.start_time <= 20190620145813).\
        filter(Trip.end_time <= 20190620151400).\
        filter(Trip.end_time >= 20190620145600)

这会返回一个像这样的列表:

(dev_id='A', accon_time='B', start_time='C',end_time='D'),(dev_id='E', accon_time='F', start_time='G',end_time='H'),(dev_id='I', accon_time='J', start_time='K',end_time='L')

将我的查询结果转换为 pandas 数据框

df = pd.DataFrame(data=qry)
print(df)

【问题讨论】:

  • 嘿只是为了确定你在df中的元素是一个字符串吗?即“(dev_id='A', accon_time='B', start_time='C',end_time='D')”?
  • 你应该在你的例子中尝试ser = ser.str.rsplit(pat=",", n=4, expand=True)而不是ser.str.rsplit(pat=",", n=4, expand=True)
  • 老实说,您可能应该将数据作为数据提取出来,而不是通过__repr__ 然后尝试解压您的字符串。
  • @kkawabat 刚刚尝试过,它返回一个空系列,如下所示:0 0 NaN 1 NaN

标签: python pandas


【解决方案1】:

在您的解析示例中ser.str.rsplit(pat=",", n=4, expand=True) 返回 ser 的输出,您需要捕获输出,否则它不会做任何事情

试试这个解析:

qry =   ["(dev_id='A', accon_time='B', start_time='C',end_time='D')",
"(dev_id='E', accon_time='F', start_time='G',end_time='H')",
"(dev_id='I', accon_time='J', start_time='K',end_time='L')"]
ser = pd.Series(qry)
df = ser.apply(lambda x: pd.Series([val.split('=')[1] for val in x[1:-1].split(',')]))
df.columns = ['dev_id', 'accon_time', 'start_time', 'end_time']

对于 ser .appy() 的每一行,我获取字符串并删除括号 x[1:-1],然后用逗号分隔 .split(','),这将给我一个键值文字列表(即 ["dev_id='A'", " accon_time='B'", " start_time='C'", "end_time='D'"])。然后对于每个文字,我将其拆分为 '=' 并返回第二个元素,即实际值 .split('=')[1]。

如果您不希望元素中的“'”在末尾使用.strip('\'') 将其剥离

   ser = ser.apply(lambda x:[val.split('=')[1].strip('\'') for val in x[1:-1].split(',')])

输出:

  dev_id accon_time start_time end_time
0    'A'        'B'        'C'      'D'
1    'E'        'F'        'G'      'H'
2    'I'        'J'        'K'      'L'

【讨论】:

  • 刚刚尝试过,它引发了类型错误。 “TypeError:'Trip' 对象不可下标”。
  • 我认为这与您的 Trip 课程有关。我假设 qry 是 str 输出的列表。试试 qry = [repr(x) for x in qry]?
猜你喜欢
  • 2019-09-02
  • 2022-06-11
  • 1970-01-01
  • 2016-11-08
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
相关资源
最近更新 更多