【发布时间】:2018-02-15 05:59:51
【问题描述】:
我正在尝试.split() 具有多个值的表中的单元格。然后我想将这些拆分值堆叠成一列。
我不断收到:AttributeError: 'DataFrame' object has no attribute 'str'
- 某些列将具有相同的名称/标签
- 值将在 str、flt、int 等之间混合
- 会有缺失值
- 我已将此表另存为 .csv
示例表:
(原表)
List , A, A , B , B , A , C
row 1,joey,mike,henry,albert ,sherru,tomkins
row 2, ,pig|soap , ,123, , ,
row 3,yes, , , and|5.3|7, , ,
row 4, ,new york|up, , , , ,
row 5,bubbles, ,movie, , , ,
(修改后的表格)
List | Value | Category
row 1,joey, A
row 1,mike,A
row 1,henry,B
row 1,albert,B
row 1,sherru,A
row 1,tomkins,C
row 2,pig,A
row 2,soap,A
row 2,123,B
row 3,yes,A
row 3,and,B
row 3,5.3,B
...
row 5,movie,B
这是我正在使用的代码,我是 python/pandas 的新手,所以它不是很好:
import pandas as pd
df = pd.read_csv('test.csv')
df2 = df.A.str.split('|').apply(pd.series)
df2.index = df.set_index([List]).index
df2.stack().reset_index([List])
【问题讨论】:
-
您的 CSV 文件无效:不同的行有不同的列数。你能适当地格式化它吗?
标签: python pandas object split attributes