【发布时间】:2018-09-23 23:36:36
【问题描述】:
我认为这应该很简单,但现在是星期五下午,我的大脑不太清楚。
我正在编写一个小文件解析,下面的代码将一组字符串转换为一个数据框,将字符串拆分。
以下是一些示例字符串:
1. NC_002523_1 Serratia entomophila plasmid pADAP, complete sequence.
2. NZ_CM003366_0 Pantoea ananatis strain CFH 7-1 plasmid CFH1-7plasmid2, whole genome shotgun sequence.
3. NZ_CP014491_0 Escherichia coli strain G749 plasmid pG749_3, complete sequence.
4. NC_015062_0 Rahnella sp. Y9602 plasmid pRAHAQ01, complete sequence.
我没想到在第 4 个条目中 sp 之后会出现 .,正如您在下面的代码中看到的那样,我拆分 . 以获得排名的第一个整数。因此,我得到一个 ValueError,表明列数比预期的多。
# Define the column headers for the section since the file's are too verbose and ambiguous
SigHit.Columns = ["Rank", "ID", "Description"]
# Store the table of loci and associated data (tab separated, removing last blank column.
# Use StringIO object to imitate a file, which means that we can use read_table and have the dtypes
# assigned automatically (necessary for functions like min() to work correctly on integers)
SigHit.Table = pd.read_table(
io.StringIO(u'\n'.join([row.rstrip('.') for row in sighits_section])),
sep='\.|\t',
engine='python',
names=SigHit.Columns)
我能想到的最简单的解决方案(直到其他一些边缘情况破坏它)是替换每个.,除了第一次出现。如何做到这一点?
我看到 maxreplace argument 到 .replace 有一个,但这与我想要的相反,只会替换第一个实例。
有什么建议吗? (更强大的解析方法也是一个有效的选择,但我必须更改的代码越少越好)。
【问题讨论】:
-
您的
1.等后面是制表符还是空格? -
您可以在前 2 个空格处
split您的数据。row.split(None,2)而不是row.rstrip。这会将空格和制表符分隔成长度为 3 的列表 -
1.后跟一个空格,字母数字 ID 后跟一个制表符,然后是描述(我没有编写输出文件的软件:P)