【发布时间】:2014-05-03 08:10:32
【问题描述】:
我知道这个话题已经讨论过几次,但我找不到适用于我的案例。我不是一个有经验的计算机用户,请记住这一点,虽然我可以玩 bash、R 并且可能也运行 perl 脚本。仅供参考 - 我在我的机器上运行 Ubuntu。
我想做的是将以下网页http://www.genome.jp/kegg-bin/get_htext?br08902.keg的可展开列表(请使用“一键模式”完全展开)转换为表格或csv格式,其中每个级别的缩进到一个单独的列。
对于分组在其下方的所有元素重复父类别也不会那么糟糕。类似于下面我为页面的前几行手动制作的选项卡。
Pathways and Ontologies Pathways br08901 KEGG pathway maps
Pathways and Ontologies Functional hierarchies br08902 BRITE functional hierarchies
Genes and Proteins Orthologs and modules ko00001 KEGG Orthology (KO)
Genes and Proteins Orthologs and modules ko00002 KEGG pathway modules
Genes and Proteins Orthologs and modules ko00003 KEGG modules and reaction modules
Genes and Proteins Protein families: metabolism ko01000 Enzymes
Genes and Proteins Protein families: metabolism ko01001 Protein kinases
Genes and Proteins Protein families: metabolism ko01009 Protein phosphatases and associated proteins
Genes and Proteins Protein families: metabolism ko01002 Peptidases
Genes and Proteins Protein families: metabolism ko01003 Glycosyltransferases
Genes and Proteins Protein families: metabolism ko01005 Lipopolysaccharide biosynthesis proteins
Genes and Proteins Protein families: metabolism ko01004 Lipid biosynthesis proteins
提前致谢!
【问题讨论】:
-
...到目前为止,您尝试解析什么?哪里不行?