【问题标题】:Retrieve splitted values in a XML file检索 XML 文件中的拆分值
【发布时间】:2021-11-22 14:20:21
【问题描述】:

我有一个这样的 XML 文件:

data = '''<dbReference type="PDB" id="6LVN">
<property type="method" value="X-ray"/>
<property type="resolution" value="2.47 A"/>
<property type="chains" value="A/B/C/D=1168-1203"/>
</dbReference>
<dbReference type="PDB" id="6LXT">
<property type="method" value="X-ray"/>
<property type="resolution" value="2.90 A"/>
<property type="chains" value="A/B/C/D/E/F=910-988, A/B/C/D/E/F=1162-1206"/>
</dbReference>
<dbReference type="PDB" id="6LXV">
<property type="method" value="X-ray"/>
<property type="resolution" value="4.90 A"/>
<property type="chains" value="A/B/C/=210-488, A/B/C/=510-688, A/B/C=800-960"/>
</dbReference>'''

我想检索所有长度值。我这样做的代码:

from bs4 import BeautifulSoup

xml_file = BeautifulSoup(data, 'lxml')
pdbs_xml = xml_file.find_all('dbreference', {'type': 'PDB'})
if len(pdbs_xml) != 0:
    for item in pdbs_xml:
        if item.find('property'):
            id_ = item['id']
            chains = item.find('property', {'type': 'chains'})
            chains2 = chains['value']
            if chains2.find(",")!= -1:
                count = chains2.count(',')
                if count >= 2:
                    chains = chains['value'].split('=')[count]
                    chains = chains.split(',')[0]
                    first_aa = chains.split('-')[0]
                    last_aa = chains.split('-')[1]
                    size_pdb = int(last_aa) - int(first_aa)
                else:
                    chains = chains['value'].split('=')[2]
                    first_aa = chains.split('-')[0]
                    last_aa = chains.split('-')[1]
                    size_pdb = int(last_aa) - int(first_aa)
            else:
                chains = chains['value'].split('=')[1]
                first_aa = chains.split('-')[0]
                last_aa = chains.split('-')[1]
                size_pdb = int(last_aa) - int(first_aa)

如您所见,有些值是分开的。理论上,我可以创建一个语句来预测每种可能性并检索所有情况(我知道我的代码现在并没有完全做到这一点),但是有更好的方法来实现这一点。所以,欢迎提出任何建议。

【问题讨论】:

  • 鉴于data xml - 应该是什么输出?什么是所有长度值。?
  • 你能举例说明“长度值”是什么意思吗?
  • 我的想法是拥有每个 PDB 的大小。例如,PDB '6LVN' 的长度是 35 (1203 - 1168)。 PDB '6LXT' 的长度为 122 ([988 - 910] + [1206 - 1162])

标签: python python-3.x xml


【解决方案1】:

给你[请注意,下面的代码不需要任何外部库]

import xml.etree.ElementTree as ET
from collections import defaultdict

data = '''<r><dbReference type="PDB" id="6LVN">
<property type="method" value="X-ray"/>
<property type="resolution" value="2.47 A"/>
<property type="chains" value="A/B/C/D=1168-1203"/>
</dbReference>
<dbReference type="PDB" id="6LXT">
<property type="method" value="X-ray"/>
<property type="resolution" value="2.90 A"/>
<property type="chains" value="A/B/C/D/E/F=910-988, A/B/C/D/E/F=1162-1206"/>
</dbReference>
<dbReference type="PDB" id="6LXV">
<property type="method" value="X-ray"/>
<property type="resolution" value="4.90 A"/>
<property type="chains" value="A/B/C/=210-488, A/B/C/=510-688, A/B/C=800-960"/>
</dbReference></r>'''

sizes = defaultdict(int)
root = ET.fromstring(data)
for ref in root.findall('.//dbReference'):
    pdb = ref.attrib['id']
    chains = ref.find('property[@type="chains"]')
    value = chains.attrib['value']
    parts = value.split(',')
    for part in parts:
        left,right = part.split('=')
        _left,_right = right.split('-')
        sizes[pdb] += int(_right)- int(_left)

print(sizes)

输出

defaultdict(<class 'int'>, {'6LVN': 35, '6LXT': 122, '6LXV': 616})

【讨论】:

  • 不鼓励只回答代码。还请描述defaultdict的用途。
【解决方案2】:

在单个 xpath 中的定义条件下查找所有 @id, @value 属性,并在列表中的偶数位置处理 @value。 数学直接在@value 和eval 上完成,并乘以-1,因为结果将是负数。无需拆分和交换。

获取@id xpath 部分
//dbReference[@type="PDB" and property[@type="chains" and string-length(@value)&gt;0]]/@id

获取@value xpath 部分
//dbReference[@type="PDB" and property[@type="chains" and string-length(@value)&gt;0]]/property[@type="chains"]/@value

from lxml import etree
tree = etree.parse('test.xml')

steps = tree.xpath('//dbReference[@type="PDB" and property[@type="chains" and string-length(@value)>0]]/@id | //dbReference[@type="PDB" and property[@type="chains" and string-length(@value)>0]]/property[@type="chains"]/@value')

for i in range(len(steps)):
    # @value appear on even positions
    if (i%2) != 0:
        items = steps[i].split(',')
        s=0
        for item in items:
            values = item.split('=')
            s+=eval(values[1])*(-1)
            
        print(steps[i-1],s)

结果:

6LVN 35
6LXT 122
6LXV 616

【讨论】:

    【解决方案3】:

    使用bs4 查找值,使用 regex 获取区间,然后使用内置函数对每个区间的差异求和。

    data = '''<dbReference type="PDB" id="6LVN">
    <property type="method" value="X-ray"/>
    <property type="resolution" value="2.47 A"/>
    <property type="chains" value="A/B/C/D=1168-1203"/>
    </dbReference>
    <dbReference type="PDB" id="6LXT">
    <property type="method" value="X-ray"/>
    <property type="resolution" value="2.90 A"/>
    <property type="chains" value="A/B/C/D/E/F=910-988, A/B/C/D/E/F=1162-1206"/>
    </dbReference>
    <dbReference type="PDB" id="6LXV">
    <property type="method" value="X-ray"/>
    <property type="resolution" value="4.90 A"/>
    <property type="chains" value="A/B/C/=210-488, A/B/C/=510-688, A/B/C=800-960"/>
    </dbReference>'''
    
    from bs4 import BeautifulSoup
    import re
    
    xml_file = BeautifulSoup(data, 'lxml')
    
    output = {}
    for tag in xml_file.find_all(type="chains", value=True):
        interval = re.findall(r'([0-9]+-[0-9]+)', tag['value'])
        output[tag.parent['id']] = sum(map(lambda p: abs(int(p[1])-int(p[0])), (map(lambda p: p.split('-'), interval))))
    
    print(output)
    

    输出

    {'6LVN': 35, '6LXT': 122, '6LXV': 616}
    

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2011-09-12
      • 2023-04-02
      • 1970-01-01
      • 1970-01-01
      • 2018-02-15
      相关资源
      最近更新 更多