【问题标题】:Extract multiple protein sequences from a Protein Data Bank along with Secondary Structure从蛋白质数据库中提取多个蛋白质序列以及二级结构
【发布时间】:2017-04-03 14:04:18
【问题描述】:

我想从任何蛋白质数据库(例如 RCSB)中提取蛋白质序列及其相应的二级结构。我只需要短序列和它们的二级结构。类似的,

ATRWGUVT     Helix

即使序列很长也没关系,但我希望在末尾有一个标签来表示它的二级结构。是否有任何编程工具或任何可用的东西。

如上所示,我只需要这么少的信息。我怎样才能做到这一点?

【问题讨论】:

  • 请参阅this thread,了解如何编写一个好的 StackOverflow 问题。

标签: bioinformatics biopython protein-database


【解决方案1】:

您可以使用DSSP

DSSP 的输出在“explanation”下进行了详细说明。输出的非常简短的摘要是:

H = α-helix
B = residue in isolated β-bridge
E = extended strand, participates in β ladder
G = 3-helix (310 helix)
I = 5 helix (π-helix)
T = hydrogen bonded turn
S = bend

【讨论】:

    【解决方案2】:
    from Bio.PDB import *
    from distutils import spawn
    

    提取序列:

    def get_seq(pdbfile):
     p = PDBParser(PERMISSIVE=0)
     structure = p.get_structure('test', pdbfile)
     ppb = PPBuilder()
     seq = ''
     for pp in ppb.build_peptides(structure):
      seq += pp.get_sequence()
    
     return seq
    

    如前所述,用 DSSP 提取二级结构:

    def get_secondary_struc(pdbfile):
        # get secondary structure info for whole pdb.
        if not spawn.find_executable("dssp"):
            sys.stderr.write('dssp executable needs to be in folder')
            sys.exit(1)
        p = PDBParser(PERMISSIVE=0)
        ppb = PPBuilder()
        structure = p.get_structure('test', pdbfile)
        model = structure[0]
        dssp = DSSP(model, pdbfile)
        count = 0
        sec = ''
        for residue in model.get_residues():
            count = count + 1
            # print residue,count
            a_key = list(dssp.keys())[count - 1]
            sec += dssp[a_key][2]
        print sec
        return sec
    

    这应该打印序列和二级结构。

    【讨论】:

      猜你喜欢
      • 2012-06-27
      • 2013-01-16
      • 1970-01-01
      • 2017-01-28
      • 1970-01-01
      • 2013-11-27
      • 2021-02-16
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多