【问题标题】:Graph.create_png error TypeError: sequence item 0: expected str instance, bytes foundGraph.create_png 错误类型错误:序列项 0:预期的 str 实例,找到的字节
【发布时间】:2018-08-28 19:58:18
【问题描述】:

我正在尝试根据我通过抓取收集的一些链接创建一个图表。如果我只查找 1 个标签,一切正常,但如果我尝试多个标签,我会收到以下错误:

File "c:\Users\qnour\Desktop\Programming\Python\GettingStarted\Wiki_Scraping.py", line 89, in <module>
    main()
  File "c:\Users\qnour\Desktop\Programming\Python\GettingStarted\Wiki_Scraping.py", line 32, in main
    drawGraph(graph)
  File "c:\Users\qnour\Desktop\Programming\Python\GettingStarted\Wiki_Scraping.py", line 85, in drawGraph
    graph.write_png('wiki_graph.png', prog='dot')
  File "C:\Users\qnour\AppData\Local\Programs\Python\Python36\lib\site-packages\pydot\__init__.py", line 1807, in <lambda>
    lambda path, f=frmt, prog=self.prog : self.write(path, format=f, prog=prog))
  File "C:\Users\qnour\AppData\Local\Programs\Python\Python36\lib\site-packages\pydot\__init__.py", line 1909, in write
    dot_fd.write(self.create(prog, format))
  File "C:\Users\qnour\AppData\Local\Programs\Python\Python36\lib\site-packages\pydot\__init__.py", line 2013, in create
    stderr_output = ''.join(stderr_output)
TypeError: sequence item 0: expected str instance, bytes found

这里是代码:

import bs4 as bs
import urllib.request
import pydot
import graphviz
from IPython.display import Image, display
import os 



def viewPydot(pdot):
    plt = Image(pdot.create_png())
    display(plt)

global sauce
global soup

def main():
    global sauce
    global soup
    firstElement = input("Please select the first element : ")
    bareLink = "https://en.wikipedia.org/wiki/"
    sectionNumber = calculateSection(bareLink+firstElement)
    if (sectionNumber == -1):
    print("no see also section ! ")
    exit(0)
url = "https://en.wikipedia.org/w/api.php?action=parse&prop=links&page={}&section={}".format(firstElement, sectionNumber)
sauce = urllib.request.urlopen(url).read()
soup = bs.BeautifulSoup(sauce, 'lxml')
listUrl = gatherLinks()
fullUrl = createNewLinks(listUrl)
graph = createGraph(listUrl, firstElement)
drawGraph(graph)
#TODO 
#ADD THE NEW LINKS TO THE GRAPH

def createNewLinks(listUrl):
    bareLink = "https://en.wikipedia.org/wiki/"
    fullUrl = []
    for item in listUrl:
        fullUrl.append(bareLink + item)
    return fullUrl


def gatherLinks():
    header = soup.find_all("span", class_="s2")
    found_star = False
    listUrl = []

    for item in header:
        if (found_star):
            print(item.text)
            listUrl.append(item.text.split('"')[1])
            found_star = False
        else:
            if (item.text == '"*"'):
                found_star = True
    return listUrl

def createGraph(listUrl, firstElement):
graph = pydot.Dot(graph_type='graph')

for graphEdge in listUrl:
    edge = pydot.Edge(firstElement, graphEdge)
    graph.add_edge(edge)
return graph


def calculateSection(url):
source = urllib.request.urlopen(url).read()
sectionSoup = bs.BeautifulSoup(source, 'lxml')

sections = sectionSoup.findAll(["h2", "h3", "h4"])

for number, item in enumerate(sections):
    print(item.text)
    if (item.text == "See also" or item.text == "See also[edit]"):
        print(number)
        return number
return -1



def drawGraph(graph):
graph.write_png('wiki_graph.png', prog='dot')
Image('wiki_graph.png')

    if __name__=="__main__":
main()

让我烦恼的是这种变化:

sections = sectionSoup.findAll(["h2", "h3", "h4"])

作者:

sections = sectionSoup.findAll("h2")

一切正常,但我需要检查所有 3 个标签。

【问题讨论】:

    标签: python python-3.x web-scraping graphviz pydot


    【解决方案1】:

    据我所知,您需要类似的东西:

    sections = sectionSoup.findAll("h2")
    sections += sectionSoup.findAll("h3")
    sections += sectionSoup.findAll("h4")
    

    【讨论】:

    • 问题是这样做,我们失去了订单
    • 在这种情况下 O 认为你应该循环遍历元素
    • 我的意思是,通过搜索 h2,然后是 h3,然后是 h4,html 页面中的顺序丢失了:/
    • 我确实是这样理解的,所以您可能必须浏览标签,如果感兴趣,请将其添加到 sections
    猜你喜欢
    • 2017-02-02
    • 2015-11-11
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2016-10-31
    • 1970-01-01
    • 1970-01-01
    • 2021-05-17
    相关资源
    最近更新 更多