【问题标题】:HTML scraper output stuck in utf-8HTML 刮板输出卡在 utf-8 中
【发布时间】:2017-04-10 11:05:05
【问题描述】:

我正在为一些中文文档制作刮板。作为项目的一部分,我试图将文档的正文刮到一个列表中,然后从该列表中编写文档的 html 版本(最终版本将包括元数据和文本,以及一个充满文档的单个 html 文件)。

我已经设法将文档的正文刮成一个列表,然后使用该列表的内容创建一个新的 HTML 文档。当我将列表输出到 csv 时,我什至可以查看内容(到目前为止一切都很好......)。 不幸的是,输出的 HTML 文档都是"\u6d88\u9664\u8d2b\u56f0\u3001\"。

有没有办法对输出进行编码,以免发生这种情况?我是否只需要长大并真正抓取页面(通过<p> 解析和组织它<p>,而不是按原样复制所有现有的HTML)然后逐个元素地构建新的HTML页面?

任何想法将不胜感激。

from bs4 import BeautifulSoup
import urllib
#csv is for the csv writer
import csv

#initiates the dictionary to hold the output

holder = []

#this is the target URL
target_url = "http://www.gov.cn/zhengce/content/2016-12/02/content_5142197.htm"

data = []

filename = "fullbody.html"
target = open(filename, 'w')

def bodyscraper(url):
    #opens the url for read access
    this_url = urllib.urlopen(url).read()
    #creates a new BS holder based on the URL
    soup = BeautifulSoup(this_url, 'lxml')

    #finds the body text
    body = soup.find('td', {'class':'b12c'})


    data.append(body)

    holder.append(data)

    print holder[0]
    for item in holder:
        target.write("%s\n" % item)

bodyscraper(target_url)


with open('bodyscraper.csv', 'wb') as f:
    writer = csv.writer(f)
    writer.writerows(holder)

【问题讨论】:

    标签: python html python-2.7 web-scraping character-encoding


    【解决方案1】:

    由于源 htm 是 utf-8 编码的,当使用 bs 时,只需解码 urllib 返回的内容即可。我测试过 html 和 csv 输出都会显示汉字,这里是修改后的代码:

    from bs4 import BeautifulSoup
    import urllib
    #csv is for the csv writer
    import csv
    
    #initiates the dictionary to hold the output
    
    holder = []
    
    #this is the target URL
    target_url = "http://www.gov.cn/zhengce/content/2016-12/02/content_5142197.htm"
    
    data = []
    
    filename = "fullbody.html"
    target = open(filename, 'w')
    
    def bodyscraper(url):
        #opens the url for read access
        this_url = urllib.urlopen(url).read()
        #creates a new BS holder based on the URL
        soup = BeautifulSoup(this_url.decode("utf-8"), 'lxml') #decoding urllib returns
    
        #finds the body text
        body = soup.find('td', {'class':'b12c'})
        target.write("%s\n" % body) #write the whole decoded body to html directly
    
    
        data.append(body)
    
        holder.append(data)
    
    
    bodyscraper(target_url)
    
    
    with open('bodyscraper.csv', 'wb') as f:
        writer = csv.writer(f)
        writer.writerows(holder)
    

    【讨论】:

    • 这适用于 csv,但 html 仍然给了我(一种新的)垃圾输出。但是,将 '' 行添加到 html 的开头给了我想要的结果。
    猜你喜欢
    • 1970-01-01
    • 2018-04-04
    • 2015-08-16
    • 2015-06-04
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多