【发布时间】:2017-04-10 11:05:05
【问题描述】:
我正在为一些中文文档制作刮板。作为项目的一部分,我试图将文档的正文刮到一个列表中,然后从该列表中编写文档的 html 版本(最终版本将包括元数据和文本,以及一个充满文档的单个 html 文件)。
我已经设法将文档的正文刮成一个列表,然后使用该列表的内容创建一个新的 HTML 文档。当我将列表输出到 csv 时,我什至可以查看内容(到目前为止一切都很好......)。
不幸的是,输出的 HTML 文档都是"\u6d88\u9664\u8d2b\u56f0\u3001\"。
有没有办法对输出进行编码,以免发生这种情况?我是否只需要长大并真正抓取页面(通过<p> 解析和组织它<p>,而不是按原样复制所有现有的HTML)然后逐个元素地构建新的HTML页面?
任何想法将不胜感激。
from bs4 import BeautifulSoup
import urllib
#csv is for the csv writer
import csv
#initiates the dictionary to hold the output
holder = []
#this is the target URL
target_url = "http://www.gov.cn/zhengce/content/2016-12/02/content_5142197.htm"
data = []
filename = "fullbody.html"
target = open(filename, 'w')
def bodyscraper(url):
#opens the url for read access
this_url = urllib.urlopen(url).read()
#creates a new BS holder based on the URL
soup = BeautifulSoup(this_url, 'lxml')
#finds the body text
body = soup.find('td', {'class':'b12c'})
data.append(body)
holder.append(data)
print holder[0]
for item in holder:
target.write("%s\n" % item)
bodyscraper(target_url)
with open('bodyscraper.csv', 'wb') as f:
writer = csv.writer(f)
writer.writerows(holder)
【问题讨论】:
标签: python html python-2.7 web-scraping character-encoding