【问题标题】:Converting python webcrawler to 3.4 from 2.7将 python webcrawler 从 2.7 转换为 3.4
【发布时间】:2014-09-19 17:29:49
【问题描述】:

对于这段代码,我正在将一个正常工作的 python webcrawler 从 2.7 转换为 3.4。我做了一些修改,但运行时仍然出现错误:

Traceback (most recent call last):
  File "Z:\testCrawler.py", line 11, in <module>
    for i in re.findall('''href=["'](.[^"']+)["']''', urllib.request.urlopen(myurl).read(), re.I):
  File "C:\Python34\lib\re.py", line 206, in findall
    return _compile(pattern, flags).findall(string)
TypeError: can't use a string pattern on a bytes-like object

这是代码本身,如果你看到语法错误,请告诉我。

#! C:\python34

import re
import urllib.request

textfile = open('depth_1.txt','wt')
print ("Enter the URL you wish to crawl..")
print ('Usage  - "http://phocks.org/stumble/creepy/" <-- With the double quotes')
myurl = input("@> ")
for i in re.findall('''href=["'](.[^"']+)["']''', urllib.request.urlopen(myurl).read(), re.I):
        print (i)  
        for ee in re.findall('''href=["'](.[^"']+)["']''', urllib.request.urlopen(i).read(), re.I):
                print (ee)
                textfile.write(ee+'\n')
textfile.close()

【问题讨论】:

  • 您需要将来自read 的响应解码为str。
  • 尽管请 - 使用 HTML 解析器来解析 html,而不是正则表达式。

标签: python python-2.7 web-crawler python-3.4


【解决方案1】:

改变

urllib.request.urlopen(myurl).read()

例如

urllib.request.urlopen(myurl).read().decode('utf-8')

这里发生的情况是 .read() 返回 bytes 而不是 str 就像在 python 2.7 中一样,因此必须使用某种编码对其进行解码。

【讨论】:

    猜你喜欢
    • 2016-11-19
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2014-09-13
    • 2016-12-20
    • 2016-05-16
    • 2018-01-10
    • 1970-01-01
    相关资源
    最近更新 更多