【问题标题】:Python TypeError Traceback (most recent call last)Python TypeError Traceback(最近一次调用最后一次)
【发布时间】:2021-01-29 03:37:02
【问题描述】:

我正在尝试构建一个爬虫,我想打印该页面上的所有链接 我正在使用 Python 3.5

这是我的代码

import requests
from bs4 import BeautifulSoup
def crawler(link):
    source_code = requests.get(link)
    source_code_string = str(source_code)
    source_code_soup = BeautifulSoup(source_code_string,'lxml')
    for item in source_code_soup.findAll("a"):
        title = item.string
        print(title)

crawler("https://www.youtube.com/watch?v=pLHejmLB16o")

但我得到这样的错误

TypeError                                 Traceback (most recent call last)
<ipython-input-13-9aa10c5a03ef> in <module>()
----> 1 crawler('http://archive.is/DPG9M')

TypeError: 'module' object is not callable

【问题讨论】:

  • 你试过重命名你的crawler方法吗?
  • 是的,我把“crawler”改成了“cat”,还是一样的错误

标签: python web-crawler


【解决方案1】:

如果你打算只打印链接的标题,你犯了一个小错误,替换行:

source_code_string = str(source_code)

使用

source_code_string = source_code.text 

除此之外,代码看起来很好并且正在运行。 让我们调用文件 web_crawler_v1.py

import requests
from bs4 import BeautifulSoup
def crawler(link):
    source_code = requests.get(link)
    source_code_string = source_code.text 
    source_code_soup = BeautifulSoup(source_code_string,'lxml')
    for item in source_code_soup.findAll("a"):
        title = item.string
        print(title)


crawler("https://www.youtube.com/watch?v=pLHejmLB16o")

关于那个错误,如果你像这样正确调用文件,你不应该得到那个错误

python3 wen_crawler_v1.py

【讨论】:

  • 我还有一个问题。我正在使用 Jupyter Notebook,当我在笔记本 RootPython 中键入这些代码时,它可以工作。但是当我创建一个新的py文件(test.py)并在其中放入相同的代码并输入“import test.py”然后它不起作用,你知道为什么吗?
  • 我已经对其进行了测试并且可以正常工作,但是该文件与您调用它的位置相同吗?因为如果不是,那可能是个问题。
【解决方案2】:

代替

source_code = requests.get(link)

使用:

source_code = requests.get(link, verify = False)

您将收到 HTTPS 警告,但代码将执行

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 2022-08-19
    • 2020-03-24
    • 1970-01-01
    • 2016-05-31
    • 2022-01-26
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多