【问题标题】:Getting list of urls form wikipedia page从维基百科页面获取 url 列表
【发布时间】:2019-07-28 14:44:52
【问题描述】:

我有一份财富 500 强公司的名单。 这是一个例子[Abbott Laboratories,Progressive,Arrow Electronics,Kraft Heinz Plains GP Holdings,Gilead Sciences,Mondelez International,Northrop Grumman]

现在我想从 Wikipedia 获取列表中每个元素的完整 url。

for example, after searching the name on Google or Wikipedia, 
it should give me back list of all wikipedia urls like: 

https://en.wikipedia.org/wiki/Abbott_Laboratories(这只是一个例子)

【问题讨论】:

  • 到目前为止,您所拥有的是......发布代码。如果您没有任何代码,请尝试编写一些代码然后发布。

标签: python web-scraping scrapy


【解决方案1】:

最大的问题是寻找可能的网站,并且只选择属于公司的网站。

一种有点错误的方法是只是将公司名称附加到 wiki url 并希望它有效。这导致 a) 它可以工作(如 Abbott Laboratories),b) 它产生一个页面,但不是正确的页面(Progressive,应该是 Progressive_Corporation)或 c) 它根本不产生任何结果。

companies = [
    "Abbott Laboratories", "Progressive", "Arrow Electronics", "Kraft Heinz Plains GP Holdings", "Gilead Sciences",
    "Mondelez International", "Northrop Grumman"
]

url = "https://en.wikipedia.org/wiki/%s"

for company in companies:
    print(url % company.replace(" ", "_"))

另一个(更好的)选择是使用 wikipedia 包 (https://pypi.org/project/wikipedia/) 及其内置的搜索功能。选择正确站点的问题仍然存在,因此您基本上必须手动执行此操作(或创建一个很好的自动选择,例如搜索“公司”一词)

companies = [
    "Abbott Laboratories", "Progressive", "Arrow Electronics", "Kraft Heinz Plains GP Holdings", "Gilead Sciences",
    "Mondelez International", "Northrop Grumman"
]

import wikipedia
for company in companies:
    options = wikipedia.search(company)
    print(company, options)

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2019-09-15
    相关资源
    最近更新 更多