【发布时间】:2016-06-09 17:58:30
【问题描述】:
我想知道如何将我的爬取结果导出到我爬取的每个不同城市的多个 csv 文件中。不知何故,我遇到了墙壁,没有正确的方法来解决它。
这是我的代码:
import requests
from bs4 import BeautifulSoup
import csv
user_agent = {'User-agent': 'Chrome/43.0.2357.124'}
output_file= open("TA.csv", "w", newline='')
RegionIDArray = [187147,187323,186338]
dict = {187147: 'Paris', 187323: 'Berlin', 186338: 'London'}
already_printed = set()
for reg in RegionIDArray:
for page in range(1,700,30):
r = requests.get("https://www.tripadvisor.de/Attractions-c47-g" + str(reg) + "-oa" + str(page) + ".html")
soup = BeautifulSoup(r.content)
g_data = soup.find_all("div", {"class": "element_wrap"})
for item in g_data:
header = item.find_all("div", {"class": "property_title"})
item = (header[0].text.strip())
if item not in already_printed:
already_printed.add(item)
print("POI: " + str(item) + " | " + "Location: " + str(dict[reg]))
writer = csv.writer(output_file)
csv_fields = ['POI', 'Locaton']
if g_data:
writer.writerow([str(item), str(dict[reg])])
我的目标是为巴黎、柏林和伦敦获取三个独立的 CSV 文件,而不是在一个大的 csv 文件中获取所有结果。
你们能帮帮我吗?感谢您的反馈:)
【问题讨论】:
-
您可能想查看 TripAdvisor Content API:developer-tripadvisor.com/content-api
-
感谢您的反馈。我很清楚这一点,但我想自己爬。不知何故,它比使用 API 更有动力;)
-
如果要根据内容写入三个不同的文件,则必须有三个独立的
csv.writer,测试内容并写入正确的文件,具体取决于测试结果。 -
感谢您的反馈。但我仍然不知道如何将城市分成三个不同的 csv 文件。我明白你的意思是拥有三个单独的 csv 编写器,但我应该如何拆分它
-
不要使用 dict 作为任何东西的名称。你隐藏了内置类型的字典。将其重命名为 RegionIDArray,因为您不需要它。遍历 dict 无论如何都会在您的“for reg in Region...”中为您提供密钥......”
标签: python csv beautifulsoup web-crawler