【发布时间】:2020-12-15 14:59:37
【问题描述】:
我用这个脚本来抓取这个网站上的评论:https://fr.trustpilot.com/review/jardiland.com
import requests
from requests import get
from bs4 import BeautifulSoup
import pandas as pd
import numpy as np
urls = ["https://fr.trustpilot.com/review/jardiland.com",
"https://fr.trustpilot.com/review/jardiland.com?page=2",
"https://fr.trustpilot.com/review/jardiland.com?page=3",
"https://fr.trustpilot.com/review/jardiland.com?page=4",
"https://fr.trustpilot.com/review/jardiland.com?page=5",
"https://fr.trustpilot.com/review/jardiland.com?page=6",
"https://fr.trustpilot.com/review/jardiland.com?page=7",
"https://fr.trustpilot.com/review/jardiland.com?page=8"]
comms = []
for url in urls :
results = requests.get(url)
soup = BeautifulSoup(results.text, "html.parser")
commentary = soup.find_all('p', class_='review-content__text')
for container in commentary:
comm = container.text
comms.append(comm)
data = pd.DataFrame({
'comms' : comms})
data['comms'] = data['comms'].str.replace('\n', '')
#print(movies.head())
data.to_csv('df.csv')
我得到了这个:
当我在 Excel 中打开它时,它并不漂亮,所以我 c/c 并使用 Text to columns,我得到了这个:
看起来不错,但是当我想用 Python 阅读它进行进一步分析时,出现错误:
df = pd.read_csv('datajardiland.csv')
df.head()
UnicodeDecodeError: 'utf-8' codec can't decode byte 0xe0 in position 24: invalid continuation byte
我尝试了一些在 stackoverflow 上找到的方法,例如:
with open("datajardiland.csv") as f:
print(f.encoding)
cp1252
df = pd.read_csv('datajardiland.csv', encoding='cp1252')
df.head()
但它不起作用。
我在这里做错了什么?
【问题讨论】:
标签: python pandas web-scraping beautifulsoup utf-8