【问题标题】:UnicodeDecodeError: 'utf-8'UnicodeDecodeError: 'utf-8'
【发布时间】:2020-12-15 14:59:37
【问题描述】:

我用这个脚本来抓取这个网站上的评论:https://fr.trustpilot.com/review/jardiland.com

import requests
from requests import get
from bs4 import BeautifulSoup
import pandas as pd
import numpy as np

urls = ["https://fr.trustpilot.com/review/jardiland.com",
        "https://fr.trustpilot.com/review/jardiland.com?page=2",
        "https://fr.trustpilot.com/review/jardiland.com?page=3",
        "https://fr.trustpilot.com/review/jardiland.com?page=4",
        "https://fr.trustpilot.com/review/jardiland.com?page=5",
        "https://fr.trustpilot.com/review/jardiland.com?page=6",
        "https://fr.trustpilot.com/review/jardiland.com?page=7",
        "https://fr.trustpilot.com/review/jardiland.com?page=8"]

comms = []

for url in urls : 
    results = requests.get(url)

    soup = BeautifulSoup(results.text, "html.parser")
    commentary = soup.find_all('p', class_='review-content__text')

    for container in commentary:
        comm  = container.text
        comms.append(comm)

    data = pd.DataFrame({
        'comms' : comms})

    data['comms'] = data['comms'].str.replace('\n', '')

#print(movies.head())

data.to_csv('df.csv')

我得到了这个:

当我在 Excel 中打开它时,它并不漂亮,所以我 c/c 并使用 Text to columns,我得到了这个:

看起来不错,但是当我想用 Python 阅读它进行进一步分析时,出现错误:

df = pd.read_csv('datajardiland.csv')
df.head()


UnicodeDecodeError: 'utf-8' codec can't decode byte 0xe0 in position 24: invalid continuation byte

我尝试了一些在 stackoverflow 上找到的方法,例如:

with open("datajardiland.csv") as f:
    print(f.encoding)

cp1252

df = pd.read_csv('datajardiland.csv', encoding='cp1252')
df.head()

但它不起作用。

我在这里做错了什么?

【问题讨论】:

    标签: python pandas web-scraping beautifulsoup utf-8


    【解决方案1】:

    尝试使用 ISO-8859 的编码参数。 df = pd.read_csv('datajardiland.csv', encoding='iso-8859-1')

    编辑:

    另外,在使用df.to_csv('test.csv', encoding='iso-8859-1') 保存 CSV 时尝试强制编码

    EDIT2:

    因为数据似乎是带有很多逗号和句点的“原始”文本。您可以尝试对df.to_csv('file.csv', sep=';') 和pd.read_csv('file.csv', sep=';') 使用不同的分隔符。或者您可以在所有文本数据周围设置“”以避免问题

    【讨论】:

    • ParserError: 标记数据时出错。 C 错误:第 8 行中预期有 3 个字段,看到 22
    • 嗯,我想说现在问题出在数据上。它看起来像带有逗号和句点的“原始”文本。您可以尝试使用 df.to_csv('file.csv', sep=';') 和 pd.read_csv('file.csv', sep=';') 的不同分隔符
    • 如果这不起作用,我建议您尝试在所有文本数据周围添加“”
    • 不错!我会用我在这里所说的来编辑答案。感谢您接受我的回答:)
    • 我会的!只是一件事,我有这个奇怪的列:未命名:0 如何在我的脚本中摆脱这个?使用 reset_index ?
    猜你喜欢
    • 2018-04-22
    • 1970-01-01
    • 2017-11-02
    • 2019-02-24
    • 1970-01-01
    • 1970-01-01
    • 2021-11-24
    • 2020-04-17
    • 1970-01-01
    相关资源
    最近更新 更多