【问题标题】:How to export data from a beautifulsoup scrape to a csv file如何将beautifulsoup抓取的数据导出到csv文件
【发布时间】:2017-06-18 12:50:23
【问题描述】:

我在网上找到了这段代码,想知道如何将收集到的数据导出到 csv 文件中。

html = urllib.urlopen(url).read()
soup = BeautifulSoup(html)

# kill all script and style elements
for script in soup(["script", "style"]):
    script.extract()    # rip it out

# get text
text = soup.body.get_text()

# break into lines and remove leading and trailing space on each
lines = (line.strip() for line in text.splitlines())
# break multi-headlines into a line each
chunks = (phrase.strip() for line in lines for phrase in line.split("       "))
# drop blank lines
text = '\n'.join(chunk for chunk in chunks if chunk)

print(text)

【问题讨论】:

  • 这取决于您拥有的 URL。您能否编辑问题以提供更多详细信息。
  • 您的示例只返回它在一个集中找到的所有文本,因此它没有以任何方式结构化。在 CSV 列中对齐它是没有意义的。您可能对网页的某个部分感兴趣,例如新闻条目。您需要使用汤来提取该文本,然后将其制成 CSV。

标签: python csv beautifulsoup


【解决方案1】:

您的代码只是从给定的 URL 中提取所有文本。这会丢失任何结构,因此很难确定所需文本的开始和结束位置。

例如,在您提供的页面上,您可以通过查看 HTML 源并确定 5 个故事都具有唯一的 HTML id 来提取所有标题。您可以使用soup() 找到这些并从中提取文本。现在您有了每篇文章的标题和摘要,然后可以将其写入 CSV 文件。以下内容已使用 Python 3.5.2 进行了测试:

from urllib.request import urlopen
from bs4 import BeautifulSoup
import csv

html = urlopen("http://www.thestar.com.my/news/nation/")
soup = BeautifulSoup(html, "html.parser")

# IDs found by looking at the HTML source in a browser
ids = [
    "slcontent3_3_ileft_0_hlFirstStory", 
    "slcontent3_3_ileft_0_hlSecondStory",
    "slcontent3_3_ileft_0_lvStoriesRight_ctrl0_hlStoryRight",
    "slcontent3_3_ileft_0_lvStoriesRight_ctrl1_hlStoryRight",
    "slcontent3_3_ileft_0_lvStoriesRight_ctrl2_hlStoryRight"]

with open("news.csv", "w", newline="", encoding='utf-8') as f_news:
    csv_news = csv.writer(f_news)
    csv_news.writerow(["Headline", "Summary"])

    for id in ids:
        headline = soup.find("a", id=id)
        summary = headline.find_next("p") 
        csv_news.writerow([headline.text, summary.text])

这将为您提供如下 CSV 文件:

Headline,Summary
Many say convicted serial rapist Selva still considered âa person of high riskâ,PETALING JAYA: Convicted serial rapist Selva Kumar Subbiah will be back in the country from Canada in three days and a policeman who knows him says there is no guarantee that he will not strike again.
Liow: Way too many road accidents,"PETALING JAYA: Road accidents took the lives of 7,152 and incurred a loss of about RM9.2bil in Malaysia last year, says Datuk Seri Liow Tiong Lai."
Ex-civil servant wins RM27.4mil jackpot,PETALING JAYA: It was the ang pow of his life.
"Despite latest regulation, many still puff away openly at parks and R&R;","KUALA LUMPUR: It was another cloudy afternoon when office workers hung out at the popular KLCC park, puffing away at the end of lunch hour, oblivious to the smoking ban there."
Police warn groups not to cause disturbances on Thaipusam,GEORGE TOWN: Police have warned supporters of the golden and silver chariots against provo­king each other during the Thaipusam celebration next week.

【讨论】:

  • 太好了,很高兴它能满足您的需求。不要忘记单击向上/向下箭头下方的灰色勾号以接受答案作为已接受的解决方案。
猜你喜欢
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2021-08-22
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2018-02-17
相关资源
最近更新 更多