【问题标题】:Formatting scraped data from a Website (BeautifulSoup)格式化从网站抓取的数据 (BeautifulSoup)
【发布时间】:2012-06-19 15:54:35
【问题描述】:

我正在使用 BeautifulSoup 和 Requests 创建一个抓取工具,用于抓取网站页面以获取比赛时间表(和结果,如果可用)。这是我目前所拥有的:

    def getMatches(self):
        url = 'http://icc-cricket.yahoo.net/match_zone/series/fixtures.php?seriesCode=ENG_WI_2012' # change seriesCode in URL for different series.
        page = requests.get(url)
        page_content = page.content
        soup = BeautifulSoup(page_content)

    result = soup.find('div', attrs={'class':'bElementBox'})
    tags = result.findChildren('tr')

    for elem in tags:
        x = elem.getText()
        print x

这些是我得到的结果:

Date & Time (GMT)fixture
Thu, May 17, 2012 10:00 AMEngland  vs  West Indies
3rd TESTA full scorecard will be available shortly.Venue: Edgbaston,    BirminghamResult: England won by 5 wickets
Fri, May 25, 2012 11:00 AMEngland  vs  West Indies
2nd TESTClick here for the full scorecardVenue: Trent Bridge, NottinghamResult:     England won by 9 wickets
Thu, Jun 7, 2012 10:00 AMEngland  vs  West Indies
1st TESTClick here for the full scorecardVenue: Lord'sResult: Match Drawn
Sat, Jun 16, 2012 9:45 AMEngland  vs  West Indies
1st ODIClick here for the full scorecardVenue: The Rose Bowl, SouthamptonResult:     England won by 114 runs (D/L Method)
Tue, Jun 19, 2012 9:45 AMEngland  vs  West Indies
2nd ODIVenue: KIA Oval
Fri, Jun 22, 2012 9:45 AMEngland  vs  West Indies
3rd ODIVenue: Headingley Carnegie
Sun, Jun 24, 2012 12:00 AMEngland  vs  West Indies
1st T20Venue: Trent Bridge, Nottingham

现在,我想以某种结构化格式对数据进行分类。字典列表,每个包含
有关单个匹配的信息将是理想的。但我坚持如何实现这一目标。结果中的输出字符串有&nbsp之类的字符,时间安排很奇怪,例如AMEngland。还有一个问题,如果我使用空格字符作为分隔符来分割字符串,像西印度群岛这样的国家,只有两个单词,将被分割,并且不会有任何统一的方法来解析它。

那么有没有一种方法可以统一解析这些数据,这样我就可以进入表单了。有点像:

[ {'date': match_date, 'home_team': team1, 'away_team': team2, 'venue': venue},{ same for match 2}, { match 3 }...]

我将不胜感激。 :)

【问题讨论】:

    标签: python html-parsing web-scraping beautifulsoup


    【解决方案1】:

    区分日期/时间和国家并不难。你可以对“地点”和“结果”做同样的事情。

    >>> import re
    >>> s = "Sun, Jun 24, 2012 12:00 AMEngland  vs  West Indies"
    >>> match = re.search(r"\b[AP]M", s)
    >>> s[0:match.end()]
    'Sun, Jun 24, 2012 12:00 AM'
    >>> s[match.end():]
    'England  vs  West Indies'
    

    【讨论】:

    • 非常感谢。我想整天看 HTML 让我有点忘记我可以只使用一个简单的正则表达式。 :)
    【解决方案2】:

    改为查看scrapy;这将使这项任务变得容易得多。

    您定义 items 以从该站点抓取:

    from scrapy.item import Item, Field
    
    class CricketMatch(Item):
        date = Field()
        home_team = Field()
        away_team = Field()
        venue = Field()
    

    然后定义一个loader with XPath expressions 来填充这些项目。之后你可以直接使用这些物品,或者produce JSON output or similar.

    【讨论】:

    • 我确实想使用 scrapy,但我正在开发的应用程序已经使用 BeautifulSoup 来完成现有任务,所以我被告知不要使用它。
    • 很遗憾,您没有在问题中指定这一点。另请注意,SO 旨在提供普遍有用的问题和答案,而不仅仅是针对个别情况,因此我将保留我的答案。
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2019-09-26
    • 2021-10-26
    • 2020-10-25
    • 2015-05-09
    • 1970-01-01
    • 2018-06-30
    相关资源
    最近更新 更多