【问题标题】:Associating the correct teams with the correct score values将正确的团队与正确的得分值相关联
【发布时间】:2013-08-16 06:31:01
【问题描述】:

我有一些代码可以从页面http://sports.yahoo.com/nhl/scoreboard?d=2013-04-01 输出团队及其所有得分值(不包括空格)。

from bs4 import BeautifulSoup
from urllib.request import urlopen

url = urlopen("http://sports.yahoo.com/nhl/scoreboard?d=2013-04-01")

content = url.read()

soup = BeautifulSoup(content)

listnames = ''
listscores = ''

for table in soup.find_all('table', class_='scores'):
    for row in table.find_all('tr'):
        for cell in row.find_all('td', class_='yspscores'):
            if cell.text.isdigit():
                listscores += cell.text
        for cell in row.find_all('td', class_='yspscores team'):
            listnames += cell.text

print (listnames)
print (listscores)

我无法解决的问题是我不太了解 Python 如何使用任何提取的信息并以如下格式为正确的团队提供正确的整数值:

Team X: 1, 5, 11.

网站的问题是所有分数都属于同一类;所有表都属于同一类。唯一不同的是 href。

【问题讨论】:

    标签: python python-3.x html-parsing beautifulsoup


    【解决方案1】:

    当您想将值与名称相关联时,dict 通常是要走的路。这是对代码的修改以演示原理:

    from bs4 import BeautifulSoup
    from urllib.request import urlopen
    
    url = urlopen('http://sports.yahoo.com/nhl/scoreboard?d=2013-04-01')
    
    content = url.read()
    
    soup = BeautifulSoup(content)
    
    results = {}
    
    for table in soup.find_all('table', class_='scores'):
        for row in table.find_all('tr'):
            scores = []
            name = None
            for cell in row.find_all('td', class_='yspscores'):
                link = cell.find('a')
                if link:
                    name = link.text
                elif cell.text.isdigit():
                    scores.append(cell.text)
            if name is not None:
                results[name] = scores
    
    for name, scores in results.items():
        print('%s: %s' % (name, ', '.join(scores)))
    

    ...运行时会给出这个输出:

    $ python3 get_scores.py
    St. Louis: 1, 2, 1
    San Jose: 0, 3, 0
    Colorado: 0, 0, 2
    Dallas: 0, 0, 0
    New Jersey: 0, 1, 0
    NY Islanders: 2, 0, 1
    Nashville: 0, 0, 2, 0
    Minnesota: 0, 1, 0
    Detroit: 1, 2, 0
    NY Rangers: 1, 1, 2
    Anaheim: 0, 3, 1
    Winnipeg: 2, 0, 0
    Chicago: 1, 1, 0, 0
    Calgary: 0, 0, 1
    Vancouver: 0, 1, 1
    Edmonton: 3, 0, 1
    Montreal: 1, 1, 2
    Carolina: 1, 0, 0
    

    除了使用字典之外,另一个重大变化是我们现在检查是否存在 a 元素来获取团队名称,而不是额外的 team 类。这确实是一种风格选择,但对我来说,这样的代码似乎更具表现力。

    【讨论】:

    • 只是出于好奇,假设(虽然它有效),如果我有多个具有不同名称的链接(让我们保持简单),例如 W 和 L(赢和输),我想要它像这样显示:TEAM X:W:25。L:16,我将如何接近那个? @ZeroPiraeus
    • @NathanielElder 老实说,我不太清楚你在问什么;无论如何,最好ask follow-up questions in a new question。
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2013-09-18
    相关资源
    最近更新 更多