【问题标题】:Web scraper is scraping both text and <span> text </span>. Span text not neededWeb scraper 正在抓取文本和 <span> text </span>。不需要跨度文本
【发布时间】:2014-02-11 23:36:28
【问题描述】:

基本上,我正在尝试使用 BeautifulSoup 在 python 中抓取表格。

我已经设法抓取了另一个链接数组中的所有数据,但是由于某种原因,当我添加.text 时,它会同时打印文本和 span 标记内的文本。不需要跨度文本。

我尝试过.string 和.text.text,但它似乎不起作用。

谁能发现这里的问题?

这是我的代码:

soup = BeautifulSoup(urllib2.urlopen('http://www.livefootballontv.com/').read())

for row in soup('div', {'id': 'tv-guide'})[0]('ul'):
    tds = row('li')
    print tds[0].string, tds[1].text, tds[1].span.string, tds[2].string, tds[3].img['alt'], '\n'
    db = MySQLdb.connect("127.0.0.1","root","","footballapp")
    cursor = db.cursor()
    sql = "INSERT INTO TVGuide(DATE, FIXTURE, COMPETITION, KICKOFF, CHANNELS) VALUES (%s,%s,%s,%s,%s)"
    results = (str(tds[0].string), str(tds[1]).text, str(tds[1].span.string), str(tds[2].string), str(tds[3].img['alt']))
    cursor.execute(sql, results)
    db.commit()
    db.rollback()
    db.close()

然后给我

2014 年 6 月 22 日星期日美国对葡萄牙巴西 2014 年世界杯 G 组 2014 年巴西世界杯 G 组晚上 11:00 BBC1

2014 年 6 月 24 日星期二 哥斯达黎加 vs 英格兰巴西 2014 年世界杯小组赛 D 2014 年巴西世界杯 D 组 ITV 下午 5:00

【问题讨论】:

标签: python web-scraping beautifulsoup


【解决方案1】:

使用contents,并访问您想要的条目。

例子:

from bs4 import BeautifulSoup
import urllib2

soup = BeautifulSoup(urllib2.urlopen('http://www.livefootballontv.com/').read())

for row in soup('div', {'id': 'tv-guide'})[0]('ul'):
    tds = row('li')
    print tds[1].contents[0]

输出:

SV Hamburg vs Bayern Munich
Arsenal vs Manchester United
Napoli vs Roma
...
USA vs Portugal
Costa Rica vs England

【讨论】:

  • 我找到了一个duplicate question btw。你也可以使用find(text=True, recursive=False)
  • 第一个效果很好,非常感谢:) Top Geezer
猜你喜欢
  • 1970-01-01
  • 2019-12-23
  • 2017-11-28
  • 2016-05-09
  • 1970-01-01
  • 1970-01-01
  • 2011-08-20
  • 1970-01-01
  • 1970-01-01
相关资源
最近更新 更多