【发布时间】:2020-05-30 16:26:22
【问题描述】:
我有以下代码输出从<div> 标记中提取的数据。
s = BeautifulSoup(driver.page_source, "lxml")
best_price_tags = s.findAll('div', "flt-subhead1 gws-flights-results__price gws-flights-results__cheapest-price")
best_prices = []
for tag in best_price_tags:
best_prices.append(tag.text.replace('€', '').strip())
变量best_price_tags的第一个元素包含以下内容:
<div class="flt-subhead1 gws-flights-results__price gws-flights-results__cheapest-price"> 1 820 € </div>
我希望上面的代码只输出值 1821。
上面的代码块有一个问题,它输出如下,考虑best_price_tags[0],'1\u202f821'的情况。
我尝试了以下方法,但不幸的是没有为我工作。
for tag in best_price_tags:
best_prices.append(int(tag.text.replace('€', '').strip()))
寻找不使用 NLP 模块的自动化解决方案。
注意:我已经编辑了 <div> 标签的确切值。以前是<div class='...'>1 820 €</div>,现在是<div class='...'> 1 820 € </div>。
【问题讨论】:
-
print(repr(tag.text))然后创建一个minimal reproducible example。 -
\u202f 是 unicode 用于 1 和 8 之间的小空间,你用的是 python 3 吗?
-
感谢@RoyZwambag,我在 Jupyter notebook 中使用 Python 3。
标签: python parsing web-scraping beautifulsoup