【发布时间】:2019-10-25 04:59:47
【问题描述】:
我正在尝试从以下 url (http://www.ancient-hebrew.org/m/dictionary/1000.html) 抓取数据。因此,每个希伯来语单词部分都以 img urls 开头,然后是 2 个文本,即实际的希伯来语单词及其发音。例如 url 中的第一个条目是以下“img1 img2 img3 אֶלֶף e-leph” 使用 wget 下载 html 后的希伯来语单词是 unicode
我给我的以下代码例如<img src="../../files/heb-anc-sm-pey.jpg"/>和<font face="arial" size="+1"> unicode_hebrew_text </font>和<a href="audio/ 505 .mp3"><img border="0" height="25" src="../../files/icon_audio.gif" width="25"/></a>
我只想要../../files/heb-anc-sm-pey.jpg
和unicode_hebrew_text 和audio/505.mp3 (without any spaces in between)
from bs4 import BeautifulSoup
raw_html = open('/Users/gansaikhanshur/TESTING/webScraping/1000.html').read()
html = BeautifulSoup(raw_html, 'html.parser')
# output: <img src="../../files/heb-anc-sm-pey.jpg"/>
imgs = html.findAll("img")
for image in imgs:
# print image source
if "jpg" in str(image):
print(image)
# output: <font face="arial" size="+1"> unicode_hebrew_text </font>
font = html('font', face="arial", size="+1")
for f in font:
continue
# output: <a href="audio/ 505 .mp3"><img border="0" height="25" src="../../files/icon_audio.gif" width="25"/></a>
mp3file = html.findAll(href=True)
for mp3 in mp3file:
if "mp3" in str(mp3):
continue
如您所见,我的代码并不能真正完成这项工作。最后,我想获取 URL 中每个单词的信息,并将其保存为文本文件或 json 文件,无论哪个更容易。
例如图片:URLsOfImages,希伯来语:txt,发音:txt,URLtoAudio:txt
对于下一个单词,以此类推。
【问题讨论】:
标签: python python-3.x python-2.7 web-scraping beautifulsoup