【发布时间】:2021-12-19 02:36:09
【问题描述】:
我正在尝试抓取以下网址:http://eecs.qmul.ac.uk/postgraduate/programmes
我有以下代码:
#Create loop to look for the td tag and print the rows
for row in rows:
row_td = row.find_all('td')
row_url = row.find_all('a')
print(row_td)
type(row_td)
这会产生以下结果:
[<td>Artificial Intelligence</td>, <td style="text-align: center;"><a href="https://www.qmul.ac.uk/postgraduate/taught/coursefinder/courses/artificial-intelligence-msc/" title="Use alt + click to follow the link">I4U2</a> </td>, <td style="text-align: center;"><a href="https://www.qmul.ac.uk/postgraduate/taught/coursefinder/courses/artificial-intelligence-msc/" title="Use alt + click to follow the link">I4U1</a> </td>]
[<td>Artificial Intelligence with Machine Learning (January 2022 Entry Only)</td>, <td style="text-align: center;"> </td>, <td style="text-align: center;"><a href="https://www.qmul.ac.uk/postgraduate/taught/coursefinder/courses/artificial-intelligence-with-machine-learning-msc/">I4U8</a></td>]
[<td><span>Big Data Science</span></td>, <td style="text-align: center;"><a href="https://www.qmul.ac.uk/postgraduate/taught/coursefinder/courses/big-data-science-msc/">H6J6</a></td>, <td style="text-align: center;"><a href="https://www.qmul.ac.uk/postgraduate/taught/coursefinder/courses/big-data-science-msc/">H6J7</a></td>]
[<td><span>Big Data Science with Machine Learning Systems (January 2022 Entry Only)</span></td>, <td style="text-align: center;"> </td>, <td style="text-align: center;"><a href="https://www.qmul.ac.uk/postgraduate/taught/coursefinder/courses/big-data-science-with-machine-learning-systems-msc/">I4U7</a></td>]
如您所见,每门课程都有一个课程名称、一个全日制 URL 和代码,然后是下一行的非全日制 URL 和代码。我想把这些记录下来,这样它就会变成:
人工智能,https://www.qmul.ac.uk/postgraduate/taught/coursefinder/courses/artificial-intelligence-msc,I4U2,https://www.qmul.ac.uk/postgraduate/taught/coursefinder/courses/artificial-intelligence-msc/,I4U1。
我已使用以下代码清理行,以便我可以实现这一点,但是,它不包括 URL。
INPUT:
str_cells = str(row_td)
cleantext = BeautifulSoup(str_cells).get_text()
print(cleantext)
OUTPUT:
[Digital and Technology Solutions (Apprenticeship), I4DA, ]
您能否提供帮助,以便我也可以在输出中包含 URL?
我发现我可以使用 soup.find_all("a") 选择 URL,但我不知道如何将它与上面的代码结合起来。
谢谢。
编辑:
我已使用以下建议修改了我的代码,但是,当我尝试将其添加为 pandas 数据框时,它似乎无法正常工作。请问有人能发现我的错误,并帮助我将其转换为熊猫数据框吗?
dfObj = pd.DataFrame(columns = ['C1', 'C2', 'C3', 'C4', 'C5'])
for row in rows:
courses = row.find_all("td")
# The fragments list will store things to be included in the final
string, such as the course title and its URLs
fragments = []
for course in courses:
if course.text.isspace():
continue
# Add the <td>'s text to fragments
fragments.append(course.text)
# Try and find an <a> tag
a_tag = course.find("a")
if a_tag:
# If one was found, add the URL to fragments
fragments.append(a_tag["href"])
# Make a string containing every fragment with ", " spacing them apart.
cleantext = ", ".join(fragments)
series_obj = pd.Series(cleantext,
index=dfObj.columns)
# Add a series as a row to the dataframe
mod_df = dfObj.append( series_obj,
ignore_index=True)
print(mod_df)
然后我得到以下输出:
【问题讨论】:
标签: python web-scraping beautifulsoup data-mining