【问题标题】:How can I use Beautiful Soup to webscrape also the URLs to be used in my pandas dataframe我如何使用 Beautiful Soup 来抓取我的 pandas 数据框中要使用的 URL
【发布时间】:2021-12-19 02:36:09
【问题描述】:

我正在尝试抓取以下网址:http://eecs.qmul.ac.uk/postgraduate/programmes

我有以下代码:

#Create loop to look for the td tag and print the rows
for row in rows:
   row_td = row.find_all('td')
   row_url = row.find_all('a')  
   print(row_td)
 type(row_td)

这会产生以下结果:

[<td>Artificial Intelligence</td>, <td style="text-align: center;"><a href="https://www.qmul.ac.uk/postgraduate/taught/coursefinder/courses/artificial-intelligence-msc/" title="Use alt + click to follow the link">I4U2</a> </td>, <td style="text-align: center;"><a href="https://www.qmul.ac.uk/postgraduate/taught/coursefinder/courses/artificial-intelligence-msc/" title="Use alt + click to follow the link">I4U1</a> </td>]
[<td>Artificial Intelligence with Machine Learning (January 2022 Entry Only)</td>, <td style="text-align: center;"> </td>, <td style="text-align: center;"><a href="https://www.qmul.ac.uk/postgraduate/taught/coursefinder/courses/artificial-intelligence-with-machine-learning-msc/">I4U8</a></td>]
[<td><span>Big Data Science</span></td>, <td style="text-align: center;"><a href="https://www.qmul.ac.uk/postgraduate/taught/coursefinder/courses/big-data-science-msc/">H6J6</a></td>, <td style="text-align: center;"><a href="https://www.qmul.ac.uk/postgraduate/taught/coursefinder/courses/big-data-science-msc/">H6J7</a></td>]
[<td><span>Big Data Science with Machine Learning Systems (January 2022 Entry Only)</span></td>, <td style="text-align: center;"> </td>, <td style="text-align: center;"><a href="https://www.qmul.ac.uk/postgraduate/taught/coursefinder/courses/big-data-science-with-machine-learning-systems-msc/">I4U7</a></td>]

如您所见,每门课程都有一个课程名称、一个全日制 URL 和代码,然后是下一行的非全日制 URL 和代码。我想把这些记录下来,这样它就会变成:

人工智能,https://www.qmul.ac.uk/postgraduate/taught/coursefinder/courses/artificial-intelligence-msc,I4U2,https://www.qmul.ac.uk/postgraduate/taught/coursefinder/courses/artificial-intelligence-msc/,I4U1。

我已使用以下代码清理行,以便我可以实现这一点,但是,它不包括 URL。

INPUT:
str_cells = str(row_td)
cleantext = BeautifulSoup(str_cells).get_text()
print(cleantext)

OUTPUT: 
[Digital and Technology Solutions (Apprenticeship), I4DA,  ]

您能否提供帮助,以便我也可以在输出中包含 URL?

我发现我可以使用 soup.find_all("a") 选择 URL,但我不知道如何将它与上面的代码结合起来。

谢谢。

编辑:

我已使用以下建议修改了我的代码,但是,当我尝试将其添加为 pandas 数据框时,它似乎无法正常工作。请问有人能发现我的错误,并帮助我将其转换为熊猫数据框吗?

dfObj = pd.DataFrame(columns = ['C1', 'C2', 'C3', 'C4', 'C5'])
for row in rows:
   courses = row.find_all("td")
# The fragments list will store things to be included in the final 
   string, such as the course title and its URLs
   fragments = []
   for course in courses:
      if course.text.isspace():
        continue
     # Add the <td>'s text to fragments
     fragments.append(course.text)
      # Try and find an <a> tag 
     a_tag = course.find("a")
     if a_tag:
        # If one was found, add the URL to fragments
        fragments.append(a_tag["href"])

     # Make a string containing every fragment with ", " spacing them apart.
         cleantext = ", ".join(fragments)
         series_obj = pd.Series(cleantext, 
                    index=dfObj.columns)
          # Add a series as a row to the dataframe  
         mod_df = dfObj.append(  series_obj,
                     ignore_index=True) 

  print(mod_df)

然后我得到以下输出:

【问题讨论】:

    标签: python web-scraping beautifulsoup data-mining


    【解决方案1】:

    您可以利用 BeautifulSoup4 的 Tag 对象的 .text 属性来构建字符串,并将内容存储在中间变量中(我选择了一个列表,以便将项目与逗号组合起来更容易)。

    这是一个成熟的例子:

    for row in rows:
        children = row.find_all("td")
        # The fragments list will store all of the stuff we want
        # included in the final string, such as the course title and
        # its URLs.
        fragments = []
        for child in children:
            # Ignore <td> tags without text (can happen when no
            # part-time version of a course exists)
            if child.text.isspace():
                continue
            # Add the <td>'s text to fragments. This could be the course
            # title (e.g. Internet of Things (Data)) or its ID (e.g. I1T2)
            fragments.append(child.text)
            # Try and find an <a> tag in the child <td>.
            a_tag = child.find("a")
            if a_tag:
                # If one was found, add the URL to fragments.
                fragments.append(a_tag["href"])
    
        # Make a string containing every fragment with ", " spacing them apart.
        cleantext = ", ".join(fragments)
        print(cleantext)
    

    【讨论】:

    • 如果您希望链接位于课程 ID 之前,只需将 fragments.append(child.text) 移动到 if a_tag 块之后。
    猜你喜欢
    • 1970-01-01
    • 2017-03-30
    • 2023-03-31
    • 1970-01-01
    • 2018-10-19
    • 1970-01-01
    • 2018-07-01
    • 1970-01-01
    • 2017-07-29
    相关资源
    最近更新 更多