【问题标题】:extract title from a link using BeautifulSoup使用 BeautifulSoup 从链接中提取标题
【发布时间】:2020-03-09 00:08:49
【问题描述】:

我正在使用 beautifulsoup 抓取一个网站,但需要帮助,因为我是 python 和 beautifulsoup 的新手 我如何从以下获得 VET “[[VET]]”

这是我目前的代码

import bs4 as bs
import urllib.request
import pandas as pd


#This is the Home page of the website
source = urllib.request.urlopen('file:///C:/Users/Aiden/Downloads/stocks/Stock%20Premarket%20Trading%20Activity%20_%20Biggest%20Movers%20Before%20the%20Market%20Opens.html').read().decode('utf-8')

soup = bs.BeautifulSoup(source,'lxml')


#find the Div and put all info into varTable
table = soup.find('table',{"id":"decliners_tbl"}).tbody



#find all Rows in table and puts into varTableRows
tableRows = table.find_all('tr')
print ("There is ",len(tableRows),"Rows in the Table")
print(tableRows)

columns = [tableRows[1].find_all('td')]
print(columns)

a = [tableRows[1].find_all("a")]
print(a)

So my output from print(a) is "[[<a class="mplink popup_link" href="https://marketchameleon.com/Overview/VET/">VET</a>]]"
 and I want to extract VET out 

广告

【问题讨论】:

  • 您能给我们一个minimal reproducible example吗?什么是"[[VET]]"(类名等)?您能向我们展示您目前的代码吗?
  • 给我们链接、您当前的代码以及 VET 是什么
  • 您好,感谢您的回复,我已经发布了我的代码
  • 欢迎来到 Stackoverflow!正如其他人所指出的,如果我们对源文件中的内容有更多信息,或者如果您减少该文件以证明问题,这将有所帮助。我的猜测是“VET”是您的 html 文件中表格中的链接,您的问题是它打印 [[VET]],但您不想要括号。尝试将最后一个分配更改为a = tableRows[1].find_all("a"),这将删除一组括号,然后您可以通过添加[0] 删除下一个,但这仅适用于至少有一个元素并且您不在乎是否存在是任何其他人。
  • 感谢您的回复,我的问题措辞很糟糕,我的输出是“[[marketchameleon.com/Overview/VET/">VET</a>]]" 我需要从中获得 VET。

标签: python screen-scraping


【解决方案1】:

您可以使用 a.text 或 a.get_text()。

如果您有多个元素,则需要对此函数进行列表理解

【讨论】:

    【解决方案2】:

    谢谢大家的回复,我用下面的代码就可以解决了

    source = urllib.request.urlopen('file:///C:/Users/Aiden/Downloads/stocks/Stock%20Premarket%20Trading%20Activity%20_%20Biggest%20Movers%20Before%20the%20Market%20Opens.html').read().decode('utf-8')
    
    
    soup = bs.BeautifulSoup(source,'html.parser')
    
    table = soup.find("table",id="decliners_tbl")
    
    for decliners in table.find_all("tbody"):
        rows = decliners.find_all("tr")
        for row in rows:
            ticker = row.find("a").text
            volume = row.findAll("td", class_="rightcell")[3].text
            print(ticker, volume)
    

    【讨论】:

      猜你喜欢
      • 2015-12-09
      • 2017-08-04
      • 2016-02-16
      • 2021-08-25
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2021-01-05
      • 2018-04-16
      相关资源
      最近更新 更多