【问题标题】:Scrape part of <li> from <ul> class?从 <ul> 类中刮掉 <li> 的一部分?
【发布时间】:2019-09-12 08:04:53
【问题描述】:

从web 的我的股票列表的“ul class mc-list”中获取每个“li class mc”的“bullets”。

我是 Python 新手,我想自动检查我的股票投资组合。

我有一个包含股票代码的文件 (mystocks.txt)(每行一张)。

我想每天查看一次 SA web 是否有任何关于我的股票的消息。

url = 'https://seekingalpha.com/dividends/dividend-news'
response = requests.get(url)
soup = BeautifulSoup(response.text, 'lxml')
for link in soup.find_all('li'):
...

预期的输出是:

如果 div.bullets 包含来自“mystocks.txt”的股票代码,则应创建一个名为“ticket”.txt 的文件并包含“div.bullets”文本。

【问题讨论】:

    标签: python-3.x web-scraping stock


    【解决方案1】:

    查看下面的实现。我希望它能让你到达那里:

    import requests
    from bs4 import BeautifulSoup
    
    link = "https://seekingalpha.com/dividends/dividend-news"
    
    #following are the pseudo list of tickers you might wanna check against
    for ticker in ['NWTUF','BSL','KRC']:
        res = requests.get(link,headers={'User-Agent':'Mozilla/5.0'})
        soup = BeautifulSoup(res.text,"lxml")
    
        for item in soup.select(".media-body"):
            #if there is no match, get rid of the content
            if ticker not in item.text:continue
    
            for elem in item.select(".bullets > ul > li, .bullets > ul > li > a"):
                print(elem.text)
            print("***"*20)
    

    【讨论】:

    • 您好,感谢您的帮助/建议,不幸的是,对于 ['NWTUF','BSL','KRC' 中的代码从 bs4 import BeautifulSoup link = "seekingalpha.com/dividends/dividend-news" ]: ... res = requests.get(link,headers={'User-Agent':'Mozilla/5.0'}) 文件“”,第 2 行 res = requests.get(link,headers={' User-Agent':'Mozilla/5.0'}) ^ IndentationError: 需要一个缩进块
    • 缩进是python编程中最重要和最基本的东西之一,你应该首先学习。但是,如果您按原样运行代码,则不应遇到任何错误。我刚才查了一下。谢谢。
    • 看起来你也可以使用 .bullets li, .bullets a +
    • 非常感谢。我将在星期一回到我的电脑后对其进行测试。
    【解决方案2】:

    小进步(学习中),添加逐行读取文件,但即使票在页面上也不打印divi记录:

    import requests
    from bs4 import BeautifulSoup
    
    link = "https://seekingalpha.com/dividends/dividend-news"
    fileHandler = open ("tickers.txt", "r")
    
    with open ("tickers.txt", "r") as fileHandler:
      for ticker in fileHandler:
        print(ticker.strip())
        res = requests.get(link,headers={'User-Agent':'Mozilla/5.0'})
        soup = BeautifulSoup(res.text,"lxml")
    
        for item in soup.select(".media-body"):
            #if there is no match, get rid of the content
            if ticker not in item.text:continue
    
            for elem in item.select(".bullets > ul > li, .bullets > ul > li > a"):
                print(elem.text)
            print("***"*20)
    
    # Close Close
    fileHandler.close()
    

    输出看起来像(尝试了所有可能的名称): rpi2:~$ ./divi.py 主要的 TJX 纳斯达克股票代码:NWFL 奥驰亚 拉尔夫劳伦

    【讨论】:

      猜你喜欢
      • 2021-09-26
      • 1970-01-01
      • 1970-01-01
      • 2021-10-16
      • 1970-01-01
      • 2019-01-24
      • 1970-01-01
      • 1970-01-01
      • 2019-06-16
      相关资源
      最近更新 更多