【问题标题】:Web scraping different football live scores sites网络抓取不同的足球现场比分网站
【发布时间】:2017-03-17 09:19:07
【问题描述】:

我需要一个数据库中的足球比分,以便开发一个应用程序。我发现的 api 不完整或没有我需要的某些功能,网络抓取实时比分网站是否合法?我想我可以抓取不同的网站来不创造流量,你怎么看?谢谢

【问题讨论】:

  • 我投票结束这个问题,因为这个问题是关于抓取网站的合法性
  • 你的问题不好回答。我建议阅读本网站的服务条款。网络上有很多关于网络抓取合法性的文章。

标签: python web-scraping live


【解决方案1】:

我不认为从网站解析数据是违法的。您可能想要做这样的事情,它是一个程序,它转到指定的网页并从特定行获取数据并将其保存到文件中以供另一个程序使用。

#this is to get the price of various stocks from Google's search page.     Let's hope this works.
import requests
# Example file for parsing and processing HTML
# import the HTMLParser module
from HTMLParser import HTMLParser
import time

metacount = 0;

x=0
while x==0:
# create a subclass and override the handler methods
class MyHTMLParser(HTMLParser):
    # function to handle character and text data (tag contents)
    def handle_data(self, data):
        #print data
        pos = self.getpos()
        #print "At line: ", pos[0], " position ", pos[1]
        if pos[0]==154:
            price=data
            print price
            # Open a file for writing and create it if it doesn't exist
            f = open("price.txt", "w+")
            # write some lines of data to the file
            f.write(price)
            f.close()
            # Open the file back up and read the contents
            #if = open("price.txt", "r")
            #if f.mode == 'r':  # check to make sure that the file was opened
        # use the read() function to read the entire file
            #   print('true')



def main():
    # instantiate the parser and feed it some HTML
    parser = MyHTMLParser()
    #stock=open('stocks.txt')
    stockname=raw_input('stock symbol')#stock.read()
    r=requests.get('http://stocks.tradingcharts.com/stocks/quotes/'+stockname)
    #print (r.status_code)
    stuff = r.text
    parser.feed(stuff)


if __name__ == "__main__":
    main();
#put it on a timer since the page is updated once every 5 minutes
time.sleep(300)

【讨论】:

  • 但是我需要开发一个使用该数据库的公共应用程序,我认为这不太合法,所以我想问
  • 等等,你是想从网页上抓取它还是想闯入为数据库服务的数据库?一种是合法的,但有些人不赞成,另一种是绝对非法的,如果您没有访问数据库的权限。
  • 我想从页面中抓取结果,使用条款中没有任何内容
  • 如果来自页面,应该没有任何合法性问题。如果它可以公开访问,那么任何人都可以而且应该能够访问它。
  • 它来自一个可公开访问的页面,所以我可以为我的免费应用程序获取和使用数据,谢谢。
猜你喜欢
  • 2016-02-03
  • 1970-01-01
  • 2022-01-21
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2020-03-08
  • 1970-01-01
  • 2017-10-16
相关资源
最近更新 更多