【问题标题】:What am I doing wrong in this for loop for Web Scraping with bs4?在使用 bs4 进行 Web Scraping 的 for 循环中我做错了什么?
【发布时间】:2020-04-01 13:24:48
【问题描述】:

我正在尝试遍历Transfermarkt 上的玩家列表,输入每个个人资料,获取他们的个人资料图片,然后抓取原始信息列表。后者,我已经实现了(您将在我的代码中看到),但前者我似乎无法开始工作。我不是他的专家,并且在我的代码方面得到了帮助。

我不想保存每个玩家图片的源链接,而不是图像本身,然后将该链接存储到我的数据框中的“PlayerImgURL”中。 (第 73 行)。

这是我的错误信息:

(.venv) PS C:\Users\cljkn\Desktop\Python scraper github> & "c:/Users/cljkn/Desktop/Python scraper github/.venv/Scripts/python.exe" "c:/Users/cljkn/Desktop/Python scraper github/.vscode/test.py"
  File "c:/Users/cljkn/Desktop/Python scraper github/.vscode/test.py", line 45
    for page in range(1, 21):
    ^
SyntaxError: invalid syntax

谢谢。

from bs4 import BeautifulSoup
import requests
import pandas as pd

playerID = []
playerImage = []
playerName = []
result = []

for page in range(1, 21):

    r = requests.get("https://www.transfermarkt.com/spieler-statistik/wertvollstespieler/marktwertetop?land_id=0&ausrichtung=alle&spielerposition_id=alle&altersklasse=alle&jahrgang=0&kontinent_id=0&plus=1",
        params= {"page": page},
        headers= {"User-Agent":"Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:74.0) Gecko/20100101 Firefox/74.0"}
    )
    soup = BeautifulSoup(r.content, "html.parser")

    links = soup.select('a.spielprofil_tooltip')

    for i in range(len(links)):
        playerID.append(links[i].get('id'))

    for i in range(len(playerID)):
        playerID[i] = 'https://www.transfermarkt.com/kylian-mbappe/profil/spieler/'+playerID[i]
        playerID = list(set(playerID))

    for i in range(len(playerID)):

        r = requests.get(playerID[i],
            params= {"page": page},
            headers= {"User-Agent":"Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:74.0) Gecko/20100101 Firefox/74.0"}
        )
    soup = BeautifulSoup(r.content, "html.parser")

    name = soup.find_all('h1')

    for image in soup.find_all('img'):
        playerName.append('title')

        playerImage.append[image.get('src')




    for page in range(1, 21):

        r = requests.get("https://www.transfermarkt.com/spieler-statistik/wertvollstespieler/marktwertetop?land_id=0&ausrichtung=alle&spielerposition_id=alle&altersklasse=alle&jahrgang=0&kontinent_id=0&plus=1",
            params= {"page": page},
            headers= {"User-Agent":"Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:74.0) Gecko/20100101 Firefox/74.0"}
        )
        soup = BeautifulSoup(r.content, "html.parser")


        tr = soup.find_all("tbody")[1].find_all("tr", recursive=False)

        result.extend([
            { 

            "Club": t[4].find("img")["alt"],
            "Age": t[2].text.strip(),
            "GamesPlayed": t[6].text.strip(),
            "GoalsDone": t[7].text.strip(),
            "OwnGoals": t[8].text.strip(),
            "Assists": t[9].text.strip(),
            "YellowCards": t[10].text.strip(),
            "SecondYellow": t[11].text.strip(),
            "StraightRed": t[12].text.strip(),
            "SubsOn": t[13].text.strip(),
            "SubsOff": t[14].text.strip(),
            "Nationality": t[3].find("img")["alt"], # for all nationality : [ i["alt"] for i in t[3].find_all("img")], 
            "Position": t[1].find_all("td")[2].text,
            "Value": t[5].text.strip(),
            #"PlayerImgURL":
            "ClubImgURL": t[4].find("img")["src"],
            "CountryImgURL": t[3].find("img")["src"] # for all country url: [ i["src"] for i in t[3].find_all("img")]
            }

            for t in (t.find_all(recursive=False) for t in tr)
        ])



df = pd.DataFrame(result,{'Name':playerImage, 'Source':playerImage})


#df.to_csv (r'S:\_ALL\Internal Projects\Introduction_2020\Transfermarkt\PlayerDetails.csv', index = False, header=True)

print(df)

【问题讨论】:

  • 仔细查看有问题的 for 循环正上方的代码行。有两个错误。
  • 你的意思是在“ playerImage.append[image.get('src')”之后
  • 我指的正是那一行——你刚刚复制粘贴到评论中的那一行。
  • 现在我看到了!!谢谢..

标签: python web-scraping beautifulsoup


【解决方案1】:

这一行的问题

playerImage.append[image.get('src')

尝试用这一行替换

playerImage.append(image.get('src'))

【讨论】:

  • 非常感谢!
  • mille grazie Hamza Lachi!
  • @zero 没问题
  • @Hamza Lachi : 已将您添加到 fb ;)
猜你喜欢
  • 2018-05-30
  • 1970-01-01
  • 2017-03-22
  • 1970-01-01
  • 1970-01-01
  • 2015-05-08
  • 2018-08-12
相关资源
最近更新 更多