【问题标题】:Web Scraping in Google Scholar with beautifulsoup and selenium in python使用 Python 中的 beautifulsoup 和 selenium 在 Google Scholar 中进行 Web Scraping
【发布时间】:2020-08-18 12:34:34
【问题描述】:

我正在尝试从 Google Scholar 个人资料中抓取。我需要具有我指定的特殊规格的配置文件。我在 Python 中使用 Beautifulsoup 和 selenium。例如,我需要在一所大学中从事我指定的某些学科的教授。你的想法是什么?

我的方式很慢,需要访问每个个人资料页面以检查我的特殊规格。如果你知道,请给我一个更快的方法。

如果有更快更好的方法来完成这项工作,请说出来。

【问题讨论】:

    标签: python selenium web-scraping beautifulsoup web-crawler


    【解决方案1】:

    您可以通过在双引号 "<univ. name>" 的标签后添加大学名称来做到这一点,例如:label:computer_vision "Michigan State University"。这样,您只能通过工作场所或电子邮件(例如 msu.edu)获得密歇根州立大学的作者,他们的兴趣与计算机视觉直接相关。

    注意:有时作者会写简短的大学缩写,例如 Michigan University -> U.Michigan, as Honglak Lee does。

    要包括这个例外,你可以使用管道| 符号,我相信它代表or。因此搜索查询将变为:label:computer_vision "Michigan State University"|"U.Michigan",翻译为密歇根州立大学或 U.Michigan。

    我找到的唯一一个可以get an idea of how to make such search queries on the Google Scholar Search Tips under How do I search by Title?的地方但是与搜索在某所大学工作的作者没有任何关系。所显示的结果是通过试验和错误实现的,这似乎是有效的。


    Code and example in the online IDE:

    from parsel import Selector
    import requests, json
    
    # https://docs.python-requests.org/en/master/user/quickstart/#passing-parameters-in-urls
    params = {
        "mauthors": 'label:computer_vision "Michigan State University"|"U.Michigan"', # search query 
        "hl": "en",                  # language
        "view_op": "search_authors"  # author results
    }
    
    # https://requests.readthedocs.io/en/master/user/quickstart/#custom-headers
    # Make sure you're using your user-agent: https://www.whatismybrowser.com/detect/what-is-my-user-agent
    headers = {
        "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/98.0.4758.87 Safari/537.36",
    }
    
    html = requests.get("https://scholar.google.com/citations", params=params, headers=headers, timeout=30)
    selector = Selector(html.text)
    
    profiles = []
    
    for profile in selector.css(".gs_ai_chpr"):
        profile_name = profile.css(".gs_ai_name a::text").get()
        profile_link = f'https://scholar.google.com{profile.css(".gs_ai_name a::attr(href)").get()}'
        profile_affiliation = profile.css('.gs_hlt::text').get()  # selects only university name without additional affiliation, e.g: Assistant Professor
        profile_email = profile.css(".gs_ai_eml::text").get()
        profile_interests = profile.css(".gs_ai_one_int::text").getall()
    
        profiles.append({
            "profile_name": profile_name,
            "profile_link": profile_link,
            "profile_affiliations": profile_affiliation,
            "profile_email": profile_email,
            "profile_interests": profile_interests
        })
    
    print(json.dumps(profiles, indent=2))
    
    
    # part of the output:
    '''
    [
      {
        "author_name": "Anil K. Jain",
        "author_link": "https://scholar.google.com/citations?hl=en&user=g-_ZXGsAAAAJ",
        "author_affiliations": "Michigan State University",
        "author_email": "Verified email at cse.msu.edu",
        "author_interests": [
          "Biometrics",
          "Computer vision",
          "Pattern recognition",
          "Machine learning",
          "Image processing"
        ]
      } # ...other profiles
    ]
    '''
    

    注意:我使用Parsel 库而不是最流行的解析库BeautifulSoup,但它非常相似并且支持XPath,并且有自己的CSS 伪元素支持,如::text 或::attr(<attribute>)。


    或者,您可以使用来自 SerpApi 的 Google Scholar Profiles API 来实现相同的目的。这是一个带有免费计划的付费 API。

    这种情况的不同之处在于您不必弄清楚刮擦部分,例如选择正确的选择器/XPath 来抓取数据或如何绕过搜索引擎的块,以及如何扩展请求数量。

    要集成的示例代码:

    from serpapi import GoogleSearch
    import os, json
    
    params = {
        "api_key": os.getenv("API_KEY"),     # SerpApi API key
        "engine": "google_scholar_profiles", # SerpApi profiles parsing engine
        "hl": "en",                          # language
        "mauthors": 'label:computer_vision "Michigan State University"|"U.Michigan"' # search query
    }
    
    search = GoogleSearch(params)
    results = search.get_dict()
    
    for profile in results["profiles"]:
        print(json.dumps(profile, indent=2))
    
    # part of the output:
    '''
    {
      "name": "Anil K. Jain",
      "link": "https://scholar.google.com/citations?hl=en&user=g-_ZXGsAAAAJ",
      "serpapi_link": "https://serpapi.com/search.json?author_id=g-_ZXGsAAAAJ&engine=google_scholar_author&hl=en",
      "author_id": "g-_ZXGsAAAAJ",
      "affiliations": "Michigan State University",
      "email": "Verified email at cse.msu.edu",
      "cited_by": 233876,
      "interests": [
        {
          "title": "Biometrics",
          "serpapi_link": "https://serpapi.com/search.json?engine=google_scholar_profiles&hl=en&mauthors=label%3Abiometrics",
          "link": "https://scholar.google.com/citations?hl=en&view_op=search_authors&mauthors=label:biometrics"
        },
        {
          "title": "Computer vision",
          "serpapi_link": "https://serpapi.com/search.json?engine=google_scholar_profiles&hl=en&mauthors=label%3Acomputer_vision",
          "link": "https://scholar.google.com/citations?hl=en&view_op=search_authors&mauthors=label:computer_vision"
        },
        {
          "title": "Pattern recognition",
          "serpapi_link": "https://serpapi.com/search.json?engine=google_scholar_profiles&hl=en&mauthors=label%3Apattern_recognition",
          "link": "https://scholar.google.com/citations?hl=en&view_op=search_authors&mauthors=label:pattern_recognition"
        },
        {
          "title": "Machine learning",
          "serpapi_link": "https://serpapi.com/search.json?engine=google_scholar_profiles&hl=en&mauthors=label%3Amachine_learning",
          "link": "https://scholar.google.com/citations?hl=en&view_op=search_authors&mauthors=label:machine_learning"
        },
        {
          "title": "Image processing",
          "serpapi_link": "https://serpapi.com/search.json?engine=google_scholar_profiles&hl=en&mauthors=label%3Aimage_processing",
          "link": "https://scholar.google.com/citations?hl=en&view_op=search_authors&mauthors=label:image_processing"
        }
      ],
      "thumbnail": "https://scholar.googleusercontent.com/citations?view_op=small_photo&user=g-_ZXGsAAAAJ&citpid=1"
    } ... other profiles
    
    '''
    

    免责声明,我为 SerpApi 工作。

    【讨论】:

      【解决方案2】:

      您可以像这样在 url 中添加您需要的主题:

      https://scholar.google.com/citations?hl=en&view_op=search_authors&mauthors=label:computer_vision+label:machine_learning

      我在这里寻找两个领域的作者:计算机视觉和机器学习

      【讨论】:

      • 如果我想从某所大学抓取科目,你会怎么想?
      • 我不确定大学,看起来它有特定大学的参考 ID
      猜你喜欢
      • 2017-05-14
      • 2021-07-12
      • 2020-01-16
      • 1970-01-01
      • 2018-11-11
      • 2018-02-13
      • 2018-04-02
      • 2019-01-06
      • 1970-01-01
      相关资源
      最近更新 更多