【问题标题】:scraping google search results page data python抓取谷歌搜索结果页面数据python
【发布时间】:2020-05-03 17:52:17
【问题描述】:

我想在搜索结果查询中抓取电子邮件。但是当我使用 css 选择器“select”访问类并打印时,它总是显示空列表。如何访问 .r 类或“class=g”?

    import requests
    from bs4 import BeautifulSoup
    
    url = "https://www.google.com/search?sxsrf=ACYBGNQA4leQETe0psVZPu7daLWbdsc9Ow%3A1579194494737&ei=fpggXpvRLMakwQKkqpSICg&q=%22computer+science+%22%22usa%22+%22%40yahoo.com%22&oq=%22computer+science+%22%22usa%22+%22%40yahoo.com%22&gs_l=psy-ab.12...0.0..7407...0.0..0.0.0.......0......gws-wiz.82okhpdJLYg&ved=0ahUKEwibiI_3zYjnAhVGUlAKHSQVBaEQ4dUDCAs"
    responce = requests.get(url)
    soup = BeautifulSoup(responce.text, "html.parser")
    test = soup.select('.r')
    print(test)

【问题讨论】:

    标签: python web-scraping beautifulsoup request python-requests


    【解决方案1】:

    您的程序是正确的,但要从 Google 获得正确答案,您需要指定 User-Agent 标头:

    导入请求 从 bs4 导入 BeautifulSoup

    url = "https://www.google.com/search?sxsrf=ACYBGNQA4leQETe0psVZPu7daLWbdsc9Ow%3A1579194494737&ei=fpggXpvRLMakwQKkqpSICg&q=%22computer+science+%22%22usa%22+%22%40yahoo.com%22&oq=%22computer+science+%22%22usa%22+%22%40yahoo.com%22&gs_l=psy-ab.12...0.0..7407...0.0..0.0.0.......0......gws-wiz.82okhpdJLYg&ved=0ahUKEwibiI_3zYjnAhVGUlAKHSQVBaEQ4dUDCAs"
    
    headers = {'User-Agent':'Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:72.0) Gecko/20100101 Firefox/72.0'}
    
    responce = requests.get(url, headers=headers)  # <-- specify custom header
    soup = BeautifulSoup(responce.text, "html.parser")
    test = soup.select('.r')
    print(test)
    

    打印:

    [<div class="r"><a href="https://www.yahoo.com/news/11-course-complete-computer-science-171322233.html" onmousedown="return rwt(this,'','','','1','AOvVaw2wM4TUxc_4V7s9GjeWTNAG','','2ahUKEwjt17Kk-YjnAhW2R0EAHcnsC3QQFjAAegQIAxAB','','',event)"><div class="TbwUpd"><img alt="https://...
    ...
    

    【讨论】:

    • 谢谢 .. 我的代码的目标是在 google 页面上显示的查询结果中废弃电子邮件。我在“.st”类中​​得到了文本。但是这个文本有点原始。我如何过滤文本并获得准确的电子邮件。
    【解决方案2】:

    要从 Google 搜索结果中删除电子邮件,您需要使用 regex

    # this regex needs possible modifications
    re.findall(r'[\w\.-]+@[\w\.-]+\.\w+', variable_where_to_search_from)
    

    代码:

    from bs4 import BeautifulSoup
    import requests, lxml, re
    
    headers = {
        "User-agent":
        "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko)"
        "Chrome/70.0.3538.102 Safari/537.36 Edge/18.19582"
    }
    
    html = requests.get('https://www.google.com/search?q="computer science ""usa" "@yahoo.com"', headers=headers)
    soup = BeautifulSoup(html.text, 'lxml')
    
    
    for result in soup.select('.tF2Cxc'):
        try:
            snippet = result.select_one('.lyLwlc').text
        except:
            snippet = None
    
        match_email = re.findall(r'[\w\.-]+@[\w\.-]+\.\w+', str(snippet))
        email = '\n'.join(match_email).strip()
        print(email)
    
    ----------
    '''
    ahmed_733@yahoo.com
    yjzou@uguam.uog
    yzou2002@yahoo.com
    ...
    

    或者,您可以使用来自 SerpApi 的 Google Organic Results API 来做同样的事情。这是一个带有免费计划的付费 API。

    它不会使用正则表达式提取电子邮件,尽管它可能是一个很棒的功能。主要区别在于完成工作比从头开始创建一切都更容易、更快捷。

    要集成的代码:

    from serpapi import GoogleSearch
    import re
    
    params = {
      "api_key": "YOUR_API_KEY",
      "engine": "google",
      "q": '"computer science ""usa" "@yahoo.com"',
    }
    
    search = GoogleSearch(params)
    results = search.get_dict()
    
    for result in results['organic_results']:
        try:
            snippet = result['snippet']
        except:
            snippet = None
    
        match_email = re.findall(r'[\w\.-]+@[\w\.-]+\.\w+', str(snippet))
        email = '\n'.join(match_email).strip()
        print(email)
    
    ---------
    '''
    shaikotweb@yahoo.com
    ahmed_733@yahoo.com
    RPeterson@L1id.com
    rj_peterson@yahoo.com
    '''
    

    免责声明,我为 SerpApi 工作。

    【讨论】:

      猜你喜欢
      • 2018-12-19
      • 2021-01-17
      • 2023-03-20
      • 2021-06-27
      • 2018-01-15
      • 2022-08-17
      • 2020-10-09
      • 2021-04-06
      • 2012-09-25
      相关资源
      最近更新 更多