【问题标题】:Using Python Beautifulsoup to collect data from LinkedIn使用 Python Beautifulsoup 从 LinkedIn 收集数据
【发布时间】:2019-02-27 18:33:08
【问题描述】:

我正在尝试使用 python beautifulsoup 模块导出我的 LinkedIn 联系人姓名。我的代码如下:

import requests
from bs4 import BeautifulSoup

client = requests.Session()

HOMEPAGE_URL = 'https://www.linkedin.com'
LOGIN_URL = 'https://www.linkedin.com/uas/login-submit'
CONNECTIONS_URL = 'https://www.linkedin.com/mynetwork/invite-connect/connections/'

html = client.get(HOMEPAGE_URL).content
soup = BeautifulSoup(html, "html.parser")
csrf = soup.find(id="loginCsrfParam-login")['value']

login_information = {
    'session_key':'username',
    'session_password':'password',
    'loginCsrfParam': csrf,
}
try:
    client.post(LOGIN_URL, data=login_information)
    print "Login Successful"
except:
    print "Failed to Login"

html = client.get(CONNECTIONS_URL).content
soup = BeautifulSoup(html , "html.parser")
print soup.find_all('div', attrs={'class' : 'mn-connection-card__name'})

但问题是我总是得到一个空列表。像下面这样:

Login Successful
[]

一个html结构是这样的:

<span class="mn-connection-card__name t-16 t-black t-bold">
      Sombody's name
    </span>

我认为我应该改变我的 soup.x 方法。我用了 find、select、find_all 都没有成功。

谢谢

【问题讨论】:

  • 你应该使用Linkedin的REST API来查询这类信息。
  • Ummmmm...当您应该尝试查找跨度(在 html 结构中)时,是否可以像您尝试查找 div(在您的代码中)一样简单?
  • @darksky 不幸的是,LinkedIn 对其 API 的权限有限,因此我需要您抓取他们的网站以收集我所需的信息。所以他们的 REST API 暂时没有用。 (我需要获取联系人的电子邮件地址、电话号码……)
  • @Mahdi 您是否按照 Jeff 的建议尝试在 soup.find_all('div', attrs={'class' : 'mn-connection-card__name'}) 中将 'div' 替换为 'span'?

标签: python beautifulsoup linkedin


【解决方案1】:

我知道我迟到了,但现在这适用于linkedin:

import requests
from bs4 import BeautifulSoup

#create a session
client = requests.Session()

#create url page variables
HOMEPAGE_URL = 'https://www.linkedin.com'
LOGIN_URL = 'https://www.linkedin.com/uas/login-submit'
CONNECTIONS_URL = 'https://www.linkedin.com/mynetwork/invite-connect/connections/'
ASPIRING_DATA_SCIENTIEST = 'https://www.linkedin.com/search/results/people/?keywords=Aspiring%20Data%20Scientist&origin=GLOBAL_SEARCH_HEADER'

#get url, soup object and csrf token value
html = client.get(HOMEPAGE_URL).content
soup = BeautifulSoup(html, "html.parser")
csrf = soup.find('input', dict(name='loginCsrfParam'))['value']

#create login parameters
login_information = {
    'session_key':'your_email',
    'session_password':'your_password',
    'loginCsrfParam': csrf,
}

#try and login
try:
    client.post(LOGIN_URL, data=login_information)
    print("Login Successful")
except:
    print("Failed to Login")

#open the html with soup object
# html = client.get(CONNECTIONS_URL).content #opens connections_url
html = client.get(ASPIRING_DATA_SCIENTIEST).content #opens ASPIRING_DATA_SCIENTIEST
soup = BeautifulSoup(html , "html.parser")
# print(soup.find_all('div', attrs={'class' : 'mn-connection-card__name'}))

# print(soup)
print(soup.prettify())

【讨论】:

    【解决方案2】:

    如果你想提取名字,你只需要

    from bs4 import BeautifulSoup
    soup = BeautifulSoup(html , "html.parser")
    target = soup.find_all('span', attrs={'class' : 'mn-connection-card__name'})
    target[0].text.strip()
    

    输出

    "Sombody's name"
    

    【讨论】:

    • 我这样做了,但列表索引超出范围
    • 我不知道该告诉你什么——我在一个干净的 jupyter 笔记本上又试了一次,它对我有用。只是为了确认一下:我正在使用的soup 变量中的html 是html = """ &lt;span class="mn-connection-card__name t-16 t-black t-bold"&gt; Sombody's name &lt;/span&gt; """
    • 能否请您发送您的完整代码?我的意思是身份验证等。
    • 对不起,我没有做整个事情(没有 LI 帐户);只关注数据提取部分...
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2022-01-02
    • 2021-12-17
    • 2019-09-13
    • 1970-01-01
    • 2021-11-15
    相关资源
    最近更新 更多