【问题标题】:Combining multiple generated dataframes into a single dataframe将多个生成的数据帧组合成一个数据帧
【发布时间】:2017-11-29 14:58:18
【问题描述】:

我想通过从 api 的每一页获取数据(每页限制 100 行)来构造一个数据框。目前以下代码返回所有数据,但结构错误。

有 17 个标题,因此我需要 17 列中的数据。但是,它输出 [100 行 x 1700 列] 的数据框,我需要 [10000 行 x 17 列]。

我不确定如何实现这一目标 - 任何帮助将不胜感激。

from ebaysdk.finding import Connection as finding
from bs4 import BeautifulSoup
import pandas as pd

x = []

for i in range(1,101):
    print(type(i))
    api = finding(siteid='EBAY-GB',appid='some_id',config_file=None)

    response = api.execute('findItemsByKeywords', {'keywords': 'phone', 'outputSelector' : 'SellerInfo',
    'paginationInput': {'entriesPerPage': '2','pageNumber': ' '+str(i)}})    

    soup = BeautifulSoup(response.content, 'lxml')

    items = soup.find_all('item')

    headers = ['itemid','title','categoryname','categoryid','postalcode','location','sellerusername','feedbackscore','positivefeedbackpercent','topratedseller','shippingservicecost','buyitnowavailable','currentprice','starttime','endtime','watchcount','conditionid']

    for object in headers:
        values = [element.text for element in soup.find_all(object)]
        x.append(values)
        df = pd.DataFrame(x)
        df = df.T
    print(x)
#[['152668959069', '252999725410'], ['Samsung GALAXY Ace GT-S5830i (Unlocked) Smartphone Android Phone- ALL COLOURS UK', '8GB 3G Unlocked Android 5.1 Quad Core Smartphone Mobile Phone 2 SIM GPS qHD'], ['Mobile & Smart Phones', 'Mobile & Smart Phones'], ['9355', '9355'], ['RM137PP'], ['Rainham,United Kingdom', 'United Kingdom'], ['deals4u_shop', 'smartlife2017'], ['15700', '456'], ['99.9', '98.5'], ['true', 'true'], ['0.0', '0.0'], ['false', 'false'], ['32.49', '48.9'], ['2017-08-18T18:36:28.000Z', '2017-06-19T09:04:40.000Z'], ['2017-12-16T18:36:28.000Z', '2017-12-16T09:04:40.000Z'], ['272', '134'], ['1000', '1000']]

    print(df)
             0                                                  1   \
0  152668959069  Samsung GALAXY Ace GT-S5830i (Unlocked) Smartp...   
1  252999725410  8GB 3G Unlocked Android 5.1 Quad Core Smartpho...   

                      2     3        4                       5   \
0  Mobile & Smart Phones  9355  RM137PP  Rainham,United Kingdom   
1  Mobile & Smart Phones  9355     None          United Kingdom   

              6      7     8     9   ...    24    25    26   27     28    29  \
0   deals4u_shop  15700  99.9  true  ...   456  98.5  true  0.0  false  48.9   

1  smartlife2017    456  98.5  true  ...   456  98.5  true  0.0  false  48.9   

                         30                        31   32    33  
0  2017-06-19T09:04:40.000Z  2017-12-16T09:04:40.000Z  214  1000  
1  2017-06-19T09:04:40.000Z  2017-12-16T09:04:40.000Z  182  1000  

编辑:添加更多代码并为第一页的前 2 个条目打印 x,为 2 页的前 2 个条目打印 df。

【问题讨论】:

  • df 输出有什么问题?并尝试在DataFrame's columns 参数中传递标题。
  • 我遇到的问题是 df 在每个循环中输出 [2 行 x 17 列]、[2 行 x 34 列]、[2 行 x 51 列] 等。我需要它来产生 [2 行 x 17 列]、[4 行 x 17 列]、[6 行 x 17 列] 等的结果。
  • 我之前尝试设置列,它返回以下错误。代码:df = pd.DataFrame(x, columns=headers) AssertionError: 17 columns pass, pass data has 2 columns
  • 另外,这看起来是一个 XML 响应。您只需 lxml 即可轻松做到这一点,无需 BS。

标签: python pandas dataframe beautifulsoup


【解决方案1】:

这应该会更好。

字典理解版:

data_dict = {obj: [element.text for element in soup.find_all(obj)] for obj in headers}    
df = pd.DataFrame(data_dict)

循环版本:

data_dict = {}
for obj in headers:
    data_dict[obj] = [element.text for element in soup.find_all(obj)]

df = pd.DataFrame(data_dict)

【讨论】:

  • 这很有帮助 - 谢谢
  • 很好,它有帮助。顺便说一句,如果它解决了你的问题,你可以accept 一个答案。
【解决方案2】:

考虑迭代地附加到具有最终连接的数据帧列表:

...
df_list = []
api = finding(siteid='EBAY-GB',appid='some_id',config_file=None)

for i in range(1,101):
    print(i)
    response = api.execute('findItemsByKeywords', 
                           {'keywords': 'phone',
                            'outputSelector' : 'SellerInfo',
                            'paginationInput': {'entriesPerPage': '2',
                                                'pageNumber': ' '+str(i)}})    

    soup = BeautifulSoup(response.content, 'lxml')

    headers = ['itemid','title','categoryname','categoryid','postalcode','location',
               'sellerusername','feedbackscore','positivefeedbackpercent','topratedseller',
               'shippingservicecost','buyitnowavailable','currentprice','starttime',
               'endtime','watchcount','conditionid']

    # LIST COMPREHENSION PARSING ELEMENTS OF API RESPONSE
    values = [element.text for element in soup.find_all(obj) for obj in headers]

    # DICT COMPREHENSION WITH ZIP TO DF THAT NAMES EACH COLUMN WITH VALUE & FILLS MISSING
    tmp = pd.DataFrame({h:v if len(v) > 1 else v+[None] for h,v in zip(headers, values)})

    # APPENDS TO LIST
    df_list.append(tmp)

# ROW BINDS TO FINAL DF
final_df = pd.concat(df_list, ignore_index=True)

【讨论】:

    【解决方案3】:
    from ebaysdk.finding import Connection as finding
    from bs4 import BeautifulSoup
    import pandas as pd
    
    def flatten(lst):
       for x in lst:
          if isinstance(x, list):
             for y in flatten(x):
                yield y           
          else:
                yield x
    
    full_dict = {}
    result = {}
    
    for i in range(1,101):
    print(i)
    
        api = finding(siteid='EBAY-GB',appid='some key',config_file=None)
        response = api.execute('findItemsByKeywords', {'keywords': 'phone', 'outputSelector' : 'SellerInfo',
    'paginationInput': {'entriesPerPage': '100','pageNumber': ' '+str(i)}})    
    
        soup = BeautifulSoup(response.content, 'lxml')
    
        items = soup.find_all('item')
    
        headers_tuple = ('itemid','title','categoryname','categoryid','postalcode','location','sellerusername','feedbackscore','positivefeedbackpercent','topratedseller','shippingservicecost','buyitnowavailable','currentprice','starttime','endtime','watchcount','conditionid')
    
        data_dict = {}
    
        for obj in headers_tuple:
            x = [element.text for element in soup.find_all(obj)]
            data_dict[obj] = x
        for key in (data_dict.keys() | full_dict.keys()):
            if key in data_dict: result.setdefault(key, []).append(data_dict[key])
            if key in full_dict: result.setdefault(key, []).append(full_dict[key])
    
    final_dict = {k: list(flatten(v)) for k, v in result.items()}
    df = pd.DataFrame.from_dict(final_dict, orient='index')
    df = df.T
    

    如果有人感兴趣,这是我的答案。它可以工作,但是由于某种原因列顺序发生了变化,我不确定为什么。感谢您的所有帮助!

    【讨论】:

      猜你喜欢
      • 2023-03-20
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多