【问题标题】:Loop on pandas dataframe over unique values only仅在唯一值上循环熊猫数据框
【发布时间】:2018-08-31 14:53:56
【问题描述】:

我有以下熊猫数据框:

DB      Table   Column  Format

Retail  Orders  ID      INTEGER
Retail  Orders  Place   STRING
Dept    Sales   ID      INTEGER
Dept    Sales   Name    STRING

我想在表格上循环,同时生成用于创建表格的 SQL。例如

create table Retail.Orders ( ID INTEGER, Place STRING)
create table Dept.Sales ( ID INTEGER, Name STRING)

我已经完成的是使用 drop_duplicate 获得不同的数据库和表,然后为每个表应用过滤器并连接字符串以创建 sql。

def generate_tables(df_cols):
    tables = df_cols.drop_duplicates(subset=[KEY_DB, KEY_TABLE])[[KEY_DB, KEY_TABLE]]

    for index, row in tables.iterrows():
        db = row[KEY_DB]
        table = row[KEY_TABLE]

        print("DB: " + db)
        print("Table: " + table)

        sql = "CREATE TABLE " + db + "." + table + " ("
        cols = df_cols.loc[(df_cols[KEY_DB] == db) & (df_cols[KEY_TABLE] == table)]
        for index, col in cols.iterrows():
            sql += col[KEY_COLUMN] + " " + col[KEY_FORMAT] + ", "

        sql += ")"

        print(sql)

是否有更好的方法来迭代数据框?

【问题讨论】:

    标签: python pandas


    【解决方案1】:

    这就是我要做的方式。首先通过df.itertuples创建字典[比df.iterrows更高效],然后使用str.format无缝包含值。

    使用set 保证字典构造的唯一性。

    我还转换为生成器,以便您可以根据需要有效地迭代它;始终可以通过list 将发电机耗尽,如下所示。

    from collections import defaultdict
    
    d = defaultdict(set)
    for row in df.itertuples():
        d[(row[1], row[2])].add((row[3], row[4]))
    
    def generate_tables_jp(d):
        for k, v in d.items():
            yield 'CREATE TABLE {0}.{1} ({2})'\
                  .format(k[0], k[1], ', '.join([' '.join(i) for i in v]))
    
    list(generate_tables_jp(d))
    

    结果:

    ['CREATE TABLE Retail.Orders (ID INTEGER, Place STRING)',
     'CREATE TABLE Dept.Sales (ID INTEGER, Name STRING)']
    

    【讨论】:

      【解决方案2】:

      如果循环是您想要的,那么是的 .iterrows() 是通过 pandas 框架的最有效方法。编辑:来自其他答案,并在此处链接 - Does iterrows have performance issues? - 我相信 .itertuples() 实际上是性能更好的生成器。

      但是,根据数据框的大小,您最好使用一些 pandas groupby 函数来提供帮助

      考虑这样的事情

      # Add a concatenation of the column name and format
      df['col_format'] =  df['Column'] + ' ' + df['Format']
      
      # Now create a frame which is the groupby of the DB/Table rows and 
      # concatenates the tuples of col_format correctly
      y1 = (df.groupby(by=['DB', 'Table'])['col_format']
              .apply(lambda x: '(' + ', '.join(x) + ')'))
      
      # Reset the index to bring the keys/indexes back in as columns
      y2 = y1.reset_index()
      
      # Now create a Series of all of the SQL statements
      all_outs = 'Create Table ' + y2['DB'] + '.' + y2['Table'] + ' ' + y2['col_format']
      
      # Look at them!
      all_outs.values
      Out[44]: 
      array(['Create Table Dept.Sales (ID INTEGER, Name STRING)',
             'Create Table Retail.Orders (ID INTEGER, Place STRING)'], dtype=object)
      

      希望这会有所帮助!

      【讨论】:

      • 我喜欢这个答案(并赞成),但要小心。对于大型数据帧,groupby 并不总是比 itertuples 快,例如看看this answer。请记住,groupby 应该具有 O(n log n) 复杂度,但其性能往往由函数驱动,而lambda 效率低下。
      【解决方案3】:

      您可以先将每行的信息组合到一个额外的列中,然后使用groupby.sum

      queries = df[KEY_COLUMN] + ' ' + df[KEY_FORMAT] + ', '
      queries.index = df.set_index(index_labels).index
      
      DB      Table 
      Retail  Orders      ID INTEGER, 
              Orders    Place STRING, 
      Dept    Sales       ID INTEGER, 
              Sales      Name STRING, 
      dtype: object
      
      queries = queries.groupby(index_labels).sum().str.strip(', ')
      
      DB      Table 
      Dept    Sales      ID INTEGER, Name STRING
      Retail  Orders    ID INTEGER, Place STRING
      dtype: object
      
      def format_queries(queries):
          query_pattern = 'CREATE TABLE %s.%s (%s)'
          for (db, table), text in queries.items():# idx, table, text
              query = query_pattern % (db, table, text)
              yield query
      list(format_queries(queries))
      
      ['CREATE TABLE Dept.Sales (ID INTEGER, Name STRING)',
       'CREATE TABLE Retail.Orders (ID INTEGER, Place STRING)']
      

      这样您就不需要lambda。我不知道这种方法或itertuples 是否会最快

      【讨论】:

        猜你喜欢
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 2016-10-27
        • 2019-02-12
        • 1970-01-01
        • 1970-01-01
        • 2020-11-20
        • 2020-07-05
        相关资源
        最近更新 更多