【问题标题】:How to replace lambda and grouping to increase performance with Pandas DataFrame如何使用 Pandas DataFrame 替换 lambda 和分组以提高性能
【发布时间】:2019-04-01 06:10:07
【问题描述】:

也许我的问题看起来很复杂,但本质上很简单。我是 Python 的新手,现在我面临代码太慢的问题。下面是代码的优化版本。我会很感激一个小的代码审查和关于如何加快它的建议。我认为最慢的操作是.apply(lambda和分组,但我不知道如何替换它们。

...
for raw_file in raw_files:
    reader = pd.read_csv(raw_file, chunksize=100000)
    for chunk in reader:
        processed_data = task(chunk)
        for name, data in processed_data:
            save_data(name, data) # some method which saves DataFrame correctly
...


def task(data):
    data = data[data['Quantity'] != 0] # remove zero items
    # add date parts as columns
    data[['dt_year', 'dt_month', 'dt_day', 'dt_day_of_year', 'dt_day_of_week', 'dt_hour']] = \
                data.apply(lambda df: to_date_parts(df['SalesDate']), axis=1)
    # group by location-item to aggregate in different files
    grouped = data.groupby(['LocationID','ItemID'])
    result = []
    for name, group in grouped:
        result += [(name, group)]
    return result



def to_date_parts(str_date):
    date = dt.datetime.strptime(str_date.split(".")[0], '%Y-%m-%d %H:%M:%S')
    dt_year = date.year
    dt_month = date.month
    dt_day = date.day
    dt_day_of_year = date.toordinal() - dt.datetime(date.year, 1, 1).toordinal() + 1
    dt_day_of_week = date.weekday()
    dt_hour = date.hour
    return pd.Series([dt_year, dt_month, dt_day, dt_day_of_year, dt_day_of_week, dt_hour])

【问题讨论】:

  • 这是XY Problem,您可以在其中寻求有关 y 解决方案的帮助,但未描述 x 问题。请提供您的情况的背景和总体情况,包括示例数据和输出以进行说明。

标签: python pandas performance datetime dataframe


【解决方案1】:

Python datetime vs Pandas datetime

您发现性能不佳有两个相互关联的原因:

  1. 您使用 Python 内置的 datetime 对象,而不是高效的 Pandas datetime 系列来存储日期。
  2. 您使用 Python 级别的 for 循环,而不是 Pandas datetime 系列支持的矢量化操作。

所以首先将您的系列转换为 Pandas datetime 系列:

date_format = '%Y-%m-%d %H:%M:%S'
df['SalesDate'] = pd.to_datetime(df['SalesDate'], format=date_format, errors='coerce')

然后直接从您的系列中提取属性:

from operator import attrgetter

# list attributes
fields = ['year', 'month', 'day', 'dayofyear', 'dayofweek', 'hour']

# extract attributes
attributes = pd.concat(attrgetter(*fields)(df['SalesDate'].dt), axis=1, keys=fields)

# join attributes to dataframe
df = df.join(attributes)

熊猫GroupBy对象

这种将项目串联到list 是不必要的:

grouped = data.groupby(['LocationID','ItemID'])
result = []
for name, group in grouped:
    result += [(name, group)]
return result

由于data.groupby(...) 是一个可迭代对象,你可以只使用return GroupBy 对象:

return data.groupby(['LocationID','ItemID'])

【讨论】:

  • 我需要试试这个。谢谢。我会更新结果
  • @AlexanderGoida,当然,我刚刚修正了一个错字并添加了一些关于 GroupBy 的信息。
猜你喜欢
  • 2016-11-01
  • 2020-03-31
  • 2012-03-25
  • 1970-01-01
  • 2011-12-29
  • 1970-01-01
  • 1970-01-01
  • 2022-11-23
  • 1970-01-01
相关资源
最近更新 更多