【问题标题】:Need help in extracting data from list of dictionaries with time lines and counts需要帮助从具有时间线和计数的字典列表中提取数据
【发布时间】:2021-09-22 04:20:04
【问题描述】:

您好,我希望从 python 字典列表中计算一些东西,其中数据如下所示。当 Name="TOM" 我想要城市、国家键和国家的计数。我想从下面的格式表中计算的文件在数据后明确提到,请查看并提出最佳的计算方法

people = [
    {"name": "Tom", "age": 10, "city": "NewYork", "Date":2021-01-04 08:37:19Z},
    {"name": "Mark", "age": 5, "country": "Japan", "Date": 2021-01-06 08:37:24Z},
    {"name": "Pam", "age": 7, "city": "London", "Date": 2021-01-04 09:26:38Z},
    {"name": "Tom", "hight": 163, "city": "California", "Date": 2021-01-08 12:50:17Z},
    {"name": "Lena", "weight": 45, "country": "Italy", "Date": 2021-01-08 12:50:17Z},
    {"name": "Ben", "age": 17, "city": "Colombo", "Date": 2021-01-21 09:56:04Z},
    {"name": "Lena", "gender": "Female", "country": "Italy", "Date": 2021-01-21 09:56:04Z},
    {"name": "Ben", "gender": "Male", "city": "Colombo", "Date": 2021-02-09 08:47:26Z},
    {"name": "Tom", "age": 10, "country": "Italy", "Date": 2021-02-09 08:47:26Z},
    {"name": "Mark", "age": 5, "country": "Japan", "Date": 2021-02-23 09:10:59Z},
    {"name": "Tom", "age": 7, "city": "London", "Date": 2021-03-08 09:39:28Z},
    {"name": "Tom", "hight": 163, "country": "Japan", "Date": 2021-03-08 09:39:28Z},
]

期望以下格式的输出,

姓名:12

汤姆:5

城市:3

时间线的数据计数,

城市:纽约

1 月 1 日、2 月 2 日、3 月 2 日

城市:伦敦

1 月 2 日、2 月 1 日、3 月 4 日

I have 1000s of these data in list of dictionaries format with lot of other parameters. I am new to this and can some please help getting this solved 

【问题讨论】:

    标签: python pandas list dictionary datetime


    【解决方案1】:

    这只是一步一步思考“我有什么”和“我需要什么”的问题。

    people = [
        {"name": "Tom", "age": 10, "city": "NewYork", "Date": '01/01/2021'},
        {"name": "Mark", "age": 5, "country": "Japan", "Date": '05/01/2021'},
        {"name": "Pam", "age": 7, "city": "London", "Date": '03/06/2021'},
        {"name": "Tom", "hight": 163, "city": "California", "Date": '04/06/2021'},
        {"name": "Lena", "weight": 45, "country": "Italy", "Date": '12/12/2020'},
        {"name": "Ben", "age": 17, "city": "Colombo", "Date": '11/12/2020'},
        {"name": "Lena", "gender": "Female", "country": "Italy", "Date": '8/01/2020'},
        {"name": "Ben", "gender": "Male", "city": "Colombo", "Date": '7/01/2020'},
        {"name": "Tom", "age": 10, "country": "Italy", "Date": '01/01/2021'},
        {"name": "Mark", "age": 5, "country": "Japan", "Date": '05/01/2021'},
        {"name": "Tom", "age": 7, "city": "London", "Date": '03/06/2021'},
        {"name": "Tom", "hight": 163, "country": "Japan", "Date": '04/06/2021'}
    ]
    
    def groupby( fld ):
        vals = { fld: 0 }
        for row in people:
            if fld in row:
                vals[fld] += 1
                if row[fld] not in vals:
                    vals[row[fld]] = 1
                else:
                    vals[row[fld]] += 1
        return vals
    
    months = ('Jan','Feb','Mar','Apr','May','Jun','Jul','Aug','Sep','Oct','Nov','Dec')
    def groupbydate( fld ):
        vals = {}
        for row in people:
            if fld in row and 'Date' in row:
                month = months[int(row['Date'].lstrip('0').split('/')[0])-1]
                if row[fld] not in vals:
                    vals[row[fld]] = {}
                if month not in vals[row[fld]]:
                    vals[row[fld]][month] = 1
                else:
                    vals[row[fld]][month] += 1
        return vals
    
    print( groupby( 'name' ) )
    print( groupby( 'city' ) )
    print( groupby( 'country' ) )
    print( )
    print( groupbydate( 'city' ) )
    

    输出:

    {'name': 12, 'Tom': 5, 'Mark': 2, 'Pam': 1, 'Lena': 2, 'Ben': 2}
    {'city': 6, 'NewYork': 1, 'London': 2, 'California': 1, 'Colombo': 2}
    {'country': 6, 'Japan': 3, 'Italy': 3}
    
    {'NewYork': {'Jan': 1}, 'London': {'Mar': 2}, 'California': {'Apr': 1}, 'Colombo': {'Nov': 1, 'Jul': 1}}
    

    使用collections.defaultdict 会缩短一点:

    from collections import defaultdict
    
    def groupby( fld ):
        vals = defaultdict(int)
        for row in people:
            if fld in row:
                vals[fld] += 1
                vals[row[fld]] += 1
        return dict(vals)
    
    months = ('Jan','Feb','Mar','Apr','May','Jun','Jul','Aug','Sep','Oct','Nov','Dec')
    def groupbydate( fld ):
        vals = {}
        for row in people:
            if fld in row and 'Date' in row:
                if row[fld] not in vals:
                    vals[row[fld]] = defaultdict(int)
                month = months[int(row['Date'].lstrip('0').split('/')[0])-1]
                vals[row[fld]][month] += 1
        return vals
    
    print( groupby( 'name' ) )
    print( groupby( 'city' ) )
    print( groupby( 'country' ) )
    print( groupbydate( 'city' ) )
    

    跟进添加年份

    def groupbyyear( fld ):
        vals = {}
        for row in people:
            if fld in row and 'Date' in row:
                if row[fld] not in vals:
                    vals[row[fld]] = defaultdict(int)
                year = int(row['Date'].split('-')[0])
                vals[row[fld]][year] += 1
        return vals
    
    print( groupby( 'name' ) )
    print( groupby( 'city' ) )
    print( groupby( 'country' ) )
    print( groupbydate( 'city' ) )
    print( groupbyyear( 'city' ) )
    

    【讨论】:

    • 谢谢...因为日期是 2021-01-04 08:35:58Z ...无法获取 groupdate 函数值...我该如何解决它
    • 这就是为什么您应该始终让您的样本数据与您的实际数据相匹配。您可以通过将.lstrip('0').split('/')[0] 替换为.split('-')[1].lstrip('0') 来将其作为字符串。
    • 哇...解决了这个问题。最后一件事,如果我想包括年份,我可以添加吗
    • 我将把它作为练习留给读者。你可以看到我做了什么作为一个模式。
    • 试过但做不到
    猜你喜欢
    • 1970-01-01
    • 2021-09-21
    • 1970-01-01
    • 2021-09-21
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2021-11-20
    • 1970-01-01
    相关资源
    最近更新 更多