【问题标题】:How can I create a dynamic dictionaries in Python using a CSV file如何使用 CSV 文件在 Python 中创建动态字典
【发布时间】:2016-07-10 05:32:00
【问题描述】:

问题很简单,我有一个包含四列的 CSV 文件,我想将第一列的值设置为我的 python 脚本中的字典。我不想将值添加到已完成任务的字典日期中。

为此的 CSV 数据可在名为 VC.csv 的文件中找到,例如:

24M Technologies,Series A,8/19/10
24M Technologies,Grant,8/16/10
2B Energy,Private Equity,3/18/14
2B Energy,Series B,3/18/14
2B Energy,Unattributed VC,5/1/08
3GSolar Photovoltaics,Series A,12/17/12
3sun Group,Growth Equity,3/3/14
3Tier Group,Series C,11/17/08

我想要的最终结果是我打印时的字典。

例如

>>> print 3TierGroup
>>>
>>>{'company': '2B Energy', 'Private Equity': '3/18/14', 'Series B': '3/18/14', 'Unattributed VC': '05/01/08'}

我的问题是尝试循环槽并向已定义的字典添加更多内容。我猜我不是在追加,而是在每次通过时都在重新创建和覆盖循环。结果我得到的是{'company': '2B Energy', 'Private Equity': '3/18/14'} 我需要代码的最后一行来测试字典是否已经存在;如果是这样,它将附加额外的圆形日期。

这是我的代码...

import csv

companyList =[]
transactionDates=[]
dictNames=[]

def fileNameCleaner(namer):
    namer = namer.replace(' ', '')
    namer = namer.replace(',','')
    namer = namer.replace('-','')
    namer = namer.replace('.','')
    namer = namer.replace('_','')
    namer = namer.replace('@','')
    namer = namer.replace('&','')
    return namer

with open('VC.csv', 'rb') as rawData:
    timelineData = csv.reader(rawData, delimiter=',', quotechar='"')      # Open CSV file and snag data
    for line in timelineData:  # Run through each row in csv
        companyList.append(fileNameCleaner(line[0])) # Create list and remove some special charcters
    companyList = list(set(companyList))    # Remove duplicates and Sort

for companyListRow in companyList:
    with open('VC.csv', 'rb') as rawDataTwo:
        timelineDataTwo = csv.reader(rawDataTwo, delimiter=',', quotechar='"')
        for TList in timelineDataTwo:
            company = TList[0]
            finRound = TList[1]
            tranDate = TList[2]
            if companyListRow == fileNameCleaner(TList[0]):
                companyListRow = {'company':TList[0], finRound:tranDate }
                print companyListRow

【问题讨论】:

  • 我认为这种数据最好在 SQL 数据库中表示和查询(想想 SQLite),因为融资类型(seriesA、seriesB 等)会经常重复并用掉很多不必要的存储。还要查询数据,找出哪些资金、以什么顺序最好在 SQL 数据库中提供服务(一张表用于公司,一张用于资金类型,一张用于外键为公司和资金类型的日期)。跨度>

标签: python python-2.7 csv dictionary append


【解决方案1】:

读取您的数据:

str1='''24M Technologies,Series A,8/19/10
24M Technologies,Grant,8/16/10
2B Energy,Private Equity,3/18/14
2B Energy,Series B,3/18/14
2B Energy,Unattributed VC,5/1/08
3GSolar Photovoltaics,Series A,12/17/12
3sun Group,Growth Equity,3/3/14
3Tier Group,Series C,11/17/08'''

list1= str1.split('\n')
print list1

我认为您只需要一本字典(不多),因此您可以按公司名称查找数据,例如像这样:

comps={}
for abc in list1:
    a,b,c=abc.split(',')
    if not a in comps:
        comps[a]= [[b,c]]
    else:
        comps[a].append([b,c])

for k,v in comps.iteritems():
    print k,v

输出:

3sun Group [['Growth Equity', '3/3/14']]
2B Energy [['Private Equity', '3/18/14'], ['Series B', '3/18/14'], ['Unattributed VC', '5/1/08']]
3Tier Group [['Series C', '11/17/08']]
3GSolar Photovoltaics [['Series A', '12/17/12']]
24M Technologies [['Series A', '8/19/10'], ['Grant', '8/16/10']]

您的字典条目的值将是一个事件列表,每个事件都是一个列表,首先是类型,然后是日期。

【讨论】:

    【解决方案2】:

    我认为这段代码将总结您的公司数据,只需通过 CSV:

    # define dict (to be keyed by company name) to accumulate company attributes from CSV file
    company_data = {}
    
    with open('VC.csv', 'rb') as rawData:
        # Open CSV file and snag data
        timelineData = csv.DictReader(rawData, delimiter=',', quotechar='"',
                                      fieldnames=['company','key','value'])
    
        # Run through each row in csv
        for line in timelineData:  
            name = filenameCleaner(line['company'])
            # get record for previously seen company, or get a new one with just the name in it
            rec = company_data.get(name, {'company': name})
    
            # add this line's key-value to the rec for this company
            rec[line['key']] = line['value']
    
            # stuff updated rec back into the overall summarizing dict
            company_data[name] = rec
    
    # now get the assembled records by getting just the values from the summarizing dict
    company_recs = company_data.values()
    

    【讨论】:

      猜你喜欢
      • 2019-01-06
      • 2021-05-15
      • 1970-01-01
      • 2020-04-17
      • 2013-11-18
      • 1970-01-01
      • 2020-03-12
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多