【问题标题】:Combine two python data processing scripts into a single work flow将两个 python 数据处理脚本组合成一个工作流
【发布时间】:2016-11-30 19:11:28
【问题描述】:

我目前正在处理一项数据处理任务。

我有两个 python 脚本,每个脚本实现一个单独的功能,但它们对相同的数据进行操作,我认为它们可以组合成一个工作流程,但我想不出最合乎逻辑的方式来实现这一点。

数据文件是here,它是 JSON,但它有两个不同的组件。

第一部分如下所示:

{
    "links": {
        "self": "http://localhost:2510/api/v2/jobs?skills=data%20science"
    },
    "data": [
        {
            "id": 121,
            "type": "job",
            "attributes": {
                "title": "Data Scientist",
                "date": "2014-01-22T15:25:00.000Z",
                "description": "Data scientists are in increasingly high demand amongst tech companies in London. Generally a combination of business acumen and technical skills are sought. Big data experience ..."
            },
            "relationships": {
                "location": {
                    "links": {
                        "self": "http://localhost:2510/api/v2/jobs/121/location"
                    },
                    "data": {
                        "type": "location",
                        "id": 3
                    }
                },
                "country": {
                    "links": {
                        "self": "http://localhost:2510/api/v2/jobs/121/country"
                    },
                    "data": {
                        "type": "country",
                        "id": 1
                    }
                },

这是由第一个 python 脚本处理的,这里:

import json
from collections import defaultdict
from pprint import pprint

with open('data-science.txt') as data_file:
    data = json.load(data_file)

locations = defaultdict(int)

for item in data['data']:
    location = item['relationships']['location']['data']['id']
    locations[location] += 1

pprint(locations)

呈现这种形式的数据:

         1: 6,
         2: 20,
         3: 2673,
         4: 126,
         5: 459,
         6: 346,
         8: 11,
         9: 68,
         10: 82,

这些是位置"id"s 和分配给该位置的记录数。

JSON 对象的另一部分如下所示:

"included": [
    {
        "id": 3,
        "type": "location",
        "attributes": {
            "name": "Victoria",
            "coord": [
                51.503378,
                -0.139134
            ]
        }
    },

并由这个python文件处理:

import json
from collections import defaultdict
from pprint import pprint

with open('data-science.txt') as data_file:
    data = json.load(data_file)

locations = defaultdict(int)

for record in data['included']:
    id = record.get('id', None)
    name = record.get('attributes', {}).get('name', None)
    coord = record.get('attributes', {}).get('coord', None)
    print(id, name, coord)

它以这种格式输出数据:

3 Victoria [51.503378, -0.139134]
1 United Kingdom None
71 data science None
32 None None
3 Victoria [51.503378, -0.139134]
1 United Kingdom None
1 data mining None
22 data analysis None
33 sdlc None
38 artificial intelligence None
39 machine learning None
40 software development None
71 data science None
93 devops None
63 None None
52 Cubitt Town [51.505199, -0.018848]

我真正想要的是最终输出看起来像这样:

3, Victoria, [51.503378, -0.139134], 2673

2673 引用第一个脚本中的作业计数。

如果它没有任何坐标,例如[51.503378, -0.139134]我可以扔了。

我确信将这些脚本组合在一起并获得输出是可能的,但我不是一个全面的思考者,我不知道如何去做。

所有真实的项目文件live here

【问题讨论】:

    标签: python json


    【解决方案1】:

    使用functions 是组合这两个脚本的一种方法,毕竟它们处理相同的数据。所以,你应该为每个处理逻辑块做一个函数,然后将结果组合到最后:

    import json
    from collections import defaultdict
    from pprint import pprint
    
    def process_locations_data(data):
        # processes the 'data' block
        locations = defaultdict(int)
        for item in data['data']:
            location = item['relationships']['location']['data']['id']
            locations[location] += 1
        return locations
    
    def process_locations_included(data):
        # processes the 'included' block
        return_list = []
        for record in data['included']:
            id = record.get('id', None)
            name = record.get('attributes', {}).get('name', None)
            coord = record.get('attributes', {}).get('coord', None)
            return_list.append((id, name, coord))
        return return_list    # return list of tuples
    
    # load the data from file once
    with open('data-science.txt') as data_file:
        data = json.load(data_file)
    
    # use the two functions on same data
    locations = process_locations_data(data)
    records = process_locations_included(data)
    
    # combine the data for printing
    for record in records:
        id, name, coord = record
        references = locations[id]   # lookup the references in the dict
        print id, name, coord, references
    

    该函数可以有更好的名称,但这应该实现您正在寻找的统一。

    【讨论】:

    • 脚本可以运行,但是当您尝试将其通过管道传输到输出文件时,您会收到错误 UnicodeEncodeError: 'ascii' codec can't encode character u'\xfc' in position 8: ordinal not in range(128)
    • 这与输入数据有关。它可以被处理,例如在这里阅读:stackoverflow.com/questions/5760936/… 但是你也可以只保存到文件而不是在最后一个循环中使用print
    猜你喜欢
    • 2016-09-03
    • 2014-08-26
    • 2014-10-02
    • 2021-02-13
    • 1970-01-01
    • 2021-03-02
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多