【问题标题】:Pyspark dataframe corrupted record when reading from python dictionary(json) got from requests, encoding problem从请求中获取的python字典(json)读取时,Pyspark数据帧损坏记录,编码问题
【发布时间】:2020-09-30 19:46:09
【问题描述】:

我正在使用 Requests 库进行 REST api 调用。

response = requests.get("https://urltomaketheapicall", headers={'authorization': 'bearer {0}'.format("7777777777777777777777777777")}, timeout=5)

当我做response.json()

我得到一个带有这些值的键

{'devices': '....iPhone\xa05S, iPhone\xa06, iPhone\xa06\xa0Plus, iPhone\xa06S'}

当我执行print(response.encoding) 时,我得到None

当我执行print(type(data[devices])) 时,我得到<class 'str'>

如果我使用print(data[devices]),我会得到不带特殊字符的'....iPhone 5S, iPhone 6, iPhone 6 Plus, iPhone 6S'。

现在如果这样做

new_dict={}
new_val = data[devices]
new_dict["devices"] = new_val
print(new_dict["devices"])

我也会得到新字典中的特殊字符。

有什么想法吗?

我想摆脱特殊字符,因为我需要读取这些 json 并将其放入 pyspark 数据帧中,使用这些字符我会得到 _corrupted_record

rd= spark.sparkContext.parallelize([data])
df = spark.read.json(rd)

我想避免像.replace("\\xa0"," ")这样的解决方案

【问题讨论】:

    标签: python apache-spark encoding pyspark python-requests


    【解决方案1】:

    A0 是一个不间断的空格。它只是字符串的一部分。它只是这样打印,因为您正在倾倒整个 dict 的 repr。如果您打印单个字符串,它将简单地打印为正确的不间断空格:

    >>> print({'a': '\xa0'})
    {'a': '\xa0'}
    >>> print('\xa0')
     
    >>>
    

    【讨论】:

    • 检查我的编辑,特殊字符我不能把它放在 pyspark 数据帧中
    • 然后专门集中另一个问题。我不知道 pyspark 以及您是否只是做错了,或者它是否根本无法处理不间断空间。
    猜你喜欢
    • 2021-05-09
    • 1970-01-01
    • 2021-08-21
    • 1970-01-01
    • 2019-08-12
    • 2019-08-02
    • 1970-01-01
    • 2016-12-05
    • 1970-01-01
    相关资源
    最近更新 更多