【问题标题】:TypeError: normalize() argument 2 must be str, not Series with a dataframe of stringsTypeError: normalize() 参数 2 必须是 str,而不是带有字符串数据框的 Series
【发布时间】:2019-06-28 02:10:17
【问题描述】:

我有一个包含每天新闻的数据框,我尝试分析当天的感受强度,也就是说,从新闻中获得的总体感受是积极的、消极的还是中性的。这是df_news 数据框:

    Date    name
0   2017-10-20  Gucci debuts art installation at its Ginza sto...
1   2018-08-01  Gucci Joins Paris Fashion Week for Its Spring ...
2   2018-04-20  Gucci launches its new creative hub Gucci ArtL...
3   2017-10-20  Gucci to launch homeware line Gucci Decor - CP...
4   2017-12-07  GUCCI opens new store at Miami Design District...
5   2018-01-12  Gucci opens Gucci Garden in Florence - LUXUO
6   2018-02-26  GUCCI's wild experiment with the Fall Winter 2...
7   2018-08-09  Gucci Revamped London Flagship Store | The Imp...
8   2018-08-01  Alessandro Michele Announces new Gucci Home co...
9   2017-10-20  Before He Picks Up the CFDA’s International Aw...

我试图用他使用的以下代码SentimentIntensityAnalyzer de nltk.sentiment.vader 来获得这种感觉的强度:

from nltk.sentiment.vader import SentimentIntensityAnalyzer
import unicodedata
sid = SentimentIntensityAnalyzer()
for date, row in df_news.T.iteritems():
    try:
        sentence = unicodedata.normalize('NFKD', df_news.loc[date, 'name']).encode('ascii','ignore')
        #print((sentence))
        ss = sid.polarity_scores(str(sentence))
        df_news.set_value(date, 'compound', ss['compound'])
        df_news.set_value(date, 'neg', ss['neg'])
        df_news.set_value(date, 'neu', ss['neu'])
        df_news.set_value(date, 'pos', ss['pos'])
    except TypeError:
        print(df_news.loc[date, 'name'])
        print(date)

但是,我在某些日期收到 TypeError。感谢try catch您不考虑并绘制下表:

    name    compound    neg neu pos
Date                    
2017-10-20  Gucci debuts art installation at its Ginza sto...               
2018-08-01  Gucci Joins Paris Fashion Week for Its Spring ...               
2018-04-20  Gucci launches its new creative hub Gucci ArtL...   0.4404  0   0.756   0.244
2017-10-20  Gucci to launch homeware line Gucci Decor - CP...               
2017-12-07  GUCCI opens new store at Miami Design District...   0   0   1   0
2018-01-12  Gucci opens Gucci Garden in Florence - LUXUO    0   0   1   0
2018-02-26  GUCCI's wild experiment with the Fall Winter 2...   0   0   1   0
2018-08-09  Gucci Revamped London Flagship Store | The Imp...   0.3182  0   0.602   0.398
2018-08-01  Alessandro Michele Announces new Gucci Home co...               
2017-10-20  Before He Picks Up the CFDA’s International Aw...               

但是当我删除 try catch 以了解它失败的原因时,我收到以下错误:

---------------------------------------------------------------------------
TypeError                                 Traceback (most recent call last)
<ipython-input-26-2e9dbfc62bce> in <module>
      4 for date, row in df_news.T.iteritems():
      5 #    try:
----> 6     sentence = unicodedata.normalize('NFKD', df_news.loc[date, 'name']).encode('ascii','ignore')
      7     #print((sentence))
      8     ss = sid.polarity_scores(str(sentence))

TypeError: normalize() argument 2 must be str, not Series

然后我认为问题在于不是字符串的行,而是例如第一行:

>>>type(df_news['name'][0])
str

获取数据

doc_data = {
  "size": 10,
  "query": {
    "bool": {
      "must" : [
       {"term":{"text":"gucci"}}
     ]
    }
  }
 }

docs = create_doc("https://elastic:rKzWu2WbXI@db.luxurynsight.com/luxurynsight_v2/news/_search",doc_data)


information_df = pd.DataFrame.from_dict(docs.json()["hits"]["hits"])

# Reading the JSON file
df_news = pd.read_json('data.json')

# Converting the element wise _source feature datatype to dictionary
df_news._source = df_news._source.apply(lambda x: dict(x))

# Creating name column
df_news['name'] = df_news._source.apply(lambda x: x['name'])

# Creating createdAt column
df_news['createdAt'] = df_news._source.apply(lambda x: x['createdAt'])

df_news['createdAt'] =  pd.to_datetime(df_news['createdAt'], unit='ms')

df_news['createdAt'] = pd.DatetimeIndex(df_news.createdAt).normalize()
#df_news.createdAt.dt.normalize()

df_news['Date'] = df_news['createdAt']

df_news = df_news[['name','Date']]
df_news = df_news.set_index('Date')
information_df._source = information_df.apply(lambda x: dict(x))
df_news.reset_index()

它应该回馈:

    Date    name
0   2017-10-20  Gucci debuts art installation at its Ginza sto...
1   2018-08-01  Gucci Joins Paris Fashion Week for Its Spring ...
2   2018-04-20  Gucci launches its new creative hub Gucci ArtL...
3   2017-10-20  Gucci to launch homeware line Gucci Decor - CP...
4   2017-12-07  GUCCI opens new store at Miami Design District...
5   2018-01-12  Gucci opens Gucci Garden in Florence - LUXUO
6   2018-02-26  GUCCI's wild experiment with the Fall Winter 2...
7   2018-08-09  Gucci Revamped London Flagship Store | The Imp...
8   2018-08-01  Alessandro Michele Announces new Gucci Home co...
9   2017-10-20  Before He Picks Up the CFDA’s International Aw...

编辑:

我将同一天出现的文章分组并放入列表中

# get date out of the index to column    
df_news = df_news.reset_index()
# optional
df_news['Date'] = pd.to_datetime(df_news['Date'])
# groupby and output group rows as list
df_news = df_news.groupby('Date')['name'].apply(list)
df_news.head()

它给了我回报:

Date
2017-10-20    [Gucci debuts art installation at its Ginza st...
2017-12-07    [GUCCI opens new store at Miami Design Distric...
2018-01-12       [Gucci opens Gucci Garden in Florence - LUXUO]
2018-02-26    [GUCCI's wild experiment with the Fall Winter ...
2018-04-20    [Gucci launches its new creative hub Gucci Art...
2018-08-01    [Gucci Joins Paris Fashion Week for Its Spring...
2018-08-09    [Gucci Revamped London Flagship Store | The Im...
Name: name, dtype: object

因此,当我尝试应用 Stael 的答案时:

sentence = df_news.loc[date, 'name'].apply(lambda x: unicodedata.normalize('NFKD', x).encode('ascii','ignore'))

也就是说对系列中的每一项进行归一化

我收到以下错误:

---------------------------------------------------------------------------
IndexingError                             Traceback (most recent call last)
<ipython-input-173-1bc93a0a065c> in <module>
      5     try:
      6         #sentence = unicodedata.normalize('NFKD', df_news.loc[date, 'name']).encode('ascii','ignore')
----> 7         sentence = df_news.loc[date, 'name'].apply(lambda x: unicodedata.normalize('NFKD', x).encode('ascii','ignore'))
      8         ss = sid.polarity_scores(str(sentence))
      9         df_news.set_value(date, 'compound', ss['compound'])

C:\ProgramData\Anaconda3\lib\site-packages\pandas\core\indexing.py in __getitem__(self, key)
   1470             except (KeyError, IndexError):
   1471                 pass
-> 1472             return self._getitem_tuple(key)
   1473         else:
   1474             # we by definition only have the 0th axis

C:\ProgramData\Anaconda3\lib\site-packages\pandas\core\indexing.py in _getitem_tuple(self, tup)
    873 
    874         # no multi-index, so validate all of the indexers
--> 875         self._has_valid_tuple(tup)
    876 
    877         # ugly hack for GH #836

C:\ProgramData\Anaconda3\lib\site-packages\pandas\core\indexing.py in _has_valid_tuple(self, key)
    218         for i, k in enumerate(key):
    219             if i >= self.obj.ndim:
--> 220                 raise IndexingError('Too many indexers')
    221             try:
    222                 self._validate_key(k, i)

IndexingError: Too many indexers

当我尝试只在索引中使用 Date 时,我得到了句子 = df_news.loc[date].aply(x: ...:

---------------------------------------------------------------------------
AttributeError                            Traceback (most recent call last)
<ipython-input-176-308d1f6c6644> in <module>
      5     try:
      6         #sentence = unicodedata.normalize('NFKD', df_news.loc[date, 'name']).encode('ascii','ignore')
----> 7         sentence = df_news.loc[date].apply(lambda x: unicodedata.normalize('NFKD', x).encode('ascii','ignore'))
      8         ss = sid.polarity_scores(str(sentence))
      9         df_news.set_value(date, 'compound', ss['compound'])

AttributeError: 'list' object has no attribute 'apply'

【问题讨论】:

    标签: python python-3.x string nltk typeerror


    【解决方案1】:

    看到您在一篇帖子中出现了第三个错误,我想我会再做一次。

    首先,我觉得您对自己的很多代码都不了解。像AttributeError: 'list' object has no attribute 'apply' 这样的错误意味着你在操作变量时不知道你的变量是什么,所以我认为你需要更慢更仔细地工作,以了解代码的每一位在你继续之前在做什么到下一节。

    也就是说,你的问题并不像你做的那么复杂——你正在尝试应用这两行代码

        sentence = unicodedata.normalize('NFKD', df_news.loc[date, 'name']).encode('ascii','ignore')
    
        ss = sid.polarity_scores(str(sentence))
    

    到数据框中“名称”列中的每个条目,这并不难。

    你可以像这样轻松地做到这一点:

    scores = []
    for entry in df['name']:
        sentence = unicodedata.normalize('NFKD', entry).encode('ascii','ignore')
        scores.append(sid.polarity_scores(str(sentence)))
    

    这将为您提供您调用ss的分数词典列表

    您可以像这样将它们应用为数据框中的列:

    df['new_column'] = [i['example_key'] for i in scores]
    

    这不是最好或最有效的方法,但它是一种非常简单的方法,可以让您实现您想要做的事情。

    祝你好运。


    如果您以前按天分组,并列出了您的字符串(顺便说一句,我认为您不应该这样做),那么您需要另一层迭代

    scores = []
    for sentence_list in df['name']:
        for entry in sentence_list:
            sentence = unicodedata.normalize('NFKD', entry).encode('ascii','ignore')
            scores.append(sid.polarity_scores(str(sentence)))
    

    【讨论】:

    • 感谢您的帮助。但是我仍然有一个TypeError 和sentence = unicodedata.normalize('NFKD', entry).encode('ascii','ignore'),因为条目是一个列表。但是,当我执行 sentence = df_news.loc[date, 'name'] 时,我可以在某个句子上应用 ss = sid.polarity_scores(str(sentence)),而另一个句子似乎不起作用,因为 neu 的返回分数为 1。
    • 我认为您没有任何理由将字符串分组到列表中。我认为您这样做是因为您在日期索引中有重复项,但这并不是一个真正的问题 - 这种方法应该可以很好地处理重复的日期。
    【解决方案2】:

    在我看来就是这一行:

    sentence = unicodedata.normalize('NFKD', df_news.loc[date, 'name']).encode('ascii','ignore')
    

    您正在尝试对 df.news.loc[...] 系列中的每个项目调用 normalize

    但是 pandas 不会为您在整个系列中应用该功能 - 我认为您想要做的是这样的:

    sentence = df_news.loc[date, 'name'].apply(lambda x: unicodedata.normalize('NFKD', x).encode('ascii','ignore')
    

    这是对系列中的每个项目应用函数(规范化)的 pandas 方式。


    编辑:

    理论 2 - 当您调用 df_news.loc[date, 'name'] 时,您正在选择 index==date 和 column=='name' 所在的项目 - 但从您的问题来看,某些日期在您的索引中重复,这意味着有时而不是获取您可以在其中调用 unicodedata.normalize 的单个记录,您将得到一个系列,这会导致您的错误。

    您会注意到,当您使用 'try: except:` 子句时未填写的记录是具有重复日期的记录。

    您需要以某种方式处理它,也许通过使用 row out of iteritems 而不是日期,但这需要您自己解决。

    【讨论】:

    • 嗯,然后在ss = sid.polarity_scores(str(sentence))上创建一个SyntaxError: invalid syntaxe
    • 你期望变量句是什么?在第一种情况下,您正在对系列进行操作,因此您会期望从中产生类似系列的内容-您不能使用 str 或系列,这是没有意义的。
    • 对不起!!我缺少括号,实际错误是 AttributeError: 'str' object has no attribute 'apply' on df_news.loc[date, 'name'].apply(lambda ...
    • 好的,我想我开始明白了——我认为df_news.loc[date, 'name'] 有时会给你一个字符串,有时会给你一个Series。我从您的问题中看到日期'2017-10-20' 在索引中出现了两次。在这种情况下,你会得到一个系列而不是一个字符串。您需要以某种方式处理它,然后才能normalize 它。 (我认为)
    • @ThePassenger 我已经编辑了我的答案以使其更清晰。
    猜你喜欢
    • 2021-02-14
    • 2020-10-12
    • 2021-05-17
    • 2018-08-02
    • 2013-12-16
    • 1970-01-01
    • 2020-12-03
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多