【问题标题】:Creating n-grams word cloud using python使用 python 创建 n-gram 词云
【发布时间】:2017-07-20 00:12:01
【问题描述】:

我正在尝试使用二元语法生成词云。我能够生成前 30 个判别词,但无法在绘图时一起显示单词。我的词云图像看起来仍然像一个 uni-gram 云。我使用了以下脚本和 sci-kit 学习包。

def create_wordcloud(pipeline): 
"""
Create word cloud with top 30 discriminative words for each category
"""

class_labels = numpy.array(['Arts','Music','News','Politics','Science','Sports','Technology'])

feature_names =pipeline.named_steps['vectorizer'].get_feature_names() 
word_text=[]

for i, class_label in enumerate(class_labels):
    top30 = numpy.argsort(pipeline.named_steps['clf'].coef_[i])[-30:]

    print("%s: %s" % (class_label," ".join(feature_names[j]+"," for j in top30)))

    for j in top30:
        word_text.append(feature_names[j])
    #print(word_text)
    wordcloud1 = WordCloud(width = 800, height = 500, margin=10,random_state=3, collocations=True).generate(' '.join(word_text))

    # Save word cloud as .png file
    # Image files are saved to the folder "classification_model" 
    wordcloud1.to_file(class_label+"_wordcloud.png")

    # Plot wordcloud on console
    plt.figure(figsize=(15,8))
    plt.imshow(wordcloud1, interpolation="bilinear")
    plt.axis("off")
    plt.show()
    word_text=[]

这是我的管道代码

pipeline = Pipeline([

# SVM using TfidfVectorizer
('vectorizer', TfidfVectorizer(max_features = 25000, ngram_range=(2, 2),sublinear_tf=True, max_df=0.95, min_df=2,stop_words=stop_words1)),
('clf',       LinearSVC(loss='squared_hinge', penalty='l2', dual=False, tol=1e-3))
])

这些是我为“艺术”类别获得的一些功能

Arts: cosmetics businesspeople, television personality, reality television, television presenters, actors london, film producers, actresses television, indian film, set index, actresses actresses, television actors, century actors, births actors, television series, century actresses, actors television, stand comedian, television personalities, television actresses, comedian actor, stand comedians, film actresses, film actors, film directors

【问题讨论】:

    标签: python scikit-learn n-gram word-cloud


    【解决方案1】:

    我认为您需要以某种方式将您的 n-gramms 加入 feature_names 与除空格之外的任何其他符号。例如,我建议使用下划线。 现在,这部分让你的 n-grams 再次分离单词,我认为:

    ' '.join(word_text)
    

    我认为您必须在这里用下划线替换空格:

    word_text.append(feature_names[j])
    

    改成这样:

    word_text.append(feature_names[j].replace(' ', '_'))
    

    【讨论】:

    • 它没有用。它用 (_) 替换所有单词,没有任何中断。
    • 我编辑了我的答案。你试过这样的事情吗?
    猜你喜欢
    • 2017-09-21
    • 2019-05-14
    • 2013-10-06
    • 1970-01-01
    • 2018-09-09
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2021-03-14
    相关资源
    最近更新 更多