【问题标题】:Two python loops that look like they should do the same thing, but output different results?两个看起来应该做同样事情但输出不同结果的python循环?
【发布时间】:2019-06-16 10:31:12
【问题描述】:

昨天我试图完成 Udacity 的第 11 课,关于文本矢量化。我检查了代码,一切似乎都运行良好 - 我接收了一些电子邮件,打开它们,删除一些签名词并将每封电子邮件的词干词返回到一个列表中。

这是循环 1:

for name, from_person in [("sara", from_sara), ("chris", from_chris)]:
    for path in from_person:
        ### only look at first 200 emails when developing
        ### once everything is working, remove this line to run over full dataset
#        temp_counter += 1
    if temp_counter < 200:
        path = os.path.join('/xxx', path[:-1])
        email = open(path, "r")

        ### use parseOutText to extract the text from the opened email

        email_stemmed = parseOutText(email)

        ### use str.replace() to remove any instances of the words
        ### ["sara", "shackleton", "chris", "germani"]

        email_stemmed.replace("sara","")
        email_stemmed.replace("shackleton","")
        email_stemmed.replace("chris","")
        email_stemmed.replace("germani","")

    ### append the text to word_data

    word_data.append(email_stemmed.replace('\n', ' ').strip())

    ### append a 0 to from_data if email is from Sara, and 1 if email is from Chris
        if from_person == "sara":
            from_data.append(0)
        elif from_person == "chris":
            from_data.append(1)

    email.close()

这是循环 2:

for name, from_person in [("sara", from_sara), ("chris", from_chris)]:
    for path in from_person:
        ### only look at first 200 emails when developing
        ### once everything is working, remove this line to run over full dataset
#        temp_counter += 1
        if temp_counter < 200:
            path = os.path.join('/xxx', path[:-1])
            email = open(path, "r")

            ### use parseOutText to extract the text from the opened email
            stemmed_email = parseOutText(email)

            ### use str.replace() to remove any instances of the words
            ### ["sara", "shackleton", "chris", "germani"]
            signature_words = ["sara", "shackleton", "chris", "germani"]
            for each_word in signature_words:
                stemmed_email = stemmed_email.replace(each_word, '')         #careful here, dont use another variable, I did and broke my head to solve it

            ### append the text to word_data
            word_data.append(stemmed_email)

            ### append a 0 to from_data if email is from Sara, and 1 if email is from Chris
            if name == "sara":
                from_data.append(0)
            else: # its chris
                from_data.append(1)


            email.close()

代码的下一部分按预期工作:

print("emails processed")
from_sara.close()
from_chris.close()

pickle.dump( word_data, open("/xxx/your_word_data.pkl", "wb") )
pickle.dump( from_data, open("xxx/your_email_authors.pkl", "wb") )


print("Answer to Lesson 11 quiz 19: ")
print(word_data[152])


### in Part 4, do TfIdf vectorization here

from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.feature_extraction import stop_words
print("SKLearn has this many Stop Words: ")
print(len(stop_words.ENGLISH_STOP_WORDS))

vectorizer = TfidfVectorizer(stop_words="english", lowercase=True)
vectorizer.fit_transform(word_data)

feature_names = vectorizer.get_feature_names()

print('Number of different words: ')
print(len(feature_names))

但是当我使用循环 1 计算单词总数时,我得到了错误的结果。当我使用循环 2 执行此操作时,我得到了正确的结果。

我查看这段代码太久了,但我无法发现其中的区别 - 我在循环 1 中做错了什么?

为了记录,我一直得到的错误答案是 38825。正确答案应该是 38757。

非常感谢您的帮助,好心的陌生人!

【问题讨论】:

    标签: python-3.x machine-learning tfidfvectorizer


    【解决方案1】:

    这些行没有做任何事情:

    email_stemmed.replace("sara","")
    email_stemmed.replace("shackleton","")
    email_stemmed.replace("chris","")
    email_stemmed.replace("germani","")
    

    replace 返回一个新字符串并且不修改email_stemmed。相反,您应该将返回值设置为 email_stemmed:

    email_stemmed = email_stemmed.replace("sara", "")
    

    以此类推。

    循环二确实在for循环中设置了返回值:

    for each_word in signature_words:
        stemmed_email = stemmed_email.replace(each_word, '')
    

    上面的代码 sn-ps 不等价,因为在第一个 sn-p 的末尾 email_stemmed 完全不变,因为 replace 被正确使用,而在第二个的末尾 @987654329 @ 实际上已经被剥夺了每个单词。

    【讨论】:

    • 太棒了,速度快得离谱。非常感谢,你救了我的理智。我会尽可能将其标记为完成 - 有一个 10 分钟的窗口。
    • 哈哈乐于助人:)
    • 虽然我在这里,为什么最后会引起这么多新词?我猜它们只是从我的源数据中多次出现,因此出现在列表中?
    • 你的意思是为什么循环一比循环二有更多的单词?我的假设是循环二中替换的单词没有在循环一中替换,因此它们出现在最后。
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 2015-08-31
    • 1970-01-01
    • 1970-01-01
    • 2019-03-07
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多