【问题标题】:PRAW Adding reddit users with 3+ comments in a subreddit to a listPRAW 将 subreddit 中具有 3 条以上评论的 reddit 用户添加到列表
【发布时间】:2017-07-14 05:48:19
【问题描述】:

我目前有一个使用 PRAW 编写的 Python reddit 机器人,它从特定的 subreddit 获取所有 cmets,查看他们的作者并确定该作者在 subreddit 中是否至少有 3 个 cmets。如果他们有 3 个以上的 cmets,则将它们添加到已批准的提交提交文本文件中。 我的代码目前“有效”,但老实说,它太糟糕了,我什至不确定如何。实现我的目标的更好方法是什么? 我目前拥有的:

def get_approved_posters(reddit):

   subreddit = reddit.subreddit('Enter_Subreddit_Here')
   subreddit_comments = subreddit.comments(limit=1000)
   dictionary_of_posters = {}
   unique_posters = {}
   commentCount = 1
   duplicate_check = False
   unique_authors_file = open("unique_posters.txt", "w")

   print("Obtaining comments...")
   for comment in subreddit_comments:
      dictionary_of_posters[commentCount] = str(comment.author)
      for key, value in dictionary_of_posters.items():
        if value not in unique_posters.values():
            unique_posters[key] = value
    for key, value in unique_posters.items():
        if key >= 3:
            commentCount += 1
        if duplicate_check is not True:
            commentCount += 1
            print("Adding author to dictionary of posters...")
            unique_posters[commentCount] = str(comment.author)
            print("Author added to dictionary of posters.")
            if commentCount >= 3:
                duplicate_check = True

   for x in unique_posters:
      unique_authors_file.write(str(unique_posters[x]) + '\n')

   total_comments = open("total_comments.txt", "w")
   total_comments.write(str(dictionary_of_posters))

   unique_authors_file.close()
   total_comments.close()

   unique_authors_file = open("unique_posters.txt", "r+")
   total_comments = open("total_comments.txt", "r")
   data = total_comments.read()
   approved_list = unique_authors_file.read().split('\n')
   print(approved_list)
   approved_posters = open("approved_posters.txt", "w")
   for username in approved_list:
      count = data.count(username)
      if(count >= 3):
        approved_posters.write(username + '\n')
      print("Count for " + username + " is " + str(count))

   approved_posters.close()
   unique_authors_file.close()
   total_comments.close()

【问题讨论】:

  • 整段代码都在一个函数中吗?你应该缩进,这样我们就可以知道函数中有什么,什么不是。
  • 这是一个功能;我会编辑它
  • 您希望获得什么样的改进?速度?准确性?这段代码对您来说有什么“错误”?关于我们如何为您提供帮助的指导并不多
  • 我目前正在使用 3 个不同的文本文件来确定谁在 subreddit 中发布了 3 次。它也不总是准确的。我最初将它用于特定的 subreddit,它工作正常,然后我在第二个 subreddit 上测试了代码,并在我的输出文件中连续 3 次使用相同的名称,而不是它是唯一的
  • 对于您的文本文件,我会将您的数据存储在 json 文件中。因此,您可以将数据添加到一个字典中,该字典按键细分为您需要跟踪的三种类型的数据,然后将其转换为 json 字符串并将其全部转储到一个文件中。有很多教程,我个人import json 并使用json.load() 和json.dump()。我会继续查看你的代码,看看还有什么可以帮助你

标签: python praw


【解决方案1】:

也许只是我今天早上速度慢,但我很难理解/理解您对 commentCount 和 unique_posters 的使用。其实应该是我吧。

我会像你一样从 subreddit 中获取所有 cmets,对于每条评论,请执行以下操作:

for comment in subreddit_comments:
    try:
        dictionary_of_posters[comment.author] += 1
    except KeyError:
        dictionary_of_posters[comment.author] = 1

for username, comment_count in dictionary_of_posters.items():
    if comment_count >= 3:
        approved_authors.append(username)

此方法利用了字典不能有两个相同键值的事实。这样,您不必进行重复检查或任何其他操作。如果它让你感觉更好,你可以去list(set(approved_authors)),这样可以消除任何杂散的重复。

【讨论】:

    猜你喜欢
    • 2020-05-04
    • 2015-09-15
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2020-10-17
    相关资源
    最近更新 更多