【问题标题】:Using multiple "blog name"/command line arguments with Tumblr API通过 Tumblr API 使用多个“博客名称”/命令行参数
【发布时间】:2018-03-25 09:25:27
【问题描述】:

不久前,我在此处发布了使用 API 从 Tumblr 博客下载数据的帮助。 birryree (https://stackoverflow.com/users/297696/birryree) 非常友好地帮助我更正了我的脚本并找出我哪里出了问题,并且从 (Print more than 20 posts from Tumblr API) 开始我一直在使用他的脚本,没有任何问题。

此脚本要求我每次手动输入我要下载的博客名称。但是,我需要下载数百个博客,所以这导致我使用相同脚本的数百个版本并且非常耗时。我做了一些谷歌搜索,发现可以编写 Python 脚本,您可以在其中从命令行输入参数,然后它们将被一一处理(如果这是正确的术语)。

我尝试编写一个脚本,让我从命令提示符运行命令,然后下载我在命令提示符中请求的三个博客。 (在这种情况下,“prettythingsicantafford.tumblr.com;theficrecfairy.tumblr.com;和 staff.tumblr.com)。

所以我要运行的脚本是:

import pytumblr
import sys


def get_all_posts(client, blog):
offset = 0
while True:
    response = client.posts(blog, limit=20, offset=offset, reblog_info=True, notes_info=True)

    # Get the 'posts' field of the response        
    posts = response['posts']

    if not posts: return

    for post in posts:
        yield post

    # move to the next offset
    offset += 20



client = pytumblr.TumblrRestClient('SECRET')
blog = (sys.argv[1], sys.argv[2], sys.argv[3])


# use our function
with open('{}-posts.txt'.format(blog), 'w') as out_file:
for post in get_all_posts(client, blog):
    print >>out_file, post

我正在从命令提示符运行以下命令

tumblr_test2.py theficrecfairy prettythingsicantafford staff

但是,我收到以下错误消息:

Traceback (most recent call last):
File "C:\Users\izzy\test\tumblr_test2.py", line 29, in <module>
for post in get_all_posts(client, blog):
File "C:\Users\izzy\test\tumblr_test2.py", line 8, in get_all_posts
response = client.posts(blog, limit=20, offset=offset, reblog_info=True, notes_info=True)
File "C:\Python27\lib\site-packages\pytumblr\helpers.py", line 46, in add_dot_tumblr
args[1] += ".tumblr.com"
TypeError: can only concatenate tuple (not "str") to tuple

为了响应这个错误,我已经尝试修改我的脚本大约两周了,但我无法纠正我无疑非常明显的错误,非常感谢任何帮助或建议。

根据 vishes_shell 的建议进行编辑:

我现在正在使用以下脚本:

import pytumblr
import sys


def get_all_posts(client, blogs):
for blog in blogs:
    offset = 0
while True:
    response = client.posts(blog, limit=20, offset=offset, reblog_info=True, notes_info=True, filter='raw')

    # Get the 'posts' field of the response        
    posts = response['posts']

    if not posts: return

    for post in posts:
        yield post

    # move to the next offset
    offset += 20


client = pytumblr.TumblrRestClient('SECRET')
blog = sys.argv

# use our function
with open('{}-postsredux.txt'.format(blog), 'w') as out_file:
for post in get_all_posts(client, blog):
    print >>out_file, post

但是,我现在收到以下错误消息:

Traceback (most recent call last):
File "C:\Users\izzy\test\tumblr_test2.py", line 27, in <module>
with open('{}-postsredux.txt'.format(blog), 'w') as out_file:
IOError: [Errno 22] invalid mode ('w') or filename: " 
['C:\\\\Users\\\\izzy\\\\test\\\\tumblr_test2.py', 
'prettythingsicantafford', 'theficrecfairy']-postsredux.txt"

【问题讨论】:

    标签: python command-line tumblr sys pytumblr


    【解决方案1】:

    当blog 是tuple 对象时,您尝试client.posts(blog, ...) 的问题,声明为:

    blog = (sys.argv[1], sys.argv[2], sys.argv[3])
    

    您需要重构您的方法以分别浏览每个博客。

    def get_all_posts(client, blogs):
        for blog in blogs:
            offset = 0
            ...
    
            while True:
                response = client.posts(blog, ...)
            ...
        ...
    
    blog = sys.argv
    ...
    

    【讨论】:

    • 谢谢你!但是,我想我可能误解了您的指示,因为我仍然收到不同的错误消息(如上面编辑的帖子中所述)。是不是因为我所做的事情意味着它试图同时处理所有三个参数(包括脚本的名称)?再次感谢您的帮助!
    • 首先,正如我从错误中看到的那样,我修复了你的错误,因为你得到了你正在获取的所有数据。新错误仅仅是因为您试图将数据写入错误的文件。我建议你在现有问题中找到解决方案,因为这样的问题很多。
    • 啊,太棒了!我将尝试解决这个问题,然后在我验证后立即接受您的回答。谢谢!
    猜你喜欢
    • 2023-04-10
    • 2014-02-27
    • 2018-06-12
    • 1970-01-01
    • 1970-01-01
    • 2012-12-28
    • 1970-01-01
    • 1970-01-01
    • 2011-02-12
    相关资源
    最近更新 更多