【问题标题】:Using Python to edit the timestamps in a list? Convert POSIX to readable format using a function使用 Python 编辑列表中的时间戳?使用函数将 POSIX 转换为可读格式
【发布时间】:2016-08-20 01:35:45
【问题描述】:

第二次编辑:

完成了用于调整时区和转换格式的 sn-p。有关导致此解决方案的详细信息,请参阅下面的正确答案。

tzvar = int(input("Enter the number of hours you'd like to add to the timestamp:"))
tzvarsecs = (tzvar*3600)
print (tzvarsecs)

def timestamp_to_str(timestamp):
    return datetime.fromtimestamp(timestamp).strftime('%H:%M:%S %m/%d/%Y')

timestamps = soup('span', {'class': '_timestamp js-short-timestamp '})
dtinfo = [timestamp["data-time"] for timestamp in timestamps]
times = map(int, dtinfo)
adjtimes = [x+tzvarsecs for x in times]
adjtimesfloat = [float(i) for i in adjtimes]
dtinfofloat = [float(i) for i in dtinfo]
finishedtimes = [x for x in map(timestamp_to_str, adjtimesfloat)]
originaltimes = [x for x in map(timestamp_to_str, dtinfofloat)]

结束第二次编辑


编辑:

此代码允许我从 HTML 文件中抓取 POSIX 时间,然后将用户输入的小时数添加到原始值。负数也可以用来减去小时数。用户将在整个小时内工作,因为更改专门针对时区进行调整。

tzvar = int(input("Enter the number of hours you'd like to add to the timestamp:"))
tzvarsecs = (tzvar*3600)
print (tzvarsecs)

timestamps = soup('span', {'class': '_timestamp js-short-timestamp '})
dtinfo = [timestamp["data-time"] for timestamp in timestamps]
times = map(int, dtinfo)
adjtimes = [x+tzvarsecs for x in times]

剩下的就是下面建议的函数的反转。如何使用函数将列表中的每个 POSIX 时间转换为可读格式?

结束编辑



以下代码创建一个 csv 文件,其中包含从已保存的 Twitter HTML 文件中抓取的数据。

Twitter 将所有时间戳转换为浏览器中用户的本地时间。我想为用户提供一个输入选项,以将时间戳调整一定的小时数,以便推文的数据反映推文的本地时间。

我目前正在抓取一个名为 'title' 的元素,它是每个永久链接的一部分。相反,我可以轻松地从每条推文中刮取 POSIX 时间。

title="2:29 PM - 28 Sep 2015"

对

data-time="1443475777" data-time-ms="1443475777000"

我将如何编辑以下部分,以便将用户输入的变量添加到每个时间戳?我不需要请求输入的帮助,我只需要知道如何在将输入传递给 python 后将其应用于时间戳列表。

timestamps = soup('a', {'class': 'tweet-timestamp js-permalink js-nav js-tooltip'})
datetime = [timestamp["title"] for timestamp in timestamps]

与此代码/项目相关的其他问题。

Fix encoding error with loop in BeautifulSoup4?

Focusing in on specific results while scraping Twitter with Python and Beautiful Soup 4?

Using Python to Scrape Nested Divs and Spans in Twitter?


完整代码。

from bs4 import BeautifulSoup
import requests
import sys
import csv
import re
from datetime import datetime
from pytz import timezone

url = input("Enter the name of the file to be scraped:")
with open(url, encoding="utf-8") as infile:
    soup = BeautifulSoup(infile, "html.parser")

#url = 'https://twitter.com/search?q=%23bangkokbombing%20since%3A2015-08-10%20until%3A2015-09-30&src=typd&lang=en'
#headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 6.1) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/41.0.2228.0 Safari/537.36'}
#r = requests.get(url, headers=headers)
#data = r.text.encode('utf-8')
#soup = BeautifulSoup(data, "html.parser")

names = soup('strong', {'class': 'fullname js-action-profile-name show-popup-with-id'})
usernames = [name.contents for name in names]

handles = soup('span', {'class': 'username js-action-profile-name'})
userhandles = [handle.contents[1].contents[0] for handle in handles]  
athandles = [('@')+abhandle for abhandle in userhandles]

links = soup('a', {'class': 'tweet-timestamp js-permalink js-nav js-tooltip'})
urls = [link["href"] for link in links]
fullurls = [permalink for permalink in urls]

timestamps = soup('a', {'class': 'tweet-timestamp js-permalink js-nav js-tooltip'})
datetime = [timestamp["title"] for timestamp in timestamps]

messagetexts = soup('p', {'class': 'TweetTextSize  js-tweet-text tweet-text'}) 
messages = [messagetext for messagetext in messagetexts]  

retweets = soup('button', {'class': 'ProfileTweet-actionButtonUndo js-actionButton js-actionRetweet'})
retweetcounts = [retweet.contents[3].contents[1].contents[1].string for retweet in retweets]

favorites = soup('button', {'class': 'ProfileTweet-actionButtonUndo u-linkClean js-actionButton js-actionFavorite'})
favcounts = [favorite.contents[3].contents[1].contents[1].string for favorite in favorites]

images = soup('div', {'class': 'content'})
imagelinks = [src.contents[5].img if len(src.contents) > 5 else "No image" for src in images]

#print (usernames, "\n", "\n", athandles, "\n", "\n", fullurls, "\n", "\n", datetime, "\n", "\n",retweetcounts, "\n", "\n", favcounts, "\n", "\n", messages, "\n", "\n", imagelinks)

rows = zip(usernames,athandles,fullurls,datetime,retweetcounts,favcounts,messages,imagelinks)

rownew = list(rows)

#print (rownew)

newfile = input("Enter a filename for the table:") + ".csv"

with open(newfile, 'w', encoding='utf-8') as f:
    writer = csv.writer(f, delimiter=",")
    writer.writerow(['Usernames', 'Handles', 'Urls', 'Timestamp', 'Retweets', 'Favorites', 'Message', 'Image Link'])
    for row in rownew:
        writer.writerow(row)

【问题讨论】:

    标签: python web-scraping timestamp list-comprehension


    【解决方案1】:

    以您的代码为例,var datetime 存储字符串日期列表。因此,让我们分 3 个步骤来剖析这个过程,只是为了便于理解。

    例子

    >>> datetime = [timestamp["title"] for timestamp in timestamps]
    >>> print(datetime)
    ['2:13 AM - 29 Sep 2015', '2:29 PM - 28 Sep 2015', '8:04 AM - 28 Sep 2015']
    

    第一步:将其转换为 Python datetime object。

    >>> datetime_obj = datetime.strptime('2:13 AM - 29 Sep 2015', '%H:%M %p - %d %b %Y')
    >>> print(datetime_obj)
    datetime.datetime(2015, 9, 29, 2, 13)
    

    第二步:将日期时间对象转换为Python structured time object

    >>> to_time = struct_date.timetuple()
    >>> print(to_time)
    time.struct_time(tm_year=2015, tm_mon=9, tm_mday=29, tm_hour=2, tm_min=13, tm_sec=0, tm_wday=1, tm_yday=272, tm_isdst=-1)
    

    第三步:使用time.mktime将结构化时间对象转换为time

    >>> timestamp = time.mktime(to_time)
    >>> print(timestamp)
    1443503580.0
    

    现在都在一起了。

    import time
    from datetime import datetime
    
    ...
    def str_to_ts(str_date):
        return time.mktime(datetime.strptime(str_date, '%H:%M %p - %d %b %Y').timetuple())
    
    datetimes = [timestamp["title"] for timestamp in timestamps]
    times = [i for i in map(str_to_ts, datetimes)]
    

    PS:datetime 是变量名的错误选择。特别是在这种情况下。 :-)

    更新

    将函数应用于列表的每个值:

    def add_time(timestamp, hours=0, minutes=0, seconds=0):
        return timestamp + seconds + (minutes * 60) + (hours * 60 * 60)
    
    datetimes = [timestamp["title"] for timestamp in timestamps]
    times = [add_time(i, 5, 0, 0) for i in datetimes]
    

    更新 2

    将时间戳转换为字符串格式的日期:

    def timestamp_to_str(timestamp):
        return datetime.fromtimestamp(timestamp).strftime('%H:%M:%S %m/%d/%Y')
    

    例子:

    >>> from time import time
    >>> from datetime import datetime
    
    >>> timestamp_to_str(time())
    '17:01:47 08/29/2016'
    

    【讨论】:

    • 我可以很容易地从 twitter 上抓取 POSIX 时间,如上所述。我不需要将时间戳从一种格式转换为另一种格式。我需要知道如何为列表中的每个时间戳添加一定的小时数。使用 POSIX 格式似乎更容易。那么,例如,如何将 17 小时或 61200 秒添加到列表中的每个项目?
    • str_to_ts 函数与我所寻找的相反。现在我已经将 61200 秒添加到每个 POSIX 时间,如何将其转换回 hh:mm:ss MM/DD/YYYY 格式?
    • 感谢 Mauro,将其应用于 POSIX 时间时出现错误。 csv 中的每个字段都显示<function timestamp_to_str at 0x0000000000177F28>,而不是更正/转换的时间。
    • finishedtimes = [x for x in map(timestamp_to_str, adjtimes)] originaltimes = [x for x in map(timestamp_to_str, dtinfo)] 这给了我一个错误,它需要一个浮点数。
    • 我就是这样做的,上面编辑的代码。感谢您对 Mauro 的所有帮助!
    【解决方案2】:

    这就是我的想法,但不确定这是不是你想要的:

    >>> timestamps = ["1:00 PM - 28 Sep 2015", "2:00 PM - 28 Sep 2016", "3:00 PM - 29 Sep 2015"]
    >>> datetime = dict(enumerate(timestamps))
    >>> datetime
    {0: '1:00 PM - 28 Sep 2015',
     1: '2:00 PM - 28 Sep 2016',
     2: '3:00 PM - 29 Sep 2015'}
    

    【讨论】:

      【解决方案3】:

      您似乎正在寻找datetime.timedelta (documentation here)。您可以通过多种方式将输入转换为datetime.datetime 对象,例如,

      timestamp = datetime.datetime.fromtimestamp(1443475777)
      

      然后您可以使用timedelta 对象对它们进行算术运算。 timedelta 只是代表时间的变化。您可以使用hours 参数构造一个,如下所示:

      delta = datetime.timedelta(hours=1)
      

      然后timestamp + delta 将在未来一小时给你另一个datetime。减法也可以,其他任意时间间隔也可以。

      【讨论】:

        猜你喜欢
        • 2018-03-07
        • 2014-09-21
        • 1970-01-01
        • 2016-02-13
        • 2018-06-09
        • 1970-01-01
        • 2022-08-12
        • 1970-01-01
        • 1970-01-01
        相关资源
        最近更新 更多