【问题标题】:Using Regex with BeautifulSoup to parse a string in Python在 Python 中使用 Regex 和 BeautifulSoup 来解析字符串
【发布时间】:2015-02-24 16:23:19
【问题描述】:

我有一系列类似于“2014 年 12 月 27 日星期六”的字符串,我想扔掉“星期六”并保存名称为“141227”的文件,即年 + 月 + 日。到目前为止,一切正常,除了我无法让 daypos 或 yearpos 的正则表达式工作。他们都给出了同样的错误:

Traceback(最近一次调用最后一次):文件“scrapewaybackblog.py”,行 17、在 daypos = byline.find(re.compile("[A-Z][a-z]*\s")) TypeError: expected a character buffer object

什么是字符缓冲区对象?这是否意味着我的表达有问题?这是我的脚本:

for i in xrange(3, 1, -1):
       page = urllib2.urlopen("http://web.archive.org/web/20090204221349/http://www.americansforprosperity.org/nationalblog?page={}".format(i))
       soup = BeautifulSoup(page.read())
       snippet = soup.find_all('div', attrs={'class': 'blog-box'})
       for div in snippet:
           byline =  div.find('div', attrs={'class': 'date'}).text.encode('utf-8')
           text = div.find('div', attrs={'class': 'right-box'}).text.encode('utf-8')

           monthpos = byline.find(",")
           daypos = byline.find(re.compile("[A-Z][a-z]*\s"))
           yearpos = byline.find(re.compile("[A-Z][a-z]*\D\d*\w*\s"))
           endpos = monthpos + len(byline)

           month = byline[monthpos+1:daypos]
           day = byline[daypos+0:yearpos]
           year = byline[yearpos+2:endpos]

           output_files_pathname = 'Data/'  # path where output will go
           new_filename = year + month + day + ".txt"
           outfile = open(output_files_pathname + new_filename,'w')
           outfile.write(date)
           outfile.write("\n")
           outfile.write(text)
           outfile.close()
       print "finished another url from page {}".format(i)

我还没有弄清楚如何使 12 月 = 12,但那是另一次了。请帮我找到合适的职位。

【问题讨论】:

    标签: python html regex beautifulsoup html-parsing


    【解决方案1】:

    不要用正则表达式解析日期字符串,而是用dateutil解析它:

    from dateutil.parser import parse
    
    for div in soup.select('div.blog-box'):
        byline = div.find('div', attrs={'class': 'date'}).text.encode('utf-8')
        text = div.find('div', attrs={'class': 'right-box'}).text.encode('utf-8')
    
        dt = parse(byline)
        new_filename = "{dt.year}{dt.month}{dt.day}.txt".format(dt=dt)
        ...
    

    或者,你可以用datetime.strptime()解析字符串,但是你需要注意suffixes

    byline = re.sub(r"(?<=\d)(st|nd|rd|th)", "", byline)
    dt = datetime.strptime(byline, '%A, %B %d %Y')
    

    re.sub() 在这里找到stndrdth 字符串after a digit 并用空字符串替换后缀。之后,日期字符串将匹配 '%A, %B %d %Y' 格式,请参阅:


    一些补充说明:

    固定版本:

    import os
    import urllib2
    
    from bs4 import BeautifulSoup
    from dateutil.parser import parse
    
    
    for i in xrange(3, 1, -1):
        page = urllib2.urlopen("http://web.archive.org/web/20090204221349/http://www.americansforprosperity.org/nationalblog?page={}".format(i))
        soup = BeautifulSoup(page)
    
        for div in soup.select('div.blog-box'):
            byline = div.find('div', attrs={'class': 'date'}).text.encode('utf-8')
            text = div.find('div', attrs={'class': 'right-box'}).text.encode('utf-8')
    
            dt = parse(byline)
    
            new_filename = "{dt.year}{dt.month}{dt.day}.txt".format(dt=dt)
            with open(os.path.join('Data', new_filename), 'w') as outfile:
                outfile.write(byline)
                outfile.write("\n")
                outfile.write(text)
    
        print "finished another url from page {}".format(i)
    

    【讨论】:

    • 你太棒了!!谢谢。
    猜你喜欢
    • 2020-06-21
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2011-07-14
    相关资源
    最近更新 更多