【问题标题】:Python - How can I open a file and specify the offset in bytes?Python - 如何打开文件并以字节为单位指定偏移量?
【发布时间】:2010-07-21 12:33:30
【问题描述】:

我正在编写一个程序,它会定期解析 Apache 日志文件以记录其访问者、带宽使用情况等。

问题是,我不想打开日志并解析我已经解析过的数据。例如:

line1
line2
line3

如果我解析该文件,我将保存所有行,然后保存该偏移量。这样,当我再次解析它时,我得到:

line1
line2
line3 - The log will open from this point
line4
line5

第二次,我会得到line4和line5。希望这是有道理的......

我需要知道的是,我该如何做到这一点? Python具有用于指定偏移量的seek()函数......那么我是否只是在解析后获取日志的文件大小(以字节为单位),然后在第二次记录它时将其用作偏移量(在seek()中)?

我似乎想不出一种编码方式>.

【问题讨论】:

    标签: python file-io byte offset


    【解决方案1】:

    借助file 类的seek 和tell 方法,您可以管理文件中的位置,请参阅 https://docs.python.org/2/tutorial/inputoutput.html

    tell 方法会告诉你下次打开时在哪里寻找

    【讨论】:

    【解决方案2】:
    log = open('myfile.log')
    pos = open('pos.dat','w')
    print log.readline()
    pos.write(str(f.tell())
    log.close()
    pos.close()
    
    log = open('myfile.log')
    pos = open('pos.dat')
    log.seek(int(pos.readline()))
    print log.readline()
    

    当然你不应该那样使用它——你应该将操作封装在像save_position(myfile)和load_position(myfile)这样的函数中,但是功能就在那里。

    【讨论】:

      【解决方案3】:

      如果您的日志文件很容易放入内存中(也就是说,您有合理的轮换策略),您可以轻松地执行以下操作:

      log_lines = open('logfile','r').readlines()
      last_line = get_last_lineprocessed() #From some persistent storage
      last_line = parse_log(log_lines[last_line:])
      store_last_lineprocessed(last_line)
      

      如果你不能这样做,你可以使用类似的东西(请参阅接受的答案对寻找和告诉的使用,以防你需要与他们一起做)Get last n lines of a file with Python, similar to tail

      【讨论】:

      • 日志用于虚拟主机,因此目前没有日志轮换。我想我应该考虑设置它......这将使您的解决方案相当有用。干杯。
      【解决方案4】:

      如果您每行解析日志行,您可以只保存上次解析的行号。下次你就得从好行开始读了。

      当您必须位于文件中非常特定的位置时,查找会更有用。

      【讨论】:

        【解决方案5】:

        注意你可以在python中从文件末尾开始seek():

        f.seek(-3, os.SEEK_END)
        

        将读取位置从 EOF 放置 3 行。

        但是,为什么不使用 diff,无论是从 shell 还是 difflib?

        【讨论】:

        • 这实际上会将读取位置从 EOF 放置 3 个字符,而不是 3 行。
        【解决方案6】:

        简单但不推荐:):

        last_line_processed = get_last_line_processed()    
        with open('file.log') as log
            for record_number, record in enumerate(log):
                if record_number >= last_line_processed:
                    parse_log(record)
        

        【讨论】:

          【解决方案7】:

          这是使用您的长度建议和告诉方法的代码证明:

          beginning="""line1
          line2
          line3"""
          
          end="""- The log will open from this point
          line4
          line5"""
          
          openfile= open('log.txt','w')
          openfile.write(beginning)
          endstarts=openfile.tell()
          openfile.close()
          
          open('log.txt','a').write(end)
          print open('log.txt').read()
          
          print("\nAgain:")
          end2 = open('log.txt','r')
          end2.seek(len(beginning))
          
          print end2.read()  ## wrong by two too little because of magic newlines in Windows
          end2.seek(endstarts)
          
          print "\nOk in Windows also"
          print end2.read()
          end2.close()
          

          【讨论】:

            【解决方案8】:

            这是一个高效且安全的 sn-p,可以将读取的偏移量保存在并行文件中。基本上是python中的logtail。

            with open(filename) as log_fd:
                offset_filename = os.path.join(OFFSET_ROOT_DIR,filename)
                if not os.path.exists(offset_filename):
                    os.makedirs(os.path.dirname(offset_filename))
                    with open(offset_filename, 'w') as offset_fd:
                        offset_fd.write(str(0))
                with open(offset_filename, 'r+') as offset_fd:
                    log_fd.seek(int(offset_fd.readline()) or 0)
                    new_logrows_handler(log_fd.readlines())
                    offset_fd.seek(0)
                    offset_fd.write(str(log_fd.tell()))
            

            【讨论】:

              猜你喜欢
              • 1970-01-01
              • 1970-01-01
              • 1970-01-01
              • 1970-01-01
              • 2012-11-30
              • 1970-01-01
              • 1970-01-01
              • 2013-07-14
              • 1970-01-01
              相关资源
              最近更新 更多