【问题标题】:Writing writing a large text file in python3在python3中编写一个大文本文件
【发布时间】:2018-01-13 23:20:27
【问题描述】:

虽然我看过一些有关该主题的文献,但我不太了解如何实现一个代码块,该代码块将写入大型文本文件而不会崩溃。

据我所知,它应该是逐行完成的,但是从我所看到的实现中,这只对已经存在的文件完成,而不是我想在每次迭代时在块中创建和写入文件循环。

这是代码块(它被 try catch 包围):

fileW = open(str(articleDate.title)+"-WC.txt", 'wb')
fileW.write(getText.encode('utf-8', errors='replace').strip()+ str(articleDate.publish_date).encode('utf-8').strip())
fileW.close()

我知道我需要另一种方式来写入文件的原因是因为我看到这个异常不断被引发,不断弹出的“块”关键字表明 write() 方法无法处理数量文字:

    File "/Users/Adrian/anaconda3/lib/python3.6/http/client.py", line 546, in _get_chunk_left
    chunk_left = self._read_next_chunk_size()
  File "/Users/Adrian/anaconda3/lib/python3.6/http/client.py", line 513, in _read_next_chunk_size
    return int(line, 16)
ValueError: invalid literal for int() with base 16: b''

During handling of the above exception, another exception occurred:

Traceback (most recent call last):
  File "/Users/Adrian/anaconda3/lib/python3.6/http/client.py", line 563, in _readall_chunked
    chunk_left = self._get_chunk_left()
  File "/Users/Adrian/anaconda3/lib/python3.6/http/client.py", line 548, in _get_chunk_left
    raise IncompleteRead(b'')
http.client.IncompleteRead: IncompleteRead(0 bytes read)

During handling of the above exception, another exception occurred:

Traceback (most recent call last):
  File "webcrawl.py", line 102, in <module>
    writeFiles()
  File "webcrawl.py", line 83, in writeFiles
    extractor = Extractor(extractor='ArticleExtractor', url=urls)
  File "/Users/Adrian/anaconda3/lib/python3.6/site-packages/boilerpipe/extract/__init__.py", line 39, in __init__
    connection  = urllib2.urlopen(request)
  File "/Users/Adrian/anaconda3/lib/python3.6/urllib/request.py", line 223, in urlopen
    return opener.open(url, data, timeout)
  File "/Users/Adrian/anaconda3/lib/python3.6/urllib/request.py", line 532, in open
    response = meth(req, response)
  File "/Users/Adrian/anaconda3/lib/python3.6/urllib/request.py", line 642, in http_response
    'http', request, response, code, msg, hdrs)
  File "/Users/Adrian/anaconda3/lib/python3.6/urllib/request.py", line 564, in error
    result = self._call_chain(*args)
  File "/Users/Adrian/anaconda3/lib/python3.6/urllib/request.py", line 504, in _call_chain
    result = func(*args)
  File "/Users/Adrian/anaconda3/lib/python3.6/urllib/request.py", line 753, in http_error_302
    fp.read()
  File "/Users/Adrian/anaconda3/lib/python3.6/http/client.py", line 456, in read
    return self._readall_chunked()
  File "/Users/Adrian/anaconda3/lib/python3.6/http/client.py", line 570, in _readall_chunked
    raise IncompleteRead(b''.join(value))
http.client.IncompleteRead: IncompleteRead(0 bytes read)

虽然我知道底部的异常名称通常是由于库名称“httplibs”更改为“urllibs”从 python 2 到 python 3,但是我使用的包是 python 3 兼容的,所以我相当肯定这是一个写作问题,任何帮助将不胜感激。

【问题讨论】:

  • 您的代码和问题标题是关于写入文件。但是您的异常是由读取 HTTP 响应引起的!两者完全没有关系!
  • @phihag 你是对的,但我指的是http响应处理程序中的_get_chunk_left,这似乎与传入的站点数据太大有关
  • @phihag 实际上在获取网络数据的行周围进行了尝试,也提供了一个解决方案,也感谢您指出这一点。

标签: python


【解决方案1】:

您可以使用上下文管理器来确保在每次操作结束时关闭文件:

import contextlib
@contextlib.contextmanager
def write_to(filename, ops = 'a'):  
    f = open(filename, ops)
    yield f
    f.close()

for chunk in data:
  with write_to('filename.txt') as f:
     f.write(chunk)

【讨论】:

  • @AdrianCoutsoftides 很高兴为您提供帮助!如果此答案对您有所帮助,请考虑接受。谢谢。
猜你喜欢
  • 1970-01-01
  • 2019-04-07
  • 2011-11-04
  • 2016-01-31
  • 2013-12-27
  • 2023-03-17
  • 2020-10-20
  • 2020-10-06
  • 2014-04-02
相关资源
最近更新 更多