【问题标题】:How to detect CRC errs in an file when extracting from a zip从zip中提取时如何检测文件中的CRC错误
【发布时间】:2018-09-20 18:32:23
【问题描述】:

我在一家拥有云基础应用程序的公司工作,该应用程序允许用户将数据文件从 PC 传输到他们的移动设备并再次传输回来,从而允许用户在现场更新和编辑他们的文件。

我会随机遇到一个问题,即在通过互联网传输文件的过程中,zip 结构被损坏,并且 zip 中的随机文件会出现 crc 错误。

通常它将是一个图像,但并非总是如此。这会导致我们的软件出现问题并阻止文件加载。我写了一个工具,它可以扫描 zip 文件并查找和修复 zip 中 xml 文件中的错误;但是,System.IO.Compression.ZipFile.ExtractToDirectory 函数不会引发错误,它会提取文件并将其写出。

我稍后调用 System.IO.Compression.ZipFile.CreateFromDirectory 然后压缩这个文件,再次没有错误导致压缩文件仍然有 crc 错误并且仍然无法在我们的软件中打开。

所以这是我的问题,我如何检查从 zip 中提取的文件是否存在 crc 错误,而不是在完成其他工作后将它们放入新的 zip 中?

我已经尝试了一些方法,但没有将每种文件类型都加载到某种阅读器中(zip 中可能有十几种文件类型)我什么也没得到。

附:我对 C# 很陌生,没有正式的编码培训,所以我可能会问一些愚蠢的问题,抱歉。 :D

一个例子可以找到here

好的,我尝试了建议的内容,这就是我得到的

使用以下代码:

//https://johnlnelson.com/tag/zip-archive-c/ using ZipArchive
string zipFile = FileName;
string extractPath = @"C:\Temp\XAP\" + xapName;
ZipArchive zipArchive = ZipFile.OpenRead(zipFile);

if (zipArchive.Entries != null && zipArchive.Entries.Count > 0)
{
foreach (ZipArchiveEntry entry in zipArchive.Entries)
{
entry.ExtractToFile(System.IO.Path.Combine(extractPath, entry.FullName));
}
}

它会产生以下错误

ERROR:
System.IO.DirectoryNotFoundException: Could not find a part of the path 'C:\Temp\XAP\CRC in bmp\assignmentmap.bmp'.
   at System.IO.__Error.WinIOError(Int32 errorCode, String maybeFullPath)
   at System.IO.FileStream.Init(String path, FileMode mode, FileAccess access, Int32 rights, Boolean useRights, FileShare share, Int32 bufferSize, FileOptions options, SECURITY_ATTRIBUTES secAttrs, String msgPath, Boolean bFromProxy, Boolean useLongPath, Boolean checkHost)
   at System.IO.FileStream..ctor(String path, FileMode mode, FileAccess access, FileShare share)
   at System.IO.Compression.ZipFile

代码无法提取 zip 中的任何文件,而不仅仅是损坏的文件。

【问题讨论】:

  • System.IO.Compression.ZipFile.CreateFromDirectory 如果遇到损坏的条目应抛出 InvalidDataException。但是由于 ZIP 文件条目允许以不同的方式存储多个元数据(例如 CRC),它可能是 ZipFile 类的一个限制,它无法检测到有问题的条目。但我只是在这里猜测,我不太确定 ZipFile 的行为。 (1/2)
  • (2/2) 出于测试目的,也许您可​​以尝试使用 ZipArchive 从损坏的 ZIP 存档中读取条目,以检查 (A) ZipArchive 是否真的能够从 ZIP 条目中获取 CRC 值标头,如果是,则 (B) 从 ZipArchiveEntries 读取未压缩的数据流时是否会引发异常。或者,如果使用 ZipArchive 的测试没有结果,并且您的项目和业务要求/约束​​允许使用 3rd-party 库,您可以使用 SharpZipLib 或类似库(当然,许可条款允许)。
  • (3/2) 我的第二条评论中的小但可能重要的遗漏:如果 ZipArchive/ZipArchiveEntry 可以从每个 ZIP 条目标头中提供 CRC32,但是从头到尾读取损坏条目的未压缩数据流end 不会抛出,您仍然可以从未压缩数据流中的数据计算 CRC32 校验和,并自己将其与 ZIP 条目标头中的 CRC32 值进行比较(这将避免对某些第 3 方库的依赖)
  • 好的,最后的评论,我保证! (也许;)您说 System.IO.Compression.ZipFile.CreateFromDirectory 创建了一个 ZIP 存档,其中包含有 CRC 错误的条目。那不应该发生。这让我相信这里可能存在误解。您是在谈论存储在 ZIP 条目元数据中的 CRC32 值(它们是 ZIP 文件格式的一部分),还是在谈论与 ZIP 文件格式本身无关的其他一些 CRC 值/类型?
  • 这是一个有趣的行为。在标准情况下,不应该提取具有不正确 CRC (损坏的压缩数据) 的文件(至少不应该在损坏点之后)。创建新的 zip 文件后,应该没有损坏的文件(可能不完整的文件除外)。由于我没有真正的代码来测试您的案例,因此只有一个理论:尝试使用 7zip 测试文件。当 7zip 不报错时,zip 文件正常,问题出在其他地方。否则,您可能需要自定义 UnzipToDir 或 TestZip 函数来处理单独的文件。

标签: c# .net zip extract crc


【解决方案1】:

.NET Framework 4.5 的 ZipArchive 中的构建似乎存在问题。损坏的文件不会以任何方式发出信号。 SharpZipLib 似乎还有另一个问题。以下示例的完整代码,包括 .NET Framework 的损坏检测:

using System;
using System.IO;
using System.IO.Compression;

using ICSharpCode.SharpZipLib.Zip;

using DotNetZipFile = System.IO.Compression.ZipFile;

namespace Stackoverflow_unzip_test
{
  internal static class Program
  {
    private static void Main()
    {
      Extract45Framework("CRCerror.zip", ".\\Unpack_1");
      Console.WriteLine();
      Console.WriteLine();
      ExtractSharpZipLib("CRCerror.zip", ".\\Unpack_2");
    }

    private static void Extract45Framework(string zipFile, string extractPath)
    {
      ZipArchive zipArchive = DotNetZipFile.OpenRead(zipFile);

      if (zipArchive.Entries != null && zipArchive.Entries.Count > 0)
      {
        Console.WriteLine("Extracting...");
        foreach (ZipArchiveEntry entry in zipArchive.Entries)
        {
          try
          {
            if (string.IsNullOrEmpty(entry.Name))
            {
              // skip directory
              continue;
            }

            string file = Path.Combine(extractPath, entry.FullName);
            string path = Path.GetDirectoryName(file);
            if (!Directory.Exists(path))
            {
              Directory.CreateDirectory(path);
            }

            Console.WriteLine(" - '" + entry.FullName + "'...");
            if (File.Exists(file))
            {
              //Console.WriteLine("   - delete previous version...");
              File.Delete(file);
            }

            entry.ExtractToFile(file);
            long length = (new FileInfo(file)).Length;
            if (entry.Length != length)
            {
              Console.WriteLine($"   - Failed to extract! Extracted only {length} out of {entry.Length} bytes");
            }
          }
          catch (Exception ex)
          {
            Console.WriteLine("   - Failed to extract: " + ex.Message);
          }
        }
      }
    }

    private static void ExtractSharpZipLib(string zipFileName, string extractPath)
    {
      using (var file = File.OpenRead(zipFileName))
      using (ZipInputStream zip = new ZipInputStream(file))
      {
        ZipEntry entry;
        try
        {
          while ((entry = zip.GetNextEntry()) != null)
          {
            SaveFile(zip, entry, extractPath);
          }
        }
        catch (Exception ex)
        {
          Console.WriteLine("   - Failed to parse zip: " + ex.Message);
        }
      }
    }

    private static void SaveFile(ZipInputStream zip, ZipEntry entry, string extractPath)
    {
      if (entry.IsDirectory)
      {
        return;
      }

      try
      {
        string file = Path.Combine(extractPath, entry.Name);
        string path = Path.GetDirectoryName(file);
        if (!Directory.Exists(path))
        {
          Directory.CreateDirectory(path);
        }

        Console.WriteLine(" - '" + entry.Name + "'...");
        if (File.Exists(file))
        {
          //Console.WriteLine("   - delete previous version...");
          File.Delete(file);
        }

        byte[] data = new byte[1024];
        long length = 0;
        using (var fs = File.OpenWrite(file))
        {
          int readLength;
          while ((readLength = zip.Read(data, 0, data.Length)) != 0)
          {
            fs.Write(data, 0, readLength);
            length += readLength;
          }
        }

        if (entry.Size != length)
        {
          Console.WriteLine($"   - Failed to extract! Extracted only {length} out of {entry.Size} bytes");
        }
      }
      catch (Exception ex)
      {
        Console.WriteLine("   - Failed to extract: " + ex.Message);
      }
    }
  }
}

应用程序输出:

Extracting...
 - 'tadd1.tx'...
 - 'assignmentmap.bmp'...
   - Failed to extract! Extracted only 34216 out of 34830 bytes
 - 'sketch.dsk'...


 - 'tadd1.tx'...
 - 'assignmentmap.bmp'...
   - Failed to extract: Index was outside the bounds of the array.
   - Failed to parse zip: Specified argument was out of the range of valid values.
Parameter name: count

【讨论】:

  • 很好,这看起来会很有帮助。您的 Extract45Framework 进程将识别任何有问题的文件,并让我在其周围放置一些代码,以确保一旦我准备好重新打包文件,任何损坏的文件都不包括在内。
猜你喜欢
  • 2017-08-17
  • 2013-05-08
  • 1970-01-01
  • 2023-03-03
  • 1970-01-01
  • 1970-01-01
  • 2020-10-25
  • 2022-01-05
  • 2011-02-19
相关资源
最近更新 更多