【问题标题】:Is there any quick way to do the following in sql or python?有什么快速的方法可以在 sql 或 python 中执行以下操作吗?
【发布时间】:2022-11-18 09:48:00
【问题描述】:

我有一个大小为 1TB 的数据集,其中包含 3 列和大约 200 亿行。我想以某种随机顺序将这些数据分成大约 80/20 块的两个子数据。但是,这两个数据应该是非重叠的,这意味着一个块中的条目不应出现在另一个块中。一个块的一列中的条目不应出现在另一块的任何列中。例如,假设示例数据是:

fruit apple seeds
vegetable carrot yellow
crops fruit lettuce
green onion vegetable
lettuce red health

两个子数据可以是

fruit apple seeds
crops fruit lettuce
lettuce red health

和

vegetable carrot yellow
green onion vegetable

对于如此大的数据,有什么有效的方法可以做到这一点吗?

【问题讨论】:

    标签: python sql


    【解决方案1】:

    您可以遍历文件并根据您布置的比例将行随机分配给 sub-data-1 和 sub-data-2。

    import random
    with open('large_file', 'r') as lf, 
    open('s1', 'w') as s1, open('s2', 'w') as s2:
        for line in lf:
            if random.random() < 0.8:
                s1.write(line)
            else:
                s2.write(line)
    

    【讨论】:

      猜你喜欢
      • 2012-05-12
      • 2020-07-01
      • 2020-12-03
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2016-09-02
      相关资源
      最近更新 更多