【问题标题】:Grouping items into subsets (power set)将项目分组为子集(幂集)
【发布时间】:2014-11-18 22:07:38
【问题描述】:

假设我有以下内容:

john: [a, b, c, d]
bob:  [a, c, d, e]
mary: [a, b, e, f]

或稍微重新格式化,以便您可以轻松查看分组:

john: [a, b, c, d]
bob:  [a,    c, d, e]
mary: [a, b,       e, f]

生成以下分组的最常见或最有效的算法是什么?

[john, bob, mary]: [a]
[john, mary]:      [b]
[john, bob]:       [c,d]
[bob, mary]:       [e]
[mary]:            [f]
[john]:            []
[bob]:             []

快速谷歌搜索后,上面的键似乎代表“电源组”。所以我正在计划以下实现:

1) 生成幂集 {{j, b, m}, {j, m}, {j, b} {b, m}, {m}, {j}, {b}} // j =约翰,b = 鲍勃,m = 玛丽

2) 生成所有字母的集合:{a, b, c, d, e, f}

3) 遍历子集,对于每个字母,查看字母是否存在于子集的所有元素中

所以...

subset = {j, b, m}

letter = a
    j contains a? true
    b contains a? true
    m contains a? true
        * add a to subset {j, b, m}

letter = b
    j contains b? true
    b contains b? false, continue

letter = c
    j contains c? true
    b contains c? true
    m contains c? false, continue
.....

subset = {j, m}
.....

有没有更好的解决方案?

编辑:上述算法有缺陷。例如,{j, m} 也会包含“a”,这是我不想要的。我想我可以简单地修改它,以便在每次迭代中,我还检查这个字母是否“不在”这个集合中的元素。所以在这种情况下,我也会检查:

if b does not contain a

【问题讨论】:

    标签: algorithm combinations powerset


    【解决方案1】:

    您可以使用两个地图/字典来实现这一点,一个是另一个的“逆”。对于第一个地图,“键”是名称,“值”是字符列表。第二个映射将字母作为键,将与其关联的名称列表作为值。

    在 Python 中

    nameDict = {'john' : ['a', 'b', 'c', 'd'], 'bob' : ['a', 'c', 'd', 'e'], 'mary' : ['a', 'b', 'e', 'f']}
    
    reverseDict = {}
    for key,values in nameDict.items():
        for v in values:
            if v in reverseDict.keys():
                reverseDict[v].append(key)
            else:
                reverseDict[v] = [key] # If adding v to dictionary for the first time it needs to be as a list element
    
    # Aggregation
    finalDict = {}
    for key,values in reverseDict.items():
        v = frozenset(values)
        if v in finalDict.keys():
            finalDict[v].append(key)
        else:
            finalDict[v] = [key] 
    

    这里,reverseDict 包含你想要的映射 a -> [john, bob, mary], b -> [john, mary] 等等。你也可以通过检查 reverseDict[' 返回的列表来检查if john does not contain a a'] 包含 john。

    [编辑] 将聚合添加到 finalDict。

    您可以使用 freezesets 作为字典键,因此 finalDict 现在包含正确的结果。打印字典:

    frozenset({'bob', 'mary'})
    ['e']
    
    frozenset({'mary'})
    ['f']
    
    frozenset({'john', 'bob'})
    ['c', 'd']
    
    frozenset({'john', 'mary'})
    ['b']
    
    frozenset({'john', 'bob', 'mary'})
    ['a']
    

    【讨论】:

    • 感谢这已接近但缺少聚合步骤。例如,您的代码将输出 c-> {b,j} 和 d-> {b,j},而我需要 {b,j}-> [c,d]
    • @nogridbag 我明白了。我认为聚合可以通过制作 another 逆映射来完成,这次是 reverseDict。这样,键 {b,j} 将具有相应的值 [c,d]。我已将此添加到我的答案中。
    • 蒂姆,现在还早,我还没有喝咖啡,但我相信你的聚合步骤只是返回原始数据集:nameDict :)
    • @nogridbag 这会教我不要测试我的答案!我已经使用 freezesets 更新了我的答案,它现在会产生正确的结果。
    【解决方案2】:

    第 3 步(迭代子集)效率低下,因为它对幂集中的每个元素都执行“j 包含 a”或“a 不在 j 中”。

    以下是我的建议:

    1) 生成幂集 {{j, b, m}, {j, m}, {j, b} {b, m}, {m}, {j}, {b}}。您不需要这一步,因为您不关心最终输出中的空映射。

    2) 遍历原始数据结构中的所有元素并构造以下内容:

    [a] : [j, b, m]
    [b] : [j, m]
    [c] : [j, b]
    [d] : [j, b]
    [e] : [b, m]
    [f] : [m]
    

    3) 反转上述结构并聚合(使用 [j, b,...] 的映射到 [a,b...] 的列表应该可以解决问题)来得到这个:

    [j, b, m] : [a]
    [j, m] : [b]
    [j, b] : [c, d]
    [b, m] : [e]
    [m] : [f]
    

    4) 比较 3 和 1 以填充剩余的空映射。

    编辑: Hers 是 Java 中的完整代码

        // The original data structure. mapping from "john" to [a, b, c..] 
        HashMap<String, HashSet<String>> originalMap = new HashMap<String, HashSet<String>>();
    
        // The final data structure. mapping from power set to [a, b...]
        HashMap<HashSet<String>, HashSet<String>> finalMap = new HashMap<HashSet<String>, HashSet<String>>();
    
        // Intermediate data structure. Used to hold [a] to [j,b...] mapping
        TreeMap<String, HashSet<String>> tmpMap = new TreeMap<String, HashSet<String>>();
    
        // Populate the original dataStructure
        originalMap.put("john", new HashSet<String>(Arrays.asList("a", "b", "c", "d")));
        originalMap.put("bob", new HashSet<String>(Arrays.asList("a", "c", "d", "e")));
        originalMap.put("mary", new HashSet<String>(Arrays.asList("a", "b", "e", "f")));
    
        // Hardcoding the powerset below. You can generate the power set using the algorithm used in googls guava library.
        // powerSet function in https://code.google.com/p/guava-libraries/source/browse/guava/src/com/google/common/collect/Sets.java
        // If you don't care about empty mappings in the finalMap, then you don't even have to create the powerset
        finalMap.put(new HashSet<String>(Arrays.asList("john", "bob", "mary")), new HashSet<String>());
        finalMap.put(new HashSet<String>(Arrays.asList("john", "bob")), new HashSet<String>());
        finalMap.put(new HashSet<String>(Arrays.asList("bob", "mary")), new HashSet<String>());
        finalMap.put(new HashSet<String>(Arrays.asList("john", "mary")), new HashSet<String>());
        finalMap.put(new HashSet<String>(Arrays.asList("john")), new HashSet<String>());
        finalMap.put(new HashSet<String>(Arrays.asList("bob")), new HashSet<String>());
        finalMap.put(new HashSet<String>(Arrays.asList("mary")), new HashSet<String>());
    
        // Iterate over the original map to prepare the tmpMap.
        for(Entry<String, HashSet<String>> entry : originalMap.entrySet()) {
            for(String value : entry.getValue()) {
                HashSet<String> set = tmpMap.get(value);
                if(set == null) {
                    set = new HashSet<String>();
                    tmpMap.put(value, set);
                }
                set.add(entry.getKey());
            }
        }
    
        // Iterate over the tmpMap result and add the values to finalMap
        for(Entry<String, HashSet<String>> entry : tmpMap.entrySet()) {
            finalMap.get(entry.getValue()).add(entry.getKey());
        }
    
        // Print the output
        for(Entry<HashSet<String>, HashSet<String>> entry : finalMap.entrySet()) {
            System.out.println(entry.getKey() +" : "+entry.getValue());
        }
    

    上面代码的输出是:

    [bob] : []
    [john] : []
    [bob, mary] : [e]
    [bob, john] : [d, c]
    [bob, john, mary] : [a]
    [mary] : [f]
    [john, mary] : [b]
    

    【讨论】:

    • 谢谢 我相信我的困惑是聚合步骤。我不确定是否使用 Set 作为 Map 的键。这有什么问题吗?我相信我可以使用 guava 和 apache commons 中的多键映射,但我最初计划在 JS 中的浏览器中执行此操作。
    • @nogridbag Set 作为 Map 的键是可以的,只要 Set 在作为键添加到 Map 后不改变。我认为在多键映射中键的顺序很重要。如果键的顺序很重要,那么您可能需要在生成键之前进行排序。
    猜你喜欢
    • 1970-01-01
    • 2011-02-10
    • 1970-01-01
    • 2018-09-06
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多