【问题标题】:How to lookup an HBase table efficiently如何高效地查找 HBase 表
【发布时间】:2023-03-20 14:04:01
【问题描述】:

我有一个 HBase 查找表,用于存储一些信息。我有一个 MapReduce 程序,它运行一些 Pentaho KTR,在 MapReduce 作业中我捕获了输出。从 KTR 输出中的某些字段中,我检索了一些键并使用它们我必须在 HBase 中查找一些值。我的情况是:

1. The rowkey is of format <Table Code>-<CRC>, ex- DDVC-XXX

For each output of the KTRs:    

2. If no result is found for a particular key(which I get from the Pentaho KTRs), 
    then increment a column value which has the rowkey of format
    <Table Code>-last, ex: DDVC-last
3. Take this incremented value and put it in the HBase table with the specific key.

所以,如果我找不到 rowkey 的值,我将在这里执行一次 Get、一次 Increment 和一次 Put 操作。有人可以给我一些建议,告诉我如何有效地做到这一点,而无需再次点击 HBase。因为,我可以看到作业所需的大部分时间是执行上述算法,该算法针对单行多次命中 HBase。

先谢谢了!!

【问题讨论】:

    标签: hbase


    【解决方案1】:

    虽然架构设计可能值得关注,但您所描述的问题可能不会在性能方面得到进一步改进。 Get、Increment 和 Put 是独立的操作,确实需要三个独立的 HBase 调用。

    【讨论】:

    • 是的。我同意你的看法。但是有什么方法可以让我批量执行,而不是为每个键执行 Get、Incr 和 Put?或者我可以像 postGet 方法那样使用 Observer 协处理器吗?
    • 啊!好问题 - 事实上我现在正在投票;)已经晚了,我会考虑的..
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 2020-01-16
    • 2021-12-11
    • 2013-10-17
    • 1970-01-01
    • 1970-01-01
    • 2012-12-16
    • 1970-01-01
    相关资源
    最近更新 更多