【问题标题】:How do I shorten and expand a uuid to a 15 or less characters如何将 uuid 缩短和扩展为 15 个或更少的字符
【发布时间】:2017-12-21 10:02:11
【问题描述】:

给定一个不带破折号的 uuid(v4),如何将其缩短为 15 个或少于 15 个字符的字符串?我应该也可以从15个字符的字符串回到原来的uuid。

我正在尝试缩短它以将其发送到平面文件中,并且文件格式指定此字段为 15 个字符的字母数字字段。鉴于缩短的 uuid,我应该能够将其映射回原始 uuid。

这是我尝试过的,但绝对不是我想要的。

export function shortenUUID(uuidToShorten: string, length: number) {
  const uuidWithoutDashes = uuidToShorten.replace(/-/g , '');
  const radix = uuidWithoutDashes.length;
  const randomId = [];

  for (let i = 0; i < length; i++) {
    randomId[i] = uuidWithoutDashes[ 0 | Math.random() * radix];
  }
  return randomId.join('');
}

【问题讨论】:

  • 你已经尝试了什么?
  • 对不起,我的错。我应该添加我正在做的事情。那完全没用,因此没有添加。
  • 我真的不明白你要做什么。您的算法是不确定的(即Math.random()),因此每次您在相同的输入 UUID 上运行 shortenUUID() 时,您都会生成一个不同的缩短版本。 如何您希望以这种方式进行 reverse 映射吗?制作更短的 UUID 的最终目标是什么?您的问题可能有完全不同的解决方案。
  • 正是我的意思,这就是我提到的原因,我所做的毫无用处。
  • @msanford 我正在尝试缩短它以将其发送到平面文件中,并且文件格式将此字段指定为 15 个字符的字母数字字段。鉴于缩短的 uuid,我应该能够将其映射回原始 uuid。

标签: javascript node.js typescript


【解决方案1】:

正如 AuxTaco 指出的那样,如果您实际上是指“字母数字”,因为它匹配“/^[A-Za-z0-9]{0,15}/”(给出 26 + 26 + 10 的位数= 62),那么这真的是不可能的。你不可能在不丢失任何东西的情况下将 3 加仑的水放入一加仑桶中。 UUID 是 128 位,因此要将其转换为 62 个字符空间,您至少需要 22 个字符 (log[base 62](2^128) == ~22)。

如果您的字符集更灵活,并且只需要 15 个 unicode 字符即可放入文本文档,那么我的回答会有所帮助。


注意:这个答案的第一部分,我认为它说的长度是 16,而不是 15。更简单的答案是行不通的。下面更复杂的版本仍然会。


为此,您需要使用某种双向压缩算法(类似于用于压缩文件的算法)。

但是,尝试压缩 UUID 之类的东西的问题是,您可能会遇到很多冲突。

UUID v4 的长度为 32 个字符(没有破折号)。它是十六进制的,所以它的字符空间是 16 个字符 (0123456789ABCDEF)

这为您提供了16^32、大约3.4028237e+38 或340,282,370,000,000,000,000,000,000,000,000,000,000 的多种可能组合。要使其在压缩后可恢复,您必须确保没有任何冲突(即,没有 2 个 UUID 变成相同的值)。这是很多可能的值(这正是我们使用这么多 UUID 的原因,2 个随机 UUID 的机会只是该数字中的 1 个)。

要将这么多可能性压缩到 16 个字符,您必须至少拥有尽可能多的可能值。如果有 16 个字符,则必须有 256 个字符(那个大数字的根 16,256^16 == 16^32`)。这是假设你有一个永远不会产生冲突的算法。

确保永远不会发生冲突的一种方法是将其从以 16 为底的数字转换为以 256 为底的数字。这将为您提供一对一的关系,确保没有碰撞并使其完全可逆。通常,在 JavaScript 中切换基数很容易:parseInt(someStr, radix).toString(otherRadix)(例如,parseInt('00FF', 16).toString(20)。不幸的是,JavaScript 最多只能处理 36 的基数,所以我们必须自己进行转换。

基数如此之大的渔获就代表了它。您可以任意选择 256 个不同的字符,将它们放入一个字符串中,然后将其用于手动转换。但是,即使您将大写和小写视为不同的字形,我也不认为标准美式键盘上有 256 个不同的符号。

更简单的解决方案是使用0 到255 和String.fromCharCode() 之间的任意字符代码。

另一个小问题是,如果我们试图将所有这些都视为一个大数字,我们就会遇到问题,因为它是一个非常大的数字,而 JavaScript 无法正确准确地表示它。

取而代之的是,由于我们已经有了十六进制,我们可以将其拆分为成对的小数,转换它们,然后将它们吐出。 32 个十六进制数字 = 16 对,所以(巧合地)这将是完美的。 (如果你必须解决这个任意大小的问题,你必须做一些额外的数学运算并转换为将数字分成几部分,转换,然后重新组合。)

const uuid = '1234567890ABCDEF1234567890ABCDEF';
const letters = uuid.match(/.{2}/g).map(pair => String.fromCharCode(parseInt(pair, 16)));
const str = letters.join('');
console.log(str);

请注意,其中有一些随机字符,因为并非每个字符代码都映射到“正常”符号。如果您要发送的内容无法处理它们,则需要使用数组方法:找到它可以处理的 256 个字符,将它们组成一个数组,然后使用 String.fromCharCode(num) 而不是 charset[num]。

要将其转换回来,您只需执行相反的操作:获取字符代码,转换为十六进制,然后将它们加在一起:

const uuid = '1234567890ABCDEF1234567890ABCDEF';

const compress = uuid => 
  uuid.match(/.{2}/g).map(pair => String.fromCharCode(parseInt(pair, 16))).join('');
  
const expand = str =>
  str.split('').map(letter => ('0' + letter.charCodeAt(0).toString(16)).substr(-2)).join('');
  
const str = compress(uuid);
const original = expand(str);

console.log(str, original, original.toUpperCase() === uuid.toUpperCase());

为了好玩,这里是您可以为任意输入基数和输出基数执行此操作的方法。

这段代码有点乱,因为它确实被扩展以使其更加不言自明,但它基本上完成了我上面描述的操作。

由于 JavaScript 没有无限的精度,如果你最终转换了一个非常大的数字(看起来像 2.00000000e+10),那么在此之后没有显示的每个数字 e 基本上都被切掉和替换了零。为了解决这个问题,你必须以某种方式分解它。

在下面的代码中,有一种“简单”的方式没有考虑到这一点,因此只适用于较小的字符串,然后是一种将其分解的正确方式。我选择了一种简单但效率较低的方法,即根据字符串变成多少位数来分解字符串。这不是最好的方法(因为数学并不是真的那样工作),但它确实有效(以需要更小的字符集为代价)。

如果您确实需要将字符集大小保持在最低限度,则可以采用更智能的拆分机制。

const smallStr = '1234';
const str = '1234567890ABCDEF1234567890ABCDEF';
const hexCharset = '0123456789ABCDEF'; // could also be an array
const compressedLength = 16;
const maxDigits = 16; // this may be a bit browser specific. You can make it smaller to be safer.

const logBaseN = (num, n) => Math.log(num) / Math.log(n);
const nthRoot = (num, n) => Math.pow(num, 1/n);
const digitsInNumber = num => Math.log(num) * Math.LOG10E + 1 | 0;
const partitionString = (str, numPartitions) => {
  const partsSize = Math.ceil(str.length / numPartitions);
  let partitions = [];
  for (let i = 0; i < numPartitions; i++) {
    partitions.push(str.substr(i * partsSize, partsSize));
  }
  return partitions;
}

console.log('logBaseN test:', logBaseN(256, 16) === 2);
console.log('nthRoot test:', nthRoot(256, 2) === 16);
console.log('partitionString test:', partitionString('ABCDEFG', 3));

// charset.length should equal radix
const toDecimalFromCharset = (str, charset) => 
    str.split('')
      .reverse()
      .map((char, index) => charset.indexOf(char) * Math.pow(charset.length, index))
      .reduce((sum, num) => (sum + num), 0);
      
const fromDecimalToCharset = (dec, charset) => {
  const radix = charset.length;
  let str = '';
  
  for (let i = Math.ceil(logBaseN(dec + 1, radix)) - 1; i >= 0; i--) {
    const part = Math.floor(dec / Math.pow(radix, i));
    dec -= part * Math.pow(radix, i);
    str += charset[part];
  }
  
  return str;
};

console.log('toDecimalFromCharset test 1:', toDecimalFromCharset('01000101', '01') === 69);
console.log('toDecimalFromCharset test 2:', toDecimalFromCharset('FF', hexCharset) === 255);
console.log('fromDecimalToCharset test:', fromDecimalToCharset(255, hexCharset) === 'FF');

const arbitraryCharset = length => new Array(length).fill(1).map((a, i) => String.fromCharCode(i));

// the Math.pow() bit is the possible number of values in the original
const simpleDetermineRadix = (strLength, originalCharsetSize, compressedLength) => nthRoot(Math.pow(originalCharsetSize, strLength), compressedLength);

// the simple ones only work for values that in decimal are so big before lack of precision messes things up
// compressedCharset.length must be >= compressedLength
const simpleCompress = (str, originalCharset, compressedCharset, compressedLength) =>
  fromDecimalToCharset(toDecimalFromCharset(str, originalCharset), compressedCharset);

const simpleExpand = (compressedStr, originalCharset, compressedCharset) =>
  fromDecimalToCharset(toDecimalFromCharset(compressedStr, compressedCharset), originalCharset);

const simpleNeededRadix = simpleDetermineRadix(str.length, hexCharset.length, compressedLength);
const simpleCompressedCharset = arbitraryCharset(simpleNeededRadix);
const simpleCompressed = simpleCompress(str, hexCharset, simpleCompressedCharset, compressedLength);
const simpleExpanded = simpleExpand(simpleCompressed, hexCharset, simpleCompressedCharset);

// Notice, it gets a little confused because of a lack of precision in the really big number.
console.log('Original string:', str, toDecimalFromCharset(str, hexCharset));
console.log('Simple Compressed:', simpleCompressed, toDecimalFromCharset(simpleCompressed, simpleCompressedCharset));
console.log('Simple Expanded:', simpleExpanded, toDecimalFromCharset(simpleExpanded, hexCharset));
console.log('Simple test:', simpleExpanded === str);

// Notice it works fine for smaller strings and/or charsets
const smallCompressed = simpleCompress(smallStr, hexCharset, simpleCompressedCharset, compressedLength);
const smallExpanded = simpleExpand(smallCompressed, hexCharset, simpleCompressedCharset);
console.log('Small string:', smallStr, toDecimalFromCharset(smallStr, hexCharset));
console.log('Small simple compressed:', smallCompressed, toDecimalFromCharset(smallCompressed, simpleCompressedCharset));
console.log('Small expaned:', smallExpanded, toDecimalFromCharset(smallExpanded, hexCharset));
console.log('Small test:', smallExpanded === smallStr);

// these will break the decimal up into smaller numbers with a max length of maxDigits
// it's a bit browser specific where the lack of precision is, so a smaller maxDigits
//  may make it safer
//
// note: charset may need to be a little bit bigger than what determineRadix decides, since we're
//  breaking the string up
// also note: we're breaking the string into parts based on the number of digits in it as a decimal
//  this will actually make each individual parts decimal length smaller, because of how numbers work,
//  but that's okay. If you have a charset just barely big enough because of other constraints, you'll
//  need to make this even more complicated to make sure it's perfect.
const partitionStringForCompress = (str, originalCharset) => {
    const numDigits = digitsInNumber(toDecimalFromCharset(str, originalCharset));
    const numParts = Math.ceil(numDigits / maxDigits);
    return partitionString(str, numParts);
}

const partitionedPartSize = (str, originalCharset) => {
    const parts = partitionStringForCompress(str, originalCharset);
    return Math.floor((compressedLength - parts.length - 1) / parts.length) + 1;
}

const determineRadix = (str, originalCharset, compressedLength) => {
    const parts = partitionStringForCompress(str, originalCharset);
    return Math.ceil(nthRoot(Math.pow(originalCharset.length, parts[0].length), partitionedPartSize(str, originalCharset)));
}

const compress = (str, originalCharset, compressedCharset, compressedLength) => {
    const parts = partitionStringForCompress(str, originalCharset);
    const partSize = partitionedPartSize(str, originalCharset);
    return parts.map(part => simpleCompress(part, originalCharset, compressedCharset, partSize)).join(compressedCharset[compressedCharset.length-1]);
}

const expand = (compressedStr, originalCharset, compressedCharset) =>
    compressedStr.split(compressedCharset[compressedCharset.length-1])
     .map(part => simpleExpand(part, originalCharset, compressedCharset))
     .join('');

const neededRadix = determineRadix(str, hexCharset, compressedLength);
const compressedCharset = arbitraryCharset(neededRadix);
const compressed = compress(str, hexCharset, compressedCharset, compressedLength);
const expanded = expand(compressed, hexCharset, compressedCharset);

console.log('String:', str, toDecimalFromCharset(str, hexCharset));
console.log('Neded radix size:', neededRadix); // bigger than normal because of how we're breaking it up... this could be improved if needed
console.log('Compressed:', compressed);
console.log('Expanded:', expanded);
console.log('Final test:', expanded === str);

要专门使用上述内容来回答问题,您可以使用:

const hexCharset = '0123456789ABCDEF';
const compressedCharset = arbitraryCharset(determineRadix(uuid, hexCharset));

// UUID to 15 characters
const compressed = compress(uuid, hexCharset, compressedCharset, 15);

// 15 characters to UUID
const expanded = expanded(compressed, hexCharset, compressedCharset);

如果任意字符中有问题,您必须做一些事情来过滤掉这些字符,或者硬编码一个特定的字符。只需确保所有函数都是确定性的(即每次结果都相同)。

【讨论】:

  • 有些问题。这不适用于表示一位十六进制数的对:尝试uuid = '01020304050607080102030405060708';。如果 UUID 包含 00,则您发送到的系统可能会在“15 个字符的字母数字字段”中阻止 U+0000 NULL。这压缩到 16 个字符,而不是 15 个。
  • 糟糕,出于某种原因,我认为它说的是 16。我实际上用一种更复杂的方法更新了我的答案,该方法适用于任意字符集和长度,因此可以用于更小的内容。您还可以过滤掉任何有问题的字符。不知道为什么您认为它不适用于表示单个数字的十六进制,除了它们可能是某些系统可能不喜欢的愚蠢字符。当将其视为相对任意时,它基本上都是二进制的。
  • 其实,明白为什么了。转换回来时它不会填充零。解决这个问题。 =p
  • 如果“15 个字符的字母数字字段”确实意味着 /^[a-zA-Z0-9\s]{15}$/ 正如您所指出的,那么是的,这实际上是不可能的。将其添加到我的答案中。
  • 这解决了我遇到的所有问题!不错!
猜你喜欢
  • 2012-07-14
  • 2020-06-26
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2013-12-26
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
相关资源
最近更新 更多