【问题标题】:Phrase/multi-word and counting matching across large data sets跨大型数据集的短语/多词和计数匹配
【发布时间】:2017-07-18 07:55:57
【问题描述】:

我有大量的数字和字母数字集,我想在其中找到常用的单词/短语,其中包含 python 2.7。

示例数据,与我的真实数据没有什么相似之处,但这可以很好地表示它。

'this is a test of the hosting',
'test is a test',
'we have more tests to run before we can trust it',
'if it true,  can trust it',
'tom is on time for ounce',
'what do you mean tom is out sick again'

我正在寻找以下类型的匹配

'is' x 5
'test' x 3
'is a test' x 2
'is a' x2
'we' x2
'trust it' x 2
'tom' x 2
..etc..

是否有一个通用的库或者我需要编写一个?我可以用蛮力做到这一点,但在我的一些较大的文件上,这可能需要数年时间。我“假设”这是一个常见问题,一些智能 cookie 已经找到了解决方案。希望这不是一个旅行推销员。

【问题讨论】:

  • 您在寻找一元、二元、三元等计数吗?
  • 我不得不承认,我不知道你所说的 unigram、bigram、trigram 是什么意思......但是快速查找让我想到了单词级别的 bigram/trigram/etc.. 匹配。我认为 4 字匹配的任何匹配集都是我想要处理的最大集。

标签: algorithm python-2.7


【解决方案1】:

我认为您正在寻找 unigram、bigram、trigram 计数。你可以在 Python 中使用 NLTK 库来做你想做的事。

另外,请查看link。

【讨论】:

  • 当我看到你的 unigram、bigram、trigram 并搜索“python unigram bigram trigram”时,我发现了很多关于它的信息。谢谢!
  • @JustBroken: 随时 :) 很多时候,只要一点提示就能得到你想要的!
猜你喜欢
  • 1970-01-01
  • 2021-05-11
  • 1970-01-01
  • 1970-01-01
  • 2013-03-15
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2013-01-21
相关资源
最近更新 更多