我制作这个脚本是为了让所有人都了解如何标记化,这样他们就可以自己构建自然语言处理的引擎。
import re
from contextlib import redirect_stdout
from io import StringIO
example = 'Mary had a little lamb, Jack went up the hill, Jill followed suit, i woke up suddenly, it was a really bad dream...'
def token_to_sentence(str):
f = StringIO()
with redirect_stdout(f):
regex_of_sentence = re.findall('([\w\s]{0,})[^\w\s]', str)
regex_of_sentence = [x for x in regex_of_sentence if x is not '']
for i in regex_of_sentence:
print(i)
first_step_to_sentence = (f.getvalue()).split('\n')
g = StringIO()
with redirect_stdout(g):
for i in first_step_to_sentence:
try:
regex_to_clear_sentence = re.search('\s([\w\s]{0,})', i)
print(regex_to_clear_sentence.group(1))
except:
print(i)
sentence = (g.getvalue()).split('\n')
return sentence
def token_to_words(str):
f = StringIO()
with redirect_stdout(f):
for i in str:
regex_of_word = re.findall('([\w]{0,})', i)
regex_of_word = [x for x in regex_of_word if x is not '']
for word in regex_of_word:
print(regex_of_word)
words = (f.getvalue()).split('\n')
我做了一个不同的过程,我从段落重新开始这个过程,让大家更了解文字处理。要处理的段落是:
example = 'Mary had a little lamb, Jack went up the hill, Jill followed suit, i woke up suddenly, it was a really bad dream...'
将段落标记为句子:
sentence = token_to_sentence(example)
结果:
['Mary had a little lamb', 'Jack went up the hill', 'Jill followed suit', 'i woke up suddenly', 'it was a really bad dream']
标记为单词:
words = token_to_words(sentence)
结果:
['Mary', 'had', 'a', 'little', 'lamb', 'Jack', 'went, 'up', 'the', 'hill', 'Jill', 'followed', 'suit', 'i', 'woke', 'up', 'suddenly', 'it', 'was', 'a', 'really', 'bad', 'dream']
我将解释这是如何工作的。
首先,我使用正则表达式搜索所有分隔单词的单词和空格并停止直到找到标点符号,正则表达式是:
([\w\s]{0,})[^\w\s]{0,}
所以计算将采用括号中的单词和空格:
'(Mary had a little lamb),( Jack went up the hill, Jill followed suit),( i woke up suddenly),( it was a really bad dream)...'
结果仍不清楚,包含一些“无”字符。所以我用这个脚本删除了“无”字符:
[x for x in regex_of_sentence if x is not '']
因此该段落将标记为句子,但不清楚句子的结果是:
['Mary had a little lamb', ' Jack went up the hill', ' Jill followed suit', ' i woke up suddenly', ' it was a really bad dream']
如您所见,结果显示了一些以空格开头的句子。所以为了在不开始空格的情况下制作一个清晰的段落,我制作了这个正则表达式:
\s([\w\s]{0,})
它会做出一个清晰的句子,如:
['Mary had a little lamb', 'Jack went up the hill', 'Jill followed suit', 'i woke up suddenly', 'it was a really bad dream']
所以,我们必须做两个过程才能取得好的结果。
你的问题的答案从这里开始……
为了将句子标记为单词,我进行段落迭代并使用正则表达式来捕获单词,同时它正在使用这个正则表达式进行迭代:
([\w]{0,})
然后再次清除空字符:
[x for x in regex_of_word if x is not '']
所以结果真的很清楚只有单词列表:
['Mary', 'had', 'a', 'little', 'lamb', 'Jack', 'went, 'up', 'the', 'hill', 'Jill', 'followed', 'suit', 'i', 'woke', 'up', 'suddenly', 'it', 'was', 'a', 'really', 'bad', 'dream']
以后要做好NLP,你需要有自己的词组数据库,如果词组在句子中,搜索一下,做成词组列表后,剩下的词就是一个词了。
使用这种方法,我可以用我的语言(印度尼西亚语)构建我自己的 NLP,它真的非常缺乏模块。
编辑:
我没有看到您想要比较单词的问题。所以你还有一句话要比较....我给你的奖金不仅仅是奖金,我给你怎么算。
mod_example = ["'Mary' 'had' 'a' 'little' 'lamb'" , 'Jack' 'went' 'up' 'the' 'hill']
在这种情况下,您必须执行的步骤是:
1. 迭代 mod_example
2. 将第一句话与 mod_example 中的单词进行比较。
3. 计算一下
所以脚本将是:
import re
from contextlib import redirect_stdout
from io import StringIO
example = 'Mary had a little lamb, Jack went up the hill, Jill followed suit, i woke up suddenly, it was a really bad dream...'
mod_example = ["'Mary' 'had' 'a' 'little' 'lamb'" , 'Jack' 'went' 'up' 'the' 'hill']
def token_to_sentence(str):
f = StringIO()
with redirect_stdout(f):
regex_of_sentence = re.findall('([\w\s]{0,})[^\w\s]', str)
regex_of_sentence = [x for x in regex_of_sentence if x is not '']
for i in regex_of_sentence:
print(i)
first_step_to_sentence = (f.getvalue()).split('\n')
g = StringIO()
with redirect_stdout(g):
for i in first_step_to_sentence:
try:
regex_to_clear_sentence = re.search('\s([\w\s]{0,})', i)
print(regex_to_clear_sentence.group(1))
except:
print(i)
sentence = (g.getvalue()).split('\n')
return sentence
def token_to_words(str):
f = StringIO()
with redirect_stdout(f):
for i in str:
regex_of_word = re.findall('([\w]{0,})', i)
regex_of_word = [x for x in regex_of_word if x is not '']
for word in regex_of_word:
print(regex_of_word)
words = (f.getvalue()).split('\n')
def convert_to_words(str):
sentences = token_to_sentence(str)
for i in sentences:
word = token_to_words(i)
return word
def compare_list_of_words__to_another_list_of_words(from_strA, to_strB):
fromA = list(set(from_strA))
for word_to_match in fromA:
totalB = len(to_strB)
number_of_match = (to_strB).count(word_to_match)
data = str((((to_strB).count(word_to_match))/totalB)*100)
print('words: -- ' + word_to_match + ' --' + '\n'
' number of match : ' + number_of_match + ' from ' + str(totalB) + '\n'
' percent of match : ' + data + ' percent')
#prepare already make, now we will use it. The process start with script below:
if __name__ == '__main__':
#tokenize paragraph in example to sentence:
getsentences = token_to_sentence(example)
#tokenize sentence to words (sentences in getsentences)
getwords = token_to_words(getsentences)
#compare list of word in (getwords) with list of words in mod_example
compare_list_of_words__to_another_list_of_words(getwords, mod_example)