【问题标题】:How do I feed an array of Tokenized Sentences to Word2vec to get embeddings?如何将一组标记化句子提供给 Word2vec 以获取嵌入?
【发布时间】:2021-11-18 23:57:54
【问题描述】:

大家好:我不知道从 word2vec 模型中获取嵌入所需的代码。

这是我的 df 的结构(它是一些基于 android 的日志):

日志日期时间 |行号 |进程ID |线程ID |优先 |应用 |留言 |事件模板 |事件ID ts int int int str str str str str

基本上,我从日志消息中创建了一个独特的事件子集,并分配了一个具有关联 ID 的模板:

def eventCreation(df):
    df['eventTemplate'] = df['message'].str.replace('\d+', '*')
    df['eventTemplate'] = df['eventTemplate'].str.replace('true', '*')
    df['eventTemplate'] = df['eventTemplate'].str.replace('false', '*')
    df['eventID'] = df.groupby(df.eventTemplate.tolist(), sort=False).ngroup() + 1
    df['eventID'] = 'E'+df['eventID'].astype(str)

def seqGen(arr, k):
    for i in range(len(arr)-k+1):
        yield arr[i:i+k]

#define the variables here
cwd = os.getcwd()
#create a dataframe of the logs concatenated
df = pd.DataFrame.from_records(process_files(cwd,getFiles))
# call functions to establish df
cleanDf(df)
featureEng(df)
eventCreation(df)
df['eventToken'] = df.eventTemplate.apply(lambda x: word_tokenize(x))
seq = []
eventArray = df[["eventToken"]].to_numpy()
for sequence in seqGen(eventArray, 9):
    seq.append(eventArray)

所以,'seq' 最终看起来像这样:

[array([['[*,com.blah.blach.blahMainblach] '],
        ['[*,*,*,com.blah.blah/.permission.ui.blah,finish-imm] '],
        ['[*,*,*,*,startingNewTask] '],
        ...,
        ['mfc, isSoftKeyboardVisible in WMS : * '],
        ['mfc, isSoftKeyboardVisible in WMS : * '],
        ['Calling a method in the system process without a qualified user: android.app.ContextImpl.startService:* android.content.ContextWrapper.startService:* android.content.ContextWrapper.startService:* com.blahblah.usbmountreceiver.USBMountReceiver.onReceive:* android.app.ActivityThread.handleReceiver:* ']],
       dtype=object),

序列是带有标记化日志消息列表的数组。计划是在训练模型之后,我可以通过将 onehot 向量和权重矩阵相乘来获得日志事件的嵌入……还有更多工作要做,但我一直在获取嵌入。

我是一个试图开发异常检测解决方案的新手。

【问题讨论】:

    标签: python-3.x logging word2vec word-embedding anomaly-detection


    【解决方案1】:

    如果您使用Python中的Gensim库以其Word2Vec实现,它需要其语言作为重新迭代序列 em>,其中每个项目本身都是字符串令牌列表 em>。

    本身具有每个项目作为字符串令牌的列表将工作。

    你的seq是关闭的,但是:

    1. 它不需要(且可能不应该是)A numpy armet的对象。
    2. 你的每一个object项目是list(好),但每个@ 987654325,但是每个list您需要将这些字符串分解为您希望模型学习的个人“单词”。

    【讨论】:

      猜你喜欢
      • 2015-06-27
      • 2020-12-26
      • 1970-01-01
      • 2019-09-04
      • 2021-05-31
      • 2020-11-14
      • 2020-04-07
      • 1970-01-01
      • 2021-11-29
      相关资源
      最近更新 更多