Python NLTK:Bigrams三字组fourgrams [英] Python NLTK: Bigrams trigrams fourgrams

查看：132 发布时间：2020/5/18 1:13:48 python nltk n-gram

本文介绍了Python NLTK:Bigrams三字组fourgrams的处理方法，对大家解决问题具有一定的参考价值，需要的朋友们下面随着小编来一起学习吧！

问题描述

我有这个例子，我想知道如何得到这个结果.我有文字，并且将其标记化，然后收集了像这样的双字母组和三字母组和四字母组

I have this example and i want to know how to get this result. I have text and I tokenize it then I collect the bigram and trigram and fourgram like that

import nltk
from nltk import word_tokenize
from nltk.util import ngrams
text = "Hi How are you? i am fine and you"
token=nltk.word_tokenize(text)
bigrams=ngrams(token,2)

字母组合:[('Hi', 'How'), ('How', 'are'), ('are', 'you'), ('you', '?'), ('?', 'i'), ('i', 'am'), ('am', 'fine'), ('fine', 'and'), ('and', 'you')]

trigrams=ngrams(token,3)

语法图:[('Hi', 'How', 'are'), ('How', 'are', 'you'), ('are', 'you', '?'), ('you', '?', 'i'), ('?', 'i', 'am'), ('i', 'am', 'fine'), ('am', 'fine', 'and'), ('fine', 'and', 'you')]

bigram [(a,b) (b,c) (c,d)]
trigram [(a,b,c) (b,c,d) (c,d,f)]
i want the new trigram should be [(c,d,f)]
which mean 
newtrigram = [('are', 'you', '?'),('?', 'i','am'),...etc

任何想法都会有所帮助

推荐答案

如果应用一些集合论(如果我正确地解释了您的问题)，您会看到想要的三元组只是元素[2:5 ]，[4:7]，[6:8]等.

If you apply some set theory (if I'm interpreting your question correctly), you'll see that the trigrams you want are simply elements [2:5], [4:7], [6:8], etc. of the token list.

您可以这样生成它们:

>>> new_trigrams = []
>>> c = 2
>>> while c < len(token) - 2:
...     new_trigrams.append((token[c], token[c+1], token[c+2]))
...     c += 2
>>> print new_trigrams
[('are', 'you', '?'), ('?', 'i', 'am'), ('am', 'fine', 'and')]

这篇关于Python NLTK:Bigrams三字组fourgrams的文章就介绍到这了，希望我们推荐的答案对大家有所帮助，也希望大家多多支持IT屋！

查看全文

Python NLTK:Bigrams三字组fourgrams [英] Python NLTK: Bigrams trigrams fourgrams

问题描述

推荐答案

相关文章

Python最新文章

热门教程

热门工具

登录关闭

Python NLTK:Bigrams三字组fourgrams [英] Python NLTK: Bigrams trigrams fourgrams

问题描述

推荐答案

相关文章

Python最新文章

热门教程

热门工具

登录 关闭

登录关闭