Skip to main content

Natural Language Processing with Python NLTK part 2 - Stop Words

Natural Language Processing


Stop words are the words which we ignore due to the fact that they do not generate any specific meaning to the sentence. Words like the, is, at etc. can be removed to extract the meaning of the sentence more easily. So NLTK has introduced us a stop words filter we can easily use. Let's see how it works.


from nltk.corpus import stopwords
from nltk.tokenize import word_tokenize

sent = "As you can see this is the blog of myself which is written by Anjula"

w = word_tokenize(sent)

# set English stop words
stop_words = set(stopwords.words('english'))

# list of standard stop words in English
print(stop_words)

# making empty arrays to store stop words and others
stop_words_in_sent = []
non_stop_words = []

# Loop through to get the stop words
for x in w:
    if x not in stop_words:
        non_stop_words.append(x)
    else:
        stop_words_in_sent.append(x)

# print result
print(non_stop_words)
print(stop_words_in_sent)

The code is simple as that the output will be as follows:




Popular posts from this blog

Natural Language Processing with Python NLTK part 5 - Chunking and Chinking

Natural Language Processing Using regular expression modifiers we can chunk out the PoS tagged words from the earlier example. The chunking is done with regular expressions defining a chunk rule. The Chinking defines what we need to exclude from the selection. Here are list of modifiers for Python: {1,3} = for digits, u expect 1-3 counts of digits, or "places" + = match 1 or more ? = match 0 or 1 repetitions. * = match 0 or MORE repetitions $ = matches at the end of string ^ = matches start of a string | = matches either/or. Example x|y = will match either x or y [] = range, or "variance" {x} = expect to see this amount of the preceding code. {x,y} = expect to see this x-y amounts of the preceding code source: https://pythonprogramming.net/regular-expressions-regex-tutorial-python-3/ Chunking import nltk from nltk.tokenize import word_tokenize # POS tagging sent = "This will be chunked. This is for Test. World is awesome. Hello world....

Natural Language Processing with Python NLTK part 1 - Tokenizer

Natural Language Processing Starting with the NLP articles first we will try the  tokenizer  in the NLTK package. Tokenizer breaks a paragraph into the relevant sub strings or sentences based on the tokenizer you used. In this I will use the Sent tokenizer, word_tokenizer and TweetTokenizer which has its specific work to do. import nltk from nltk.tokenize import sent_tokenize, word_tokenize, TweetTokenizer para = "Hello there this is the blog about NLP. In this blog I have made some posts. " \ "I can come up with new content." tweet = "#Fun night. :) Feeling crazy #TGIF" # tokenizing the paragraph into sentences and words sent = sent_tokenize(para) word = word_tokenize(para) # printing the output print ( "this paragraph has " + str(len(sent)) + " sentences and " + str(len(word)) + " words" ) # print each sentence k = 1 for i in sent: print ( "sentence ...

Unexpected Success in Filtering Repeters

So thought the previous post was the last one, but guess what I was just having a typo in the controller. This is huge relief for me. Here what I followed from the tutorial was that you can actually create a controller for your app, and Include it using ng-controller. Controller is actually a function that use $scope . Through this the model and controller work together. So here is the Controller; 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 /* Controllers */ var myApp = angular.module( 'myApp' ,[]); myApp.controller( 'namectrl' , function ($scope){ $scope.names = [ { 'Fname' : 'anna' , 'Lname' : 'kendrick' }, { 'Fname' : 'jason' , 'Lname' : 'mraz' }, { 'Fname' : 'jennifer' , 'Lna...