Non-ASCII character error using nltk in Python

Question

I'm trying to use the solution code given in the following link: Unicode Tagging in Python NLTK

In the solution given by omerbp:

from nltk.corpus import indian
from nltk.tag import tnt

train_data = indian.tagged_sents('hindi.pos')
tnt_pos_tagger = tnt.TnT()
tnt_pos_tagger.train(train_data) #Training the tnt Part of speech tagger with hindi data

print tnt_pos_tagger.tag(nltk.word_tokenize(word_to_be_tagged))

I'm getting the following error:

'SyntaxError: Non-ASCII character '\xe0' in file q12.py on line 1, but no encoding declared; see http://www.python.org/peps/pep-0263.html for details' in line 1.

This error message seems to come up a lot here - are any of those links helpful? — halfer
– halfer, Commented Mar 5, 2016 at 11:55
Possible duplicate of Python NLTK: SyntaxError: Non-ASCII character '\xc3' in file (Senitment Analysis -NLP) — CaptSolo
– CaptSolo, Commented Mar 5, 2016 at 12:12

Alex · Accepted Answer · 2016-03-10 17:25:48Z

1

Add these two lines on the top of your file:

#!/usr/bin/python
# -*- coding: utf-8 -*-

They will instruct the interpreter to encode every charater as UTF-8 instead of ASCII.

answered Mar 10, 2016 at 17:25

Alex

7,0676 gold badges22 silver badges36 bronze badges

Sign up to request clarification or add additional context in comments.

Collectives™ on Stack Overflow

Non-ASCII character error using nltk in Python

1 Answer 1

Comments

Your Answer

Linked

Hot Network Questions

Collectives™ on Stack Overflow

1 Answer 1

Comments

Your Answer

Sign up or log in

Post as a guest

Linked

Related