Test for str:
isinstance(unicode_or_bytestring, str)
or, if you must handle bytestrings, test for bytes separately:
isinstance(unicode_or_bytestring, bytes)
The two types are deliberately not exchangible; use explicit encoding (for str -> bytes) and decoding (bytes -> str) to convert between the types.
In Python 2, where the modern Python 3 str type is called unicode and str is the precursor of the Python 3 bytes type, you could use basestring to test for both:
isinstance(unicode_or_bytestring, basestring)
basestring is only available in Python 2, and is the abstract base type of both str and unicode.
If you wanted to test for just unicode, then do so explicitly:
isinstance(unicode_tring, unicode)
Answer from Martijn Pieters on Stack OverflowTest for str:
isinstance(unicode_or_bytestring, str)
or, if you must handle bytestrings, test for bytes separately:
isinstance(unicode_or_bytestring, bytes)
The two types are deliberately not exchangible; use explicit encoding (for str -> bytes) and decoding (bytes -> str) to convert between the types.
In Python 2, where the modern Python 3 str type is called unicode and str is the precursor of the Python 3 bytes type, you could use basestring to test for both:
isinstance(unicode_or_bytestring, basestring)
basestring is only available in Python 2, and is the abstract base type of both str and unicode.
If you wanted to test for just unicode, then do so explicitly:
isinstance(unicode_tring, unicode)
Is there a Unicode string object type?
Yes, it is called unicode:
>>> s = u'hello'
>>> isinstance(s, unicode)
True
>>>
Note that in Python 3.x, this type was removed because all strings are now Unicode.
python - How do I check if a string is unicode or ascii? - Stack Overflow
Python 2: Says isinstance(string_obj, six.text_types) is unicode
python - What is the difference between isinstance('aaa', basestring) and isinstance('aaa', str)? - Stack Overflow
Unicode check
In Python 3, all strings are sequences of Unicode characters. There is a bytes type that holds raw bytes.
In Python 2, a string may be of type str or of type unicode. You can tell which using code something like this:
def whatisthis(s):
if isinstance(s, str):
print "ordinary string"
elif isinstance(s, unicode):
print "unicode string"
else:
print "not a string"
This does not distinguish "Unicode or ASCII"; it only distinguishes Python types. A Unicode string may consist of purely characters in the ASCII range, and a bytestring may contain ASCII, encoded Unicode, or even non-textual data.
How to tell if an object is a unicode string or a byte string
You can use type or isinstance.
In Python 2:
>>> type(u'abc') # Python 2 unicode string literal
<type 'unicode'>
>>> type('abc') # Python 2 byte string literal
<type 'str'>
In Python 2, str is just a sequence of bytes. Python doesn't know what
its encoding is. The unicode type is the safer way to store text.
If you want to understand this more, I recommend http://farmdev.com/talks/unicode/.
In Python 3:
>>> type('abc') # Python 3 unicode string literal
<class 'str'>
>>> type(b'abc') # Python 3 byte string literal
<class 'bytes'>
In Python 3, str is like Python 2's unicode, and is used to
store text. What was called str in Python 2 is called bytes in Python 3.
How to tell if a byte string is valid utf-8 or ascii
You can call decode. If it raises a UnicodeDecodeError exception, it wasn't valid.
>>> u_umlaut = b'\xc3\x9c' # UTF-8 representation of the letter 'Ü'
>>> u_umlaut.decode('utf-8')
u'\xdc'
>>> u_umlaut.decode('ascii')
Traceback (most recent call last):
File "<stdin>", line 1, in <module>
UnicodeDecodeError: 'ascii' codec can't decode byte 0xc3 in position 0: ordinal not in range(128)
In Python versions prior to 3.0 there are two kinds of strings "plain strings" and "unicode strings". Plain strings (str) cannot represent characters outside of the Latin alphabet (ignoring details of code pages for simplicity). Unicode strings (unicode) can represent characters from any alphabet including some fictional ones like Klingon.
So why have two kinds of strings, would it not be better to just have Unicode since that would cover all the cases? Well it is better to have only Unicode but Python was created before Unicode was the preferred method for representing strings. It takes time to transition the string type in a language with many users, in Python 3.0 it is finally the case that all strings are Unicode.
The inheritance hierarchy of Python strings pre-3.0 is:
object
|
|
basestring
/ \
/ \
str unicode
'basestring' introduced in Python 2.3 can be thought of as a step in the direction of string unification as it can be used to check whether an object is an instance of str or unicode
>>> string1 = "I am a plain string"
>>> string2 = u"I am a unicode string"
>>> isinstance(string1, str)
True
>>> isinstance(string2, str)
False
>>> isinstance(string1, unicode)
False
>>> isinstance(string2, unicode)
True
>>> isinstance(string1, basestring)
True
>>> isinstance(string2, basestring)
True
All strings are basestrings, but unicode strings are not of type str. Try this instead:
>>> a=u'aaaa'
>>> print isinstance(a, basestring)
True
>>> print isinstance(a, str)
False