Test for str:

isinstance(unicode_or_bytestring, str)

or, if you must handle bytestrings, test for bytes separately:

isinstance(unicode_or_bytestring, bytes)

The two types are deliberately not exchangible; use explicit encoding (for str -> bytes) and decoding (bytes -> str) to convert between the types.

In Python 2, where the modern Python 3 str type is called unicode and str is the precursor of the Python 3 bytes type, you could use basestring to test for both:

isinstance(unicode_or_bytestring, basestring)

basestring is only available in Python 2, and is the abstract base type of both str and unicode.

If you wanted to test for just unicode, then do so explicitly:

isinstance(unicode_tring, unicode)
Answer from Martijn Pieters on Stack Overflow
🌐
GitHub
github.com › jdunck › python-unicodecsv › issues › 51
No isinstance unicode in Python3 · Issue #51 · jdunck/python-unicodecsv
March 22, 2015 - The isinstance statement with unicode does not exist in python3: if isinstance(s, unicode): Here's how I fixed it in my versions, but I still don't know if it's actually the best way to test unicod...
Author: jdunck
Discussions

python - How do I check if a string is unicode or ascii? - Stack Overflow
In Python 2, a string may be of type str or of type unicode. You can tell which using code something like this: def whatisthis(s): if isinstance(s, str): print "ordinary string" elif isinstance(s, unicode): print "unicode string" else: print "not a string" More on stackoverflow.com
🌐 stackoverflow.com
February 14, 2011
Python 2: Says isinstance(string_obj, six.text_types) is unicode
mypy is claiming that the variable should be of type 'unicode' when it should allow type 'str' Sample code: #!/usr/bin/python -ttu import sys from typing import Any import six def m... More on github.com
🌐 github.com
5
January 27, 2018
python - What is the difference between isinstance('aaa', basestring) and isinstance('aaa', str)? - Stack Overflow
'basestring' introduced in Python 2.3 can be thought of as a step in the direction of string unification as it can be used to check whether an object is an instance of str or unicode · >>> string1 = "I am a plain string" >>> string2 = u"I am a unicode string" >>> isinstance(string1, str) True ... More on stackoverflow.com
🌐 stackoverflow.com
December 30, 2009
Unicode check
I wrote an if statement to determine ... was a unicode string or not, then perform some encoding. It would not work. I figured out a workaround, but could anyone explain why my original code failed? ... I started with Python 3, but I had a go at it anyway. ... You have to use a built-in instead. So: ... Ahh, I've never used the isinstance module ... More on reddit.com
🌐 r/learnpython
6
12
October 15, 2013
🌐
Python-future
python-future.org › isinstance.html
isinstance — Python-Future documentation
>>> from __future__ import unicode_literals >>> assert isinstance('unicode string 2', str)
Top answer
1 of 13
325

In Python 3, all strings are sequences of Unicode characters. There is a bytes type that holds raw bytes.

In Python 2, a string may be of type str or of type unicode. You can tell which using code something like this:

def whatisthis(s):
    if isinstance(s, str):
        print "ordinary string"
    elif isinstance(s, unicode):
        print "unicode string"
    else:
        print "not a string"

This does not distinguish "Unicode or ASCII"; it only distinguishes Python types. A Unicode string may consist of purely characters in the ASCII range, and a bytestring may contain ASCII, encoded Unicode, or even non-textual data.

2 of 13
136

How to tell if an object is a unicode string or a byte string

You can use type or isinstance.

In Python 2:

>>> type(u'abc')  # Python 2 unicode string literal
<type 'unicode'>
>>> type('abc')   # Python 2 byte string literal
<type 'str'>

In Python 2, str is just a sequence of bytes. Python doesn't know what its encoding is. The unicode type is the safer way to store text. If you want to understand this more, I recommend http://farmdev.com/talks/unicode/.

In Python 3:

>>> type('abc')   # Python 3 unicode string literal
<class 'str'>
>>> type(b'abc')  # Python 3 byte string literal
<class 'bytes'>

In Python 3, str is like Python 2's unicode, and is used to store text. What was called str in Python 2 is called bytes in Python 3.


How to tell if a byte string is valid utf-8 or ascii

You can call decode. If it raises a UnicodeDecodeError exception, it wasn't valid.

>>> u_umlaut = b'\xc3\x9c'   # UTF-8 representation of the letter 'Ü'
>>> u_umlaut.decode('utf-8')
u'\xdc'
>>> u_umlaut.decode('ascii')
Traceback (most recent call last):
  File "<stdin>", line 1, in <module>
UnicodeDecodeError: 'ascii' codec can't decode byte 0xc3 in position 0: ordinal not in range(128)
🌐
GitHub
github.com › python › mypy › issues › 4516
Python 2: Says isinstance(string_obj, six.text_types) is unicode · Issue #4516 · python/mypy
January 27, 2018 - mypy is claiming that the variable should be of type 'unicode' when it should allow type 'str' Sample code: #!/usr/bin/python -ttu import sys from typing import Any import six def main(): foo("hello") def foo(a): # type: (Any) -> None reveal_type(a) b = "hello" reveal_type(b) if isinstance(a, six.string_types): reveal_type(a) b = a print "Type of 'a' is {}".format(type(a)) if '__main__' == __name__: sys.exit(main()) Output is: $ mypy --py2 tt.py tt.py:14: error: Revealed type is 'Any' tt.py:16: error: Revealed type is 'builtins.str' tt.py:18: error: Revealed type is 'builtins.unicode' tt.py:19: error: Incompatible types in assignment (expression has type "unicode", variable has type "str") If I run the program (commenting out the 'reveal_type()' lines) I see: $ ./tt.py Type of 'a' is <type 'str'> Reactions are currently unavailable ·
Author: python
🌐
Ufkapano
ufkapano.github.io › scicomppy › week02 › strings.html
Strings in Python 2
# One-character Unicode strings can also be created with the # unichr() built-in function. chr(97) # return 'a', 8-bit chr(255) # return '\xff', 8-bit ord('\xff') # return 255 unichr(97) # return u'a' = u'\x61' = u'\u0061' = u'\U00000061' unichr(256) # return u'\u0100' unichr(40960) # return u'\ua000' ord(u'\ua000') # return 40960 u = unichr(40960) + u'abcd' + unichr(1972) # u'\ua000abcd\u07b4' [ord(c) for c in u] # [40960, 97, 98, 99, 100, 1972] isinstance(u, unicode) # return True isinstance(u, str) # return False # encode() returns an 8-bit string version of the Unicode string u_utf8 = u.encode('utf-8') # return '\xea\x80\x80abcd\xde\xb4' (no u'...') isinstance(u_utf8, unicode) # return False isinstance(u_utf8, str) # return True # decode() interprets the 8-bit string using the given encoding u_utf8.decode('utf-8') # return u'\ua000abcd\u07b4' u'\u0061'.encode('utf-8') # return 'a'
🌐
John A. Bachman
johnbachman.net › building-a-python-23-compatible-unicode-sandwich.html
John A. Bachman – Building a Python 2/3 compatible Unicode Sandwich
March 14, 2017 - The best solution is to use Unicode everywhere in Python 2, importing from builtins import str (as recommended above) and then using isinstance(foo, str), where you would have previously used basestring.
Find elsewhere
🌐
Python-future
python-future.org › what_else.html
What else you need to know — Python-Future documentation
>>> from __future__ import unicode_literals >>> assert isinstance('unicode string 2', str)
🌐
Gitbooks
hacktec.gitbooks.io › effective-python › content › en › Chapter1 › Item3.html
Item 3: Know the Differences Between bytes, str, and unicode · Effective Python
You want to operate on Unicode characters that have no specific encoding. You’ll often need two helper functions to convert between these two cases and to ensure that the type of input values matches your code’s expectations. In Python 3, you’ll need one method that takes a str or bytes and always returns a str. def to_str(bytes_or_str): if isinstance(bytes_or_str, bytes): value = bytes_or_str.decode(‘utf-8’) else: value = bytes_or_str return value # Instance of str
Top answer
1 of 4
405

In Python versions prior to 3.0 there are two kinds of strings "plain strings" and "unicode strings". Plain strings (str) cannot represent characters outside of the Latin alphabet (ignoring details of code pages for simplicity). Unicode strings (unicode) can represent characters from any alphabet including some fictional ones like Klingon.

So why have two kinds of strings, would it not be better to just have Unicode since that would cover all the cases? Well it is better to have only Unicode but Python was created before Unicode was the preferred method for representing strings. It takes time to transition the string type in a language with many users, in Python 3.0 it is finally the case that all strings are Unicode.

The inheritance hierarchy of Python strings pre-3.0 is:

          object
             |
             |
         basestring
            / \
           /   \
         str  unicode

'basestring' introduced in Python 2.3 can be thought of as a step in the direction of string unification as it can be used to check whether an object is an instance of str or unicode

>>> string1 = "I am a plain string"
>>> string2 = u"I am a unicode string"
>>> isinstance(string1, str)
True
>>> isinstance(string2, str)
False
>>> isinstance(string1, unicode)
False
>>> isinstance(string2, unicode)
True
>>> isinstance(string1, basestring)
True
>>> isinstance(string2, basestring)
True
2 of 4
9

All strings are basestrings, but unicode strings are not of type str. Try this instead:

>>> a=u'aaaa'
>>> print isinstance(a, basestring)
True
>>> print isinstance(a, str)
False
🌐
Readthedocs
portingguide.readthedocs.io › en › latest › strings.html
Strings — Conservative Python 3 Porting Guide 1.0 documentation
For these cases, Python 2 provided the class basestring, from which both str and unicode derived: if isinstance(value, basestring): print("It's stringy!")
🌐
Sagemath
trac.sagemath.org › ticket › 16064
#16064 (Python 3 preparation: Handle basestring (Py2) vs. str (Py3)) – Sage
This abstract type is the superclass for str and unicode. It cannot be called or instantiated, but it can be used to test whether an object is an instance of str or unicode. isinstance(obj, basestring) is equivalent to isinstance(obj, (str, unicode)). New in version 2.3. This is not available ...
🌐
Reddit
reddit.com › r/learnpython › unicode check
r/learnpython on Reddit: Unicode check
October 15, 2013 - However, you do often need to use isinstance() when dealing with strings. Specifically, strings are hard to tell apart from other iterables, so it's common (in Python 2.x) to check isinstance(mystr, basestring) to tell if you've got an actual string. basestring is the superclass for both str and unicode.
🌐
Python
bugs.python.org › issue38003
Issue 38003: Change 2to3 to replace 'basestring' with '(str,bytes)' - Python tracker
September 1, 2019 - This issue tracker has been migrated to GitHub, and is currently read-only. For more information, see the GitHub FAQs in the Python's Developer Guide · This issue has been migrated to GitHub: https://github.com/python/cpython/issues/82184
🌐
Python Reference
python-reference.readthedocs.io › en › latest › docs › functions › isinstance.html
isinstance — Python Reference (The Right Way) 0.1 documentation
>>> isinstance(u'foo', (basestring, str, unicode)) True >>> isinstance(u'foo', (basestring, str)) True >>> isinstance(u'foo', (basestring)) True >>> isinstance(u'foo', (str)) False ·
🌐
Python-future
python-future.org › compatible_idioms.html
Cheat Sheet: Writing Python 2-3 compatible code — Python-Future documentation
# Python 2 and 3: option 1 from builtins import map myiter = map(func, myoldlist) assert isinstance(myiter, iter)
🌐
GitHub
github.com › python-pillow › Pillow › issues › 304
isinstance(filename, "utf-8") · Issue #304 · python-pillow/Pillow
July 25, 2013 - This line: https://github.com/python-imaging/Pillow/blob/master/PIL/ImageFont.py#L264 causes error as described here: http://stackoverflow.com/questions/17863735/python-pillow-better-pil-encoding-check-bug Please, fix it.
Author: python-pillow
🌐
Python
python.org › search
Welcome to Python.org
The official home of the Python Programming Language
🌐
Armin Ronacher
lucumr.pocoo.org › 2013 › 7 › 2 › the-updated-guide-to-unicode
The Updated Guide to Unicode on Python | Armin Ronacher's Thoughts and Writings
July 2, 2013 - In Python 2 that problem solved itself because bytestrings were promoted to unicode strings automatically. On Python 3 that is no longer the case which makes it much harder to implement with APIs that do both. Werkzeug and Flask use the following helpers to provide (or work with) APIs that deal with both strings and bytes: def normalize_string_tuple(tup): """Ensures that all types in the tuple are either strings or bytes. """ tupiter = iter(tup) is_text = isinstance(next(tupiter, None), str) for arg in tupiter: if isinstance(arg, str) != is_text: raise TypeError('Cannot mix str and bytes arguments (got %s)' % repr(tup)) return tup def make_literal_wrapper(reference): """Given a reference string it returns a function that can be used to wrap ASCII native-string literals to coerce it to the given string type.