The official node.js documentation for Buffer is the best place to check for something like this. As previously noted, Buffer currently supports these encodings: 'ascii', 'utf8', 'utf16le'/'ucs2', 'base64', 'base64url', 'latin1'/'binary', and 'hex'.
The official node.js documentation for Buffer is the best place to check for something like this. As previously noted, Buffer currently supports these encodings: 'ascii', 'utf8', 'utf16le'/'ucs2', 'base64', 'base64url', 'latin1'/'binary', and 'hex'.
As is always the way, I spent a while Googling but found nothing until after I posted the question:
http://www.w3resource.com/node.js/nodejs-buffer.php has the answer. You can use the following types in .toString() on a buffer:
asciiutf8utf16leucs2(alias ofutf16le)base64binaryhex
The correct answer is:
p.communicate(b"insert into egg values ('egg')")
Note the leading b, telling you that it's a string of bytes, not a string of unicode characters. Also, if you are reading this from a file:
value = open('thefile', 'rt').read()
p.communicate(value)
The change that to:
value = open('thefile', 'rb').read()
p.communicate(value)
Again, note the 'b'.
Now if your value is a string you get from an API that only returns strings no matter what, then you need to encode it.
p.communicate(value.encode('latin-1'))
Latin-1, because unlike ASCII it supports all 256 bytes. But that said, having binary data in unicode is asking for trouble. It's better if you can make it binary from the start.
You can convert it to bytes with encode method:
>>> "insert into egg values ('egg');".encode('ascii') # ascii is just an example
b"insert into egg values ('egg');"
Decode the bytes object to produce a string:
>>> b"abcde".decode("utf-8")
'abcde'
The above example assumes that the bytes object is in UTF-8, because it is a common encoding. However, you should use the encoding your data is actually in!
Decode the byte string and turn it in to a character (Unicode) string.
Python 3:
encoding = 'utf-8'
b'hello'.decode(encoding)
or
str(b'hello', encoding)
Python 2:
encoding = 'utf-8'
'hello'.decode(encoding)
or
unicode('hello', encoding)
The most concise I could come up with is the following, which you may be able to make more concise with a few convenience functions (or even replacing/overriding the print function):
# -*- coding=utf-8 -*-
import codecs
import os
import sys
# if you include the -*- coding line, you can use this
output = 'bar' + u'โ'
# otherwise, use this
output = 'bar' + b'\xe2\x86\x92'.decode('utf-8')
if sys.stdout.encoding == 'UTF-8':
print(output)
else:
output += os.linesep
if sys.version_info[0] >= 3:
sys.stdout.buffer.write(bytes(output.encode('utf-8')))
else:
codecs.getwriter('utf-8')(sys.stdout).write(output)
The best option is using the -*- encoding line, which allows you to use the actual character in the file. But if for some reason, you can't use the encoding line, it's still possible to accomplish without it.
This (both with and without the encoding line) works on Linux (Arch) with python 2.7.7 and 3.4.1. It also works if the terminal's encoding is not UTF-8. (On Arch Linux, I just change the encoding by using a different LANG environment variable.)
LANG=zh_CN python test.py
It also sort of works on Windows, which I tried with 2.6, 2.7, 3.3, and 3.4. By sort of, I mean I could get the 'โ' character to display only on a mintty terminal. On a cmd terminal, that character would display as 'ฮรฅร'. (There may be something simple I'm missing there.)
If you don't need to print to sys.stdout.buffer, then the following should print fine to sys.stdout. I tried it in both Python 2.7 and 3.4, and it seemed to work fine:
# -*- coding=utf-8 -*-
print("bar" + u"โ")
I think struct is what you are looking for. Anyhow you have to find out the serialized data format. It is not a C/C++ struct that is sent but (hopefilly) a well described serialized format which may be binaray. Important topics are little/big endianess, size of double, size of int and padding.
from struct import *
arg1, arg2, arg3, arg4, arg5 = unpack('!ddddi', buffer)
standard sizes would require 36 bytes
>>> from struct import *
>>> a='\x00\x00\x00\x01'*9
>>> unpack('>ddddi', a)
(2.1219957915e-314, 2.1219957915e-314, 2.1219957915e-314, 2.1219957915e-314, 1)
>>> unpack('<ddddi', a)
(7.291122046717944e-304, 7.291122046717944e-304, 7.291122046717944e-304, 7.291122046717944e-304, 16777216)
>>> unpack('!ddddi', a)
(2.1219957915e-314, 2.1219957915e-314, 2.1219957915e-314, 2.1219957915e-314, 1)
>>>
EDIT: unpacking to variables, named tuple and class
unpack directly to your variables
from struct import unpack
b = "\x00\x00\x00\x05\x00\x00\x00\x2a"
x, y = unpack('!ii', b)
print x
print y
unpack to named tuple
from struct import unpack
from collections import namedtuple
b = "\x00\x00\x00\x05\x00\x00\x00\x2a"
Point = namedtuple('Point', 'x, y')
p = Point._make(unpack("!ii", b))
print p.x
print p.y
remember: tuple attributes are read only
unpack and initalize a class
from struct import unpack
b = "\x00\x00\x00\x05\x00\x00\x00\x2a"
class Point(object):
def __init__(self, x, y):
self.x = x
self.y = y
p = Point(*unpack('!ii', b))
print p.x
print p.y
I too used struct unpacking (and had to search the character codes every time) until I learned to do the same using ctypes. I have no idea which one performs faster, but I do know which one I can translate without problems just looking at the C definition.
from ctypes import LittleEndianStructure, BigEndianStructure, c_int16, c_double, sizeof
#there are bit explicit types if you want to fix the size, like c_int8/16/32/64
class Packet(LittleEndianStructure): # replace by BigEndian as needed
_fields_ = [
("arg1", c_double),
("arg1", c_double),
("arg3", c_double),
("arg4", c_double),
("arg5", c_int16)
]
_pack_=1 # comment this or not depending on padding
mypacket = Packet.from_buffer_copy(payload)
It is for me much easier to read, and it has the advantage of providing me with an object that collects all the data and can be accessed by name
print(mypacket.arg1)
