You need to pass in a bytes object, rather than a str. The typical way to go from str (a unicode string in Python 3) to bytes is to use the .encode() method on the string and specify the encoding you wish to use.
my_bytes = my_string.encode('utf-8')
Answer from Amber on Stack OverflowYou need to pass in a bytes object, rather than a str. The typical way to go from str (a unicode string in Python 3) to bytes is to use the .encode() method on the string and specify the encoding you wish to use.
my_bytes = my_string.encode('utf-8')
Just call fileinput.input(...,mode='rb') to open files in binary mode. Such files produce binary strings instead of Unicode strings as files opened in text mode do.
It allows you to skip an unnecessary (implicit) decoding of bytes read from disk followed by immediate encoding them back to bytes using .encode() before passing them to md5().
Your categorical variable has two levels, so there is no actual difference between dummy-coding vs. simply entering the variable into the analysis. That is, to dummy code you would create one new variable with two values but your original variable is already one variable with two values. Dummy-coding is important for variables with more than two possible values. So, in this case the computer won't consider Pave > Grvl.
But if you have more than two variables then you should use dummy variables.
For your data, you can use pandas.get_dummies() or sklearn's one hot encoder to achieve your result.
- How to encode?
sklearn.preprocessing provides various classes for this purpose, LabelBinarizer is one of them.
- Wouldn't the computer interpret this as Pave > Grvl?
Consider an example, where people prefer Paved house in comparison to Graveled. Then their is a relationship between values and hence it should be treated as something like you have mentioned, other wise it should be independent values(refer next answer).
- How does the computer differentiate between binary and integer encoding?
As I mentioned above, if the categorical values have some relationship(as mentioned above), then in such a case it should be integer values(0,1,2 and so on), otherwise it should be binary. Binary representation will help us in presenting as an independent value to ML model (however it doesn't make much sense in this case as you just have 2 values). But consider an example where a feature have more than 2 categorical values. If they all are independent then it should be represented as binary value i.e in the form of OneHotEncoding(refer sklearn.preprocessing classes).
I need to read file in binary format-->
with open("tex.pdf", mode='br') as file:
fileContent = file.read()
for i in fileContent:
print(i,end=" ")so this provide decimal integers which i think it in ascii format but ascii values only have 0-127
and this output display integers greater that 127 (like 225 108 180 193)
Can someone tell me what encoding/machanism used?