The problem is that Javascript strings are encoded in UTF-16, and browsers do not offer very many good tools to deal with encodings. A great resource specifically for dealing with Base64 encodings and Unicode can be found at MDN: https://developer.mozilla.org/en-US/docs/Web/API/WindowBase64/Base64_encoding_and_decoding
Their suggested solution for encoding strings without using the often cited solution involving the deprecated unescape function is:
function b64EncodeUnicode(str) {
return btoa(encodeURIComponent(str).replace(/%([0-9A-F]{2})/g, function(match, p1) {
return String.fromCharCode('0x' + p1);
}));
}
For further solutions and details I highly recommend you read the entire page.
Answer from deceze on Stack Overflowpython - How to decode text with special symbols using base64 in python3? - Stack Overflow
android - base64 decode string and encode all special characters lost - Stack Overflow
Base64 Encoding not handlling special characters
Base64decode(x).decode("utf-8") doesn't handle multiple escape characters properly
The problem is that Javascript strings are encoded in UTF-16, and browsers do not offer very many good tools to deal with encodings. A great resource specifically for dealing with Base64 encodings and Unicode can be found at MDN: https://developer.mozilla.org/en-US/docs/Web/API/WindowBase64/Base64_encoding_and_decoding
Their suggested solution for encoding strings without using the often cited solution involving the deprecated unescape function is:
function b64EncodeUnicode(str) {
return btoa(encodeURIComponent(str).replace(/%([0-9A-F]{2})/g, function(match, p1) {
return String.fromCharCode('0x' + p1);
}));
}
For further solutions and details I highly recommend you read the entire page.
The following works for me:
JavaScript:
// Traian Băsescu encodes to VHJhaWFuIELEg3Nlc2N1
var base64 = btoa(unescape(encodeURIComponent( $("#Contact_description").val() )));
PHP:
$utf8 = base64_decode($base64);
These are url-quoted strings, so url-unquoting is the correct procedure. The first step is unquote them with urllib.parse.unquote. Only after that should you attempt base64-decoding and there's no need to manually mess around with the base64 padding character =.
The website you reference ignores invalid base64 characters and also infers the padding from the length of the base64-encoded data. So you give the website MTA0MzI%3D and it throws away the % because it's not valid base64 char, then processes MTA0MzI3D and returns 104327. Base64 padding is redundant and I'm not sure why some base64 encoding standards specify to have it in there but many do.
Example:
import base64
import urllib.parse
# List of string which we are trying to decode
encoded_text_list = ['MTA0MDI0', 'MTA0MDYw', 'MTA0MDgz', 'MTA0MzI%3D']
# Iterating and decoding string using base64
for k in encoded_text_list:
url_unquoted = urllib.parse.unquote(k)
print(k, base64.b64decode(url_unquoted).decode('utf-8'))
Output
MTA0MDI0 104024
MTA0MDYw 104060
MTA0MDgz 104083
MTA0MzI%3D 10432
and 10432 is the correct output, not 104327.
The problem is % is not a valid base64 character. The decoder expects an = sign there, and instead find %3D, which happens to be the URL encoding of =. This likely means the value is url encoded somewhere upstream from your code. Depending on requirements, you have some options:
- Call
k = parse(k); see builtin parse function - Call
k = k.replace('%3D', '=')to clean up this error - Change the inputs to not be url encoded
What's likely here is that you're in a UTF-8 locale but your terminal or terminal emulator is using a charset other than UTF-8.
To get that particular error, you must have the iconv from GNU libiconv as opposed to the one from GNU libc, and that Ç can't be encoded as 0xC7 as it is in a lot of single-byte encodings that have that character or you'd get iconv: (stdin):1:10: incomplete character or shift sequence instead.
I'd bet for IBM850 aka CP850 as that's still found on Microsoft Windows (and IIRC it's even the default in the "console" of English flavours of Windows believe it or not). In that charset, Ç is encoded as byte 0x80 (0200 in octal).
To confirm whether that's the case, enter:
printf %s 'Ç' | od -An -vto1
Which if I'm right should output 200.
$ locale charmap
UTF-8
$ printf 'FILE_DATE.\200' | iconv -t utf-8
FILE_DATE.
iconv: (stdin):1:10: cannot convert
Your locale uses UTF-8, so expects characters encoded in UTF-8. That iconv -t utf-8 would be a no-op as it's encoding from the locale's charmap (which is UTF-8) to UTF-8, except it still tries to decode that input expecting it to be UTF-8 encoded. But that 0x80 byte from that IBM850-encoded Ç is invalid in UTF-8, hence the error.
Here you'd want to fix your terminal emulator's encoding to UTF-8 (if it's indeed Microsoft Windows console, then maybe https://superuser.com/q/269818 is relevant) so it matches that of the locale. Then, when you type Ç, it would send the two bytes 0xc3 0x87 instead of 0xc7.
Then you'd just do:
printf %s 'FILE_DATE.Ç' | base64
To get the base64 encoding of the UTF-8 encoding of that string.
If you cannot change the encoding of your terminal, and you know which it is (here maybe IBM850 if that's Microsoft Windows), then:
printf %s 'FILE_DATE.Ç' | iconv -f IBM850 | base64
Would again give you the base64 encoding of the UTF-8 encoding of that string.
Understand that base64 does not deal with characters. It deals with bytes. It has absolutely no interaction with UTF-8 or any other locale. It can encode arbitrary bytes (e.g. an executable ELF file, a .JPG, a .MOV, ...) into bytes that will be in the range A-Z a-z 0-9 + /. (It pads the output with = to compensate for inputs that are not a multiple of three bytes.)
Specifically, base64 encoding splits each three (consecutive) 8-bit bytes into four 6-bit values, and then encodes each 6-bit with the appropriate plain character in the list above. So the input can be any 8-bit value, and the output is guaranteed to be plain ASCII text with no extended characters. This process is 100% reversible and lossless.
All this was originally designed to transmit binary data via email and RS232 protocols which were only designed to convey plain text, and would freak out if the user data had any control characters.
Try this on your system.
$ echo FILE_DATE.Ç > foo1
$ od -t x1ac foo1
0000000 46 49 4c 45 5f 44 41 54 45 2e c3 87 0a
F I L E _ D A T E . C bel nl
F I L E _ D A T E . 303 207 \n
0000015
$ #.. Note the accented character is UTF-8 octal 303 207.
$ base64 foo1 > foo2
$ od -t x1ac foo2
0000000 52 6b 6c 4d 52 56 39 45 51 56 52 46 4c 73 4f 48
R k l M R V 9 E Q V R F L s O H
R k l M R V 9 E Q V R F L s O H
0000020 43 67 3d 3d 0a
C g = = nl
C g = = \n
0000025
$ cat foo2
RklMRV9EQVRFLsOHCg==
$ base64 -d foo2 > foo3
$ cat foo3
FILE_DATE.Ç
$ od -t x1ac foo3
0000000 46 49 4c 45 5f 44 41 54 45 2e c3 87 0a
F I L E _ D A T E . C bel nl
F I L E _ D A T E . 303 207 \n
0000015
$ iconv -t utf-8 foo1 > foo4
$ #.. No error given.
$ od -t x1ac foo4
0000000 46 49 4c 45 5f 44 41 54 45 2e c3 87 0a
F I L E _ D A T E . C bel nl
F I L E _ D A T E . 303 207 \n
0000015
$ #.. iconv makes no changes, and shows no error.
$ echo FILE_DATE.Ç | iconv -t utf-8 | base64 > foo5
$ #.. Your original pipeline, no error.
$ cmp foo2 foo5
$ #.. Shows no difference in the base64 versions.
$ cat foo5
RklMRV9EQVRFLsOHCg==
$ file foo{1..5}
foo1: UTF-8 Unicode text
foo2: ASCII text
foo3: UTF-8 Unicode text
foo4: UTF-8 Unicode text
foo5: ASCII text
$