Remove the regular space that you have first in the pattern:

 str = str.replace(/[\u00A0\u1680โ€‹\u180e\u2000-\u2009\u200aโ€‹\u200bโ€‹\u202f\u205fโ€‹\u3000]/g,'');
Answer from Guffa on Stack Overflow
Discussions

regex - Replace unicode matches in javascript - Stack Overflow
I would like to replace the matched words/characters found in a search, for example if I search for a and I get the result รกdรกm, I would like to highlight the รก's. Something like: "รกdรกm".replace(/... More on stackoverflow.com
๐ŸŒ stackoverflow.com
Javascript Replace Unicode Characters - JavaScript - SitePoint Forums | Web Development & Design Community
Hi there, Iโ€™ve been trying to write a JavaScript method that replaces instances of RTF double quotes and single quotes or apostrophes. It seems some of our users are pasting RTF from MS Word and thatโ€™s causing some problems in Oracle. The RTF characters look like this: โ€œ,โ€, โ€˜,โ€™ ... More on sitepoint.com
๐ŸŒ sitepoint.com
0
January 19, 2010
How to replace Unicode characters in the following scenario using javascript?
How to replace Unicode characters in the following scenario using javascript? Using javascript I want to replace Unicode characters with a wrapper according to their style. If possible include a range of Unicode characters([a-z]) in regex for styles other than regular. input = abc๐‘ข๐‘ฃ๐‘ค... More on javascript.tutorialink.com
๐ŸŒ javascript.tutorialink.com
1
node.js - remove/replace Unicode characters javascript - Stack Overflow
It shows as "(-500" but when I copy it to the notepad it comes with the a invisible Unicode, how can I get the innerHTML just as simple text "-500" but without the Unicode and without the "(". More on stackoverflow.com
๐ŸŒ stackoverflow.com
๐ŸŒ
Reddit
reddit.com โ€บ r/learnjavascript โ€บ replace unicode in a string with its symbol...
r/learnjavascript on Reddit: Replace Unicode in a string with its symbol...
August 9, 2018 -

Hey, so I should start by saying, I donโ€™t know if the title is actually what Iโ€™m trying to do... hence my problem.

Iโ€™m getting data from a wordpress api and the titles have, what I believe are Unicode or ASCII codes for symbols in them, โ€˜&โ€™ and โ€˜-โ€˜ for example are a string of numbers, within a string that makes up the title.

Iโ€™m not familiar enough with the wordpress api to know how to get it without that if thatโ€™s even possible and Iโ€™ve tried many different ways to just change these codes within the string into their symbols... please... someone put me out my misery and tell me how to do this?

I donโ€™t seem to be able to google the correct thing, trust me, Iโ€™ve tried! I think itโ€™s because Iโ€™m not phrasing the question correctly, but I donโ€™t know how to phrase it :S

Cheers for any help ๐Ÿ‘๐Ÿผ

Edit - Iโ€™m using React Native. Would decodeURIComponent(); be an option? Iโ€™m not looking to work with a uri but it seems to have the effect Iโ€™m after from looking at some google results :S

Top answer
1 of 5
12

Adapted from Semplice, following link from here.

[^\x00-\x80] matches any character not in the ASCII range.
Note that some of the characters may not be encoded correctly from the copy and paste.

var latin_map = {"ร":"A","ฤ‚":"A","แบฎ":"A","แบถ":"A","แบฐ":"A","แบฒ":"A","แบด":"A","ว":"A","ร‚":"A","แบค":"A","แบฌ":"A","แบฆ":"A","แบจ":"A","แบช":"A","ร„":"A","วž":"A","ศฆ":"A","ว ":"A","แบ ":"A","ศ€":"A","ร€":"A","แบข":"A","ศ‚":"A","ฤ€":"A","ฤ„":"A","ร…":"A","วบ":"A","แธ€":"A","ศบ":"A","รƒ":"A","๊œฒ":"AA","ร†":"AE","วผ":"AE","วข":"AE","๊œด":"AO","๊œถ":"AU","๊œธ":"AV","๊œบ":"AV","๊œผ":"AY","แธ‚":"B","แธ„":"B","ฦ":"B","แธ†":"B","ษƒ":"B","ฦ‚":"B","ฤ†":"C","ฤŒ":"C","ร‡":"C","แธˆ":"C","ฤˆ":"C","ฤŠ":"C","ฦ‡":"C","ศป":"C","ฤŽ":"D","แธ":"D","แธ’":"D","แธŠ":"D","แธŒ":"D","ฦŠ":"D","แธŽ":"D","วฒ":"D","ว…":"D","ฤ":"D","ฦ‹":"D","วฑ":"DZ","ว„":"DZ","ร‰":"E","ฤ”":"E","ฤš":"E","ศจ":"E","แธœ":"E","รŠ":"E","แบพ":"E","แป†":"E","แป€":"E","แป‚":"E","แป„":"E","แธ˜":"E","ร‹":"E","ฤ–":"E","แบธ":"E","ศ„":"E","รˆ":"E","แบบ":"E","ศ†":"E","ฤ’":"E","แธ–":"E","แธ”":"E","ฤ˜":"E","ษ†":"E","แบผ":"E","แธš":"E","๊ช":"ET","แธž":"F","ฦ‘":"F","วด":"G","ฤž":"G","วฆ":"G","ฤข":"G","ฤœ":"G","ฤ ":"G","ฦ“":"G","แธ ":"G","วค":"G","แธช":"H","ศž":"H","แธจ":"H","ฤค":"H","โฑง":"H","แธฆ":"H","แธข":"H","แธค":"H","ฤฆ":"H","ร":"I","ฤฌ":"I","ว":"I","รŽ":"I","ร":"I","แธฎ":"I","ฤฐ":"I","แปŠ":"I","ศˆ":"I","รŒ":"I","แปˆ":"I","ศŠ":"I","ฤช":"I","ฤฎ":"I","ฦ—":"I","ฤจ":"I","แธฌ":"I","๊น":"D","๊ป":"F","๊ฝ":"G","๊ž‚":"R","๊ž„":"S","๊ž†":"T","๊ฌ":"IS","ฤด":"J","ษˆ":"J","แธฐ":"K","วจ":"K","ฤถ":"K","โฑฉ":"K","๊‚":"K","แธฒ":"K","ฦ˜":"K","แธด":"K","๊€":"K","๊„":"K","ฤน":"L","ศฝ":"L","ฤฝ":"L","ฤป":"L","แธผ":"L","แธถ":"L","แธธ":"L","โฑ ":"L","๊ˆ":"L","แธบ":"L","ฤฟ":"L","โฑข":"L","วˆ":"L","ล":"L","ว‡":"LJ","แธพ":"M","แน€":"M","แน‚":"M","โฑฎ":"M","ลƒ":"N","ล‡":"N","ล…":"N","แนŠ":"N","แน„":"N","แน†":"N","วธ":"N","ฦ":"N","แนˆ":"N","ศ ":"N","ว‹":"N","ร‘":"N","วŠ":"NJ","ร“":"O","ลŽ":"O","ว‘":"O","ร”":"O","แป":"O","แป˜":"O","แป’":"O","แป”":"O","แป–":"O","ร–":"O","ศช":"O","ศฎ":"O","ศฐ":"O","แปŒ":"O","ล":"O","ศŒ":"O","ร’":"O","แปŽ":"O","ฦ ":"O","แปš":"O","แปข":"O","แปœ":"O","แปž":"O","แป ":"O","ศŽ":"O","๊Š":"O","๊Œ":"O","ลŒ":"O","แน’":"O","แน":"O","ฦŸ":"O","วช":"O","วฌ":"O","ร˜":"O","วพ":"O","ร•":"O","แนŒ":"O","แนŽ":"O","ศฌ":"O","ฦข":"OI","๊Ž":"OO","ฦ":"E","ฦ†":"O","ศข":"OU","แน”":"P","แน–":"P","๊’":"P","ฦค":"P","๊”":"P","โฑฃ":"P","๊":"P","๊˜":"Q","๊–":"Q","ล”":"R","ล˜":"R","ล–":"R","แน˜":"R","แนš":"R","แนœ":"R","ศ":"R","ศ’":"R","แนž":"R","ษŒ":"R","โฑค":"R","๊œพ":"C","ฦŽ":"E","ลš":"S","แนค":"S","ล ":"S","แนฆ":"S","ลž":"S","ลœ":"S","ศ˜":"S","แน ":"S","แนข":"S","แนจ":"S","ลค":"T","ลข":"T","แนฐ":"T","ศš":"T","ศพ":"T","แนช":"T","แนฌ":"T","ฦฌ":"T","แนฎ":"T","ฦฎ":"T","ลฆ":"T","โฑฏ":"A","๊ž€":"L","ฦœ":"M","ษ…":"V","๊œจ":"TZ","รš":"U","ลฌ":"U","ว“":"U","ร›":"U","แนถ":"U","รœ":"U","ว—":"U","ว™":"U","ว›":"U","ว•":"U","แนฒ":"U","แปค":"U","ลฐ":"U","ศ”":"U","ร™":"U","แปฆ":"U","ฦฏ":"U","แปจ":"U","แปฐ":"U","แปช":"U","แปฌ":"U","แปฎ":"U","ศ–":"U","ลช":"U","แนบ":"U","ลฒ":"U","ลฎ":"U","ลจ":"U","แนธ":"U","แนด":"U","๊ž":"V","แนพ":"V","ฦฒ":"V","แนผ":"V","๊ ":"VY","แบ‚":"W","ลด":"W","แบ„":"W","แบ†":"W","แบˆ":"W","แบ€":"W","โฑฒ":"W","แบŒ":"X","แบŠ":"X","ร":"Y","ลถ":"Y","ลธ":"Y","แบŽ":"Y","แปด":"Y","แปฒ":"Y","ฦณ":"Y","แปถ":"Y","แปพ":"Y","ศฒ":"Y","ษŽ":"Y","แปธ":"Y","ลน":"Z","ลฝ":"Z","แบ":"Z","โฑซ":"Z","ลป":"Z","แบ’":"Z","ศค":"Z","แบ”":"Z","ฦต":"Z","ฤฒ":"IJ","ล’":"OE","แด€":"A","แด":"AE","ส™":"B","แดƒ":"B","แด„":"C","แด…":"D","แด‡":"E","๊œฐ":"F","ษข":"G","ส›":"G","สœ":"H","ษช":"I","ส":"R","แดŠ":"J","แด‹":"K","สŸ":"L","แดŒ":"L","แด":"M","ษด":"N","แด":"O","ษถ":"OE","แด":"O","แด•":"OU","แด˜":"P","ส€":"R","แดŽ":"N","แด™":"R","๊œฑ":"S","แด›":"T","โฑป":"E","แดš":"R","แดœ":"U","แด ":"V","แดก":"W","ส":"Y","แดข":"Z","รก":"a","ฤƒ":"a","แบฏ":"a","แบท":"a","แบฑ":"a","แบณ":"a","แบต":"a","วŽ":"a","รข":"a","แบฅ":"a","แบญ":"a","แบง":"a","แบฉ":"a","แบซ":"a","รค":"a","วŸ":"a","ศง":"a","วก":"a","แบก":"a","ศ":"a","ร ":"a","แบฃ":"a","ศƒ":"a","ฤ":"a","ฤ…":"a","แถ":"a","แบš":"a","รฅ":"a","วป":"a","แธ":"a","โฑฅ":"a","รฃ":"a","๊œณ":"aa","รฆ":"ae","วฝ":"ae","วฃ":"ae","๊œต":"ao","๊œท":"au","๊œน":"av","๊œป":"av","๊œฝ":"ay","แธƒ":"b","แธ…":"b","ษ“":"b","แธ‡":"b","แตฌ":"b","แถ€":"b","ฦ€":"b","ฦƒ":"b","ษต":"o","ฤ‡":"c","ฤ":"c","รง":"c","แธ‰":"c","ฤ‰":"c","ษ•":"c","ฤ‹":"c","ฦˆ":"c","ศผ":"c","ฤ":"d","แธ‘":"d","แธ“":"d","ศก":"d","แธ‹":"d","แธ":"d","ษ—":"d","แถ‘":"d","แธ":"d","แตญ":"d","แถ":"d","ฤ‘":"d","ษ–":"d","ฦŒ":"d","ฤฑ":"i","ศท":"j","ษŸ":"j","ส„":"j","วณ":"dz","ว†":"dz","รฉ":"e","ฤ•":"e","ฤ›":"e","ศฉ":"e","แธ":"e","รช":"e","แบฟ":"e","แป‡":"e","แป":"e","แปƒ":"e","แป…":"e","แธ™":"e","รซ":"e","ฤ—":"e","แบน":"e","ศ…":"e","รจ":"e","แบป":"e","ศ‡":"e","ฤ“":"e","แธ—":"e","แธ•":"e","โฑธ":"e","ฤ™":"e","แถ’":"e","ษ‡":"e","แบฝ":"e","แธ›":"e","๊ซ":"et","แธŸ":"f","ฦ’":"f","แตฎ":"f","แถ‚":"f","วต":"g","ฤŸ":"g","วง":"g","ฤฃ":"g","ฤ":"g","ฤก":"g","ษ ":"g","แธก":"g","แถƒ":"g","วฅ":"g","แธซ":"h","ศŸ":"h","แธฉ":"h","ฤฅ":"h","โฑจ":"h","แธง":"h","แธฃ":"h","แธฅ":"h","ษฆ":"h","แบ–":"h","ฤง":"h","ฦ•":"hv","รญ":"i","ฤญ":"i","ว":"i","รฎ":"i","รฏ":"i","แธฏ":"i","แป‹":"i","ศ‰":"i","รฌ":"i","แป‰":"i","ศ‹":"i","ฤซ":"i","ฤฏ":"i","แถ–":"i","ษจ":"i","ฤฉ":"i","แธญ":"i","๊บ":"d","๊ผ":"f","แตน":"g","๊žƒ":"r","๊ž…":"s","๊ž‡":"t","๊ญ":"is","วฐ":"j","ฤต":"j","ส":"j","ษ‰":"j","แธฑ":"k","วฉ":"k","ฤท":"k","โฑช":"k","๊ƒ":"k","แธณ":"k","ฦ™":"k","แธต":"k","แถ„":"k","๊":"k","๊…":"k","ฤบ":"l","ฦš":"l","ษฌ":"l","ฤพ":"l","ฤผ":"l","แธฝ":"l","ศด":"l","แธท":"l","แธน":"l","โฑก":"l","๊‰":"l","แธป":"l","ล€":"l","ษซ":"l","แถ…":"l","ษญ":"l","ล‚":"l","ว‰":"lj","ลฟ":"s","แบœ":"s","แบ›":"s","แบ":"s","แธฟ":"m","แน":"m","แนƒ":"m","ษฑ":"m","แตฏ":"m","แถ†":"m","ล„":"n","ลˆ":"n","ล†":"n","แน‹":"n","ศต":"n","แน…":"n","แน‡":"n","วน":"n","ษฒ":"n","แน‰":"n","ฦž":"n","แตฐ":"n","แถ‡":"n","ษณ":"n","รฑ":"n","วŒ":"nj","รณ":"o","ล":"o","ว’":"o","รด":"o","แป‘":"o","แป™":"o","แป“":"o","แป•":"o","แป—":"o","รถ":"o","ศซ":"o","ศฏ":"o","ศฑ":"o","แป":"o","ล‘":"o","ศ":"o","รฒ":"o","แป":"o","ฦก":"o","แป›":"o","แปฃ":"o","แป":"o","แปŸ":"o","แปก":"o","ศ":"o","๊‹":"o","๊":"o","โฑบ":"o","ล":"o","แน“":"o","แน‘":"o","วซ":"o","วญ":"o","รธ":"o","วฟ":"o","รต":"o","แน":"o","แน":"o","ศญ":"o","ฦฃ":"oi","๊":"oo","ษ›":"e","แถ“":"e","ษ”":"o","แถ—":"o","ศฃ":"ou","แน•":"p","แน—":"p","๊“":"p","ฦฅ":"p","แตฑ":"p","แถˆ":"p","๊•":"p","แตฝ":"p","๊‘":"p","๊™":"q","ส ":"q","ษ‹":"q","๊—":"q","ล•":"r","ล™":"r","ล—":"r","แน™":"r","แน›":"r","แน":"r","ศ‘":"r","ษพ":"r","แตณ":"r","ศ“":"r","แนŸ":"r","ษผ":"r","แตฒ":"r","แถ‰":"r","ษ":"r","ษฝ":"r","โ†„":"c","๊œฟ":"c","ษ˜":"e","ษฟ":"r","ล›":"s","แนฅ":"s","ลก":"s","แนง":"s","ลŸ":"s","ล":"s","ศ™":"s","แนก":"s","แนฃ":"s","แนฉ":"s","ส‚":"s","แตด":"s","แถŠ":"s","ศฟ":"s","ษก":"g","แด‘":"o","แด“":"o","แด":"u","ลฅ":"t","ลฃ":"t","แนฑ":"t","ศ›":"t","ศถ":"t","แบ—":"t","โฑฆ":"t","แนซ":"t","แนญ":"t","ฦญ":"t","แนฏ":"t","แตต":"t","ฦซ":"t","สˆ":"t","ลง":"t","แตบ":"th","ษ":"a","แด‚":"ae","ว":"e","แตท":"g","ษฅ":"h","สฎ":"h","สฏ":"h","แด‰":"i","สž":"k","๊ž":"l","ษฏ":"m","ษฐ":"m","แด”":"oe","ษน":"r","ษป":"r","ษบ":"r","โฑน":"r","ส‡":"t","สŒ":"v","ส":"w","สŽ":"y","๊œฉ":"tz","รบ":"u","ลญ":"u","ว”":"u","รป":"u","แนท":"u","รผ":"u","ว˜":"u","วš":"u","วœ":"u","ว–":"u","แนณ":"u","แปฅ":"u","ลฑ":"u","ศ•":"u","รน":"u","แปง":"u","ฦฐ":"u","แปฉ":"u","แปฑ":"u","แปซ":"u","แปญ":"u","แปฏ":"u","ศ—":"u","ลซ":"u","แนป":"u","ลณ":"u","แถ™":"u","ลฏ":"u","ลฉ":"u","แนน":"u","แนต":"u","แตซ":"ue","๊ธ":"um","โฑด":"v","๊Ÿ":"v","แนฟ":"v","ส‹":"v","แถŒ":"v","โฑฑ":"v","แนฝ":"v","๊ก":"vy","แบƒ":"w","ลต":"w","แบ…":"w","แบ‡":"w","แบ‰":"w","แบ":"w","โฑณ":"w","แบ˜":"w","แบ":"x","แบ‹":"x","แถ":"x","รฝ":"y","ลท":"y","รฟ":"y","แบ":"y","แปต":"y","แปณ":"y","ฦด":"y","แปท":"y","แปฟ":"y","ศณ":"y","แบ™":"y","ษ":"y","แปน":"y","ลบ":"z","ลพ":"z","แบ‘":"z","ส‘":"z","โฑฌ":"z","ลผ":"z","แบ“":"z","ศฅ":"z","แบ•":"z","แตถ":"z","แถŽ":"z","ส":"z","ฦถ":"z","ษ€":"z","๏ฌ€":"ff","๏ฌƒ":"ffi","๏ฌ„":"ffl","๏ฌ":"fi","๏ฌ‚":"fl","ฤณ":"ij","ล“":"oe","๏ฌ†":"st","โ‚":"a","โ‚‘":"e","แตข":"i","โฑผ":"j","โ‚’":"o","แตฃ":"r","แตค":"u","แตฅ":"v","โ‚“":"x"};

function embolden( str, chr ){
    return str.replace( /[^\x00-\x80]/g,
        function (a) { 
            return chr == latin_map[a] ? '<b>' + a + '</b>' : a;
        } 
    );
}

embolden( 'รกdรกm', 'a' );    // "<b>รก</b>d<b>รก</b>m"
2 of 5
3

I've tried this code, see if it's what you're looking for:

'รกdรกm'.replace(/./g,function(char){
    switch(char.toLowerCase()){
        case 'รก':
        case 'ร ':
        case 'รข':
        case 'รฃ':
            return '*';
        break;
    }
    return char;
});

EDIT:

To replace all chars that don't belong to the ASCII table, just check if the char has a char code up to 127, since the ASCII table char codes are defined between 0 and 127 (notice that รก doesn't belong to the Unicode table, but to the Extended ASCII table, that comes from 0 up to 255):

'รกdรกm'.replace(/./g,function(char){
    return char.charCodeAt(0)<=127 ? char : '<b>' + char + '</b>';
});
๐ŸŒ
SitePoint
sitepoint.com โ€บ javascript
Javascript Replace Unicode Characters - JavaScript - SitePoint Forums | Web Development & Design Community
January 19, 2010 - Hi there, Iโ€™ve been trying to write a JavaScript method that replaces instances of RTF double quotes and single quotes or apostrophes. It seems some of our users are pasting RTF from MS Word and thatโ€™s causing some probโ€ฆ
๐ŸŒ
Metring
ricardometring.com โ€บ javascript-replace-special-characters
JavaScript: Replacing Special Characters - The Clean Way
const str = 'รร‰รร“รšรกรฉรญรณรบรขรชรฎรดรปร รจรฌรฒรนร‡รง'; const parsed = str.normalize('NFD').replace(/[\u0300-\u036f]/g, ''); console.log(parsed); ... The normalize method was introduced in the ES6 version of JavaScript in 2015.
๐ŸŒ
WebDeveloper.com
webdeveloper.com โ€บ community โ€บ 223152-javascript-replace-unicode-characters
Javascript Replace Unicode Characters
Normally using unformatted text str.replace(โ€˜โ€œโ€™, โ€˜โ€โ€˜) would work, but thatโ€™s not working. ... nnn matches an ASCII character (octal) dnn matches an ASCII character (hex) unnnn matches a Unicode character.
Find elsewhere
๐ŸŒ
Stack Overflow
stackoverflow.com โ€บ questions โ€บ 65820649 โ€บ how-to-replace-unicode-characters-in-the-following-scenario-using-javascript
regex - How to replace Unicode characters in the following scenario using javascript? - Stack Overflow
If possible include a range of Unicode characters([a-z]) in regex for styles other than regular. input = abc๐‘ข๐‘ฃ๐‘ค๐‘ฅ๐š๐›๐œ๐๐’‡๐’ˆ๐’‰๐’Š ยท expected ouput = <span class="regular>abc</span><i> ๐‘ข๐‘ฃ๐‘ค๐‘ฅ</i><b>๐š๐›๐œ๐</b><b><i>๐’‡๐’ˆ๐’‰๐’Š</i></b> Copytext = 'abc๐‘ข๐‘ฃ๐‘ค๐‘ฅ๐š๐›๐œ๐๐’‡๐’ˆ๐’‰๐’Š'; text = text.replace(/([a-z]+)/,'<span class="regular">$1</span>'); text = text.replace(/([๐‘Ž๐‘๐‘๐‘‘๐‘’๐‘“๐‘”โ„Ž๐‘–๐‘—๐‘˜๐‘™๐‘š๐‘›๐‘œ๐‘๐‘ž๐‘Ÿ๐‘ ๐‘ก๐‘ข๐‘ฃ๐‘ค๐‘ฅ๐‘ฆ๐‘ง]+)/,'<i>$1</i>'); text = text.replace(/([๐š๐›๐œ๐๐ž๐Ÿ๐ ๐ก๐ข๐ฃ๐ค๐ฅ๐ฆ๐ง๐จ๐ฉ๐ช๐ซ๐ฌ๐ญ๐ฎ๐ฏ๐ฐ๐ฑ๐ฒ๐ณ]+)/,'<b>$1</b>'); text = text.replace(/([๐’‚๐’ƒ๐’„๐’…๐’†๐’‡๐’ˆ๐’‰๐’Š๐’‹๐’Œ๐’๐’Ž๐’๐’๐’‘๐’’๐’“๐’”๐’•๐’–๐’—๐’˜๐’™๐’š๐’›]+)/,'<b><i>$1</i></b>');
๐ŸŒ
YouTube
youtube.com โ€บ watch
How to Remove or Replace Unicode Characters in Text with JavaScript - YouTube
Learn how to effectively `remove` or `replace` unwanted Unicode characters from your text using JavaScript with this easy-to-follow guide.---This video is ba...
Published ย  September 25, 2025
Views ย  4
๐ŸŒ
tanaike
tanaikech.github.io โ€บ home โ€บ "replacing u+00a0 with u+0020 as unicode using google apps script"
Replacing U+00A0 with U+0020 as Unicode using Google Apps Script | tanaike - Google Apps Script, Gemini API, and Developer Tips
January 23, 2023 - GasTips / jstips ยท Google Apps Script / Javascript ยท Gists ยท This is a sample script for checking and replacing a character of U+00A0 (no-break space) with U+0020 (space) as Unicode using Google Apps Script. When Iโ€™m seeing the questions on Stackoverflow, I sometimes saw the situation that the script doesnโ€™t work while the script is correct.
๐ŸŒ
LearnByExample
learnbyexample.github.io โ€บ learn_js_regexp โ€บ unicode.html
Unicode - Understanding JavaScript RegExp
Similar to escape sequence character sets, the \p{} construct offers various predefined sets to work with Unicode.
๐ŸŒ
Stack Overflow
stackoverflow.com โ€บ questions โ€บ 39851985
regex js: Replace character by unicode
The thing is I don't manage to replace something by a unicode character. var text = $(this).text(); $(this).text(text.replace(/.\?/g, '\u202F\?')); <script src="https://ajax.googleapis.com/ajax/libs/jquery/2.1.1/jquery.min.js"></script> <p>Contrary to popular belief, Lorem Ipsum is not simply; random text !
Top answer
1 of 2
11

RegEx Improvements

The regex can be shortened by using case-insensitive match with i flag. We can remove the characters which are added as both lowercase and uppercase in the regex.

After removing lowercase characters regex will be as below

\?ร•รŒ_|ล |ลฝ|ร€|ร|ร‚|รƒ|ร„|ร…|ร†|ร‡|รˆ|ร‰|รŠ|ร‹|รŒ|ร|รŽ|ร|ร‘|ร’|ร“|ร”|ร•|ร–|ร˜|ร™|รš|ร›|รœ|ร|รž|รŸ|รฐ|รฟ|_ล’โ€š|__|_

Here's live demo of regex

The regex can be further improved by using character class which will make the matches faster than OR conditions

\?ร•รŒ_|_ล’โ€š|[ล ลฝร€รร‚รƒร„ร…ร†ร‡รˆร‰รŠร‹รŒรรŽรร‘ร’ร“ร”ร•ร–ร˜ร™รšร›รœรรžรŸรฐรฟ_]+

Adding + quantifier also has positive effect on the number of steps taken to match characters when the characters in the character class are consecutive/adjacent to each other.

Here's the demo on RegEx101, without + quantifierScreenshot and with + quantifierscreenshot applied on the same data. Note that in these demos, PHP is selected as the steps taken to match is not shown for JavaScript. Also, the regex is different, it also contains lowercase counterparts of those special characters as i flag is not working with PHP and don't want to apply u(Unicode) flag as it is not supported in JavaScript.

These demos are created only to show difference when + is applied on character class. The effect should be similar in JavaScript.

Note that the __(two underscores) are redundant as _ is already added in character class and with g flag it'll remove all occurrences.

Method Chaining

As replace returns a string, any other string method can be called on it. Multiple calls to replace can be chained.

str.replace(someRegexOrString, someString)
    .replace(someOtherRegexOrString, someOtherString);

This is equivalent to

var temp = str.replace(someRegexOrString, someString);
var result = temp.replace(someOtherRegexOrString, someOtherString);

Replacing HTML

jQuery html() accepts a function which will receive the current innerHTML of the element on which the method is called as parameter and replaces the returned content to the element.

The code can be written as

$('.rte').html(function(index, currentHTML) {
    return doSomeOperationOn(currentHTML);
});

Complete Code

With above changes, the code will be

$(document).ready(function() {
    var regex = /\?ร•รŒ_|_ล’โ€š|[ล ลฝร€รร‚รƒร„ร…ร†ร‡รˆร‰รŠร‹รŒรรŽรร‘ร’ร“ร”ร•ร–ร˜ร™รšร›รœรรžรŸรฐรฟ_]+/gi;

    $('.rte').html(function(i, oldHTML) {
        return oldHTML.replace(regex, ' ')
            .replace(/[^\x00-\x7F]|\?/g, '');
    });
});

$(document).ready(function() { is more readable than $(function() {. So, you may also consider using more expressive form.

2 of 2
2

The code may be correct in itself, but it does the wrong thing.

If by strange you mean unknown to someone who only knows English, that's no excuse for removing any letters you don't know. Would you really want to look at street signs for Cafs (which were legitimate Cafรฉs before)?

If you get strange character sequences like รƒยถ, that's an encoding problem and you need to fix it properly instead of hiding it.

If you really have to keep your code, at least be honest and replace each unknown character with a question mark or the Unicode replacement character so that it is clearly visible that something unexpected happened here.

๐ŸŒ
Dmitri Pavlutin
dmitripavlutin.com โ€บ what-every-javascript-developer-should-know-about-unicode
What every JavaScript developer should know about Unicode
November 15, 2021 - Unicode in JavaScript: basic concepts, escape sequences, normalization, surrogate pairs, combining marks and how to avoid pitfalls
๐ŸŒ
GitHub
gist.github.com โ€บ mathiasbynens โ€บ 1243213
Escape all characters in a string using both Unicode and hexadecimal escape sequences ยท GitHub
Demo combining Unicode escapes with hexadecimal escapes, returning the smallest possible string: http://mothereff.in/js-escapes Demo using Unicode escapes only: http://jsfiddle.net/mathias/BjwyC/ ... function unicodeEscape(str) { return str.replace(/[\s\S]/g, function (escape) { return '\\u' + ('0000' + escape.charCodeAt().toString(16)).slice(-4); }); }