You can use .replace. For example:
>>> df = pd.DataFrame({'col2': {0: 'a', 1: 2, 2: np.nan}, 'col1': {0: 'w', 1: 1, 2: 2}})
>>> di = {1: "A", 2: "B"}
>>> df
col1 col2
0 w a
1 1 2
2 2 NaN
>>> df.replace({"col1": di})
col1 col2
0 w a
1 A 2
2 B NaN
or directly on the Series, i.e. df["col1"].replace(di, inplace=True).
python - Remap values in pandas column with a dict, preserve NaNs - Stack Overflow
python - using dict to replace values in a pandas column - Stack Overflow
python - Use dictionary to replace a string within a string in Pandas columns - Stack Overflow
Replacing words in Pandas series using a dictionary
Hi there,
I'm trying to replace values in a dataframe column with others using a dictionary and replace, but when I try the code below, the replaced values look completely wrong!
data_humidity = data_season.copy()
replace_vals={'30%-70%':50,'<30%':15,'>70%':85,'NaN':'NaN'}
data_humidity['Humidity(%)'] = data_humidity.replace({'Humidity(%): replace_vals'})
data_humidity['Humidity(%)']
Can someone point me to where I'm going wrong?
Many thanks
You can use .replace. For example:
>>> df = pd.DataFrame({'col2': {0: 'a', 1: 2, 2: np.nan}, 'col1': {0: 'w', 1: 1, 2: 2}})
>>> di = {1: "A", 2: "B"}
>>> df
col1 col2
0 w a
1 1 2
2 2 NaN
>>> df.replace({"col1": di})
col1 col2
0 w a
1 A 2
2 B NaN
or directly on the Series, i.e. df["col1"].replace(di, inplace=True).
map can be much faster than replace
If your dictionary has more than a couple of keys, using map can be much faster than replace. There are two versions of this approach, depending on whether your dictionary exhaustively maps all possible values (and also whether you want non-matches to keep their values or be converted to NaNs):
Exhaustive Mapping
In this case, the form is very simple:
df['col1'].map(di) # note: if the dictionary does not exhaustively map all
# entries then non-matched entries are changed to NaNs
Although map most commonly takes a function as its argument, it can alternatively take a dictionary or series: Documentation for Pandas.series.map
Non-Exhaustive Mapping
If you have a non-exhaustive mapping and wish to retain the existing variables for non-matches, you can add fillna:
df['col1'].map(di).fillna(df['col1'])
as in @jpp's answer here: Replace values in a pandas series via dictionary efficiently
Benchmarks
Using the following data with pandas version 0.23.1:
di = {1: "A", 2: "B", 3: "C", 4: "D", 5: "E", 6: "F", 7: "G", 8: "H" }
df = pd.DataFrame({ 'col1': np.random.choice( range(1,9), 100000 ) })
and testing with %timeit, it appears that map is approximately 10x faster than replace.
Note that your speedup with map will vary with your data. The largest speedup appears to be with large dictionaries and exhaustive replaces. See @jpp answer (linked above) for more extensive benchmarks and discussion.
Try with map:
import pandas as pd
df = pd.DataFrame({
'column_1': ['Gm', 'Boi', 'Xv', 'Gm', 'Zl', 'Oe'],
'column_2': ['tri', 'bi', 'tri', 'bi', 'uni', 'uni']
})
dict_a = {'Gm': 'tri', 'Boi': 'bi', 'Xv': 'uni', 'Zl': 'uni', 'Oe': 'bi'}
df['column_2'] = df['column_1'].map(dict_a)
print(df)
df:
column_1 column_2
0 Gm tri
1 Boi bi
2 Xv uni
3 Gm tri
4 Zl uni
5 Oe bi
You could try this (however mapping would likely be preferred):
import pandas as pd #using version 1.1.3
# build the df as described
occurrences=['Gm','Boi','Xv','Gm','Zl','Oe']
df = pd.DataFrame(occurrences,columns=['col1'])
df['col2'] = pd.Series(['tri','bi','tri','bi','uni','uni'] )
#set dict
dict_a = {'Gm':'tri', 'Boi':'bi', 'Xv':'uni', 'Zl':'uni', 'Oe':'bi'}
#replace col2 values
df['col2']=df.replace({'col1': dict_a})
#review results
df.head()
dict:
- For a DataFrame nested dictionaries, e.g., ``{'a': {'b': np.nan}}``, are read as follows: look in column 'a' for the value 'b' and replace it with NaN. The `value` parameter should be ``None`` to use a nested dict in this way. You can nest regular expressions as well. Note that column names (the top-level dictionary keys in a nested dictionary) **cannot** be regular expressions.
You can create dictionary and then replace:
ids = {'Id':['NYC','LA','UK'],
'City':['New York City','Los Angeles','United Kingdom']}
ids = dict(zip(ids['Id'], ids['City']))
print (ids)
{'UK': 'United Kingdom', 'LA': 'Los Angeles', 'NYC': 'New York City'}
df['commentTest'] = df['Comment'].replace(ids, regex=True)
print (df)
Categories Comment Type \
0 animal The NYC tree is very big tree
1 plant The cat from the UK is small dog
2 object The rock was found in LA. rock
commentTest
0 The New York City tree is very big
1 The cat from the United Kingdom is small
2 The rock was found in Los Angeles.
It's actually much faster to use str.replace() than replace(), even though str.replace() requires a loop:
ids = {'NYC': 'New York City', 'LA': 'Los Angeles', 'UK': 'United Kingdom'}
for old, new in ids.items():
df['Comment'] = df['Comment'].str.replace(old, new, regex=False)
# Categories Type Comment
# 0 animal tree The New York City tree is very big
# 1 plant dog The cat from the United Kingdom is small
# 2 object rock The rock was found in Los Angeles
The only time replace() outperforms a str.replace() loop is with small dataframes:

The timing functions for reference:
def Series_replace(df):
df['Comment'] = df['Comment'].replace(ids, regex=True)
return df
def Series_str_replace(df):
for old, new in ids.items():
df['Comment'] = df['Comment'].str.replace(old, new, regex=False)
return df
Note that if ids is a dataframe instead of dictionary, you can get the same performance with itertuples():
ids = pd.DataFrame({'Id': ['NYC', 'LA', 'UK'], 'City': ['New York City', 'Los Angeles', 'United Kingdom']})
for row in ids.itertuples():
df['Comment'] = df['Comment'].str.replace(row.Id, row.City, regex=False)
I have a tokenized pandas Series that looks like this:
reviews = [['bad', 'movie', 'it', 'was', 'turrible'],['bad', 'acting', 'in' 'it'], ['ok', 'experience'],...]
I have a dictionary like this:
d = {'turrible':'terrible', 'ok':'okay',...}
Any words in the reviews that appear in the dictionary keys should be replaced with the dictionary values. So the expected output is:
reviews = [['bad', 'movie', 'it', 'was', 'terrible'],['bad', 'acting', 'in', 'it'], ['okay', 'experience'],...]
I've searched for hours, and I've tried these solutions, but I am not getting the expected output.
Trial 1:
pattern = re.compile(r'\b(' + '|'.join(d.keys()) + r')\b') result = pattern.sub(lambda x: d[x.group()], reviews)
Output: error: incomplete escape \u
Trial 2:
def replaceWords(text,wdict): return ''.join(wdict.get(word,word) for word in text) replaceWords(docs,d) Output: TypeError unhashable type: 'list'
Trial 3 - no error message but did not get expected output:
reviews = reviews.replace(d)
Trial 4:
reviews = reviews.replace(d, regex=True) error: missing ), unterminated subpattern
Any help would be appreciated.
When map isn't an option, there's always replace;
df['col_1'] = df['col_1'].replace(value2fixed)
df
col_1 col_2
0 dada 500
1 mel 650
2 hoodie 750
The difference between map and replace is that map replaces "invalid" keys with NaNs - in contrast, replace does not touch them.
You can use replace with a nested dictionary:
df.replace({'col_1':value2fixed}, inplace=True)
>>> df
col_1 col_2
0 dada 500
1 mel 650
2 hoodie 750
The nested dictionary syntax reads as:
Nested dictionaries, e.g., {‘a’: {‘b’: nan}}, are read as follows: look in column ‘a’ for the value ‘b’ and replace it with nan.
From the docs