I like re module solutions

if not re.search(r'(?<!X)AAA', myString):
    ...
Answer from Gribouillis on Stack Overflow
๐ŸŒ
Python.org
discuss.python.org โ€บ python help
Trying to replace input string with another string with multiple conditions - Python Help - Discussions on Python.org
March 18, 2022 - Python noob here! Iโ€™m trying to replace an input string with another string, depending on which condition is met. My problem is that it only works when the first condition is met. When I run the program it prompts me for an input of A, B, C, D or E. If I enter โ€œaโ€ or โ€œAโ€ the string gets converted to โ€œClass Aโ€, which shows when the print statement is executed.
๐ŸŒ
Note.nkmk.me
note.nkmk.me โ€บ home โ€บ python
Extract and Replace Elements That Meet the Conditions of a List of Strings in Python | note.nkmk.me
May 19, 2023 - Extract, replace, convert elements ... although their usage is grammatically optional. ... The startswith() method returns True if the string starts with the specific string....
๐ŸŒ
Stack Overflow
stackoverflow.com โ€บ questions โ€บ 27297360 โ€บ python-string-replace-conditional
python string replace conditional - Stack Overflow
May 22, 2017 - notice the use of brackets which are required due to operator precedence and you need to index the df itself, what you compared was a list with a single entry which was a string. spot.ix('programa'=='CLASSIFICADOES' & ['espec']=='', 'tipo') = 'N' ... Sign up to request clarification or add additional context in comments. ... The spot['tipo'] = np.where((spot['programa']=='CLASSIFICADOES') & (spot['espec']==''), 'N') seems to be right, but now python tells me np.where should have both x,y or none.
๐ŸŒ
W3Schools
w3schools.com โ€บ python โ€บ ref_string_replace.asp
Python String replace() Method
Remove List Duplicates Reverse ... Study Plan Python Interview Q&A Python Training ... The replace() method replaces a specified phrase with another specified phrase....
Top answer
1 of 2
1

Check the following code with regex solution:

import re

# set up the regex pattern
# the words which should be skipped, must be whole word and case-insensitive
ptn_to_skip = re.compile(r'\b(?:no|none)\b', re.IGNORECASE)

# the pattern for mapping
# Note: any regex meta charaters need to be escaped, or it will fail.
ptn_to_map = re.compile(r'\b(' + '|'.join(replace_terms_df.Text.tolist()) + r')\b')

# map from text to Replace_item
terms_map = replace_terms_df.set_index('Text').Replace_item

def adjust_text(x):
    # if 1 - 3 ptn_to_skip found, return x, 
    # otherwise, map the matched group \1 with terms_map
    if 0 < len(ptn_to_skip.findall(x)) <= 3:
        return x
    else:
        return ptn_to_map.sub(lambda y: terms_map[y.group(1)], x)

# do the conversion:
text_df['new_text'] = text_df.Text.apply(adjust_text)

Some Notes:

  • I converted the texts in replace_terms_df.Text into a regex. default the texts are all plain-text without regex meta characters.
  • if there are any regex meta characters like '$', ']' etc, you will have to escape them. regex tends to be slow especially with meta characters, if you have large chuck of data, don't suggest this solution to you.

Update:

A new logic is added to check the excluded-words ['no', 'none'] first, if matches, then find the next 0-3 words which are not themselves excluded-words, save them to \1, the actual matched search-word will be saved in \2. then in the regex replacement part, handle them differently.

Below are the new code:

import re

# pattern to excluded words (must match whole-word and case insensitive)
ptn_to_excluded = r'\b(?i:no|none)\b'

# ptn_1 to match the excluded-words ['no', 'none'] and the following maximal 3 words which are not excluded-words
# print(ptn_1)  -->    \b(?i:no|none)\b\s*(?:(?!\b(?i:no|none)\b)\S+\s*){,3}
# where (?:(?!\b(?i:no|none)\b)\S+\s*) matches any words '\S+' which is not in ['no', 'none'] followed by optional white-spaces
# {,3} to specify matches up to 3 words 
ptn_1 = r'{0}\s*(?:(?!{0})\S+\s*){{,3}}'.format(ptn_to_excluded)

# ptn_2 is the list of words you want to convert with your terms_map
# print(ptn_2)    -->    \b(?:random|here|some)\b
ptn_2 = r'\b(?:' + '|'.join(replace_terms_df.Text.tolist()) + r')\b'

# new pattern based on the alternation using ptn_1 and ptn_2
# regex:  (ptn_1)|(ptn_2)
new_ptn = re.compile('({})|({})'.format(ptn_1, ptn_2))

# map from text to Replace_item
terms_map = replace_terms_df.set_index('Text').Replace_item

# regex function to do the convertion
def adjust_map(x):
    return new_ptn.sub(lambda m:  m.group(1) or terms_map[m.group(2)], x)

# do the conversion:
text_df['new_text'] = text_df.Text.apply(adjust_map)

Explanation:

I defined two sub-patterns:

  • ptn_1: try to match the words you want to be excluded, i.e., the words 'no', 'none' followed by at most 3 more words which are not in ['no', 'none']
  • ptn_2: try to match one of the words you want to convert based on the replace_terms_df.

How it works:

  • with the alternation '|', the regex engine will make sure ptn_1 matches before ptn_2, if neither matches, the original text is kept.
  • The matched ptn_1 text will be saved in m.group(1) and ptn_2 result to m.group(2)
  • In the replacement part. If m.group(1) is not Empty(meaning ptn_1 is matched) then return m.group(1) (thus this part of matches is untouched), otherwise return terms_map[y.group(2)]

Some tests below:

In []: print(new_ptn)
re.compile('(\\b(?i:no|none)\\b\\s*(?:(?!\\b(?i:no|none)\\b)\\S+\\s*){,3})|(\\b(random|here|some)\\b)')

In[]: for i in [
    'yes, no such a random text'
  , 'yes, no such a a random text'
  , 'no no no such a random text no such here here here no'
 ]: print('{}:\n  [{}]'.format(i, adjust_map(i)))
...:
yes, no such a random text:
  [yes, no such a random text]
yes, no such a a random text:
  [yes, no such a a <RANDOM_REPLACED> text]
no no no such a random text no such here here here no:
  [no no no such a random text no such here here <HERE_REPLACED> no]

Let me know if this works.

More to consider:

  • in ptn_1, '\S+' is used to define a WORD, this will have issue if one of the words is something like ',none', this preceding 'comma' will let it skip the (?!\b(?:no|none)) test.
  • In fact, should ',no', '"none"' be excluded? this will impact how words are counted. modifying ptn_to_excluded could be enough.
2 of 2
1

Starting by creating a dictionary of replaced items will help. You can do the following:

# create a dict
make_dict = replace_terms_df.set_index('Text')['Replace_item'].to_dict()

# this function does the replacement work
def g_val(strin, dic):

    d = []
    if 'none' in strin or 'no' in strin:
        return strin
    else:
        for i in strin.split():
            if i not in dic:
                d.append(i)
            else:
                d.append(dic[i])
        return ' '.join(d)

## apply the function
text_df['new_text'] = text_df['Text'].apply(lambda x: g_val(x, dic=make_dict))

## check output
print(text_df['new_text'])

0    <HERE_REPLACED> is <SOME_REPLACED> <RANDOM_REP...
1                       no such random text, none here
2                          more <RANDOM_REPLACED> text

Explanation

In the function, we are doing:
1. If the string contains none or no, we return the string as is.
2. If it doesn't contain none or no, we check if the word is available in the dictionary, if yes, we return the replaced value else the existing value.

Find elsewhere
๐ŸŒ
Programiz
programiz.com โ€บ python-programming โ€บ methods โ€บ string โ€บ replace
Python String replace()
If the old substring is not found, it returns a copy of the original string. ... # replacing 'cold' with 'hurt' print(song.replace('cold', 'hurt')) song = 'Let it be, let it be, let it be, let it be'
Top answer
1 of 1
3

Use a regular expression:

term = 'test(s|ing)?'
df_1['color_value'] = df_1['color_value'].str.replace(term, '', regex=True)
print(df_1)

Output

    id color_value
0  001       blue_
1  002         red
2  003     yellow_
3  004      orange
4  005        blue
5  006         red
6  007       blue_
7  008      orange

From the documentation on str.replace:

pat str or compiled regex
String can be a character sequence or regular expression.

UPDATE

For including "new", "origin" you could do use another regex:

term = 'test(s|ing)?|new|orig'
df_1['color_value'] = df_1['color_value'].str.replace(term, '', regex=True)
print(df_1)

Output

    id color_value
0  001       blue_
1  002         red
2  003     yellow_
3  004     orange_
4  005       blue_
5  006         red
6  007       blue_
7  008      orange

General Solution

If you have many words I suggest you use a library such as trrex it will build a regular expression from a set of words:

import pandas as pd
import trrex as tx

df_1 = pd.DataFrame({'id': ['001', '002', '003', '004', '005', '006', '007', '008'],
                     'color_value': ['blue_test', 'red', 'yellow_tests', 'orange_orig',
                                     'blue_new', 'red', 'blue_testing', 'orange']})

term = tx.make(['test', 'tests', 'testing', 'orig', 'new'], prefix="", suffix="")
df_1['color_value'] = df_1['color_value'].str.replace(term, '', regex=True)
print(df_1)

Output

    id color_value
0  001       blue_
1  002         red
2  003     yellow_
3  004     orange_
4  005       blue_
5  006         red
6  007       blue_
7  008      orange

The pattern for the given example is:

term = tx.make(['test', 'tests', 'testing', 'orig', 'new'], prefix="", suffix="")
print(term)

Output (pattern build by trrex)

(?:test(?:ing|s)?|new|orig)

DISCLAIMER

I'm the author of trrex

๐ŸŒ
Stack Overflow
stackoverflow.com โ€บ questions โ€บ 66754171 โ€บ how-to-replace-character-in-string-upon-a-condition-python
pandas - How to replace character in string upon a condition - python - Stack Overflow
... Show activity on this post. ... s[index-1:index+1] to match a substring. But you don't need a loop, you can just use replace() to replace each substring....
๐ŸŒ
Stack Overflow
stackoverflow.com โ€บ questions โ€บ 55486297 โ€บ replace-string-and-store-on-condition
python - Replace string and store on condition - Stack Overflow
Connect and share knowledge within a single location that is structured and easy to search. Learn more about Teams ... I have a python script in which I defined a working directory Directory_1 as a hardcoded string. If it's not present in the machine, it prompts to select some other directory Directory_2. Directory_2 should replace ...
Top answer
1 of 16
385

Here is a short example that should do the trick with regular expressions:

import re

rep = {"condition1": "", "condition2": "text"} # define desired replacements here

# use these three lines to do the replacement
rep = dict((re.escape(k), v) for k, v in rep.items()) 
pattern = re.compile("|".join(rep.keys()))
text = pattern.sub(lambda m: rep[re.escape(m.group(0))], text)

For example:

>>> pattern.sub(lambda m: rep[re.escape(m.group(0))], "(condition1) and --condition2--")
'() and --text--'
2 of 16
200

You could just make a nice little looping function.

def replace_all(text, dic):
    for i, j in dic.iteritems():
        text = text.replace(i, j)
    return text

where text is the complete string and dic is a dictionary โ€” each definition is a string that will replace a match to the term.

Note: in Python 3, iteritems() has been replaced with items()


Careful: Python dictionaries don't have a reliable order for iteration. This solution only solves your problem if:

  • order of replacements is irrelevant
  • it's ok for a replacement to change the results of previous replacements

Update: The above statement related to ordering of insertion does not apply to Python versions greater than or equal to 3.6, as standard dicts were changed to use insertion ordering for iteration.

For instance:

d = { "cat": "dog", "dog": "pig"}
my_sentence = "This is my cat and this is my dog."
replace_all(my_sentence, d)
print(my_sentence)

Possible output #1:

"This is my pig and this is my pig."

Possible output #2

"This is my dog and this is my pig."

One possible fix is to use an OrderedDict.

from collections import OrderedDict
def replace_all(text, dic):
    for i, j in dic.items():
        text = text.replace(i, j)
    return text
od = OrderedDict([("cat", "dog"), ("dog", "pig")])
my_sentence = "This is my cat and this is my dog."
replace_all(my_sentence, od)
print(my_sentence)

Output:

"This is my pig and this is my pig."

Careful #2: Inefficient if your text string is too big or there are many pairs in the dictionary.

๐ŸŒ
Regular-Expressions.info
regular-expressions.info โ€บ replaceconditional.html
Replacement String Conditionals
You also need to escape literal closing parentheses (Boost) or curly braces (PCRE2) with backslashes inside conditionals. In replacement string flavors that support conditionals, you can escape colons, parentheses, curly braces, and even question marks with backslashes to make sure they are interpreted as literals anywhere in the replacement string.