You can use str.contains to mask the rows that contain 'ball' and then overwrite with the new value:
In [71]:
df.loc[df['sport'].str.contains('ball'), 'sport'] = 'ball sport'
df
Out[71]:
name sport
0 Bob tennis
1 Jane ball sport
2 Alice ball sport
To make it case-insensitive pass `case=False:
df.loc[df['sport'].str.contains('ball', case=False), 'sport'] = 'ball sport'
Answer from EdChum on Stack OverflowYou can use str.contains to mask the rows that contain 'ball' and then overwrite with the new value:
In [71]:
df.loc[df['sport'].str.contains('ball'), 'sport'] = 'ball sport'
df
Out[71]:
name sport
0 Bob tennis
1 Jane ball sport
2 Alice ball sport
To make it case-insensitive pass `case=False:
df.loc[df['sport'].str.contains('ball', case=False), 'sport'] = 'ball sport'
You can use apply with a lambda. The x parameter of the lambda function will be each value in the 'sport' column:
df.sport = df.sport.apply(lambda x: 'ball sport' if 'ball' in x else x)
You are not keeping the result of x.replace(). Try the following instead:
for tag in tags:
x = x.replace(tag, '')
print x
Note that your approach matches any substring, and not just full words. For example, it would remove the LOL in RUN LOLA RUN.
One way to address this would be to enclose each tag in a pair of r'\b' strings, and look for the resulting regular expression. The r'\b' would only match at word boundaries:
for tag in tags:
x = re.sub(r'\b' + tag + r'\b', '', x)
The method str.replace() does not change the string in place -- strings are immutable in Python. You have to bind x to the new string returned by replace() in each iteration:
for tag in tags:
x = x.replace(tag, "")
Note that the if statement is redundant; str.replace() won't do anything if it doesn't find a match.
python - Replacing string if it contains specified pattern - Stack Overflow
How to replace a cell value if it contains a string? Pandas
Check if certain Strings are present in a Sentence and Replace them with another String using Python 3.6 - Stack Overflow
python - Should I check if a substring exists before trying to replace it? - Stack Overflow
I have a DataFrame where customers entered their job title's manually instead of from a drop down menu so now I have what could be 50 different ways to describe "Director". I want to replace any cell value with "Director" if that cell contains the string "Director". So it could be "Senior Director Engineering Design" or "Director of R&D" etc but I want to rename them all "Director".
import re
s = 'The day is at not all bad'
pattern=r'(not)(?(1).+(bad))'
match=re.search(pattern,s)
new_string=re.sub(pattern,"good",s)
print(new_string)
output:
The day is at good
Regex explanation :
I used if else condition regex here :
How if else in regex works , well this is very simple if else regex syntax:
(condition1)(?(1)(do something else))
(?(A)X|Y)
This means "if proposition A is true, then match pattern X; otherwise, match pattern Y."
so in this regex :
(not)(?(1).+(bad))
it matches 'bad' if 'not' in the string, the condition is 'not' must present in the string.
Second Regex :
if you want you can also use this regex:
(not.+)(bad)
In this group(2) is matching 'bad'.
Your string :
>>> s = 'The day is not at all bad' #input
>>> print(output)
>>> 'The day is good' # output
There are a couple of ways you could approach this. One way is to convert the sentence to a list of words, locate "not" and "bad" in the list, remove them and all the elements in between and then insert "good".
>>> s = 'the day is not at all bad'
>>> start, stop = 'not', 'bad'
>>> words = s.split()
>>> words
['the', 'day', 'is', 'not', 'at', 'all', 'bad']
>>> words.index(start)
3
>>> words.index(stop)
6
>>> del words[3:7] # add 1 to stop index to delete "bad"
>>> words
['the', 'day', 'is']
>>> words.insert(3, 'good')
>>> words
['the', 'day', 'is', 'good']
>>> output = ' '.join(words)
>>> print(output)
the day is good
Another method is to use regular expressions to find a pattern that matches "not" followed by zero or more words, followed by "bad". The re.sub function finds strings that match a given pattern and replaces them with a string that you provide:
>>> import re
>>> pattern = r'not\w+bad'
>>> re.search(pattern, s)
>>> pattern = r'not(\s+\w+)* bad' # pattern matches "not <words> bad"
>>> re.sub(pattern, 'good', s)
'the day is good'
Simple is better than complex. Doing the replacement necessitates checking anyway, so no reason to write it out yourself. Special cases aren't special enough; you don't need any special handling to replace zero instances of the substring - it works the same as replacing any other number of instances.
Don't check.
Both options you have there have the same results. However in terms of performance, the second one is better than the first.
The first one will make you go through the string twice, once to check and once to replace. So it will always go through the string once, possibly twice.
The second one, will always go through the string only once to try and replace.
For the replace function, if the first argument (substring1) isnโt found in the string, nothing will happen. So the second one is perfectly safe to use.
Always remember that simpler is better.