I like re module solutions
if not re.search(r'(?<!X)AAA', myString):
...
I don't want my script to notice any presence of
"XAAA"when it doesif 'AAA' not in myString.
You can first remove any occurrences of the sub-string XAAA from myString, and then test if AAA is not a sub-string of myString:
if 'AAA' not in myString.replace('XAAA', ''):
# do action
>>> import re
>>> string = '85.2147....85.2150..85.2152..85.2166...85.2180.85.2190.'
>>> result = re.findall('(\d+\.\d+)\.+',string)
>>> result
['85.2147', '85.2150', '85.2152', '85.2166', '85.2180', '85.2190']
You should read up on Python regular expressions and book mark a regular expression test page to help you solve these problems. Also, it's good to remember that Python can be run in interactive mode to allow you to test things out by hand very quickly.
>>> import re
>>> test = "85.2147....85.2150..85.2152..85.2166...85.2180.85.2190."
>>> target = re.compile(r'(\d+\.\d+)(\.+)')
>>> match = re.findall(target, test)
>>> match
[('85.2147', '....'), ('85.2150', '..'), ('85.2152', '..'), ('85.2166', '...'), ('85.2180', '.'), ('85.2190', '.')]
>>> res
['85.2147', '85.2150', '85.2152', '85.2166', '85.2180', '85.2190']
>>>
Check the following code with regex solution:
import re
# set up the regex pattern
# the words which should be skipped, must be whole word and case-insensitive
ptn_to_skip = re.compile(r'\b(?:no|none)\b', re.IGNORECASE)
# the pattern for mapping
# Note: any regex meta charaters need to be escaped, or it will fail.
ptn_to_map = re.compile(r'\b(' + '|'.join(replace_terms_df.Text.tolist()) + r')\b')
# map from text to Replace_item
terms_map = replace_terms_df.set_index('Text').Replace_item
def adjust_text(x):
# if 1 - 3 ptn_to_skip found, return x,
# otherwise, map the matched group \1 with terms_map
if 0 < len(ptn_to_skip.findall(x)) <= 3:
return x
else:
return ptn_to_map.sub(lambda y: terms_map[y.group(1)], x)
# do the conversion:
text_df['new_text'] = text_df.Text.apply(adjust_text)
Some Notes:
- I converted the texts in
replace_terms_df.Textinto a regex. default the texts are all plain-text without regex meta characters. - if there are any regex meta characters like '$', ']' etc, you will have to escape them. regex tends to be slow especially with meta characters, if you have large chuck of data, don't suggest this solution to you.
Update:
A new logic is added to check the excluded-words ['no', 'none'] first, if matches, then find the next 0-3 words which are not themselves excluded-words, save them to \1, the actual matched search-word will be saved in \2. then in the regex replacement part, handle them differently.
Below are the new code:
import re
# pattern to excluded words (must match whole-word and case insensitive)
ptn_to_excluded = r'\b(?i:no|none)\b'
# ptn_1 to match the excluded-words ['no', 'none'] and the following maximal 3 words which are not excluded-words
# print(ptn_1) --> \b(?i:no|none)\b\s*(?:(?!\b(?i:no|none)\b)\S+\s*){,3}
# where (?:(?!\b(?i:no|none)\b)\S+\s*) matches any words '\S+' which is not in ['no', 'none'] followed by optional white-spaces
# {,3} to specify matches up to 3 words
ptn_1 = r'{0}\s*(?:(?!{0})\S+\s*){{,3}}'.format(ptn_to_excluded)
# ptn_2 is the list of words you want to convert with your terms_map
# print(ptn_2) --> \b(?:random|here|some)\b
ptn_2 = r'\b(?:' + '|'.join(replace_terms_df.Text.tolist()) + r')\b'
# new pattern based on the alternation using ptn_1 and ptn_2
# regex: (ptn_1)|(ptn_2)
new_ptn = re.compile('({})|({})'.format(ptn_1, ptn_2))
# map from text to Replace_item
terms_map = replace_terms_df.set_index('Text').Replace_item
# regex function to do the convertion
def adjust_map(x):
return new_ptn.sub(lambda m: m.group(1) or terms_map[m.group(2)], x)
# do the conversion:
text_df['new_text'] = text_df.Text.apply(adjust_map)
Explanation:
I defined two sub-patterns:
- ptn_1: try to match the words you want to be excluded, i.e., the words 'no', 'none' followed by at most 3 more words which are not in ['no', 'none']
- ptn_2: try to match one of the words you want to convert based on the replace_terms_df.
How it works:
- with the alternation '|', the regex engine will make sure ptn_1 matches before ptn_2, if neither matches, the original text is kept.
- The matched ptn_1 text will be saved in m.group(1) and ptn_2 result to m.group(2)
- In the replacement part. If m.group(1) is not Empty(meaning ptn_1 is matched) then return m.group(1) (thus this part of matches is untouched), otherwise return terms_map[y.group(2)]
Some tests below:
In []: print(new_ptn)
re.compile('(\\b(?i:no|none)\\b\\s*(?:(?!\\b(?i:no|none)\\b)\\S+\\s*){,3})|(\\b(random|here|some)\\b)')
In[]: for i in [
'yes, no such a random text'
, 'yes, no such a a random text'
, 'no no no such a random text no such here here here no'
]: print('{}:\n [{}]'.format(i, adjust_map(i)))
...:
yes, no such a random text:
[yes, no such a random text]
yes, no such a a random text:
[yes, no such a a <RANDOM_REPLACED> text]
no no no such a random text no such here here here no:
[no no no such a random text no such here here <HERE_REPLACED> no]
Let me know if this works.
More to consider:
- in ptn_1, '\S+' is used to define a WORD, this will have issue if one of the words is something like ',none', this preceding 'comma' will let it skip the (?!\b(?:no|none)) test.
- In fact, should ',no', '"none"' be excluded? this will impact how words are counted. modifying
ptn_to_excludedcould be enough.
Starting by creating a dictionary of replaced items will help. You can do the following:
# create a dict
make_dict = replace_terms_df.set_index('Text')['Replace_item'].to_dict()
# this function does the replacement work
def g_val(strin, dic):
d = []
if 'none' in strin or 'no' in strin:
return strin
else:
for i in strin.split():
if i not in dic:
d.append(i)
else:
d.append(dic[i])
return ' '.join(d)
## apply the function
text_df['new_text'] = text_df['Text'].apply(lambda x: g_val(x, dic=make_dict))
## check output
print(text_df['new_text'])
0 <HERE_REPLACED> is <SOME_REPLACED> <RANDOM_REP...
1 no such random text, none here
2 more <RANDOM_REPLACED> text
Explanation
In the function, we are doing:
1. If the string contains none or no, we return the string as is.
2. If it doesn't contain none or no, we check if the word is available in the dictionary, if yes, we return the replaced value else the existing value.
Looks like you need np.where
Ex:
test = ["No","No Error"]
data = pd.DataFrame(test, columns=["text"])
data['text'] = np.where(data['text'] == 'No', "Data not available", data['text'])
# OR data.text[data.text=='No'] = "Data not available"
print(data)
Output:
text
0 Data not available
1 No Error
You can also use, Series.replace
data.text.replace({"No":"Data not available"})
0 Data not available
1 No Error
Name: text, dtype: object
First problem:
readin = new.read
You're not calling the method, you won't get the file contents in readin
Second problem:
if "docket =3ghi" == True:
You're comparing if a string is True - it's never True.
Let's break down your current statement:
if "docket = 3ghi" == True:
Non-empty strings evaluate like True, but are not exactly True. True is a boolean, so you are asking "is this string a boolean?" That is always False:
"somestr" == True
# False
Fix it to check if the string is in a part of a file. For example:
with open(("test.txt",'r') as new:
for line in new.read(): # read in the file and iterate
if "somestr" in line:
# do something
Note I've also added the parens to new.read() so that you don't get exceptions like function doesn't support iteration
If you want to replace all occurrences of e that have a t after them, you can use the following:
>>> import re
>>> re.sub(r'e(?=.*t)', 'i', 'forest')
'forist'
>>> re.sub(r'e(?=.*t)', 'i', 'jungle')
'jungle'
>>> re.sub(r'e(?=.*t)', 'i', 'greet')
'griit'
This uses a lookahead, which is a zero-width assertion that checks to see if t is anywhere later in the string without consuming any characters. This allows the regex to find each e that fits this condition instead of just the first one.
>>> import re
>>> re.sub(r'e(.*t)', r'i\1', "forest")
'forist'
>>> re.sub(r'e(.*t)', r'i\1', "jungle")
'jungle'
using regular expression word boundary:
import re
print(re.sub(r"\bis\b","was","This is my sentence"))
Better than a mere split because works with punctuation as well:
print(re.sub(r"\bis\b","was","This is, of course, my sentence"))
gives:
This was, of course, my sentence
Note: don't skip the r prefix, or your regex would be corrupt: \b would be interpreted as backspace.
A simple but not so all-round solution (as given by Jean-Francios Fabre) without using regular expressions.
' '.join(x if x != word else new_word for x in string.split())
You can use regex with lambda function that receives the match object:
0 qqq--++www++
1 1234+5678-
dtype: object
s.str.replace(pat=r"\+|-", repl= lambda mo: "+" if mo.group()=="-" else "-", regex=True)
0 qqq++--www--
1 1234-5678+
dtype: object
You can do something like this:
Char 1 What-What
Char 2 What+What
Char 3 0
Char 4 0
Char 5 0
Char 6 0
Char 7 0
Char 8 0
mySeries.loc[mySeries.str.contains(r'[-+]') == True] = mySeries.str.translate(str.maketrans("+-", "-+"))
Char 1 What+What
Char 2 What-What
Char 3 0
Char 4 0
Char 5 0
Char 6 0
Char 7 0
Char 8 0
If it's not a series you have to do it this way:
A B C D E F G H
Char 1 1 0 0 0 0 0 0 What-What
Char 2 1 0 0 0 0 0 0 What+What
Char 3 0 1 0 0 0 0 0 0
Char 4 0 0 1 0 0 0 0 0
Char 5 0 0 0 1 0 0 0 0
Char 6 0 0 0 0 1 0 0 0
Char 7 0 0 0 0 1 0 0 0
Char 8 0 0 0 0 0 1 0 0
df.H.loc[a.str.contains(r'[-+]') == True] = df.H.str.translate(str.maketrans("+-", "-+"))
A B C D E F G H
Char 1 1 0 0 0 0 0 0 What+What
Char 2 1 0 0 0 0 0 0 What-What
Char 3 0 1 0 0 0 0 0 0
Char 4 0 0 1 0 0 0 0 0
Char 5 0 0 0 1 0 0 0 0
Char 6 0 0 0 0 1 0 0 0
Char 7 0 0 0 0 1 0 0 0
Char 8 0 0 0 0 0 1 0 0
As i told in the comment section, You don't really need to use index.str.endswith until strictly it needs to be rather use anchors like for start ^ and endswith $ that should do a Job for you.
Just taking @Scott's sample for consideration.
df.index.str.replace(r'_s$', '_sp', regex=True)
I'm retaining this answer here for the sake of posterity ..
IIUC,
import pandas as pd
import numpy as np
โ
df = pd.DataFrame(np.random.randint(0,100,(5,5)), columns=['a_s','b','c_s','d','e'], index=['A','B_s','C','D_s','E_s'])
โ
df.columns = df.columns.str.replace('_s','_sp')
df.index = df.index.str.replace('_s','_sp')
โ
print(df)
Output:
a_sp b c_sp d e
A 51 80 48 93 34
B_sp 96 16 73 15 29
C 27 85 35 93 69
D_sp 92 79 90 71 85
E_sp 4 63 2 77 14
How about
>> df.loc[(df['date'] == '2016-09-10') & (df['value'] == 'value1'), 'value'] = 'value7'
You may want to read Indexing and Selecting Data section and this for more info
If you want to keep everything else same and change what you need, try this,
df.loc[df['date'] == '2016-09-10' , 'value'] = 'value7'
You should use:
word = word.replace(word[i],'7', 1)
to indicate that you want to make one character replacement. Calling replace() without indicating how many replacements you wish to make will replace any occurrence of the character "e" (as found at word[i]) by "7".
the answer above has a little bug for example: when your word = 'ebcdeefghiijkl' the result of replace_letter(word) will be '7abcdeefgh7ijkl' you can try this:
def replace_letter(word):
result=[]
for i in range(len(word)):
if i!=len(word)-1 and word[i] == word[i+1]:
result.append('7')
else:
result.append(word[i])
return ''.join(result)