You need to manually replace newlines with \n using the replace method.
Set the lineterminator option to the desired character sequence. More info on what else is available is in the docs.
with open('csvfile.csv', 'w') as csvOutput:
writer = csv.writer(csvOutput, delimiter='|', escapechar=' ', quoting=csv.QUOTE_NONE, lineterminator='\n')
for row in data:
writer.writerow([s.replace('\n', '\\n').encode('utf-8') for s in row])
Answer from Jeff Mercado on Stack OverflowYou need to manually replace newlines with \n using the replace method.
Set the lineterminator option to the desired character sequence. More info on what else is available is in the docs.
with open('csvfile.csv', 'w') as csvOutput:
writer = csv.writer(csvOutput, delimiter='|', escapechar=' ', quoting=csv.QUOTE_NONE, lineterminator='\n')
for row in data:
writer.writerow([s.replace('\n', '\\n').encode('utf-8') for s in row])
I'm guessing it is to do with the \ being an escape character in Python. \\ is interpreted to be the actual \ string. So your code should be as follows:
with open('csvfile.csv', 'w') as csvOutput:
testData = csv.writer(csvOutput, delimiter='|', escapechar=' ', quoting = csv.QUOTE_NONE)
for row in data:
newRow = [row[x].encode('utf-8') for x in xrange(len(row))]
testData.writerow(newRow)
testData.writerow('\\n') #Note the final row with \\
python - How to store strings in CSV with new line characters? - Data Science Stack Exchange
newline in text fields with python csv writer - Stack Overflow
csv writer: append doesn't go to new line
What does the `newline=" "` argument do?
I am running the script on a Windows 7 machine using Python 3.4.1. The script runs correctly and produces a csv file, the only problem is at the end of each line is a '\r\n' which causes an extra blank line to appear when displayed in Excel. How do get remove the extra blank line?
import pyodbc
import csv
connect = pyodbc.connect('driver={SQL Server Native Client 10.0};SERVER=xxxxx-7- VM;DATABASE=DEMO_xxxx_TEST;UID=DEMO_xxxx_TEST;PWD=password')
cursor = connect.cursor()
cursor.execute('''select Report_name
,UserID
,FieldName
,XPos
,YPos
,Hidden
,Picture
from Rpt_Nudge_2
where Report_Name = 'UB04' and UserID = 'myid' and YPos > 8886 and YPos < 9888
order by YPos, XPos''')
col_names = [i[0] for i in cursor.description]
print(col_names)
nudge = cursor.fetchall()
with open('UB04_nudge.csv', 'w') as csvfile:
fileout = csv.writer(csvfile)
row = fileout.writerow(col_names)
for line in nudge:
fileout.writerow(line)
connect.close()
The strip() method removes whitespace, including newlines.
fileout.writerow(line.strip())
In Python 2, you could write to CSV files with the 'wb' option on the file and avoid this.
In Python 3, it's a little different - here's the documentation, take a look at the footnote.
Since you're opening the csv file as a file, you should replace line 25 with this:
with open('UB04_nudge.csv', 'w', newline='') as csvfile:
Basically, since Windows uses \r\n line endings, file() is already planning to write a newline ending out after each line. CSV does this as well - so you're getting the duplicate newlines after each row. By setting newline='', you're telling file() to not terminate new lines - which works since csv() will terminate the lines on its own.
Here is a simple solution: Replace all \n with \\n before saving to CSV. This will preserve the newline characters.
df.loc[:, "Column_Name"] = df["Column_Name"].apply(lambda x: x.replace('\n', '\\n'))
df.to_csv("df.csv", index=False)
I assume that you want to keep the newlines in the strings for some reason after you have loaded the csv files from disk. Also that this is done again in Python. My solution will require Python 3, although the principle could be applied to Python 2.
The main trick
This is to replace the \n characters before writing with a weird character that otherwise wouldn't be included, then to swap that weird character back for \n after reading the file back from disk.
For my weird character, I will use the Icelandic thorn: Þ, but you can choose anything that should otherwise not appear in your text variables. Its name, as defined in the standardised Unicode specification is: LATIN SMALL LETTER THORN. You can use it in Python 3 a couple of ways:
weird_literal = 'þ'
weird_name = '\N{LATIN SMALL LETTER THORN}'
weird_char = '\xfe' # hex representation
weird_literal == weird_name == weird_char # True
That \N is pretty cool (and works in python 3.6 inside formatted strings too)... it basically allows you to pass the Name of a character, as per Unicode's specification.
An alternative character that may serve as a good standard is '\u2063' (INVISIBLE SEPARATOR).
Replacing \n
Now we use this weird character to replace '\n'. Here are the two ways that pop into my mind for achieving this:
using a list comprehension on your list of lists:
data:new_data = [[sample[0].replace('\n', weird_char) + weird_char, sample[1]] for sample in data]putting the data into a dataframe, and using replace on the whole
textcolumn in one godf1 = pd.DataFrame(data, columns=['text', 'category']) df1.text = df.text.str.replace('\n', weird_char)
The resulting dataframe looks like this, with newlines replaced:
text category
0 some text in one line 1
1 text withþnew line character 0
2 another newþline character 1
Writing the results to disk
Now we write either of those identical dataframes to disk. I set index=False as you said you don't want row numbers to be in the CSV:
FILE = '~/path/to/test_file.csv'
df.to_csv(FILE, index=False)
What does it look like on disk?
text,category
some text in one line,1
text withþnew line character,0
another newþline character,1
Getting the original data back from disk
Read the data back from file:
new_df = pd.read_csv(FILE)
And we can replace the Þ characters back to \n:
new_df.text = new_df.text.str.replace(weird_char, '\n')
And the final DataFrame:
new_df
text category
0 some text in one line 1
1 text with\nnew line character 0
2 another new\nline character 1
If you want things back into your list of lists, then you can do this:
original_lists = [[text, category] for index, text, category in old_df_again.itertuples()]
Which looks like this:
[['some text in one line', 1],
['text with\nnew line character', 0],
['another new\nline character', 1]]
The created file is a valid CSV file - parsers should be able to identify that a pair of quotes is open when they find a newline character.
If you want to avoid these characters for being able to see then with a normal, non CSV aware, text editor, then you have to escape the newlines, so that the data output is transformed from the real newline character (a single byte with decimal value 10 (\x0a or \n) ) to two printable character sequence: \ and n (2 bytes with decimal values 92 and 110) - or any other sequence of your choice.
On the Python side, that is simply achievable with a str.replace call. However, although you will then see the CSV data rows in the same "physical TXT data rows", in a similar manner applications that will later read this file as data, like a spreadsheet or other Python scripts, won't recognize these sequences as "newlines": you will have to replace them again for newlines after being read (or just keep the modified data and work with it).
tbl_name= 'testnewlines'
with open(tbl_name+'.csv','w', newline='', encoding='utf-8') as f:
writer=csv.DictWriter(f,fieldnames=field_names,delimiter='|')
writer.writeheader()
for d_row in data_rows:
# the next three lines could be written as a comprehension.
# I am unwinding them for clarity
new_row = {}
for key, value in data_rows.items():
new_row[key] = value.replace("\n", "\\n") if isinstance(value, str) else value
# "\\n" escapes the "\" itself so it is a literal "\"
# character, and not a character escaping the "n"
writer.writerow(new_row)
Just to emphasize with other words: this will create a file that is more neat to look at with a text editor, but will not preserve the recorded data for a round-trip: it will require a custom-step after reading to undo this replacement.
I would do it by simply escaping the (\n).
Replace this :
for d_row in data_rows:
writer.writerow(d_row)
By this :
for d_row in data_rows:
writer.writerow({k: v.replace('\n', '\\n') if isinstance(v, str)
else v for k, v in d_row.items()})
Output (.csv) :
id|number|scope
100001|a01|row1 test \n\nrow1 test\n\nrow1 test
100002|a02|row2 test \n\nrow2 test\n\nrow2 test
So, basically, I've got this code:
new_list_csv = []
def add_it():
#gets the values from Entry
#three different tk Entrys generate three different values
name= self.e.get()
ex= self.e_x.get()
ey= self.e_y.get()
#adds them to the new list
new_list_csv.append(name)
new_list_csv.append(ex)
new_list_csv.append(ey)
#appends new characters to csv file
with open ('chr_list_copy.txt', 'a', newline='') as write_obj:
csv_writer = csv.writer(write_obj)
csv_writer.writerow(new_list_csv)and the csv file looks like this:
name,x,y Ergo,1,1 Sum,5,5 Name12,1,2 Name34,3,4 Name56,5,6 Name78,7,8 Name910,9,10 Name1112,11,12 Name1314,13,14
if I try and add [Otto,6,9], I get:
name,x,y [...] Name1314,13,14Otto,6,9
instead of:
name,x,y [...] Name1314,13,14 Otto,6,9
I've used this very same structure for the rest of my code, but in this specific instance, it doesn't work.
For a uni assignment, I need to process a csv file. However these csv files were written such that the EOL character in every case is a carriage return (\r) character. The assignment prompt says that I cannot read in the file using the csv module as it will not recognise the \r character and thus the requirement of the assignment is that, in my Python script, I change the carriage return to a newline or any other EOL character that Python's csv module can recognise as a valid EOL. How would I do this? Also why doesn't the csv module recognise the carriage return? (Sorry for poor English and/or bad formatting)
Python 3:
The official csv documentation recommends opening the file with newline='' on all platforms to disable universal newlines translation:
with open('output.csv', 'w', newline='', encoding='utf-8') as f:
writer = csv.writer(f)
...
The CSV writer terminates each line with the lineterminator of the dialect, which is '\r\n' for the default excel dialect on all platforms because that's what RFC 4180 recommends.
Python 2:
On Windows, always open your files in binary mode ("rb" or "wb"), before passing them to csv.reader or csv.writer.
Although the file is a text file, CSV is regarded a binary format by the libraries involved, with \r\n separating records. If that separator is written in text mode, the Python runtime replaces the \n with \r\n, hence the \r\r\n observed in the file.
See this previous answer.
While @john-machin gives a good answer, it's not always the best approach. For example, it doesn't work on Python 3 unless you encode all of your inputs to the CSV writer. Also, it doesn't address the issue if the script wants to use sys.stdout as the stream.
I suggest instead setting the 'lineterminator' attribute when creating the writer:
import csv
import sys
doc = csv.writer(sys.stdout, lineterminator='\n')
doc.writerow('abc')
doc.writerow(range(3))
That example will work on Python 2 and Python 3 and won't produce the unwanted newline characters. Note, however, that it may produce undesirable newlines (omitting the LF character on Unix operating systems).
In most cases, however, I believe that behavior is preferable and more natural than treating all CSV as a binary format. I provide this answer as an alternative for your consideration.
Based on your comments, the data you're being served doesn't actually include carriage returns or newlines, it includes the text representing the escapes for carriage returns and newlines (so it really has a backslash, r, backslash, n in the data). It's otherwise already in the form you want, so you don't need to involve the csv module at all, just interpret the escapes to their correct value, then write the data directly.
This is relatively simple using the unicode-escape codec (which also handles ASCII escapes):
import codecs # Needed for text->text decoding
# ... retrieve data here, store to res ...
# Converts backslash followed by r to carriage return, by n to newline,
# and so on for other escapes
decoded = codecs.decode(res, 'unicode-escape')
# newline='' means don't perform line ending conversions, so you keep \r\n
# on all systems, no adding, no removing characters
# You may want to explicitly specify an encoding like UTF-8, rather than
# relying on the system default, so your code is portable across locales
with open(title, 'w', newline='') as f:
f.write(decoded)
If the strings you receive are actually wrapped in quotes (so print(repr(s)) includes quotes on either end), it's possible they're intended to be interpreted as JSON strings. In that case, just replace the import and creation of decoded with:
import json
decoded = json.loads(res)
If I understand your question correctly, can't you just replace the string?
with open(title, 'w') as f: f.write(res.replace("¥r¥n","¥n"))