If they are all within the same line, that is there are no line breaks between "1." and "2." then you can iterate over the lines of the file like this:
for line in open("myfile.txt"):
#do stuff
The line will be disposed of and overwritten at each iteration meaning you can handle large file sizes with ease. If they're not on the same line:
for line in open("myfile.txt"):
if #regex to match start of new string
parsed_line = line
else:
parsed_line += line
and the rest of your code.
Answer from wheaties on Stack OverflowHello,
I would like to know how to best approach the following task.
I have a large file which is a few million lines. I want to read each line and do something with it.
Considerations:
-
I don’t want the whole file loaded into memory at once, I want this loaded in chunks
-
Threading should be used (unless there’s a better option)
My initial thought process in pseudocode would be something like:
lines_at_once = 100000
for line in read(file, lines_at_once):
do_something(line)What is the best way to read a very large file in Python
When should I use the mmap module for file reading
How can I avoid memory errors when processing large files
If they are all within the same line, that is there are no line breaks between "1." and "2." then you can iterate over the lines of the file like this:
for line in open("myfile.txt"):
#do stuff
The line will be disposed of and overwritten at each iteration meaning you can handle large file sizes with ease. If they're not on the same line:
for line in open("myfile.txt"):
if #regex to match start of new string
parsed_line = line
else:
parsed_line += line
and the rest of your code.
Why don't you just read the file char by char using file.read(1)?
Then, you could - in each iteration - check whether you arrived at the char 1. Then you have to make sure that storing the string is fast.
I have a fairly large text file which I would like to run in chunks. In order to do this with the subprocess library, one would execute following shell command:
"cat hugefile.log"
with the code:
import subprocess
task = subprocess.Popen("cat hugefile.log", shell=True, stdout=subprocess.PIPE)
data = task.stdout.read()
Using print(data) will spit out the entire contents of the file at once. How can I present the number of chunks, and then access the contents of this file by the chunk size (e.g. chunk = three lines at a time).
It must be something like:
chunksize = 1000 # break up hugefile.log into 1000 chunks
for chunk in data:
print(chunk)
The equivalent question with Python open() of course uses the code
with open('hugefile.log', 'r') as f:
read_data = f.read()
How would you read_data in chunks?
Maybe I'm missing something, but why don't you just use read()'s size argument?
To read a file’s contents, call f.read(size), which reads some quantity of data and returns it as a string (in text mode) or bytes object (in binary mode). size is an optional numeric argument. When size is omitted or negative, the entire contents of the file will be read and returned; it’s your problem if the file is twice as large as your machine’s memory. Otherwise, at most size bytes are read and returned.
chunksize = 1000
with open('hugefile.log', 'r') as f:
while True:
read_data = f.read(chunksize)
if not read_data:
break # done
print(read_data)
is the chunk supposed to be a list of 3 lines or a single string made by combining 3 lines?
It's not clear when do you want your chunk to end, tho - when it encounters a '>' at a beginning of a line or anywhere in the line, so I'll assume the first scenario:
chunk = []
with open("your_large_file.ext", "r") as f:
for _ in xrange(4): # skip 4 lines, use range() on Python 3.x instead
next(f)
for line in f:
if chunk and line.startswith(">"): # break on > if we're already collecting a chunk
break
chunk.append(line)
print("".join(chunk)) # or whatever you want to do with it
.
>TRINITY_DN63782_c0_g1_i1 len=433 path=[411:0-432] [-1, 411, -2]
ATAGACACGAACACAAACACATAAATAATTTGAGAAAATAGAAGTGATTGAACTTGTTGG
TGTGGTACAGGTGTCAAACAAACCTTCAACCAGAAGTTTTGTTGCTGCATAAATCATAGT
GACACTCTGATATGATATCAAAGAAAATCATGTAACCCAAATACATCCCTAAGTATCTAG
TTGAAGCTACAGTCCACTAATTGTAACAATATTAAGTAATTATGAAATGAACCATTTGCA
You can use this function if you know from which line the data starts:
def extract_chunk(start_line):
"""
start_line is the line number where your data starts, counting from 0
"""
lines = []
with open("data.txt") as f:
for i, line in enumerate(f):
if i == start_line:
lines.append(line)
elif not line.startswith(">") and i > start_line:
lines.append(line)
elif line.startswith(">"):
break
return "".join(lines)