Solution, if need create one big DataFrame if need processes all data at once (what is possible, but not recommended):

Then use concat for all chunks to df, because type of output of function:

df = pd.read_csv('Check1_900.csv', sep='\t', iterator=True, chunksize=1000)

isn't dataframe, but pandas.io.parsers.TextFileReader - source.

tp = pd.read_csv('Check1_900.csv', sep='\t', iterator=True, chunksize=1000)
print tp
#<pandas.io.parsers.TextFileReader object at 0x00000000150E0048>
df = pd.concat(tp, ignore_index=True)

I think is necessary add parameter ignore index to function concat, because avoiding duplicity of indexes.

EDIT:

But if want working with large data like aggregating, much better is use dask, because it provides advanced parallelism.

Answer from jezrael on Stack Overflow
🌐
Pandas
pandas.pydata.org › docs › reference › api › pandas.read_csv.html
pandas.read_csv — pandas 3.0.6 documentation
Internally process the file in chunks, resulting in lower memory use while parsing, but possibly mixed type inference. To ensure no mixed types either set False, or specify the type with the dtype parameter. Note that the entire file is read into a single DataFrame regardless, use the chunksize or iterator parameter to return the data in chunks.
Top answer
1 of 4
37

Solution, if need create one big DataFrame if need processes all data at once (what is possible, but not recommended):

Then use concat for all chunks to df, because type of output of function:

df = pd.read_csv('Check1_900.csv', sep='\t', iterator=True, chunksize=1000)

isn't dataframe, but pandas.io.parsers.TextFileReader - source.

tp = pd.read_csv('Check1_900.csv', sep='\t', iterator=True, chunksize=1000)
print tp
#<pandas.io.parsers.TextFileReader object at 0x00000000150E0048>
df = pd.concat(tp, ignore_index=True)

I think is necessary add parameter ignore index to function concat, because avoiding duplicity of indexes.

EDIT:

But if want working with large data like aggregating, much better is use dask, because it provides advanced parallelism.

2 of 4
21

You do not need concat here. It's exactly like writing sum(map(list, grouper(tup, 1000))) instead of list(tup). The only thing iterator and chunksize=1000 does is to give you a reader object that iterates 1000-row DataFrames instead of reading the whole thing. If you want the whole thing at once, just don't use those parameters.

But if reading the whole file into memory at once is too expensive (e.g., takes so much memory that you get a MemoryError, or slow your system to a crawl by throwing it into swap hell), that's exactly what chunksize is for.

The problem is that you named the resulting iterator df, and then tried to use it as a DataFrame. It's not a DataFrame; it's an iterator that gives you 1000-row DataFrames one by one.

When you say this:

My problem is I don't know how to use stuff like these below for the whole df and not for just one chunk

The answer is that you can't. If you can't load the whole thing into one giant DataFrame, you can't use one giant DataFrame. You have to rewrite your code around chunks.

Instead of this:

df = pd.read_csv('Check1_900.csv', sep='\t', iterator=True, chunksize=1000)
print df.dtypes
customer_group3 = df.groupby('UserID')

… you have to do things like this:

for df in pd.read_csv('Check1_900.csv', sep='\t', iterator=True, chunksize=1000):
    print df.dtypes
    customer_group3 = df.groupby('UserID')

Often, what you need to do is aggregate some data—reduce each chunk down to something much smaller with only the parts you need. For example, if you want to sum the entire file by groups, you can groupby each chunk, then sum the chunk by groups, and store a series/array/list/dict of running totals for each group.

Of course it's slightly more complicated than just summing a giant series all at once, but there's no way around that. (Except to buy more RAM and/or switch to 64 bits.) That's how iterator and chunksize solve the problem: by allowing you to make this tradeoff when you need to.

Discussions

How to set custom chunksize parameters to read csv into pandas data frame?
I don't think there is anything to help with that. You could parse the file once and detect when the month changes so you calculate the chunk size - but this would assume all chunk sizes are the same. There are other tools you could use for "larger than memory" datasets e.g. pyarrow and duckdb Example file: >>> print(open("time.csv").read(), end="") time,value 2022-01-23T01:52:10Z,1 2022-02-23T01:52:10Z,2 2022-01-23T01:53:10Z,2 2022-03-23T01:52:10Z,3 2022-04-23T01:52:10Z,4 2022-01-23T01:53:10Z,3 2022-05-23T01:52:10Z,5 2022-01-23T01:53:10Z,4 2022-02-23T01:52:10Z,8 2022-03-23T01:52:10Z,9 2022-01-23T01:53:10Z,5 You could parse the csv file, add a month column with duckdb and convert the result to parquet. parquet is much smaller (compressed) and faster to read than csv. It's also simple to split up data into multiple files. >>> import duckdb, pyarrow ... ... reader = duckdb.connect().execute(""" ... select month(time) as month, * from 'time.csv' ... """).fetch_record_batch() ... ... pyarrow.dataset.write_dataset( ... reader, ... base_dir="/tmp/mydata/", ... format="parquet", ... partitioning=["month"], ... partitioning_flavor="hive" ... ) ... The partitioning=["month"] specifices the list of columns to split by. >>> from pathlib import Path >>> list(Path("/tmp/mydata").rglob("*.parquet")) [PosixPath('/tmp/mydata/month=1/part-0.parquet'), PosixPath('/tmp/mydata/month=2/part-0.parquet'), PosixPath('/tmp/mydata/month=3/part-0.parquet'), PosixPath('/tmp/mydata/month=4/part-0.parquet'), PosixPath('/tmp/mydata/month=5/part-0.parquet')] pandas has a read_parquet() (as does polars) >>> pd.read_parquet("/tmp/mydata/month=1/part-0.parquet") time value 0 2022-01-23 01:52:10 1 1 2022-01-23 01:53:10 2 2 2022-01-23 01:53:10 3 3 2022-01-23 01:53:10 4 4 2022-01-23 01:53:10 5 Or depending on what exactly you're doing with the - you could implement the functionality with duckdb directly. More on reddit.com
🌐 r/learnpython
1
1
November 22, 2022
python - Refactoring pandas using an iterator via chunksize - Bioinformatics Stack Exchange
This question was also asked on Stack Overflow Bioinformatics rationale eggNOG files can be very big and sump all available RAM for regular to medium sized desktops. I am looking for advice on usi... More on bioinformatics.stackexchange.com
🌐 bioinformatics.stackexchange.com
Does pandas.read_csv's chunks have context on the entire dataset?
chunksize has nothing to do with the actual data, it's for reading the file itself as to not fill up your memory with a huge file. The rest of the file isn't read in so there's literally no way to know if there's more of a specific type of data or not. That is for your program logic You get a chunk, you do your logic on it and calculate your stuff, then in the next chunk you'll need to do your logic again so for your example you'd need to take the max of the order size in your current chunk + your running max. That way if the old max is bigger it's kept or if something in the current chunk is bigger, that becomes your new max More on reddit.com
🌐 r/learnpython
2
3
April 4, 2024
python - Pandas Chunksize iterator - Stack Overflow
I have a 1GB, 70M row file which anytime I load it all it runs out of memory. I have read in 1000 rows and been able to prototype what I'd like it to do. My problem is not knowing how to get the next More on stackoverflow.com
🌐 stackoverflow.com
🌐
Pandas
pandas.pydata.org › docs › user_guide › scale.html
Scaling to large datasets — pandas 3.0.6 documentation
Some readers, like pandas.read_csv(), offer parameters to control the chunksize when reading a single file.
🌐
GeeksforGeeks
geeksforgeeks.org › pandas › how-to-load-a-massive-file-as-small-chunks-in-pandas
How to Load a Massive File as small chunks in Pandas? - GeeksforGeeks
July 15, 2025 - This example demonstrates how to use chunksize parameter in the read_csv function to read a large CSV file in chunks, rather than loading the entire file into memory at once.
🌐
Towards AI
towardsai.com › home › publication › data science › efficient pandas: using chunksize for large datasets
Efficient Pandas: Using Chunksize for Large Datasets | Towards AI
December 10, 2020 - This shows that the chunksize acts just like the next() function of an iterator, in the sense that an iterator uses the next() function to get its’ next element, while the get_chunksize() function grabs the next specified number of rows of data from the data frame, which is similar to an iterator.
🌐
Reddit
reddit.com › r/learnpython › how to set custom chunksize parameters to read csv into pandas data frame?
r/learnpython on Reddit: How to set custom chunksize parameters to read csv into pandas data frame?
November 22, 2022 -

Hello, so I have a massive 5GB+ csv file I am trying to read into a pandas data frame in python. The csv file has over 100 million rows of data. The data is a simple timeseries data set, and so a single timestamp column and then a corresponding value column, where each row represents a single second, proceeding in chronological order. Though when trying to read this in as a pandas data frame, given the enormous size of the csv file, I run out of memory to allocate to reading in this data on my machine. To avoid this problem, I am trying to read in this csv data in chunks, using the following code:

Chunksize = 2500000 
for chunk in pd.read_csv("my_file.csv", chunksize=Chunksize):     
    print(chunk.head()) 

This works, where I am able to read in my csv file into data frame chunks of 2,500,000 rows each (the last chunk would of course be the remainder of < 2,500,000 rows).

However, I want an explicit reason for my chunk size, as opposed to just a "best judgement" selection, such as the 2,500,000 row chunk size I use above. What I want to figure out is, how can I set my chunk size to be custom based on a given parameter? Specifically, I want each of my chunks to be all of the rows corresponding to unique months in my time series data set. And so let's say this time series dataset has for example 3 years, 5 months, and 9 days of data, and so 3x12 = 36 months + 5 months = 41 months and 9 days of data = 42 chunks, where I have 41 chunks of full month-long second-resolution data and then the last chunk made up of 9 days worth of 1-second resolution data.

How can I augment the chunksize argument in pd.read_csv() to accommodate a custom parameter such as delimiting by months? I am guessing this would involve some sort of manipulation in the timestamp as a datetime object, but I am not sure how to actually specify this delineation, since the chunksize argument just requires a single value. I would appreciate any guidance on this matter, thank you!

Top answer
1 of 1
2
I don't think there is anything to help with that. You could parse the file once and detect when the month changes so you calculate the chunk size - but this would assume all chunk sizes are the same. There are other tools you could use for "larger than memory" datasets e.g. pyarrow and duckdb Example file: >>> print(open("time.csv").read(), end="") time,value 2022-01-23T01:52:10Z,1 2022-02-23T01:52:10Z,2 2022-01-23T01:53:10Z,2 2022-03-23T01:52:10Z,3 2022-04-23T01:52:10Z,4 2022-01-23T01:53:10Z,3 2022-05-23T01:52:10Z,5 2022-01-23T01:53:10Z,4 2022-02-23T01:52:10Z,8 2022-03-23T01:52:10Z,9 2022-01-23T01:53:10Z,5 You could parse the csv file, add a month column with duckdb and convert the result to parquet. parquet is much smaller (compressed) and faster to read than csv. It's also simple to split up data into multiple files. >>> import duckdb, pyarrow ... ... reader = duckdb.connect().execute(""" ... select month(time) as month, * from 'time.csv' ... """).fetch_record_batch() ... ... pyarrow.dataset.write_dataset( ... reader, ... base_dir="/tmp/mydata/", ... format="parquet", ... partitioning=["month"], ... partitioning_flavor="hive" ... ) ... The partitioning=["month"] specifices the list of columns to split by. >>> from pathlib import Path >>> list(Path("/tmp/mydata").rglob("*.parquet")) [PosixPath('/tmp/mydata/month=1/part-0.parquet'), PosixPath('/tmp/mydata/month=2/part-0.parquet'), PosixPath('/tmp/mydata/month=3/part-0.parquet'), PosixPath('/tmp/mydata/month=4/part-0.parquet'), PosixPath('/tmp/mydata/month=5/part-0.parquet')] pandas has a read_parquet() (as does polars) >>> pd.read_parquet("/tmp/mydata/month=1/part-0.parquet") time value 0 2022-01-23 01:52:10 1 1 2022-01-23 01:53:10 2 2 2022-01-23 01:53:10 3 3 2022-01-23 01:53:10 4 4 2022-01-23 01:53:10 5 Or depending on what exactly you're doing with the - you could implement the functionality with duckdb directly.
Find elsewhere
🌐
Architecture-performance
architecture-performance.fr › home › blog › reading a sql table by chunks with pandas
Reading a SQL table by chunks with Pandas | Architecture & Performance
May 16, 2022 - Note that the result of the stream_results and max_row_buffer arguments might differ a lot depending on the database, DBAPI/database adapter. Here we load a table from PostgreSQL with the psycopg2 adapter. It seems that the server side cursor is the default with psycopg2 when using chunksize in pd.read_sql().
🌐
Medium
michael-scherding.medium.com › dynamic-chunk-sizing-with-pandas-and-bigquery-daec4fbc9798
Dynamic chunk sizing with Pandas and Bigquery | by Michaël Scherding | Medium
April 30, 2024 - Dynamic chunk sizing in Pandas offers a practical solution for processing large datasets on machines with limited memory.
🌐
Python⇒Speed
pythonspeed.com › articles › chunking-pandas
Reducing Pandas memory usage #3: Reading in chunks
January 6, 2023 - Reduce Pandas memory usage by loading and then processing a file in chunks rather than all at once, using Pandas’ chunksize option.
🌐
Medium
medium.com › itversity › scalable-data-processing-with-pandas-handling-large-csv-files-in-chunks-b15c5a79a3e3
Scalable Data Processing with Pandas: Handling Large CSV Files in Chunks | by Durga Gadiraju | itversity | Medium
February 12, 2025 - Handling large CSV files can be challenging, but Pandas’ chunk processing feature offers a practical solution. By using the chunksize parameter in read_csv(), you can efficiently load, process, and analyze large datasets without running into memory issues.
🌐
Hashnode
mwass.hashnode.dev › processing-data-in-chunks-using-pandas
Processing data in chunks using pandas
May 25, 2023 - This allows us to read data that cannot otherwise fit in the memory and process it. The chunksize argument, which takes an integer value, is used to make the read_csv function read the dataset in chunks with the size of the value.
🌐
Janakiev
janakiev.com › blog › python-pandas-chunks
Reading and Writing Pandas DataFrames in Chunks - njanakiev
April 3, 2021 - In this short example you will see how to apply this to CSV files with pandas.read_csv. First, create a TextFileReader object for iteration. This won’t load the data until you start iterating over it. Here it chunks the data in DataFrames with 10000 rows each: df_iterator = pd.read_csv( 'input_data.csv.gz', chunksize=10000, compression='gzip') Now, you can use the iterator to load the chunked DataFrames iteratively.
🌐
Reddit
reddit.com › r/learnpython › does pandas.read_csv's chunks have context on the entire dataset?
r/learnpython on Reddit: Does pandas.read_csv's chunks have context on the entire dataset?
April 4, 2024 -

Don't have a sample dataset to play with right now so thought I'd just ask.

Is chunksize basically just a more automated way of setting nrows and going through all the rows in batches, or is there an advantage to chunking where you could do operations that require context on the whole dataset?

For example, for a dataset of transactions, say I need to pull all the biggest orders for each account, and the accounts are not clumped up within the dataset. So when I iterate over the partitions, taking MAX of the order-size column would not be accurate.

Assuming the dataset is too big to fit in memory at once, I can't figure out how I'd go about this (besides using completely different tools/packages, which is out of scope for now).

🌐
Saturn Cloud
saturncloud.io › blog › how-to-write-large-pandas-dataframes-to-csv-file-in-chunks
How to Write Large Pandas Dataframes to CSV File in Chunks | Saturn Cloud Blog
May 1, 2026 - In this article, we explored how to write large Pandas dataframes to CSV file in chunks. Writing a large dataframe to a CSV file in chunks can help to alleviate memory errors and make the process faster. By breaking the dataframe into smaller chunks, we can write the file in segments and avoid memory errors.
🌐
Another Dev Notes
acepor.github.io › 2017 › 08 › 03 › using-chunksize
Using Chunksize in Pandas – Another Dev Notes
August 3, 2017 - In our main task, we set chunksize as 200,000, and it used 211.22MiB memory to process the 10G+ dataset with 9min 54s. the pandas.DataFrame.to_csv() mode should be set as ‘a’ to append chunk results to a single file; otherwise, only the last chunk will be saved.
🌐
Saturn Cloud
saturncloud.io › blog › how-to-efficiently-read-large-csv-files-in-python-pandas
How to Efficiently Read Large CSV Files in Python Pandas | Saturn Cloud Blog
May 1, 2026 - import pandas as pd chunksize = 1000 for chunk in pd.read_csv('large_file.csv', chunksize=chunksize): # process each chunk here
🌐
Like Geeks
likegeeks.com › home › python › pandas › pandas read_sql with chunksize: unlock parallel processing
Pandas read_sql with chunksize: Unlock Parallel Processing
Here’s how you can use chunksize with a sample SQL query: import pandas as pd from sqlalchemy import create_engine engine = create_engine('sqlite:///sample_data.db') query = """ SELECT user_id, session_start, session_end, bytes_sent, bytes_received FROM data_usage_logs """ # Use chunksize to read the SQL query in chunks chunk_size = 5000 chunks = pd.read_read_sql(query, engine, chunksize=chunk_size) for chunk in chunks: print(chunk.head()) break
🌐
Pandas
pandas.pydata.org › docs › reference › api › pandas.DataFrame.to_csv.html
pandas.DataFrame.to_csv — pandas 3.0.6 documentation
DataFrame.to_csv(path_or_buf=None, *, sep=',', na_rep='', float_format=None, columns=None, header=True, index=True, index_label=None, mode='w', encoding=None, compression='infer', quoting=None, quotechar='"', lineterminator=None, chunksize=None, date_format=None, doublequote=True, escapechar=None, decimal='.', errors='strict', storage_options=None)[source]# Write object to a comma-separated values (csv) file.