import numpy as np
import pandas as pd
import scipy.sparse as sparse
df = pd.DataFrame(np.arange(1,10).reshape(3,3))
arr = sparse.coo_matrix(([1,1,1], ([0,1,2], [1,2,0])), shape=(3,3))
df['newcol'] = arr.toarray().tolist()
print(df)
yields
0 1 2 newcol
0 1 2 3 [0, 1, 0]
1 4 5 6 [0, 0, 1]
2 7 8 9 [1, 0, 0]
Answer from unutbu on Stack Overflowimport numpy as np
import pandas as pd
import scipy.sparse as sparse
df = pd.DataFrame(np.arange(1,10).reshape(3,3))
arr = sparse.coo_matrix(([1,1,1], ([0,1,2], [1,2,0])), shape=(3,3))
df['newcol'] = arr.toarray().tolist()
print(df)
yields
0 1 2 newcol
0 1 2 3 [0, 1, 0]
1 4 5 6 [0, 0, 1]
2 7 8 9 [1, 0, 0]
df = pd.DataFrame(np.arange(1,10).reshape(3,3))
df['newcol'] = pd.Series(your_2d_numpy_array)
What is the proper way to insert an array into a pandas data frame?
python - how to append numpy array to a pandas dataframe - Stack Overflow
How can I store a numpy matrix in a pandas dataframe?
Numpy array vs Pandas Dataframe
@BalrogOfMoira is that really faster than simply creating the dataframe to append?
df.append(pd.DataFrame(arr.reshape(1,-1), columns=list(df)), ignore_index=True)
Otherwise @Wonton you could simply concatenate arrays then write to a data frame, which could the be appended to the original data frame.
This will work:
df.append(pd.DataFrame(arr).T)
I am doing something like this.
df = pd.DataFrame({
'a':[1,1,1,1],
'b':['foo','foo','bar','bar'],
'c':[3,3,None,3],
'd':['foo','foo',None,'bar']})
df.loc[2, ['c','d']] = list([3, ['bar','foo']])Before:
a b c d 0 1 foo 3.0 foo 1 1 foo 3.0 foo 2 1 bar NaN None 3 1 bar 3.0 bar
After:
a b c d 0 1 foo 3.0 foo 1 1 foo 3.0 foo 2 1 bar 3.0 [bar, foo] 3 1 bar 3.0 bar
But I get the following warning:
VisibleDeprecationWarning: Creating an ndarray from ragged nested sequences (which is a list-or-tuple of lists-or-tuples-or ndarrays with different lengths or shapes) is deprecated. If you meant to do this, you must specify 'dtype=object' when creating the ndarray.
I've always gotten this message and just ignored it, but I want to do thing correctly so when this deprecation take place in the future I don't get caught out. How should I go about this so I am doing it the correct way so I don't get this warning?
Assign the predictions to a variable and then extract the columns from the variable to be assigned to the pandas dataframe cols. If x is the 2D numpy array with predictions,
x = sentiment_model.predict_proba(test_matrix)
then you can do,
test_data['prediction0'] = x[:,0]
test_data['prediction1'] = x[:,1]
import numpy as np
import pandas as pd
df = pd.DataFrame(
np.arange(10).reshape(5, 2), columns=['a', 'b'])
print('df:', df, sep='\n')
arr = np.arange(100, 104).reshape(2, 2)
print('array to append:', arr, sep='\n')
df = df.append(pd.DataFrame(arr, columns=df.columns), ignore_index=True)
print('df:', df, sep='\n')
output
df:
a b
0 0 1
1 2 3
2 4 5
3 6 7
4 8 9
array to append:
[[100 101]
[102 103]]
df:
a b
0 0 1
1 2 3
2 4 5
3 6 7
4 8 9
5 100 101
6 102 103