Check out this article.

# import warnings filter
from warnings import simplefilter
# ignore all future warnings
simplefilter(action='ignore', category=FutureWarning)

The simplest way is just to ignore it. The author also discusses how to fix it, so you might want to check that out.

Answer from astro_bear on Stack Overflow
Top answer
1 of 2
1

Check out this article.

# import warnings filter
from warnings import simplefilter
# ignore all future warnings
simplefilter(action='ignore', category=FutureWarning)

The simplest way is just to ignore it. The author also discusses how to fix it, so you might want to check that out.

2 of 2
-1

KNN can be done like this.

import numpy as np
import matplotlib.pyplot as plt
import pandas as pd


url = "https://archive.ics.uci.edu/ml/machine-learning-databases/iris/iris.data"

# Assign colum names to the dataset
names = ['sepal-length', 'sepal-width', 'petal-length', 'petal-width', 'Class']

# Read dataset to pandas dataframe
dataset = pd.read_csv(url, names=names)


dataset.head()


X = dataset.iloc[:, :-1].values
y = dataset.iloc[:, 4].values


from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.20)


from sklearn.preprocessing import StandardScaler
scaler = StandardScaler()
scaler.fit(X_train)

X_train = scaler.transform(X_train)
X_test = scaler.transform(X_test)


from sklearn.neighbors import KNeighborsClassifier
classifier = KNeighborsClassifier(n_neighbors=5, metric='minkowski')
classifier.fit(X_train, y_train)


y_pred = classifier.predict(X_test)


from sklearn.metrics import classification_report, confusion_matrix
print(confusion_matrix(y_test, y_pred))
print(classification_report(y_test, y_pred))

# Result:
                 precision    recall  f1-score   support

    Iris-setosa       1.00      1.00      1.00        13
Iris-versicolor       1.00      0.89      0.94         9
 Iris-virginica       0.89      1.00      0.94         8

       accuracy                           0.97        30
      macro avg       0.96      0.96      0.96        30
   weighted avg       0.97      0.97      0.97        30


error = []
# Calculating error for K values between 1 and 40
for i in range(1, 40):
    knn = KNeighborsClassifier(n_neighbors=i)
    knn.fit(X_train, y_train)
    pred_i = knn.predict(X_test)
    error.append(np.mean(pred_i != y_test))


plt.figure(figsize=(12, 6))
plt.plot(range(1, 40), error, color='red', linestyle='dashed', marker='o',
         markerfacecolor='blue', markersize=10)
plt.title('Error Rate K Value')
plt.xlabel('K Value')
plt.ylabel('Mean Error')

🌐
GitHub
github.com › aladdinpersson › Machine-Learning-Collection › blob › master › ML › algorithms › knn › knn.py
Machine-Learning-Collection/knn.py at master · aladdinpersson/Machine-Learning-Collection
X_test_squared = np.sum(X_test ** 2, axis=1, keepdims=True) X_train_squared = np.sum(self.X_train ** 2, axis=1, keepdims=True) two_X_test_X_train = np.dot(X_test, self.X_train.T) # (Taking sqrt is not necessary: min distance won't change since sqrt is monotone) return np.sqrt( self.eps + X_test_squared - 2 * two_X_test_X_train + X_train_squared.T ·
Author: aladdinpersson
🌐
OneCompiler
onecompiler.com › python › 3ybprbus8
3ybprbus8 - Python - OneCompiler
Write, Run & Share Python code online using OneCompiler's Python online compiler for free. It's one of the robust, feature-rich online compilers for python language, supporting both the versions which are Python 3 and Python 2.7. Getting started with the OneCompiler's Python editor is easy and fast.
🌐
GitHub
github.com › YosefLab › Hotspot › blob › master › hotspot › knn.py
Hotspot/hotspot/knn.py at master · YosefLab/Hotspot
wnorm = weights.sum(axis=1, keepdims=True) wnorm[wnorm == 0] = 1.0 · weights = weights / wnorm · · return weights · · · @jit(nopython=True) def compute_node_degree(neighbors, weights): · D = np.zeros(neighbors.shape[0]) · for i in range(neighbors.shape[0]): for k in range(neighbors.shape[1]): ·
Author: YosefLab
🌐
Lukasheumos
lukasheumos.com › fast-knn-imputation.html
Lukas Heumos - Fast KNN imputation
# Batch prefill NaNs with fallback ... keepdims=True) This eliminates thousands of Python loop iterations and GPU synchronization points. Real-world data rarely has uniform missingness. When too few complete rows exist to build a reliable index, we iteratively exclude the most-NaN-heavy features until reaching a configurable threshold (min_data_ratio). Features that cannot be imputed via KNN fall back ...
🌐
scikit-learn
scikit-learn.org › stable › modules › generated › sklearn.neighbors.KNeighborsClassifier.html
KNeighborsClassifier — scikit-learn 1.9.1 documentation
If None, predictions for all indexed points are used; in this case, points are not considered their own neighbors. This means that knn.fit(X, y).score(None, y) implicitly performs a leave-one-out cross-validation procedure and is equivalent to cross_val_score(knn, X, y, cv=LeaveOneOut()) but typically much faster.
Find elsewhere
🌐
Substack
jaspinders30.substack.com › jaspinder's substack › [ai ml| learning notes day 13| knn & distance-based learning ]
[AI ML| Learning Notes Day 13| kNN & Distance-Based Learning ]
February 2, 2026 - b2 = sum(X_train^2, axis=1, keepdims=True).T → (1, n) dist2 = a2 + b2 - 2 * (X_test @ X_train.T) → (m, n) dist = sqrt(max(dist2, 0)) (numerical guard) Why this is FAANG-relevant: it’s the cleanest way to show you can translate math to vectorized, memory-aware code. All-pairs distances cost: Time: O(m * n * d) Memory: O(m * n) for D · That’s why kNN is typically expensive at inference.
🌐
SciPy
docs.scipy.org › doc › scipy › reference › generated › scipy.stats.mode.html
mode — SciPy v1.18.0 Manual
>>> stats.mode(a, axis=None, keepdims=True) ModeResult(mode=[[3]], count=[[5]]) >>> stats.mode(a, axis=None, keepdims=False) ModeResult(mode=3, count=5)
Top answer
1 of 4
46

Consider a small 2d array:

In [180]: A=np.arange(12).reshape(3,4)
In [181]: A
Out[181]: 
array([[ 0,  1,  2,  3],
       [ 4,  5,  6,  7],
       [ 8,  9, 10, 11]])

Sum across rows; the result is a (3,) array

In [182]: A.sum(axis=1)
Out[182]: array([ 6, 22, 38])

But to sum (or divide) A by the sum requires reshaping

In [183]: A-A.sum(axis=1)
...
ValueError: operands could not be broadcast together with shapes (3,4) (3,) 
In [184]: A-A.sum(axis=1)[:,None]   # turn sum into (3,1)
Out[184]: 
array([[ -6,  -5,  -4,  -3],
       [-18, -17, -16, -15],
       [-30, -29, -28, -27]])

If I use keepdims, "the result will broadcast correctly against" A.

In [185]: A.sum(axis=1, keepdims=True)   # (3,1) array
Out[185]: 
array([[ 6],
       [22],
       [38]])
In [186]: A-A.sum(axis=1, keepdims=True)
Out[186]: 
array([[ -6,  -5,  -4,  -3],
       [-18, -17, -16, -15],
       [-30, -29, -28, -27]])

If I sum the other way, I don't need the keepdims. Broadcasting this sum is automatic: A.sum(axis=0)[None,:]. But there's no harm in using keepdims.

In [190]: A.sum(axis=0)
Out[190]: array([12, 15, 18, 21])    # (4,)
In [191]: A-A.sum(axis=0)
Out[191]: 
array([[-12, -14, -16, -18],
       [ -8, -10, -12, -14],
       [ -4,  -6,  -8, -10]])

If you prefer, these actions might make more sense with np.mean, normalizing the array over columns or rows. In any case it can simplify further math between the original array and the sum/mean.

2 of 4
5

You can keep the dimension with "keepdims=True" if you sum a matrix For example:

import numpy as np
x  = np.array([[1,2,3],[4,5,6]])
x.shape
# (2, 3)

np.sum(x, keepdims=True).shape
# (1, 1)
np.sum(x, keepdims=True)
# array([[21]]) <---the reault is still a 1x1 array

np.sum(x, keepdims=False).shape
# ()
np.sum(x, keepdims=False)
# 21 <--- the result is an integer with no dimesion
🌐
Kaggle
kaggle.com › general › 318045
Role of keepdims in Numpy! | Kaggle
If we use keepdims, "the result will broadcast correctly against" the array on which we are performing the operations. I came across a problem of finding the...
Top answer
1 of 2
102

@Ney @hpaulj is correct, you need to experiment, but I suspect you don't realize that summation for some arrays can occur along axes. Observe the following which reading the documentation

>>> a
array([[0, 0, 0],
       [0, 1, 0],
       [0, 2, 0],
       [1, 0, 0],
       [1, 1, 0]])
>>> np.sum(a, keepdims=True)
array([[6]])
>>> np.sum(a, keepdims=False)
6
>>> np.sum(a, axis=1, keepdims=True)
array([[0],
       [1],
       [2],
       [1],
       [2]])
>>> np.sum(a, axis=1, keepdims=False)
array([0, 1, 2, 1, 2])
>>> np.sum(a, axis=0, keepdims=True)
array([[2, 4, 0]])
>>> np.sum(a, axis=0, keepdims=False)
array([2, 4, 0])

You will notice that if you don't specify an axis (1st two examples), the numerical result is the same, but the keepdims = True returned a 2D array with the number 6, whereas, the second incarnation returned a scalar. Similarly, when summing along axis 1 (across rows), a 2D array is returned again when keepdims = True. The last example, along axis 0 (down columns), shows a similar characteristic... dimensions are kept when keepdims = True.
Studying axes and their properties is critical to a full understanding of the power of NumPy when dealing with multidimensional data.

2 of 2
9

An example showing keepdims in action when working with higher dimensional arrays. Let's see how the shape of the array changes as we do different reductions:

import numpy as np
a = np.random.rand(2,3,4)
a.shape
# => (2, 3, 4)
# Note: axis=0 refers to the first dimension of size 2
#       axis=1 refers to the second dimension of size 3
#       axis=2 refers to the third dimension of size 4

a.sum(axis=0).shape
# => (3, 4)
# Simple sum over the first dimension, we "lose" that dimension 
# because we did an aggregation (sum) over it

a.sum(axis=0, keepdims=True).shape
# => (1, 3, 4)
# Same sum over the first dimension, but instead of "loosing" that 
# dimension, it becomes 1.

a.sum(axis=(0,2)).shape
# => (3,)
# Here we "lose" two dimensions

a.sum(axis=(0,2), keepdims=True).shape
# => (1, 3, 1)
# Here the two dimensions become 1 respectively
🌐
Medium
medium.com › @akp83540 › keep-dim-argument-5ee1829c9ead
Keep Dim Argument. Easy: | by Abhishek Kumar Pandey | Medium
June 12, 2024 - In deep learning, especially when dealing with neural networks, the “KeepDim” argument is used to specify how dimensions should be treated during operations that reduce the size of tensors (multi-dimensional arrays).
🌐
Medium
medium.com › swlh › image-classification-with-k-nearest-neighbours-51b3a289280
Image Classification with K Nearest Neighbours | by Paarth Bir | The Startup | Medium
August 6, 2019 - Image Classification with K Nearest Neighbours K-Nearest Neighbours (k-NN) is a supervised machine learning algorithm i.e. it learns from a labelled training set by taking in the training data X …
🌐
Alibaba Cloud
topic.alibabacloud.com › a › the-meaning-of-keepdims-in-python-numpy_1_29_30286150.html
The meaning of keepdims in Python NumPy
November 13, 2017 - Keepdims is mainly used to maintain the two-dimensional properties of matricesImport= Np.array ([[1,2],[3,4]])# is added by line and retains the second dimension of print(Np.sum (A, Axis=1, keepdims=True)# Add by line, do not maintain the second