machine learning - Faster kNN Classification Algorithm in Python - Stack Overflow
Really simple and easy K-Nearest Neighbors algorithm from scratch in Python
The k-Nearest Neighbors (kNN) Algorithm in Python – Real Python
I wouldn't remove the sex column. I'm not a biologist but I think it affects the physical measurements and helps the algorithm.
More on reddit.comk-nearest neighbors from scratch in pure Python
Nice project! I also have a repo with "from scratch implementations", but I love to compare it with other approaches :)
More on reddit.comScikit-learn uses a KD Tree or Ball Tree to compute nearest neighbors in O[N log(N)] time. Your algorithm is a direct approach that requires O[N^2] time, and also uses nested for-loops within Python generator expressions which will add significant computational overhead compared to optimized code.
If you'd like to compute weighted k-neighbors classification using a fast O[N log(N)] implementation, you can use sklearn.neighbors.KNeighborsClassifier with the weighted minkowski metric, setting p=2 (for euclidean distance) and setting w to your desired weights. For example:
from sklearn.neighbors import KNeighborsClassifier
model = KNeighborsClassifier(metric='wminkowski', p=2,
metric_params=dict(w=weights))
model.fit(X_train, y_train)
y_predicted = model.predict(X_test)
you can take a look at this great article introducing faiss
Make kNN 300 times faster than Scikit-learn’s in 20 lines!
it is on GPU and developed in CPP behind the seen
import numpy as np
import faiss
class FaissKNeighbors:
def __init__(self, k=5):
self.index = None
self.y = None
self.k = k
def fit(self, X, y):
self.index = faiss.IndexFlatL2(X.shape[1])
self.index.add(X.astype(np.float32))
self.y = y
def predict(self, X):
distances, indices = self.index.search(X.astype(np.float32), k=self.k)
votes = self.y[indices]
predictions = np.array([np.argmax(np.bincount(x)) for x in votes])
return predictions