k-nearest neighbours
Arguments
- x
Numeric matrix or
cudasparseobject with observations in rows.- k
Number of neighbours.
- metric
Exact distance metric,
"euclidean"or"cosine".- device
One of
"auto","cuda", or"cpu".- batch_size
Maximum number of query rows in each dense distance block. Larger batches may be faster but use more memory.
Value
A cuda_knn list with index and distance matrices of size
nrow(x) by k, followed by the selected metric and actual device.
Neighbours in every row are ordered by distance and then row index.
When x has row names, both matrices retain them as query identifiers;
neighbour identities can be recovered with
rownames(result$index)[result$index].
Details
Neighbours are exact: every row is compared with every other row. The observation itself is always excluded. Equal distances are resolved deterministically in favour of the smaller row index.
The implementation constructs at most a
min(batch_size, nrow(x))-by-nrow(x) dense distance block instead of a
complete pairwise distance matrix. The native CUDA backend keeps distance
blocks and deterministic top-k selection on the GPU. A torch backend with
stable-sort support follows the same residency contract. Both transfer only
the final n-by-k index and distance matrices. Compatibility backends
without device-side selection transfer each distance block to the CPU for
stable ordering. On CPU, Euclidean blocks use the same guarded
translated-and-scaled implementation as cuda_distance().