annoy
Approximate Nearest Neighbors in C++/Python optimized for memory usage and loading/saving to disk.
Decision gist · record as of 2026-08-14
Yes, if you need approximate nearest-neighbor search and want to share indexes across processes. The library is production-stable, permissively licensed, and has no known vulnerabilities. Install friction is medium due to C++ compilation, and the package is aging (last release mid-2023) but not abandoned. Best suited for use cases where approximate results are acceptable, memory efficiency matters, and you can tolerate the build step.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires a C++ compiler and build tools to compile the native extension during installation.
- Medium install friction due to compiled C++ components requiring a build step.
- The package is aging (last release 2023-06-14, 1157 days ago) but remains actively maintained with no recent commits blocked; the repository is not archived and has substantial community adoption (14285 stars).
License · maintenance · safety
Apache License 2.0 (permissive) — Licensed under Apache License 2.0, a permissive license that allows commercial and private use with minimal restrictions, making it safe for most production deployments.
last release 2023-06-14 (1157 days) · last repo commit 2025-10-29 · 14,285 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 1,232,480 downloads/mo, #4,184 on PyPI
Alternatives
Verify before relying
pip install annoy
from annoy import AnnoyIndex
import random
f = 40
t = AnnoyIndex(f, 'angular')
for i in range(1000):
v = [random.gauss(0, 1) for z in range(f)]
t.add_item(i, v)
t.build(10)
t.save('test.ann')
u = AnnoyIndex(f, 'angular')
u.load('test.ann')
neighbors = u.get_nns_by_item(0, 100)- Whether the package still builds reliably against modern Python versions beyond 3.9 (classifiers list ends at 3.9)
- Current performance characteristics compared to newer approximate nearest neighbor libraries
- Whether memory-mapping behavior is consistent across all supported operating systems
What it is and what it does
Annoy is a nearest-neighbor search library built on C++ with Python bindings, designed to find points in space closest to a query point across many dimensions. It trades some accuracy for speed by building forests of random trees, allowing you to tune the tradeoff between search speed and precision at query time. The core innovation is that indexes are stored as static, memory-mapped files on disk—meaning you build an index once, save it, and then any number of processes can load and query it simultaneously without rebuilding or copying data.
The library supports multiple distance metrics (Euclidean, Manhattan, cosine, Hamming, and dot product) and is optimized for memory efficiency, making it practical for large datasets that would otherwise exhaust RAM. It was originally built at Spotify for music recommendation systems working with millions of high-dimensional vectors. You create an index by adding vectors with integer IDs, build a forest of trees to structure the search space, save the index to disk, and then load it (via mmap) in any process to perform fast approximate nearest-neighbor queries.
Use it for
- Build a recommendation engine that finds similar users or items by querying pre-built vector indexes across multiple application servers
- Index millions of embeddings from a machine learning model and serve nearest-neighbor queries in production without rebuilding indexes
- Share a single large nearest-neighbor index across many parallel processes or Hadoop jobs without duplicating data in memory
- Implement similarity search for music, images, or text embeddings where approximate results are acceptable and speed matters more than perfect accuracy
- Reduce memory footprint for high-dimensional vector search by using disk-based indexes that are memory-mapped on demand
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you need approximate nearest-neighbor search and want to share indexes across processes.
The library is production-stable, permissively licensed, and has no known vulnerabilities. Install friction is medium due to C++ compilation, and the package is aging (last release mid-2023) but not abandoned. Best suited for use cases where approximate results are acceptable, memory efficiency matters, and you can tolerate the build step.
Install
annoy on PyPI
Before you install
Medium install friction due to compiled C++ components requiring a build step. The package is aging (last release 2023-06-14, 1157 days ago) but remains actively maintained with no recent commits blocked; the repository is not archived and has substantial community adoption (14285 stars).
Requires a C++ compiler and build tools to compile the native extension during installation.
License in practice
Licensed under Apache License 2.0, a permissive license that allows commercial and private use with minimal restrictions, making it safe for most production deployments.
Quickstart
pip install annoy
from annoy import AnnoyIndex
import random
f = 40
t = AnnoyIndex(f, 'angular')
for i in range(1000):
v = [random.gauss(0, 1) for z in range(f)]
t.add_item(i, v)
t.build(10)
t.save('test.ann')
u = AnnoyIndex(f, 'angular')
u.load('test.ann')
neighbors = u.get_nns_by_item(0, 100)
Verify before relying
- Whether the package still builds reliably against modern Python versions beyond 3.9 (classifiers list ends at 3.9)
- Current performance characteristics compared to newer approximate nearest neighbor libraries
- Whether memory-mapping behavior is consistent across all supported operating systems
Package facts
| License | Apache License 2.0 permissive |
| Python support | Not specified |
| Install friction | Medium. Platform-specific wheel |
| Runtime dependencies | None |
| Maintenance | Aging 1,157 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 1,232,480 / month, #4,184 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 5 - Production/StableProgramming Language :: PythonProgramming Language :: Python :: 2.6Programming Language :: Python :: 2.7Programming Language :: Python :: 3.3Programming Language :: Python :: 3.4Programming Language :: Python :: 3.5Programming Language :: Python :: 3.6Programming Language :: Python :: 3.7Programming Language :: Python :: 3.8Programming Language :: Python :: 3.9 |
Evidence: annoy-1.17.3-cp310-cp310-macosx_11_0_arm64.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “nearest neighbors in high dimensions”
- annoyAnnoy searches for approximate nearest neighbors in high-dimensional…
- pynndescentPyNNDescent builds approximate nearest neighbor search indexes using…
- pyspark-hnswProvides a PySpark-compatible implementation of the Hierarchical…
Give your agent the search over MCP, or paste the wish link into any chat.
More Scientific/Engineering packages
NumPy provides an N-dimensional array object and a comprehensive suite of mathematical, linear algebra, Fourier transform, and random number functions for scientific computing in Python.
pandas provides fast, flexible data structures (Series and DataFrame) for loading, cleaning, transforming, and analyzing labeled or relational data in Python.
scipy provides numerical algorithms for mathematics, science, and engineering—including optimization, integration, linear algebra, Fourier transforms, signal and image processing, and ODE solvers—built on numpy arrays.
scikit-learn provides a comprehensive Python library for supervised and unsupervised machine learning, including classification, regression, clustering, dimensionality reduction, and model evaluation tools built on NumPy and SciPy.
Install it if you need to train, evaluate, or deploy supervised or unsupervised learning models.
dill extends Python's pickle module to serialize and deserialize a much wider range of Python objects, including functions, lambdas, classes, and interpreter sessions, to byte streams for storage or network transmission.
Multiprocess is an enhanced fork of Python's standard multiprocessing library that uses dill for better serialization, allowing you to spawn processes with a threading-like API and share complex objects between them.
Install it if you use multiprocessing and encounter pickle serialization limits with lambdas or complex objects.
See also pynndescent · voyager · nmslib · usearch · scann · pyspark-hnsw · libcuvs-cu12 · simsimd · faiss-gpu · smmap