Exact and approximate search
Queries are exact and give perfect recall by default. Add an HNSW or IVFFlat index to trade a little recall for much faster search at scale.
pgvector is an open-source PostgreSQL extension that lets you store embeddings next to your relational data and run exact or approximate nearest-neighbor search with plain SQL.
No separate vector database. Vectors live in Postgres, covered by the same transactions, backups, and joins.
pgvector is an open-source extension for PostgreSQL that adds a vector data type and distance operators so you can store embeddings alongside your relational data and run exact or approximate nearest-neighbor search with ordinary SQL.
It supports L2 (Euclidean), inner-product, cosine, L1 (taxicab), Hamming, and Jaccard distances; single-precision (vector), half-precision (halfvec), binary (bit), and sparse (sparsevec) vectors; and two approximate index types, HNSW and IVFFlat. Because it is a Postgres extension, you keep ACID compliance, point-in-time recovery, JOINs, replication, and the rest of the Postgres feature set — there is no separate vector service to operate.
You use it from any language that has a Postgres client, and you can generate vectors in one language and query them from another. It is commonly used for semantic search, retrieval-augmented generation (RAG), recommendations, image similarity, and deduplication.
Everything below is provided by the extension itself. Each capability links to the corresponding documentation section.
Queries are exact and give perfect recall by default. Add an HNSW or IVFFlat index to trade a little recall for much faster search at scale.
vector for single precision, halfvec for half precision, bit for binary vectors, and sparsevec for sparse vectors — so you can match precision and storage to your workload.
L2, negative inner product, cosine, L1, Hamming, and Jaccard. Mix and match by creating one index per distance function you actually query with.
HNSW for better query performance and IVFFlat for faster builds and lower memory, each with tunable build and query parameters such as m, ef_search, lists, and probes.
ACID transactions, point-in-time recovery, JOINs, replication, and standard backup tooling — because your vectors are rows in a normal Postgres table.
Store halfvec to shrink the working set, use binary quantization with re-ranking for smaller indexes, or index subvectors and concatenated vectors for larger models.
Enable the extension in a database, create a table with a vector column, insert some embeddings, and ask for the nearest neighbors. The full walkthrough is in the documentation.
CREATE EXTENSION vector;.vector(3).'[1,2,3]', or bulk load with COPY.ORDER BY embedding <-> '[...]' LIMIT 5.3 with the number of dimensions your model produces. Sentence-transformer and OpenAI embedding models typically output several hundred to a few thousand dimensions.CREATE EXTENSION vector;
CREATE TABLE items (id bigserial PRIMARY KEY, embedding vector(3));
INSERT INTO items (embedding) VALUES ('[1,2,3]'), ('[4,5,6]');
SELECT * FROM items ORDER BY embedding <-> '[3,1,2]' LIMIT 5;
# Adds pgvector to the official Postgres image
$ docker pull pgvector/pgvector:pg18-trixie
$ docker run --name pgvector -e POSTGRES_PASSWORD=postgres -p 5432:5432 -d pgvector/pgvector:pg18-trixie
# Then, in the database:
=> CREATE EXTENSION vector;
See every method, including Homebrew, APT, conda-forge, and building from source, on the install page.
Without an index, search is exact and returns perfect recall but scans the table. Approximate indexes trade a little recall for speed. Pick the trade-off that fits your data size and latency target.
| Approach | Recall | Build | Memory | Best for |
|---|---|---|---|---|
| No index (exact) | Perfect | None | Low | Small tables, or when you need guaranteed exact results |
| HNSW | High (tunable) | Slower | Higher | Best query performance; can be built before data is loaded |
| IVFFlat | Moderate (tunable) | Faster | Lower | Large tables where build time and memory matter; build after loading data |
Not sure which parameters to use? The index planner generates the exact CREATE INDEX statement and tuning values for your dimensions, row count, and distance metric. Full details, including filtering and iterative index scans, are in the indexing documentation.
Tell it your vector type, distance metric, dimensions, and estimated row count. It produces the correct distance operator and operator class, a ready-to-run CREATE INDEX statement, recommended query tuning values, a storage estimate, and warnings when your dimensions exceed the index limits.
It runs entirely in your browser — nothing you type is uploaded anywhere.
Open the index planner-- vector, cosine distance, 1536 dims, 1,000,000 rows
CREATE INDEX ON items USING hnsw (embedding vector_cosine_ops);
-- Query tuning (default 40; raise for better recall)
SET hnsw.ef_search = 40;
-- Verify with EXPLAIN (ANALYZE, BUFFERS)
A short summary of the most common methods. Every command, plus source builds and troubleshooting, is on the install page.
| Platform | Method | Command |
|---|---|---|
| Linux / macOS | Build from source | make && make install |
| Windows | Build with nmake | nmake /F Makefile.win install |
| Any (Docker) | Official image | docker pull pgvector/pgvector:pg18-trixie |
| macOS | Homebrew | brew install pgvector |
| Debian / Ubuntu | PostgreSQL APT | sudo apt install postgresql-18-pgvector |
| Conda | conda-forge | conda install -c conda-forge pgvector |
18 with your PostgreSQL server version. pgvector supports PostgreSQL 13 and newer.After installing, enable it once per database and confirm the version:
CREATE EXTENSION vector;
SELECT extversion FROM pg_extension WHERE extname = 'vector';
-- expected: 0.8.6
Short, self-contained answers. Deeper explanations live in the documentation.
pgvector is an open-source extension for PostgreSQL that adds a vector data type and distance operators, so you can store embeddings alongside your relational data and run exact or approximate nearest-neighbor search with ordinary SQL. It supports L2, inner-product, cosine, L1, Hamming, and Jaccard distances; single-, half-, binary-, and sparse-precision vectors; and HNSW and IVFFlat indexes.
Yes. pgvector is released under the PostgreSQL License, a permissive open-source license. The source code and issue tracker are on GitHub at github.com/pgvector/pgvector.
pgvector supports PostgreSQL 13 and newer. The current release, v0.8.6, ships Docker image tags for PostgreSQL 13 through 18. Earlier majors (12 and below) are not supported by current versions.
With HNSW or IVFFlat indexes, vector columns can be indexed up to 2,000 dimensions, halfvec up to 4,000 dimensions, and bit up to 64,000 dimensions. Sparse vectors are limited to 1,000 non-zero elements for indexing. To go beyond these limits, use half-precision indexing, binary quantization, subvector indexing, or dimensionality reduction.
HNSW generally gives better query performance for the same recall, but it builds more slowly and uses more memory; it can be created on an empty table. IVFFlat builds faster and uses less memory but has lower query performance, and it should be created after the table has data, with lists and probes tuned for good recall.
No. pgvector is not a separate database — it is a PostgreSQL extension, so vectors are stored in the same database as your other data. There is no separate vector service to run, and you keep Postgres features such as ACID transactions, point-in-time recovery, and JOINs.
You can declare the column with the vector type without a dimension. Indexes, however, require rows with the same number of dimensions, so you index one dimension set using expression and partial indexing, for example CREATE INDEX ON embeddings USING hnsw ((embedding::vector(3)) vector_l2_ops) WHERE (model_id = 123);.
pgvector is not a binary installer you download and run. It is a PostgreSQL extension you obtain as source, a package, or a container image. Get the source from the pgvector GitHub releases, or install it with Docker, Homebrew, APT, Yum, pkg, APK, PGXN, or conda-forge.
Yes. pgvector works from any language with a PostgreSQL client. The project maintains official bindings for Python, JavaScript and TypeScript, Go, Java, Rust, Ruby, and many more; Python and Ruby also work with frameworks such as Django and Rails.
Use the langchain-postgres package, which provides a PGVector vector store backed by pgvector. Install it, point it at a PostgreSQL connection string, and store and query embeddings through LangChain. The integration is listed on the ecosystem page.
Yes. pgvector is a widely used PostgreSQL extension under the PostgreSQL License. Vectors are stored in normal PostgreSQL tables, so it inherits ACID transactions, point-in-time recovery, replication, and backups. Approximate indexes (HNSW and IVFFlat) trade some recall for speed and should be tuned; see the comparison for when a dedicated vector database fits better.
Yes. Many hosted PostgreSQL providers offer pgvector, including Supabase, Neon, Timescale, Amazon RDS, Google Cloud SQL, and Azure Database for PostgreSQL, so you can enable the extension without installing anything. Check your provider's documentation for the version it supports. See the hosted providers list.
Storing, querying, indexing, tuning, filtering, and a complete reference of types, operators, and functions.
Read the docs →Copy-paste SQL for every common task, from enabling the extension to re-ranking with binary quantization.
Open the cheatsheet →Why an index is not used, why a query returns fewer rows, memory errors, and install failures — with fixes.
Find a fix →Source builds, Windows, Docker, Homebrew, APT, Yum, conda-forge, and upgrading, all in one place.
Install pgvector →How an in-database Postgres extension compares with FAISS, Pinecone, Qdrant, and Milvus.
Compare options →Official language bindings, related extensions, ORM integrations, monitoring, and hosted providers.
Browse the ecosystem →