v0.8.6 · released 2026-07-29 PostgreSQL License Postgres 13+

Vector similarity search for Postgres

pgvector is an open-source PostgreSQL extension that lets you store embeddings next to your relational data and run exact or approximate nearest-neighbor search with plain SQL.

No separate vector database. Vectors live in Postgres, covered by the same transactions, backups, and joins.

Overview

What is pgvector?

pgvector is an open-source extension for PostgreSQL that adds a vector data type and distance operators so you can store embeddings alongside your relational data and run exact or approximate nearest-neighbor search with ordinary SQL.

It supports L2 (Euclidean), inner-product, cosine, L1 (taxicab), Hamming, and Jaccard distances; single-precision (vector), half-precision (halfvec), binary (bit), and sparse (sparsevec) vectors; and two approximate index types, HNSW and IVFFlat. Because it is a Postgres extension, you keep ACID compliance, point-in-time recovery, JOINs, replication, and the rest of the Postgres feature set — there is no separate vector service to operate.

You use it from any language that has a Postgres client, and you can generate vectors in one language and query them from another. It is commonly used for semantic search, retrieval-augmented generation (RAG), recommendations, image similarity, and deduplication.

Capabilities

What you get

Everything below is provided by the extension itself. Each capability links to the corresponding documentation section.

Exact and approximate search

Queries are exact and give perfect recall by default. Add an HNSW or IVFFlat index to trade a little recall for much faster search at scale.

Four vector types

vector for single precision, halfvec for half precision, bit for binary vectors, and sparsevec for sparse vectors — so you can match precision and storage to your workload.

Six distance operators

L2, negative inner product, cosine, L1, Hamming, and Jaccard. Mix and match by creating one index per distance function you actually query with.

HNSW and IVFFlat indexes

HNSW for better query performance and IVFFlat for faster builds and lower memory, each with tunable build and query parameters such as m, ef_search, lists, and probes.

All the Postgres features

ACID transactions, point-in-time recovery, JOINs, replication, and standard backup tooling — because your vectors are rows in a normal Postgres table.

Scale with quantization

Store halfvec to shrink the working set, use binary quantization with re-ranking for smaller indexes, or index subvectors and concatenated vectors for larger models.

Quickstart

Get running in four steps

Enable the extension in a database, create a table with a vector column, insert some embeddings, and ask for the nearest neighbors. The full walkthrough is in the documentation.

  1. Enable the extension once per database with CREATE EXTENSION vector;.
  2. Create a table with a column typed to your embedding dimensions, e.g. vector(3).
  3. Insert vectors as array-like literals such as '[1,2,3]', or bulk load with COPY.
  4. Query the neighbors with ORDER BY embedding <-> '[...]' LIMIT 5.
Replace 3 with the number of dimensions your model produces. Sentence-transformer and OpenAI embedding models typically output several hundred to a few thousand dimensions.
SQL postgres
CREATE EXTENSION vector;

CREATE TABLE items (id bigserial PRIMARY KEY, embedding vector(3));

INSERT INTO items (embedding) VALUES ('[1,2,3]'), ('[4,5,6]');

SELECT * FROM items ORDER BY embedding <-> '[3,1,2]' LIMIT 5;

See every method, including Homebrew, APT, conda-forge, and building from source, on the install page.

Choose an index

Exact, HNSW, or IVFFlat?

Without an index, search is exact and returns perfect recall but scans the table. Approximate indexes trade a little recall for speed. Pick the trade-off that fits your data size and latency target.

Search approaches at a glance
Approach Recall Build Memory Best for
No index (exact) Perfect None Low Small tables, or when you need guaranteed exact results
HNSW High (tunable) Slower Higher Best query performance; can be built before data is loaded
IVFFlat Moderate (tunable) Faster Lower Large tables where build time and memory matter; build after loading data

Not sure which parameters to use? The index planner generates the exact CREATE INDEX statement and tuning values for your dimensions, row count, and distance metric. Full details, including filtering and iterative index scans, are in the indexing documentation.

Free tool

Index planner & SQL generator

Tell it your vector type, distance metric, dimensions, and estimated row count. It produces the correct distance operator and operator class, a ready-to-run CREATE INDEX statement, recommended query tuning values, a storage estimate, and warnings when your dimensions exceed the index limits.

It runs entirely in your browser — nothing you type is uploaded anywhere.

Open the index planner
Output example sql
-- vector, cosine distance, 1536 dims, 1,000,000 rows
CREATE INDEX ON items USING hnsw (embedding vector_cosine_ops);

-- Query tuning (default 40; raise for better recall)
SET hnsw.ef_search = 40;

-- Verify with EXPLAIN (ANALYZE, BUFFERS)
Install

Ways to install pgvector

A short summary of the most common methods. Every command, plus source builds and troubleshooting, is on the install page.

Common installation methods
Platform Method Command
Linux / macOSBuild from sourcemake && make install
WindowsBuild with nmakenmake /F Makefile.win install
Any (Docker)Official imagedocker pull pgvector/pgvector:pg18-trixie
macOSHomebrewbrew install pgvector
Debian / UbuntuPostgreSQL APTsudo apt install postgresql-18-pgvector
Condaconda-forgeconda install -c conda-forge pgvector
Replace 18 with your PostgreSQL server version. pgvector supports PostgreSQL 13 and newer.

After installing, enable it once per database and confirm the version:

SQL postgres
CREATE EXTENSION vector;
SELECT extversion FROM pg_extension WHERE extname = 'vector';
-- expected: 0.8.6
FAQ

Frequently asked questions

Short, self-contained answers. Deeper explanations live in the documentation.

What is pgvector?

pgvector is an open-source extension for PostgreSQL that adds a vector data type and distance operators, so you can store embeddings alongside your relational data and run exact or approximate nearest-neighbor search with ordinary SQL. It supports L2, inner-product, cosine, L1, Hamming, and Jaccard distances; single-, half-, binary-, and sparse-precision vectors; and HNSW and IVFFlat indexes.

Is pgvector free and open source?

Yes. pgvector is released under the PostgreSQL License, a permissive open-source license. The source code and issue tracker are on GitHub at github.com/pgvector/pgvector.

Which PostgreSQL versions does pgvector support?

pgvector supports PostgreSQL 13 and newer. The current release, v0.8.6, ships Docker image tags for PostgreSQL 13 through 18. Earlier majors (12 and below) are not supported by current versions.

How many dimensions can I index in pgvector?

With HNSW or IVFFlat indexes, vector columns can be indexed up to 2,000 dimensions, halfvec up to 4,000 dimensions, and bit up to 64,000 dimensions. Sparse vectors are limited to 1,000 non-zero elements for indexing. To go beyond these limits, use half-precision indexing, binary quantization, subvector indexing, or dimensionality reduction.

Should I use HNSW or IVFFlat?

HNSW generally gives better query performance for the same recall, but it builds more slowly and uses more memory; it can be created on an empty table. IVFFlat builds faster and uses less memory but has lower query performance, and it should be created after the table has data, with lists and probes tuned for good recall.

Does pgvector need a separate server or database?

No. pgvector is not a separate database — it is a PostgreSQL extension, so vectors are stored in the same database as your other data. There is no separate vector service to run, and you keep Postgres features such as ACID transactions, point-in-time recovery, and JOINs.

Can I store vectors with different dimensions in the same column?

You can declare the column with the vector type without a dimension. Indexes, however, require rows with the same number of dimensions, so you index one dimension set using expression and partial indexing, for example CREATE INDEX ON embeddings USING hnsw ((embedding::vector(3)) vector_l2_ops) WHERE (model_id = 123);.

How do I download pgvector?

pgvector is not a binary installer you download and run. It is a PostgreSQL extension you obtain as source, a package, or a container image. Get the source from the pgvector GitHub releases, or install it with Docker, Homebrew, APT, Yum, pkg, APK, PGXN, or conda-forge.

Can I use pgvector from Python or another language?

Yes. pgvector works from any language with a PostgreSQL client. The project maintains official bindings for Python, JavaScript and TypeScript, Go, Java, Rust, Ruby, and many more; Python and Ruby also work with frameworks such as Django and Rails.

How do I use pgvector with LangChain?

Use the langchain-postgres package, which provides a PGVector vector store backed by pgvector. Install it, point it at a PostgreSQL connection string, and store and query embeddings through LangChain. The integration is listed on the ecosystem page.

Is pgvector production-ready?

Yes. pgvector is a widely used PostgreSQL extension under the PostgreSQL License. Vectors are stored in normal PostgreSQL tables, so it inherits ACID transactions, point-in-time recovery, replication, and backups. Approximate indexes (HNSW and IVFFlat) trade some recall for speed and should be tuned; see the comparison for when a dedicated vector database fits better.

Does pgvector work on managed Postgres such as Supabase?

Yes. Many hosted PostgreSQL providers offer pgvector, including Supabase, Neon, Timescale, Amazon RDS, Google Cloud SQL, and Azure Database for PostgreSQL, so you can enable the extension without installing anything. Check your provider's documentation for the version it supports. See the hosted providers list.