pgvector documentation

Everything you need to store embeddings in PostgreSQL and search them with pgvector: enabling the extension, storing and querying vectors, building HNSW and IVFFlat indexes, tuning recall, and the complete reference of types, operators, and functions.

TL;DR
  • pgvector adds vector, halfvec, bit, and sparsevec types and six distance operators to PostgreSQL 13+.
  • Enable it once per database with CREATE EXTENSION vector;, then query with ORDER BY embedding <-> '[...]' LIMIT n.
  • Queries are exact by default. Add an HNSW or IVFFlat index for approximate search that trades recall for speed.
  • HNSW: tune m, ef_construction, and hnsw.ef_search. IVFFlat: build after loading data and tune lists and ivfflat.probes.
  • With an approximate index, filters are applied after the scan — enable iterative index scans when queries return too few rows.
  • Current version is v0.8.6 (released 2026-07-29) under the PostgreSQL License. See the changelog for details.

Requirements & setup

pgvector is a PostgreSQL extension and supports PostgreSQL 13 and newer. Current versions no longer support PostgreSQL 12 and below. Install the extension, then enable it once per database where you want to use it.

Installation methods include building from source, Docker, Homebrew, PGXN, APT, Yum, pkg, APK, conda-forge, and Postgres.app. The install guide covers all of them, and the index planner generates the CREATE INDEX statement for your data.

SQL postgres
CREATE EXTENSION vector;
If you see type "vector" does not exist, the extension is not installed or not enabled in the current database. Run CREATE EXTENSION vector; in that database first.

Quickstart

Create a table with a fixed-dimension vector column, insert a few vectors, and request the nearest neighbors. The distance operator in the ORDER BY clause determines the metric.

SQL postgres
CREATE TABLE items (id bigserial PRIMARY KEY, embedding vector(3));

INSERT INTO items (embedding) VALUES ('[1,2,3]'), ('[4,5,6]');

SELECT * FROM items ORDER BY embedding <-> '[3,1,2]' LIMIT 5;

The vector(3) column stores 3-dimensional vectors. Replace 3 with the output size of your embedding model — for example 768, 1536, or 3072.

Storing vectors

Create a new table with a vector column, or add one to an existing table. Besides vector, the same operations work for halfvec, bit, and sparsevec.

SQL postgres
CREATE TABLE items (id bigserial PRIMARY KEY, embedding vector(3));

ALTER TABLE items ADD COLUMN embedding vector(3);

Insert vectors as array-like literals, or bulk-load them with COPY using the binary format, which is much faster for large datasets. Upserts use the standard ON CONFLICT syntax.

SQL postgres
INSERT INTO items (embedding) VALUES ('[1,2,3]'), ('[4,5,6]');

COPY items (embedding) FROM STDIN WITH (FORMAT BINARY);

INSERT INTO items (id, embedding) VALUES (1, '[1,2,3]'), (2, '[4,5,6]')
    ON CONFLICT (id) DO UPDATE SET embedding = EXCLUDED.embedding;

UPDATE items SET embedding = '[1,2,3]' WHERE id = 1;
DELETE FROM items WHERE id = 1;

Querying

Ask for the nearest neighbors by ordering by the distance operator and limiting the result. This is the shape the query planner understands, and the same shape an index accelerates.

SQL postgres
-- nearest neighbors to a vector
SELECT * FROM items ORDER BY embedding <-> '[3,1,2]' LIMIT 5;

-- nearest neighbors to a row (excluding the row itself)
SELECT * FROM items WHERE id != 1
    ORDER BY embedding <-> (SELECT embedding FROM items WHERE id = 1) LIMIT 5;

-- rows within a distance (combine with ORDER BY + LIMIT to use an index)
SELECT * FROM items WHERE embedding <-> '[3,1,2]' < 5;

-- distance value, inner product, and cosine similarity
SELECT embedding <-> '[3,1,2]' AS distance FROM items;
SELECT (embedding <#> '[3,1,2]') * -1 AS inner_product FROM items;
SELECT 1 - (embedding <=> '[3,1,2]') AS cosine_similarity FROM items;

The inner-product operator <#> returns the negative inner product, because PostgreSQL only supports ascending order index scans on operators. Multiply by -1 to get the actual inner product. Cosine similarity is 1 - cosine distance.

Vector columns also support the aggregates AVG and SUM, individually and per group.

SQL postgres
SELECT AVG(embedding) FROM items;
SELECT category_id, AVG(embedding) FROM items GROUP BY category_id;

Indexing overview

By default, pgvector performs exact nearest-neighbor search, which provides perfect recall but scans every row. Adding an index enables approximate nearest-neighbor search, trading some recall for speed. Unlike typical indexes, you will see different results for queries after adding an approximate index.

Two index types are supported:

  • HNSW — a multilayer graph with better query performance (speed/recall trade-off), but slower builds and higher memory use. It can be created on an empty table because there is no training step.
  • IVFFlat — partitions vectors into lists; faster builds and lower memory, but lower query performance. Create it after the table has data.

Add an index for each distance function you want to use, matching the operator class to the type and metric. For example:

SQL postgres
CREATE INDEX ON items USING hnsw (embedding vector_l2_ops);
CREATE INDEX ON items USING hnsw (embedding vector_ip_ops);
CREATE INDEX ON items USING hnsw (embedding vector_cosine_ops);
CREATE INDEX ON items USING hnsw (embedding vector_l1_ops);

Use halfvec_* operator classes for halfvec, bit_hamming_ops or bit_jaccard_ops for bit, and sparsevec_* for sparsevec. The full list is in the operator class reference.

HNSW

An HNSW index creates a multilayer graph. It generally gives better query performance than IVFFlat for the same recall, at the cost of slower builds and more memory. There is no training step, so you can create it before loading data.

SQL postgres
-- Hamming and Jaccard index the bit type
CREATE INDEX ON items USING hnsw (embedding bit_hamming_ops);
CREATE INDEX ON items USING hnsw (embedding bit_jaccard_ops);

-- Index options (defaults shown); higher ef_construction = better recall, slower build
CREATE INDEX ON items USING hnsw (embedding vector_l2_ops) WITH (m = 16, ef_construction = 64);

Index options

  • m — the maximum number of connections per layer (default 16).
  • ef_construction — the size of the dynamic candidate list used to build the graph (default 64).

Query options

The search candidate list is controlled by hnsw.ef_search (default 40). A higher value gives better recall at the cost of speed. Use SET LOCAL inside a transaction to apply it to a single query.

SQL postgres
SET hnsw.ef_search = 100;

BEGIN;
SET LOCAL hnsw.ef_search = 100;
SELECT ...;
COMMIT;

Build time

HNSW builds are significantly faster when the graph fits in maintenance_work_mem. When it no longer fits, PostgreSQL emits a notice. As with other index types, building after loading your initial data is faster, and increasing parallel workers helps.

SQL postgres
SET maintenance_work_mem = '8GB';
SET max_parallel_maintenance_workers = 7;  -- plus leader

-- check progress
SELECT phase, round(100.0 * blocks_done / nullif(blocks_total, 0), 1) AS "%"
    FROM pg_stat_progress_create_index;

The HNSW phases are initializing and loading tuples.

IVFFlat

An IVFFlat index divides vectors into lists and, at query time, searches a subset of the lists closest to the query vector. It builds faster and uses less memory than HNSW but has lower query performance.

Three keys to good recall: create the index after the table has data, choose an appropriate number of lists, and set an appropriate number of probes at query time.

SQL postgres
CREATE INDEX ON items USING ivfflat (embedding vector_l2_ops) WITH (lists = 100);
CREATE INDEX ON items USING ivfflat (embedding vector_ip_ops) WITH (lists = 100);
CREATE INDEX ON items USING ivfflat (embedding vector_cosine_ops) WITH (lists = 100);
CREATE INDEX ON items USING ivfflat (embedding bit_hamming_ops) WITH (lists = 100);
Create IVFFlat indexes after the table has data. An index built on too few rows can return poor results.

Choose lists starting around rows / 1000 for up to 1 million rows, and sqrt(rows) above 1 million. At query time, start with sqrt(lists) probes. The index planner calculates these for you.

SQL postgres
SET ivfflat.probes = 10;

BEGIN;
SET LOCAL ivfflat.probes = 10;
SELECT ...;
COMMIT;

For build time, increase parallel workers as with HNSW. IVFFlat phases are initializing, performing k-means, assigning tuples, and loading tuples (the percentage is only shown during loading tuples):

SQL postgres
SELECT phase, round(100.0 * tuples_done / nullif(tuples_total, 0), 1) AS "%"
    FROM pg_stat_progress_create_index;

Filtering

There are several ways to combine nearest-neighbor search with a WHERE clause. A good starting point is a B-tree index on the filter column, which can give fast exact search in many cases.

SQL postgres
SELECT * FROM items WHERE category_id = 123 ORDER BY embedding <-> '[3,1,2]' LIMIT 5;

-- exact index on the filter column
CREATE INDEX ON items (category_id);

-- multicolumn index
CREATE INDEX ON items (location_id, category_id);

Exact indexes work well when a condition matches a low percentage of rows. With an approximate index, filtering is applied after the index is scanned. If a condition matches 10% of rows, an HNSW query with the default hnsw.ef_search of 40 returns only about 4 matching rows on average. For more results, enable iterative index scans.

If you filter by only a few distinct values, consider a partial index. If you filter by many different values, consider partitioning.
SQL postgres
-- partial index for a single value
CREATE INDEX ON items USING hnsw (embedding vector_l2_ops) WHERE (category_id = 123);

-- partitioning for many values
CREATE TABLE items (embedding vector(3), category_id int) PARTITION BY LIST(category_id);

Iterative index scans

Starting with 0.8.0, you can enable iterative index scans. When a query filters results, the scan automatically continues through the index until enough rows are found (or it reaches hnsw.max_scan_tuples or ivfflat.max_probes).

Iterative scans support strict or relaxed ordering. Strict keeps results exactly ordered by distance; relaxed allows results to be slightly out of order but gives better recall.

SQL postgres
SET hnsw.iterative_scan = strict_order;
SET hnsw.iterative_scan = relaxed_order;
SET ivfflat.iterative_scan = relaxed_order;

-- HNSW scan limits
SET hnsw.max_scan_tuples = 20000;
SET hnsw.scan_mem_multiplier = 2;

-- IVFFlat scan limit
SET ivfflat.max_probes = 100;

With relaxed ordering, you can recover strict ordering using a materialized CTE. For queries that filter by distance, place the distance filter outside a materialized CTE for best performance.

SQL postgres
WITH relaxed_results AS MATERIALIZED (
    SELECT id, embedding <-> '[1,2,3]' AS distance FROM items
    WHERE category_id = 123 ORDER BY distance LIMIT 5
) SELECT * FROM relaxed_results ORDER BY distance + 0;  -- + 0 needed for Postgres 17+

WITH nearest_results AS MATERIALIZED (
    SELECT id, embedding <-> '[1,2,3]' AS distance FROM items ORDER BY distance LIMIT 5
) SELECT * FROM nearest_results WHERE distance < 5 ORDER BY distance;

hnsw.max_scan_tuples defaults to 20,000 and does not affect the initial scan. hnsw.scan_mem_multiplier defaults to 1 and is a multiple of work_mem — try increasing it if raising max_scan_tuples does not improve recall. If ivfflat.max_probes is lower than ivfflat.probes, the latter is used.

Half-precision vectors

The halfvec type stores half-precision vectors, halving storage versus vector while allowing up to 4,000 indexed dimensions. You can also index a full-precision column at half precision.

SQL postgres
CREATE TABLE items (id bigserial PRIMARY KEY, embedding halfvec(3));

-- index a full-precision column at half precision
CREATE INDEX ON items USING hnsw ((embedding::halfvec(3)) halfvec_l2_ops);

SELECT * FROM items ORDER BY embedding::halfvec(3) <-> '[1,2,3]' LIMIT 5;

Binary vectors & quantization

The bit type stores binary vectors and supports Hamming (<~>) and Jaccard (<%>) distances. Binary quantization compresses a vector index to a fraction of its size; re-rank with the original vectors for better recall.

SQL postgres
CREATE TABLE items (id bigserial PRIMARY KEY, embedding bit(3));
INSERT INTO items (embedding) VALUES ('000'), ('111');
SELECT * FROM items ORDER BY embedding <~> '101' LIMIT 5;

-- binary quantization with re-ranking
CREATE INDEX ON items USING hnsw ((binary_quantize(embedding)::bit(3)) bit_hamming_ops);

SELECT * FROM (
    SELECT * FROM items
    ORDER BY binary_quantize(embedding)::bit(3) <~> binary_quantize('[1,-2,3]') LIMIT 20
) ORDER BY embedding <=> '[1,-2,3]' LIMIT 5;

Sparse vectors

The sparsevec type stores sparse vectors efficiently. The literal format is {index1:value1,index2:value2}/dimensions, and indices start at 1 like SQL arrays.

SQL postgres
CREATE TABLE items (id bigserial PRIMARY KEY, embedding sparsevec(5));
INSERT INTO items (embedding) VALUES ('{1:1,3:2,5:3}/5'), ('{1:4,3:5,5:6}/5');
SELECT * FROM items ORDER BY embedding <-> '{1:3,3:1,5:2}/5' LIMIT 5;

Indexing subvectors

For models that support it, you can index a subvector with expression indexing and re-rank with the full vectors.

SQL postgres
CREATE INDEX ON items USING hnsw ((subvector(embedding, 1, 3)::vector(3)) vector_cosine_ops);

SELECT * FROM (
    SELECT * FROM items
    ORDER BY subvector(embedding, 1, 3)::vector(3) <=> subvector('[1,2,3,4,5]'::vector, 1, 3) LIMIT 20
) ORDER BY embedding <=> '[1,2,3,4,5]' LIMIT 5;

Hybrid search

Combine vector search with PostgreSQL full-text search. You can fuse the two result sets with Reciprocal Rank Fusion or a cross-encoder. The upstream pgvector-python repository includes worked examples of both.

SQL postgres
SELECT id, content FROM items, plainto_tsquery('hello search') query
    WHERE textsearch @@ query ORDER BY ts_rank_cd(textsearch, query) DESC LIMIT 5;

Multitenancy

When multiple tenants share one approximate index, one tenant's vectors can affect recall and speed for the others. For tenant isolation, use list partitioning or separate tables.

SQL postgres
CREATE TABLE items (customer_id int, embedding vector(3)) PARTITION BY LIST(customer_id);

Performance & scaling

Use a tuning tool such as PgTune to set initial server parameters. shared_buffers should typically be about 25% of server memory. Check settings with SHOW config_file; and SHOW shared_buffers;, and restart PostgreSQL for changes to take effect.

  • Storing: use halfvec instead of vector for a smaller working set.
  • Loading: use COPY for bulk loads and add indexes after loading the initial data.
  • Indexing: raise maintenance_work_mem and parallel workers; create indexes CONCURRENTLY in production to avoid blocking writes.
  • Querying: use EXPLAIN (ANALYZE, BUFFERS) to debug. If vectors are normalized to length 1 (like OpenAI embeddings), use inner product for best performance.
SQL postgres
EXPLAIN (ANALYZE, BUFFERS) SELECT * FROM items ORDER BY embedding <-> '[3,1,2]' LIMIT 5;

SET max_parallel_workers_per_gather = 4;   -- speed up exact search without an index
CREATE INDEX CONCURRENTLY ...;              -- avoid blocking writes

-- vacuuming is faster after a reindex
REINDEX INDEX CONCURRENTLY index_name;
VACUUM table_name;

Scaling: shrink the working set with halfvec and use binary quantization with re-ranking for indexes. Scale vertically with more memory, CPU, and storage, or horizontally with replicas. For sharding, PostgreSQL-compatible options include Citus and PgDog.

Monitoring

Use existing tools such as pg_stat_statements or PgHero to monitor performance. Check index size with pg_relation_size, and measure recall by comparing approximate results against exact results.

SQL postgres
SELECT pg_size_pretty(pg_relation_size('index_name'));

BEGIN;
SET LOCAL enable_indexscan = off;  -- force exact search to compare recall
SELECT ...;
COMMIT;

Languages

Use pgvector from any language with a PostgreSQL client — you can even generate vectors in one language and query them from another. The project maintains bindings for many languages (Ada, C, C++, C#, Crystal, D, Dart, Elixir, Erlang, Go, Java/Kotlin/Scala, JavaScript/TypeScript, Julia, Lua, Nim, OCaml, Perl, PHP, Python, R, Ruby, Rust, Swift, and more) under the pgvector GitHub organization. See the ecosystem page for the directory.

Upgrading

Install the latest version using the same method you originally used, then run ALTER EXTENSION vector UPDATE; in each database you want to upgrade. You can check the current version in a database with the query below.

SQL postgres
ALTER EXTENSION vector UPDATE;
SELECT extversion FROM pg_extension WHERE extname = 'vector';  -- expected: 0.8.6

Installation

The install guide covers every method: building from source on Linux and macOS, building on Windows with nmake, Docker, Homebrew, PGXN, APT, Yum, pkg, APK, conda-forge, Postgres.app, the GitHub Actions action, and hosted providers. It also lists the common installation errors and fixes.

Troubleshooting

See the troubleshooting page for step-by-step fixes. The most common questions are covered there:

  • Why a query isn’t using the index.
  • Why results decrease after adding an HNSW or IVFFlat index.
  • Why a parallel table scan isn’t used.
  • Memory notices during HNSW builds and installation errors.

Reference: types & limits

Each element of a vector or halfvec must be finite (no NaN, Infinity, or -Infinity).

Vector data types
TypeStorage per rowMax dimensionsMax indexed dimensions
vector4 * dimensions + 8 bytes16,0002,000
halfvec2 * dimensions + 8 bytes16,0004,000
bitdimensions / 8 + 8 bytes—64,000
sparsevec8 * non-zero + 16 bytesup to 16,000 non-zero1,000 non-zero

Approximate index limits apply to HNSW and IVFFlat. IVFFlat does not support sparsevec. To index beyond these limits, use half-precision indexing, binary quantization, subvector indexing, or dimensionality reduction.

Mixed dimensions and higher precision

SQL postgres
-- store vectors of different dimensions in one column (index per dimension set)
CREATE TABLE embeddings (model_id bigint, item_id bigint, embedding vector,
    PRIMARY KEY (model_id, item_id));
CREATE INDEX ON embeddings USING hnsw ((embedding::vector(3)) vector_l2_ops) WHERE (model_id = 123);

-- higher precision with double precision[] or numeric[]
CREATE TABLE items2 (id bigserial PRIMARY KEY, embedding double precision[]);
ALTER TABLE items2 ADD CHECK (vector_dims(embedding::vector) = 3);
CREATE INDEX ON items2 USING hnsw ((embedding::vector(3)) vector_l2_ops);

Reference: operators

vector operators
OperatorDescriptionAdded
+element-wise addition—
-element-wise subtraction—
*element-wise multiplication0.5.0
||concatenate0.7.0
<->Euclidean distance—
<#>negative inner product—
<=>cosine distance—
<+>taxicab distance0.7.0
halfvec operators (all added in 0.7.0)
OperatorDescription
+element-wise addition
-element-wise subtraction
*element-wise multiplication
||concatenate
<->Euclidean distance
<#>negative inner product
<=>cosine distance
<+>taxicab distance
bit and sparsevec operators
OperatorDescriptionTypesAdded
<~>Hamming distancebit0.7.0
<%>Jaccard distancebit0.7.0
<-> <#> <=> <+>Euclidean, negative inner product, cosine, taxicabsparsevec0.7.0

Reference: functions

vector functions
FunctionDescriptionAdded
binary_quantize(vector) → bitbinary quantize0.7.0
cosine_distance(vector, vector) → double precisioncosine distance—
inner_product(vector, vector) → double precisioninner product—
l1_distance(vector, vector) → double precisiontaxicab distance0.5.0
l2_distance(vector, vector) → double precisionEuclidean distance—
l2_normalize(vector) → vectornormalize with Euclidean norm0.7.0
subvector(vector, integer, integer) → vectorsubvector0.7.0
vector_dims(vector) → integernumber of dimensions—
vector_norm(vector) → double precisionEuclidean norm—
halfvec functions (all added in 0.7.0)
FunctionDescription
binary_quantize(halfvec) → bitbinary quantize
cosine_distance(halfvec, halfvec) → double precisioncosine distance
inner_product(halfvec, halfvec) → double precisioninner product
l1_distance(halfvec, halfvec) → double precisiontaxicab distance
l2_distance(halfvec, halfvec) → double precisionEuclidean distance
l2_norm(halfvec) → double precisionEuclidean norm
l2_normalize(halfvec) → halfvecnormalize with Euclidean norm
subvector(halfvec, integer, integer) → halfvecsubvector
vector_dims(halfvec) → integernumber of dimensions
bit and sparsevec functions
FunctionDescriptionAdded
hamming_distance(bit, bit) → double precisionHamming distance0.7.0
jaccard_distance(bit, bit) → double precisionJaccard distance0.7.0
cosine_distance(sparsevec, sparsevec) → double precisioncosine distance0.7.0
inner_product(sparsevec, sparsevec) → double precisioninner product0.7.0
l1_distance(sparsevec, sparsevec) → double precisiontaxicab distance0.7.0
l2_distance(sparsevec, sparsevec) → double precisionEuclidean distance0.7.0
l2_norm(sparsevec) → double precisionEuclidean norm0.7.0
l2_normalize(sparsevec) → sparsevecnormalize with Euclidean norm0.7.0

Reference: aggregates

Aggregate functions
FunctionDescriptionAdded
avg(vector) → vectoraverage—
sum(vector) → vectorsum0.5.0
avg(halfvec) → halfvecaverage0.7.0
sum(halfvec) → halfvecsum0.7.0

Reference: operator classes

Match the operator class to the column type and distance metric. Create one index per operator class you query with.

Index operator classes by type and distance
TypeL2Inner productCosineL1
vectorvector_l2_opsvector_ip_opsvector_cosine_opsvector_l1_ops
halfvechalfvec_l2_opshalfvec_ip_opshalfvec_cosine_opshalfvec_l1_ops
sparsevecsparsevec_l2_opssparsevec_ip_opssparsevec_cosine_opssparsevec_l1_ops
Binary (bit) operator classes
DistanceOperator class
Hammingbit_hamming_ops
Jaccardbit_jaccard_ops

Reference: server settings

Query-time settings can be applied globally with SET or per query with SET LOCAL inside a transaction.

pgvector settings and defaults
SettingDefaultScope
hnsw.ef_search40query
hnsw.iterative_scanoffquery — strict_order / relaxed_order (0.8.0+)
hnsw.max_scan_tuples20000query
hnsw.scan_mem_multiplier1query (multiple of work_mem)
ivfflat.probes1query
ivfflat.iterative_scanoffquery — strict_order / relaxed_order (0.8.0+)
ivfflat.max_probes—query — capped by ivfflat.probes if lower
HNSW m16index build option
HNSW ef_construction64index build option
maintenance_work_memPostgres defaultserver (raise for faster HNSW builds)
max_parallel_maintenance_workers2server (index build parallelism)
max_parallel_workers8server (raise if many workers are used)
Want these values generated for your table? Use the index planner, or copy ready-made snippets from the cheatsheet.

Glossary

Plain-language definitions of the terms used throughout this documentation.

Embedding
A fixed-length list of numbers produced by a model that represents an item — text, image, audio, or other data — so that similar items are close together in vector space.
Vector
In pgvector, the vector data type, which stores single-precision (4-byte) floating-point values. Each vector takes 4 * dimensions + 8 bytes and can hold up to 16,000 dimensions.
Dimension
One position in a vector. The length of the vector (its number of dimensions) is fixed per column when you declare it, for example vector(1536).
Nearest-neighbor search
Finding the rows whose vectors are closest to a query vector, usually with ORDER BY distance_operator LIMIT n. The operator in the ORDER BY selects the distance metric.
Exact vs. approximate search
Without an index, pgvector compares the query against every row (exact search, perfect recall, slower as the table grows). An HNSW or IVFFlat index makes search approximate: faster, but it may miss some true neighbors.
Recall
The fraction of the true nearest neighbors that an approximate search actually returns. Higher recall means more accurate results, generally at the cost of speed. You can measure it by comparing approximate results with exact results.
HNSW
Hierarchical Navigable Small World — a graph-based approximate index with good query performance at the cost of slower builds and higher memory use. It can be created on an empty table and is tuned with m, ef_construction, and hnsw.ef_search.
IVFFlat
An inverted-file approximate index that divides vectors into a number of lists and, at query time, searches the closest lists. It builds faster and uses less memory than HNSW, and should be created after the table has data.
lists and probes
IVFFlat's number of clusters (lists, set when the index is built) and the number searched per query (ivfflat.probes). More probes improve recall at the cost of speed.
m, ef_construction, ef_search
HNSW tuning values: m is the maximum number of connections per layer (default 16), ef_construction is the build-time candidate-list size (default 64), and hnsw.ef_search is the query-time candidate-list size (default 40).
Quantization
Reducing the precision of stored vectors to save space and speed up search. Binary quantization converts a vector to a binary representation for a much smaller index; re-ranking with the original vectors recovers recall.
Half precision
The halfvec type, which stores 2-byte floating-point values (2 * dimensions + 8 bytes) and can index up to 4,000 dimensions — half the storage of vector.
Binary vector
The bit type for 0/1 vectors, such as image hashes, searched with Hamming (<~>) or Jaccard (<%>) distance and indexable up to 64,000 dimensions.
Sparse vector
The sparsevec type, which stores only non-zero elements. Its literal form is {index1:value1,index2:value2}/dimensions, with indices starting at 1.
Distance operator
The SQL operator that expresses a distance metric: <-> L2, <#> negative inner product, <=> cosine, <+> L1, and, for bit vectors, <~> Hamming and <%> Jaccard.
Operator class
The index binding that pairs a column's type with a distance metric, such as vector_cosine_ops for a vector column searched by cosine distance. Create one index per operator class you query with.