Blog · Architecture & benchmarks

Polign and DuckDB: sharing memory across agents

September 24, 2026 · Anup Talwalkar · Polign · Updated September 26, 2026

I put this comparison together to compare the two technologies, and how they would differentiate your use case based on your specific requirements (as they change over time). My focus is agent memory with Polign Recall, and helping multiple agents share a knowledge base through vector and semantic retrieval.

DuckDB is an excellent embedded database. Polign is a retrieval service that many agents on many machines can share. On one machine and one million vectors, Polign answered faster at the same recall with a fifth of the memory. The bigger difference is what happens when a second or third agent needs the data. The experience vastly changes when the stateless factor of this setup is needed.

5×
less memory under load
0.7 GiB vs 3.7 GiB
3×
faster to first answer
1.1 s vs 3.7 s
22%
lower latency at 0.98 recall
4.5 ms vs 5.8 ms p50
0
acknowledged writes lost
2,690 writes, server killed mid-load

The short answer

Use DuckDB when

  • one application process owns the data
  • the vector index fits in RAM
  • you want SQL joins, aggregates and Parquet next to vector search

Use Polign when

  • several agents, workers or machines read and write the same memory
  • you want to add serving nodes without copying data around
  • memory should stay small while the data lives in cheap object storage

The difference: one file vs. a shared service

A typical agent setup has a worker ingesting documents, a web server handling requests, and several agents reading and adding memories. They run as separate processes, often on separate machines.

DuckDB is a file opened by one process. While our benchmark process held the database open for writing, a second process could not open it at all, not even read-only. DuckDB's concurrency documentation confirms the rule: one writing process, or many readers with no writer. DuckDB also offers Quack, a client-server protocol in beta, and DuckLake for shared access. We did not test either.

Polign is a server. Every agent connects over HTTP or gRPC, and the data lives in an object store such as S3, GCS or Azure Blob. A node holds only caches it can rebuild, so scaling is adding nodes, not moving data.

What you need DuckDB Polign
A second process writes while the first is running Blocked by the file lock Yes. Every client talks to the server
Agents on other machines Need a copy of the file Connect to any node
More read traffic A bigger machine More nodes on the same bucket, including read-only nodes
A machine dies The data was on its disk Start another node. The data is in the bucket
Agent B reads what agent A just wrote, on another node Not applicable Pass the write token from A's write; the read waits until the node has it
Memory for 1M vectors The whole HNSW index in RAM A bounded cache; the rest stays in storage
Many agents, many nodes, one store

Agents, workers and apps
↓ HTTP / gRPC

Polign node

Reads and writes
Caches only

Polign node

Reads and writes
Caches only

↓
Your object store bucket
All durable data

The same service also does keyword (BM25) and hybrid search, metadata filters, API keys and encryption at rest. The architecture guide covers how nodes share a store.

The benchmark

Setup

Cohere 1M: 1,000,000 vectors, 768 dimensions, cosine. One laptop, 12 cores, 24 GiB RAM. Polign 0.7.1 serving a local directory, hot tier off. DuckDB 1.5.5 with the VSS extension and default HNSW settings. Medians of three rounds of 1,000 queries against exact ground truth. This measures one node on local disk. It does not measure multi-node scaling or object storage.

Search: Polign is faster per query, DuckDB serves more queries per second

Each row pairs the two engines at about the same recall. With one client, Polign had lower latency at every matched level up to 0.99. With eight clients hitting one process, DuckDB served 1.1× to 2.4× more queries a second.

Recall@10 DuckDB p50 Polign p50 DuckDB QPS, 8 clients Polign QPS, 8 clients
~0.95
DuckDB 0.940, Polign 0.948
4.4 ms 2.9 ms 1,111 967
~0.98
0.979 both
5.8 ms 4.5 ms 930 650
~0.99
DuckDB 0.989, Polign 0.990
9.6 ms 7.3 ms 710 374
Highest tested
DuckDB 0.993, Polign 0.996
10.9 ms 13.2 ms 511 212

Polign reached the highest recall of either engine, 0.996. Its default setting gives 0.990. The throughput gap is real for a single node. Polign's answer to more traffic is more nodes on the same store, which this test did not measure.

Memory and startup: Polign stays small

DuckDB Polign
Start to first answer 3.66 s 1.14 s
Memory after first answer 1.05 GiB 87 MiB
Memory under 8-client load 3.5 to 4.0 GiB 0.6 to 0.7 GiB

DuckDB's HNSW index must be fully in RAM, and its documentation notes the index is outside its memory_limit. Polign reads compressed index data from storage through a 256 MiB cache. The DuckDB figure includes its Python host; Polign's separate client used another 90 MiB, which the 5× headline counts.

Writes and durability: DuckDB writes faster, Polign never lost one

DuckDB Polign
Single-record write, p50 3.3 ms 10.4 ms
Writes per second, one client 259 56
New records returned as the top result 991 / 1,000 1,000 / 1,000
Acknowledged writes surviving a hard kill Not tested 2,690 / 2,690
Index build, 1M vectors 2 m 54 s 5 m 49 s

DuckDB is 4.6× faster for serial single writes. Polign only acknowledges a write once it is in durable storage, which is what makes it safe for another node to read. We killed the Polign server with SIGKILL three times while four clients wrote: every acknowledged record came back intact and searchable. DuckDB's VSS documentation calls HNSW persistence experimental and warns of possible index corruption after an unexpected shutdown.

Using both

Keep analysis in DuckDB and serve shared memory from Polign. Export from DuckDB as Parquet and import:

-- in DuckDB
COPY (SELECT 'doc-' || id AS id, v FROM items) TO 'items.parquet' (FORMAT parquet);

# load it into a Polign store
polign-import -store fs:./memory -collection items \
  -metric cosine -vector-column v items.parquet

We checked this on 20,000 rows: every id and vector matched byte for byte.

What this test does not cover

How we ran it
  • Versions. Polign server and importer 0.7.1 (commit 361df1009c31), Python SDK 0.7.0. DuckDB 1.5.5, VSS b833341. Python 3.14.7.
  • Index settings. DuckDB: cosine HNSW, M 16, M0 32, ef_construction 128, 12 threads, hnsw_enable_experimental_persistence = true. Polign: 1,000 IVF cells, 192-byte PQ codes, rerank pool 100, -store fs:DIR -hot-max 0 -maintain 0, telemetry off.
  • Queries. DuckDB used precomputed vector literals in SQL (bound Python parameters were about 10× slower on writes) with a verified HNSW_INDEX_SCAN plan. Polign used precomputed vectors over gRPC. Timings include client overhead.
  • Ground truth. Exact float64 cosine top 10 over the same million rows, recomputed for all 1,000 queries. All searches ran before any write test.
  • Repetition. Three fresh processes per engine, alternating which ran first, settings shuffled, 200 warmup queries each. OS caches were not cleared, so starts are warm-cache. Throughput is eight Python threads for 15 seconds. Memory is resident set size, no memory limit imposed.
  • Writes. 1,000 single-record writes per engine on top of 1,001,000 rows, one run. Stored vectors verified byte for byte in both. Crash test: three SIGKILL trials on a separate small store.
  • Other settings. DuckDB ef_search 128: 0.968 recall, 5.0 ms, 1,085 QPS. Polign nprobe 128: 0.991 recall, 7.5 ms, 367 QPS. DuckDB exact scan: 1.000 recall, 294 ms p50.

Try it

The getting started guide runs Polign against a local directory in a few minutes. Point it at a bucket when a second agent or machine needs the same memory. Polign Recall builds agent memory on top of it.