Polign and DuckDB: sharing memory across agents
I put this comparison together to compare the two technologies, and how they would differentiate your use case based on your specific requirements (as they change over time). My focus is agent memory with Polign Recall, and helping multiple agents share a knowledge base through vector and semantic retrieval.
DuckDB is an excellent embedded database. Polign is a retrieval service that many agents on many machines can share. On one machine and one million vectors, Polign answered faster at the same recall with a fifth of the memory. The bigger difference is what happens when a second or third agent needs the data. The experience vastly changes when the stateless factor of this setup is needed.
0.7 GiB vs 3.7 GiB
1.1 s vs 3.7 s
4.5 ms vs 5.8 ms p50
2,690 writes, server killed mid-load
The short answer
Use DuckDB when
- one application process owns the data
- the vector index fits in RAM
- you want SQL joins, aggregates and Parquet next to vector search
Use Polign when
- several agents, workers or machines read and write the same memory
- you want to add serving nodes without copying data around
- memory should stay small while the data lives in cheap object storage
The difference: one file vs. a shared service
A typical agent setup has a worker ingesting documents, a web server handling requests, and several agents reading and adding memories. They run as separate processes, often on separate machines.
DuckDB is a file opened by one process. While our benchmark process held the database open for writing, a second process could not open it at all, not even read-only. DuckDB's concurrency documentation confirms the rule: one writing process, or many readers with no writer. DuckDB also offers Quack, a client-server protocol in beta, and DuckLake for shared access. We did not test either.
Polign is a server. Every agent connects over HTTP or gRPC, and the data lives in an object store such as S3, GCS or Azure Blob. A node holds only caches it can rebuild, so scaling is adding nodes, not moving data.
| What you need | DuckDB | Polign |
|---|---|---|
| A second process writes while the first is running | Blocked by the file lock | Yes. Every client talks to the server |
| Agents on other machines | Need a copy of the file | Connect to any node |
| More read traffic | A bigger machine | More nodes on the same bucket, including read-only nodes |
| A machine dies | The data was on its disk | Start another node. The data is in the bucket |
| Agent B reads what agent A just wrote, on another node | Not applicable | Pass the write token from A's write; the read waits until the node has it |
| Memory for 1M vectors | The whole HNSW index in RAM | A bounded cache; the rest stays in storage |
Agents, workers and apps
↓ HTTP / gRPC
Polign node
Reads and writes
Caches only
Polign node
Reads and writes
Caches only
↓
Your object store bucket
All durable data
The same service also does keyword (BM25) and hybrid search, metadata filters, API keys and encryption at rest. The architecture guide covers how nodes share a store.
The benchmark
Cohere 1M: 1,000,000 vectors, 768 dimensions, cosine. One laptop, 12 cores, 24 GiB RAM. Polign 0.7.1 serving a local directory, hot tier off. DuckDB 1.5.5 with the VSS extension and default HNSW settings. Medians of three rounds of 1,000 queries against exact ground truth. This measures one node on local disk. It does not measure multi-node scaling or object storage.
Search: Polign is faster per query, DuckDB serves more queries per second
Each row pairs the two engines at about the same recall. With one client, Polign had lower latency at every matched level up to 0.99. With eight clients hitting one process, DuckDB served 1.1× to 2.4× more queries a second.
| Recall@10 | DuckDB p50 | Polign p50 | DuckDB QPS, 8 clients | Polign QPS, 8 clients |
|---|---|---|---|---|
| ~0.95 DuckDB 0.940, Polign 0.948 |
4.4 ms | 2.9 ms | 1,111 | 967 |
| ~0.98 0.979 both |
5.8 ms | 4.5 ms | 930 | 650 |
| ~0.99 DuckDB 0.989, Polign 0.990 |
9.6 ms | 7.3 ms | 710 | 374 |
| Highest tested DuckDB 0.993, Polign 0.996 |
10.9 ms | 13.2 ms | 511 | 212 |
Polign reached the highest recall of either engine, 0.996. Its default setting gives 0.990. The throughput gap is real for a single node. Polign's answer to more traffic is more nodes on the same store, which this test did not measure.
Memory and startup: Polign stays small
| DuckDB | Polign | |
|---|---|---|
| Start to first answer | 3.66 s | 1.14 s |
| Memory after first answer | 1.05 GiB | 87 MiB |
| Memory under 8-client load | 3.5 to 4.0 GiB | 0.6 to 0.7 GiB |
DuckDB's HNSW index must be fully in RAM, and its
documentation
notes the index is outside its memory_limit. Polign reads compressed index data from
storage through a 256 MiB cache. The DuckDB figure includes its Python host; Polign's
separate client used another 90 MiB, which the 5× headline counts.
Writes and durability: DuckDB writes faster, Polign never lost one
| DuckDB | Polign | |
|---|---|---|
| Single-record write, p50 | 3.3 ms | 10.4 ms |
| Writes per second, one client | 259 | 56 |
| New records returned as the top result | 991 / 1,000 | 1,000 / 1,000 |
| Acknowledged writes surviving a hard kill | Not tested | 2,690 / 2,690 |
| Index build, 1M vectors | 2 m 54 s | 5 m 49 s |
DuckDB is 4.6× faster for serial single writes. Polign only acknowledges a write once it is in durable storage, which is what makes it safe for another node to read. We killed the Polign server with SIGKILL three times while four clients wrote: every acknowledged record came back intact and searchable. DuckDB's VSS documentation calls HNSW persistence experimental and warns of possible index corruption after an unexpected shutdown.
Using both
Keep analysis in DuckDB and serve shared memory from Polign. Export from DuckDB as Parquet and import:
-- in DuckDB
COPY (SELECT 'doc-' || id AS id, v FROM items) TO 'items.parquet' (FORMAT parquet);
# load it into a Polign store
polign-import -store fs:./memory -collection items \
-metric cosine -vector-column v items.parquet
We checked this on 20,000 rows: every id and vector matched byte for byte.
What this test does not cover
- Multi-node throughput and failover, the part where Polign is designed to shine.
- Object storage. Both engines ran on local disk.
- Filters, keyword and hybrid search, mixed read/write traffic.
- Other dataset sizes, Polign's hot tier, DuckDB's Quack and DuckLake.
How we ran it
- Versions. Polign server and importer 0.7.1 (commit
361df1009c31), Python SDK 0.7.0. DuckDB 1.5.5, VSSb833341. Python 3.14.7. - Index settings. DuckDB: cosine HNSW, M 16, M0 32, ef_construction 128, 12 threads,
hnsw_enable_experimental_persistence = true. Polign: 1,000 IVF cells, 192-byte PQ codes, rerank pool 100,-store fs:DIR -hot-max 0 -maintain 0, telemetry off. - Queries. DuckDB used precomputed vector literals in SQL (bound Python parameters were
about 10× slower on writes) with a verified
HNSW_INDEX_SCANplan. Polign used precomputed vectors over gRPC. Timings include client overhead. - Ground truth. Exact float64 cosine top 10 over the same million rows, recomputed for all 1,000 queries. All searches ran before any write test.
- Repetition. Three fresh processes per engine, alternating which ran first, settings shuffled, 200 warmup queries each. OS caches were not cleared, so starts are warm-cache. Throughput is eight Python threads for 15 seconds. Memory is resident set size, no memory limit imposed.
- Writes. 1,000 single-record writes per engine on top of 1,001,000 rows, one run. Stored vectors verified byte for byte in both. Crash test: three SIGKILL trials on a separate small store.
- Other settings. DuckDB ef_search 128: 0.968 recall, 5.0 ms, 1,085 QPS. Polign nprobe 128: 0.991 recall, 7.5 ms, 367 QPS. DuckDB exact scan: 1.000 recall, 294 ms p50.
Try it
The getting started guide runs Polign against a local directory in a few minutes. Point it at a bucket when a second agent or machine needs the same memory. Polign Recall builds agent memory on top of it.