Security
Polign Recall keeps your agents' memory, and polign_db is the server that stores it. Both run in your own infrastructure. This page answers what a security review asks: what is stored, who can read it, what leaves your network, how it is encrypted, and how to report a problem.
At a glance
| Question | Answer |
|---|---|
| Where does memory live? | In your own S3, GCS, or Azure bucket, or a local folder. Polign runs no hosted copy. |
| Does anything reach Polign? | Only an optional heartbeat with version, platform, and two counts. It is off whenever encryption is on, and you can turn it off. Details |
| Who can read a memory? | Holders of an API key for that server. A key can be bound to one namespace so it cannot see any other. Details |
| Is it encrypted? | Optionally before it leaves the server, with keys you hold, so the bucket and the cloud provider see only ciphertext. Details |
| Can a memory be deleted? | Forgetting withdraws a fact but keeps its history. Permanent deletion means deleting its records in polign_db. Details |
| Can I prove what an agent knew? | Yes. Export an audit bundle and verify it offline, without our server. Details |
| What is the license? | Recall is open source under Apache 2.0. polign_db is a free, closed-source binary. Details |
What Recall stores
Every memory is a short record, written once and never edited:
- Subject, predicate, and value, such as
customer-4812,refund_exception,revoked - When it was observed, and how: stated by the user, inferred by the agent, or returned by a tool
- A confidence between 0 and 1
- Evidence, when the fact was proposed from text: the exact excerpt it came from (up to 2 KB), and a link to a note that holds the full text
That last point matters for a privacy review. When your agent proposes facts from a conversation, Recall keeps the words they came from, so you can trace a belief back to its source. Treat the memory store as holding conversation excerpts, and protect and retain it accordingly.
Recall only talks to the polign_db server you point it at. It runs no model and calls no embedding service unless you configure one, and it sends nothing to Polign.
Keeping memories apart
Agents share memory when they use the same server, collection, and namespace. To keep one team's, customer's, or agent's memory away from another's, bind each API key to a namespace:
polign-apikey create -namespace acme:support
The server stamps that namespace on every write made with the key and limits every read, search, list, and delete to it. A namespaced key cannot read, overwrite, or delete another namespace's records, even when both share one collection and the same record ids. This is enforced by the server, not by a filter the client could leave out. A key created without a namespace is an operator key and sees everything.
For the strongest separation, such as one customer per contract, run one deployment per customer with its own bucket and keys.
Forgetting and deleting
Forgetting is not deletion. forget adds a withdrawal:
the fact stops appearing in current answers, and the earlier statements stay in the
history. That is what lets you answer "what did the agent believe on Sep 20?" after the
fact was withdrawn.
When a record must be erased, for example under a data-subject request, delete its
records from polign_db. Use history to list the record ids behind a fact,
then delete each one through the HTTP API
(DELETE /v1/collections/{c}/vectors/{id}), along with the evidence note each
one points to, or delete the whole collection. Recall has no single erase call yet. Erasing records removes them from
the history, so audit bundles exported earlier will no longer match a fresh export.
Proving what an agent knew
Recall can export a versioned audit bundle: the records and rules behind a memory
answer at a given moment. A standalone verifier, recall-audit, replays the
bundle and checks its digest from the file alone, with no database, model, or Polign
service involved. The bundle covers what memory returned. It does not record the
agent's reasoning or its final reply.
Running on one machine
When the Python client or recall setup keeps memory in a local
folder, the server it starts is locked down by default:
- It listens on
127.0.0.1only, on ports chosen at random. - Every request needs a key. The key is generated on your machine, bound to the
recallnamespace, and stored in a file only your user can read. - Telemetry is off.
- Memories stay in that folder. Nothing is written anywhere else.
What the server sends over the network
polign_db's normal traffic goes only to the object store you configured
(-store) and that store's credential services, such as AWS STS or instance
metadata. polign_db has no license check and no crash reporting. The polign CLI
and polign-apikey connect only to the server address you give them.
The heartbeat
- When it is on. By default, unless encryption or confidential mode
is on. Turn it off with
--telemetry=false. - What it sends. Exactly six fields:
version,os,arch,install_id(128 random bits),successful_queries(since the process started), andvector_count(ornull). No bucket names, hostnames, namespaces, credentials, queries, or results. - Where and how often.
https://telemetry.polign.com/v1/ping, at startup and every ten minutes. - See it yourself.
--show-telemetryprints the exact payload and exits without connecting. - What we keep. Those six fields plus the time received. No request logs and no IP or header enrichment, though the network provider necessarily sees source IPs.
- Failure is harmless. A failed heartbeat never changes what the database can do.
Update checks are a separate opt-in: --check-updates makes one request
to api.github.com/repos/Polign/polign/releases/latest and exits. It sends no
install id or counts and never downloads anything. Turning telemetry on does not turn
update checks on.
Confidential mode (--confidential or
POLIGN_CONFIDENTIAL=true) turns telemetry off regardless of other flags,
refuses update checks, and refuses to start without an encryption keyring. There is no
plaintext fallback.
The simplest way to rely on all this is to enforce it: turn telemetry off and allow
egress only to your object store and its credential services. Everything keeps working,
which is also why air-gapped deployments work against an internal S3-compatible store or
a local fs: folder.
Verify a download
Each release attaches a checksums.txt with the SHA-256 of every archive.
The install script checks it before unpacking. By
hand:
V=$(curl -fsSL https://dl.polign.com/latest/version)
OS=$(uname -s | tr '[:upper:]' '[:lower:]')
ARCH=$(uname -m | sed 's/x86_64/amd64/;s/aarch64/arm64/')
curl -fsSLO "https://dl.polign.com/$V/polign_db_${OS}_${ARCH}.tar.gz"
curl -fsSLO "https://dl.polign.com/$V/checksums.txt"
grep "polign_db_${OS}_${ARCH}.tar.gz" checksums.txt | shasum -a 256 -c -
Archives are not signed yet. Checksums and archives are published together on GitHub
releases and fetched over TLS, and signed releases are on the roadmap. Every archive
carries its LICENSE and CHANGELOG, so you can always tell what
a host is running.
Access control on the server
- Local by default. The server binds
127.0.0.1unless you tell it otherwise. Anything exposed beyond localhost belongs behind TLS and a deliberate network boundary. - Require a key for data. The vector data plane is open by default;
start with
-require-data-keyto require an API key on every HTTP and gRPC operation. The collection API always requires one. - Separate admin keys guard the management API at
/v1/admin. - Keys are stored hashed. A key looks like
plgn_<key_id>_<secret>with a 256-bit random secret. Only the secret's SHA-256 is stored, so a leaked bucket listing yields no usable credential. Checks are constant-time and every failure returns the same error. Theplgn_prefix lets secret scanners spot a leaked key. - No static cloud keys. The server reaches its bucket as the
instance or task role. A collection bound to your own bucket is reached by assuming an
IAM role, with a server-minted
ExternalIdguarding against the confused deputy problem, and the wiring is verified before the collection activates.
Hardening steps are in Operate in production.
Known limitations
Written down here so they are chosen, not discovered:
- Rate limiting is one server-wide ceiling. There is no per-caller metering; a gateway provides fairness where it matters.
- There are no storage quotas per caller or namespace.
- Apart from a namespace binding, keys have no scopes, such as read-only, and no expiry. A key lives until you disable it.
- Renewed TLS certificates are reloaded without a restart since 0.7.0. mTLS is a mesh or proxy responsibility. See TLS renewal.
Encryption at rest
Bucket encryption (SSE-S3, SSE-KMS, and the GCS and Azure equivalents) is done by the storage service, which decrypts on every read and so handles your plaintext. It protects against a lost disk, not against the cloud provider.
For that, polign_db encrypts before data leaves the server. Give it a keyring with
-store-encryption-key-file (or POLIGN_STORE_ENCRYPTION_KEY_FILE)
and everything it writes is AES-256-GCM ciphertext: memories and their history, indexes,
API-key records, and its local disk cache. The bucket, its backups, and the cloud key
service see only ciphertext. Keys come from a file or the environment and are never
fetched from a cloud key service.
- Tampering is caught. Moving, reordering, or truncating encrypted data is reported as corruption, never silently read.
- No accidental plaintext. A server started without the keys refuses an encrypted store instead of reading garbage or writing plaintext into it.
- Rotation is adding a line to the keyring. New writes use the new key and old data keeps reading.
- Decide up front. An existing store cannot be converted in place; move to a new bucket or prefix.
- Content, not shape. Object names, sizes, and access patterns stay visible to the store.
Setup steps are in Operate in production.
Confidential computing
Confidential computing hardware keeps memory encrypted while it is in use, so nothing else on the host, the cloud operator included, can read it. polign_db fits inside such a boundary without changes: the server is one static Linux binary with no control plane and no vendor in the data path. Moving it inside the boundary is moving one process, and auditing it is hashing one file against the release checksums.
On confidential VMs (AMD SEV-SNP or Intel TDX on AWS, GCP, and Azure) it runs unmodified. Pair it with client-side encryption so plaintext exists only inside protected memory and the bucket holds ciphertext. The keyring arrives however the platform releases secrets after attestation: an environment variable on dstack, or a tmpfs file on a confidential VM. AWS Nitro Enclaves also need a vsock proxy to reach the bucket. If you are deploying into an enclave today, write to us and we will work through it with you.
Licenses
Polign Recall is open source under the Apache License 2.0. You can read, modify, and ship it. The source is at github.com/Polign/recall.
polign_db is proprietary and its source is not public. It is free to
self-host, with no license key, no activation, and no limit on nodes or data.
Database API keys are locally generated security credentials, not license keys. You may run
each release in production and build your own products and services on it, including
ones your customers use. You may not redistribute it, resell it as a hosted database
service, modify it, or reverse engineer it. The full text ships in every release archive
as LICENSE, and a later release can change terms only for that release.
Pricing covers redistribution under an OEM agreement.
Reporting a vulnerability
Email security@polign.com with a description and a reproduction if you have one. Please don't open a public issue for anything you believe is exploitable. You'll get an acknowledgement within a few days, and fixes ship in the next patch release with credit unless you ask otherwise. Only the latest release gets security fixes, so upgrading is the patch channel.