Pricing

Onboarding is free. Answers are cheap.

We read your library in at no charge, hold it in memory for a small monthly fee, and charge a flat rate per question. We never charge for input or context tokens.

Estimate your bill ↓

Estimate your bill

Your library, your volume.

Set your document count and your monthly questions to see the whole bill, meter by meter.

Your workload

Your estimated bill

Two meters. Onboarding, input, and context are always free.

One time onboarding
 
Monthly memory
 
Monthly inference
 
Monthly total
Memory plus inference, every month

Input and context tokens are free at every volume. Onboarding is free too. You pay a small amount to hold each document in memory and a flat rate per question answered.

Plans

Two ways to buy.

Start on your own, or reserve committed capacity for sensitive workloads.

Pay as you go

Self serve, up and running the same day. You pay only for the two meters, with a small monthly minimum, and you can grow your library whenever you like. Best when you want to try it on a real workload without a commitment.

  • Onboard every document free
  • Hold your library in memory month to month
  • Pay a flat rate for every question answered
Onboard today

Dedicated capacity

Committed capacity for regulated and data sensitive teams. Reserve steady performance for a large library and a busy question load, with the isolation and controls your compliance team expects. Best when memory holds material you cannot share.

  • Reserved capacity for steady response times
  • Built for regulated and data sensitive teams
  • Pricing shaped to your library and volume
Contact us

Why it stays cheap

Read once, then answer for a fraction.

Onboarding is free

The work of reading a document happens a single time, at onboarding, and we do not charge for it. Growing your library costs nothing up front.

Every answer skips the re-read

Because the document already lives in memory, answering a question never re-reads it. Each answer costs a fraction of a per token AI call, so a busy library stays affordable.

Memory is cheap to hold

Holding a document in memory is inexpensive, so a large library stays affordable month to month. The bill grows gently with your library and stays predictable.

Performance

Performance holds as usage grows.

Response times stay steady as more people and agents ask questions at once. Because answers come from memory rather than re-reading, capacity keeps opening up while retrieval based AI slows down under the same load. Measured on a single dedicated serving tier.

Engram CaaS retrieval based AI waiting avoided

Measured on a single dedicated serving tier as concurrent questions climb. Response times for Engram CaaS stay steady while retrieval based AI slows down, because every answer comes from memory instead of re-reading the source on each question.