Benchmark methodology

Evidence backed benchmarks.

Every claim is measured head to head against retrieval based AI on identical models, identical questions, and identical retrieval.

Claim by claim

How we measured each claim.

Each claim runs against retrieval based AI with identical model and retrieval on both sides, so the comparison isolates how the answer is served and nothing else.

ClaimHow it was measured
Less re-processed text per questionWith identical model and retrieval on both sides, we compared how much text each approach re-reads to answer a question. Engram reads the library once, so each new question re-processes only the question itself.
Same answer qualityAn independent judge scored the answers from both sides for meaning and found the two even, a statistical tie with retrieval based AI, confirmed by an independent judge.
Faster first tokenWith one question at a time and identical model and retrieval on both sides, we measured the real time to the first streamed word of the answer.
Steady throughput under concurrent loadAs more people and agents ask at once, we measured answers completed per second on both sides, with queueing included, on identical model and retrieval.
Lower cost per queryWe took the real hourly price of the serving hardware and divided by the answers it sustained, so the cost reflects what the hardware actually delivers under load.

Throughput under load

Measured in the busy conditions real traffic creates.

We ramp up the load and measure both sides on identical model and retrieval, in the conditions a busy service actually creates as more people and agents ask at once. Retrieval based AI re-reads its context on every question, so its lead falls further behind as the load climbs, while Engram answers from memory it has already read.

Identical on both sides

The same model and the same retrieved documents feed both approaches, so the comparison reflects how the answer is served and not a difference in retrieval quality.

Measured, not modeled

First token is the real time to the first streamed word. Throughput is answers completed per second at each level of load, with queueing counted in, measured on the running service.

Real running cost

Cost per query is the real hourly price of the serving hardware divided by the answers it sustains, so it reflects what the hardware delivers under load rather than a list rate.

The live performance curve, with the full measured ramp, is on the Pricing page.

Answer quality

Same accuracy, far less work per question.

Answer quality is a statistical tie with retrieval based AI. Across a large document haystack and hundreds of questions, Engram matches an equally retrieved comparison within noise while re-processing far less text per question.

Same retrieval, both sides

One shared retriever feeds both approaches across a large document haystack and hundreds of questions, so the comparison isolates how the answer is served rather than retrieval quality.

The tie

A paired significance test put the gap between the two well inside statistical noise. On answer quality, the two are not distinguishable.

Independent judge

The result is confirmed by an independent judge, a more capable model than the one under test, so the verdict does not depend on a lenient self grade.

See the full benchmark pack.

Members can sign in for the full benchmark pack with every figure and every configuration. Anyone can request it through the form.