Evidence & methodology

What we measured.
And what it means.

Every number below comes with the conditions it was measured under and the limits of what it proves. Where we don't have evidence, we say so instead of estimating.

Observed results

Four things we can actually show.

Each entry separates the observed result, the conditions it was measured under, what it means in ordinary language, and where it stops.

Fresh-session continuity resume

What it measures

Whether the correct goal, constraints, and decisions can be picked back up after a session reset or a switch of model.

Why it matters

This is the moment most people notice AI 'forgetting', a new session starts and the thread is gone.

Observed result
33.3% → 75% average resume (a 2.25× lift)
Test conditions: Controlled private test across OpenAI GPT-4.1 Mini, Anthropic Claude Haiku 4.5, and Gemini 2.5 Flash, comparing baseline calls with Willow-assisted calls.
Simply put

In our tests, roughly three times out of four the thread came back, instead of one in three.

Limitation

A small controlled test on our own scenarios, not an independent or public benchmark. Your results will depend on your workflow and provider.

Context compression (continuity artifact)

What it measures

How small a compact continuity artifact can be while still preserving the facts the thread needs.

Why it matters

Replaying an entire transcript on every turn is the main reason long AI sessions get slow and expensive.

Observed result
Artifact ≈ 14.6% of the source transcript, an 85.4% reduction in replayed source tokens
Test conditions: Single synthetic long-context compression test: source transcript ≈ 2,140.67 tokens, Willow artifact ≈ 313 tokens (token counts are estimates).
Simply put

Willow carried the thread forward using a small note instead of the whole conversation.

Limitation

This measures repeated context tokens only. Reducing repeated tokens may lower provider cost and unnecessary compute, but it is not a measured reduction in your bill and we have no emissions data, so we make no carbon claim.

Provider-adaptive control adherence

What it measures

Whether a model actually follows the adaptive control Willow generates for it.

Why it matters

A continuity signal is only useful if the model on the other side responds to it.

Observed result
Gemini followed-rate 37.5% → 62.5% (high-load/urgent cases reached 100% after policy refinement)
Test conditions: Focused private validation on one provider; adherence varied across providers.
Simply put

Some models take the hint more readily than others, and we tune for that.

Limitation

One provider, one focused test. We don't present this as a general provider ranking.

Live stack validation

What it measures

That the shipped API, continuity state, and multi-model path all work end to end.

Why it matters

It's the difference between a described product and a working one.

Observed result
Live API health, continuity state, and multi-model path verified against the running stack
Test conditions: Controlled launch smoke run against the live stack.
Simply put

The things we say exist, exist, and we check them.

Limitation

A smoke test proves the paths work, it doesn't measure quality or scale. Our automated test suite is passing, but a test count is build health, not performance, so we don't present it as a result.

Methodology & privacy

How these tests were run.

Enough to judge the evidence, without handing over the recipe.

What we share

The signals Willow returns, the shape of each test, the providers involved, the observed result, and the limits of that result.

What we don't share

Internal scoring formulas, thresholds, repair heuristics, orchestration prompts, private test fixtures, and infrastructure detail. These are protected, and publishing them would help others copy rather than help you evaluate.

Your data

None of these results come from customer conversations. They were produced on our own test material. See Privacy for how your own data is handled.

Live API health verified
Continuity state and re-entry paths verified
Multi-model path smoke test verified
Automated test suite passing (build health, not a performance result)
Care with claims

What we won't say.

Continuity is a hard problem. We'd rather be precise than loud.

Willow does not claim

  • zero drift
  • guaranteed cost savings or ROI
  • elimination of hallucinations
  • measured reductions in carbon emissions
  • universal benchmark leadership
  • validated performance in every domain

Willow does claim

  • controlled private tests show promising continuity and compression results
  • Willow exposes coherence signals developers can inspect for themselves
  • Willow helps applications preserve the active frame without replaying raw transcripts
  • broader and independent validation is still ahead of us
Judge for yourself

Test it against your own week.

  1. 01Pick a real workflow where your AI loses the thread today.
  2. 02Run the task as you run it now, and write down what went wrong.
  3. 03Run the same task with Willow in the loop.
  4. 04Compare how often the thread survives, how much context you resend, and how long repair takes.
  5. 05Repeat across more than one provider and more than one session reset.

See it hold the thread yourself.

Willow Observatory is live. Start a thread, then wire the same continuity engine into your own stack on hosted Willow.