# Message test — https://www.elastic.co/elasticsearch/context-engine

After reading your page, only 3 of 15 personas could name a reason to pick you over a similar option.

- **Page tested:** https://www.elastic.co/elasticsearch/context-engine
- **Audience tested against:** developers building ai agents and want to scale them efficiently
- **Personas:** 15 simulated
- **Report:** https://grader.wynter.com/r/context-engine-elastic-t4n7uJ4

> These answers are generated by AI, scored on Wynter's B2B Message
> Layers framework using behaviorally-diverse simulated personas. The
> methodology is real and the critique is directional. What a simulated
> persona cannot have is a live budget, a renewal coming up, or a boss
> asking about this quarter.

---

## 01 · The scores

Every persona answered all four questions. These are four independent
proportions of the same panel, not stages of a funnel.

| Layer | Question | Cleared the bar | Strength | Of those who passed |
| --- | --- | --- | --- | --- |
| 1. Clarity | Do they understand what you do? | 15/15 | 78% | all with reservations |
| 2. Relevance | Can they tell what it solves, and who it's for? | 12/15 | 67% | all with reservations |
| 3. Value | Do they actually want it? | 10/15 | 59% | all with reservations |
| 4. Differentiation | Is there a reason to pick you over the alternatives? | 3/15 | 33% | all with reservations |

**Brand alignment** (a side metric, not one of the four layers) — 14/15, 74% strength (all with reservations). Does the page read like the company you actually are?

**Fix first: Differentiation.** Earliest failing layer, walking the sequence in order — not simply the lowest score.

---

## 02 · What to change, layer by layer

Ordered worst-first. Specific edits, not a restatement of the score.

### Differentiation

**Add a line under "Connect" stating what data can stay outside Elasticsearch.**

"Elasticsearch indices, external data sources via Kibana connectors, or ES|QL queries" reads as a requirement to move everything into the Elastic stack. Say plainly which sources are read in place through connectors versus indexed, so buyers can judge…

*effort medium · impact high · tested against Answer the live objection*

**Move "Works with any harness" into the hero subhead with the named harnesses.**

The one thing buyers said competitors can't match, running across Claude Code, Codex, LangChain and AgentCore, is buried mid-page under a generic benefit heading. Put the named harness list in the first screen so the differentiator lands before anyone scrolls.

*effort low · impact high · tested against Front-load the meaning*

**Replace "Make any agent efficient" with a claim naming what only this does.**

"Make any agent efficient" is a sentence any context or RAG vendor could run unchanged. Lead the benefits block with the specific combination on offer: precomputed context served to agents you already run, through one query.

*effort low · impact medium · tested against Give a reason to choose you*

### Value

**Add methodology under the benchmark table: dataset, model, prompt count, who ran it.**

The 0.625 to 0.917 accuracy and 75% token figures are the reason buyers take the meeting, and they arrive with no source, so they read as a vendor test on vendor data. Name the dataset, the baseline RAG setup and the model beside the numbers.

*effort medium · impact high · tested against Proof next to the claim*

**Add a line under "Get access" explaining what Experimental Private Preview commits a team to.**

The "Experimental" footnote stops budget and pilot conversations because nobody can tell what they are signing up for. State who is eligible, whether it costs anything, what support exists, and expected GA timing.

*effort low · impact high · tested against Answer the live objection*

**Attribute the "reduced cost by 55% and latency by 40%" support agent result to a named deployment.**

That claim floats with no customer, scale or workload behind it. Say whose support agents, how many tickets or queries, and over what period, or move it next to the benchmark methodology.

*effort medium · impact medium · tested against Proof next to the claim*

### Relevance

**Add a who-it-is-for line under the H1 naming the role and stack.**

Readers have to reverse-engineer the audience from Elasticsearch, Kibana and ES|QL mentions. Say outright that it is for platform and AI engineering teams running agents on enterprise data.

*effort low · impact high · tested against Name the audience*

### Clarity

**Define "Knowledge Indicators" in plain words at first use, before the benefits list.**

The term carries four different things, schemas, entities, summaries and memories, and is never defined where it first appears. Give it a one-sentence definition on first mention so readers stop guessing.

*effort low · impact medium · tested against Plain language*

### Brand alignment (side metric)

**Rewrite "Connect" step to explain ES|QL and Kibana connectors for non-Elastic readers.**

The steps assume prior Elastic product knowledge, so the page reads as a feature update for existing customers. Gloss each Elastic term once so a buyer comparing vendors can follow the flow.

*effort medium · impact medium · tested against Plain language*

---

## 03 · What is working

### The 75% token reduction and accuracy numbers are the one thing that earns a second look

Eight respondents singled out the before/after token and accuracy figures as concrete, checkable against their own data, and the sole reason to take a meeting or evaluate. Several tied them directly to lower inference costs and fewer wrong answers.

> that 75% token drop (174.8M→42.6M) would show up directly in our inference costs, and the accuracy jump (0.625→0.917) would mean fewer embarrassing wrong answers
> 
> — Engineering Manager, Software Development, 501-1000

> accuracy 0.625→0.917, input tokens 174.8M→42.6M, F1 0.561→0.827 — that's specific and checkable in a way most vendor pages aren't
> 
> — Developer, Artificial Intelligence, 5000+

> the subhead "Context that's complete and fresh" plus the first paragraph about connecting enterprise sources to "build accurate agents" and cutting tokens/latency told me what pain it's addressing within a few lines.
> 
> — Lead Developer, Technology Services, 5000+

> that's specific enough to be checkable, and most competitors just say "faster" or "cheaper" with nothing behind it
> 
> — Lead Developer, Technology Services, 201-500

> the table claiming 0.625 to 0.917 accuracy and 174.8M to 42.6M input tokens is a concrete enough delta that I'd want it verified on our own data. That's worth a meeting
> 
> — Senior Developer, Enterprise Software, 5000+

### Working with multiple agent harnesses is the one differentiator respondents named

One respondent credited the framework integration for working across multiple agent harnesses. No other point identified a capability competitors lack.

> the "Works with any harness" integration list (LangChain, Claude Code, AWS AgentCore, Gemini) - if that's real and not aspirational, it's a genuine differentiator
> 
> — Lead Developer, Technology Services, 501-1000

---

## 04 · What the personas said

### The benchmark numbers are not believed because no methodology or source is disclosed

Six respondents flagged that the accuracy and token-reduction metrics arrive with no disclosed methodology, source, or case study, and read as a single vendor-run test on vendor data. The same numbers that attract them also fail to survive scrutiny.

> there's no methodology or benchmark source given, so I'd treat those as marketing until proven otherwise
> 
> — Lead Developer, Technology Services, 501-1000

> it's one undisclosed "Source" benchmark with no methodology — I'd want to know what counted as baseline RAG, what prompts made up those 96, and whether it holds on our own data before I'd put this ahead of other spend.
> 
> — Senior Developer, Enterprise Software, 51-200

> The numbers (55% cost cut, 75% fewer tokens) are the only thing that gave it substance — without a source or case study link though, I'd want to verify those before I believed them.
> 
> — Lead Developer, Technology Services, 201-500

### Core terms including Knowledge Indicators and context engine are used without definition

Four respondents could not parse stacked jargon offered without plain-English definitions. One said Knowledge Indicator conflates schema, entities, summaries, and memories into an undifferentiated term.

> It's the jargon stacked without definition up front — "Knowledge Indicators," "AI index," "composable ES|QL," "Elastic Workflows" — all introduced as if I already know what they do, so I'm reverse-engineering the product from feature names instead of being told plainly what it is in one sentence.
> 
> — Engineering Manager, Software Development, 1001-5000

> terms like "Knowledge Indicators," "AI index," and "harness" being used confidently without a plain-English definition up front — you have to piece together what they mean from context
> 
> — Developer, Artificial Intelligence, 5000+

> It's the term "Knowledge Indicator" doing too much work — it's used for schema metadata, extracted entities, document summaries, and agent memories all at once, so I can't tell if it's one data structure or a marketing umbrella over four different ones.
> 
> — Senior Developer, Enterprise Software, 51-200

> Terms like "Knowledge Indicators," "precompute," and "context engine" itself are doing a lot of unexplained work — they're coined-sounding terms stacked on top of each other
> 
> — Engineering Manager, Software Development, 501-1000

### The page never says who it is for

Seven respondents reported the target audience is nowhere stated and must be reverse-engineered from Elasticsearch, Kibana, and ES|QL references. One noted the problem itself lands quickly even though the reader is left to self-identify.

> the intended reader is never explicitly named - no "for data platform teams" or "for AI engineering leads"
> 
> — Lead Developer, Technology Services, 501-1000

> But the "who" is never explicit — there's no line saying "built for platform teams" or "for enterprises running agents at scale," I had to infer it from the use cases (Operations Analysis, Knowledge Base, Investigation Agents) and the integrations list (LangChain, Claude, AWS AgentCore).
> 
> — Engineering Manager, Software Development, 1001-5000

> Who it's for is inferred, not spelled out explicitly — there's no "for platform teams" or "for ML engineers" banner
> 
> — Developer, Artificial Intelligence, 5000+

> the intended reader is never spelled out explicitly; there's no "for data engineers" or "for platform teams" line anywhere. I had to infer it from the mechanics — mentions of Elasticsearch indices, Kibana connectors, ES|QL, role-based access control
> 
> — Senior Developer, Enterprise Software, 51-200

> The intended reader is never explicitly named as a job title, but it's inferable within seconds from the content — "What does a context engine do for agents?", ES|QL queries, Kibana connectors, framework integrations like LangChain and AWS AgentCore
> 
> — Senior Developer, Enterprise Software, 5000+

### Experimental Private Preview status with no pricing disqualifies the product from any…

Five respondents said the preview label, absent pricing, and lack of named customer references block budget approval, roadmap commitment, and pilot decisions, and lose the comparison to GA alternatives with production SLAs.

> The "Experimental" badge next to "Get access" would actually rule it out for near-term shortlisting over a competitor that's GA — if another tool in this space has production SLAs and this one admits it's not there, that's a real differentiator against it, not for it
> 
> — Senior Developer, Enterprise Software, 201-500

> But "Private Preview" and "currently not priced" tells me this is early, and there's not one named customer logo anywhere on this page backing those numbers — it's all generic "support agents" and "Operations Analysis Agents" use cases.
> 
> — Engineering Manager, Software Development, 1001-5000

> But what rules it out, or at least stalls it, is "Experimental" and "currently not priced separately while under Private Preview" sitting right next to those big accuracy numbers — if a competitor on my shortlist is GA, priced, and has even one named reference customer, they win by default because I'm not betting a renewal cycle on something Elastic itself is still calling experimental.
> 
> — Engineering Manager, Software Development, 1001-5000

> it's "Experimental"/Private Preview with no pricing — that's three reasons to not bet a roadmap on it yet
> 
> — Developer, Artificial Intelligence, 1001-5000

### Elasticsearch dependency reads as lock-in rather than a context layer

Four respondents said the product requires adopting the full Elastic stack rather than sitting on top of what they have, making store-agnostic competitors more attractive. One added that rivals with named case studies would win the meeting.

> The hard disqualifier risk is the ES|QL/Elasticsearch-index dependency buried in the "How it works" section — if a competing option is store-agnostic, that alone could knock this one out regardless of the numbers
> 
> — Senior Developer, Enterprise Software, 201-500

> if we're not already on Elastic, this probably means adopting their whole stack, not just a context layer, and that's a much bigger ask than the page admits to
> 
> — Developer, Artificial Intelligence, 5000+

> if a rival on my shortlist shows a named customer case study or a reproducible eval instead of an unlinked citation, they win the meeting. The other lock-in factor is the Elastic-native plumbing itself: ES|QL, Elasticsearch indices, Kibana connectors
> 
> — Senior Developer, Enterprise Software, 51-200

### The page reads as an upsell to existing Elastic customers, not a pitch to new buyers

Four respondents said the page presumes prior Elasticsearch, Kibana, and ES|QL knowledge and positions the product as a feature update rather than a neutral vendor comparison. One read it as an established vendor extending into agents with unclear maturity.

> it assumes I already know what ES|QL and Kibana are, drops "Knowledge Indicators" as a term without much hand-holding at first
> 
> — Lead Developer, Technology Services, 5000+

> this is an internal product team shipping a feature update to its installed base, not a sales org trying to win a net-new logo like mine.
> 
> — Engineering Manager, Software Development, 51-200

> it reads like an upsell to existing Elastic customers rather than a neutral pitch to someone comparing across vendors — which matters since we're not fully on Elastic
> 
> — Senior Developer, Enterprise Software, 201-500

> an established company testing a new SKU on existing customers rather than a startup desperate to prove itself - so the confidence of the copy is slightly ahead of the product's actual maturity
> 
> — Developer, Artificial Intelligence, 501-1000

---

## 05 · The hardest read

An adversarial pass over the findings. Every claim below was checked
against the panel's own answers; unsupported ones were dropped.

- **The only asset carrying the page is also the asset that collapses under scrutiny, leaving nothing standing.** *(high)*
  Eight respondents named the token and accuracy figures as the sole reason to take a meeting, while six rejected those same numbers as an undisclosed vendor-run test. The page's one load-bearing claim is the one readers refuse to believe.
- **The page cannot convert anyone because it never identifies who should read it.** *(high)*
  Seven respondents said the audience must be reverse-engineered from Elasticsearch, Kibana, and ES|QL references, and four read the page as a feature update for existing customers. Self-identification is being outsourced to the reader.
- **Every commercial next step is blocked, so even persuaded readers cannot act.** *(high)*
  Five respondents said the Private Preview label, missing pricing, and absent customer references stop budget approval, roadmap commitment, and pilots outright. Interest generated by the metrics has nowhere to go.
- **The page loses head-to-head comparisons on its own terms.** *(high)*
  Four respondents said the Elastic stack requirement reads as lock-in versus store-agnostic rivals, and five said GA alternatives with production SLAs win the comparison. Only one respondent named any capability competitors lack.
- **Differentiation is effectively absent from the page as written.** *(high)*
  A single respondent out of fifteen named multi-harness support as a differentiator, and no other point identified a capability competitors lack. Four respondents instead read the Elastic dependency as a reason to look elsewhere.
- **Unexplained vocabulary compounds the credibility gap rather than signaling sophistication.** *(medium)*
  Four respondents could not parse Knowledge Indicators or context engine without definitions, one noting the term conflates schema, entities, summaries, and memories. Jargon stacked on unverifiable numbers gives readers two reasons to disengage.

---

## 06 · Who answered

| # | Role | Industry | Company size |
| --- | --- | --- | --- |
| 1 | Senior Developer | Enterprise Software | 201-500 |
| 2 | Lead Developer | Technology Services | 501-1000 |
| 3 | Engineering Manager | Software Development | 1001-5000 |
| 4 | Developer | Artificial Intelligence | 5000+ |
| 5 | Senior Developer | Enterprise Software | 51-200 |
| 6 | Lead Developer | Technology Services | 201-500 |
| 7 | Engineering Manager | Software Development | 501-1000 |
| 8 | Developer | Artificial Intelligence | 1001-5000 |
| 9 | Senior Developer | Enterprise Software | 5000+ |
| 10 | Lead Developer | Technology Services | 51-200 |
| 11 | Engineering Manager | Software Development | 201-500 |
| 12 | Developer | Artificial Intelligence | 501-1000 |
| 13 | Senior Developer | Enterprise Software | 1001-5000 |
| 14 | Lead Developer | Technology Services | 5000+ |
| 15 | Engineering Manager | Software Development | 51-200 |

---

## 07 · Before you act on this

The methodology is real, and the critique is directional. What a
simulated persona cannot have is a live budget, a renewal coming up, or
a boss asking about this quarter. **Validate anything you're betting on
with real ICPs who are actually in-market.** Being wrong is more
expensive than you think. Finding out is cheaper than you'd guess.

Wynter runs message testing with verified B2B professionals — trusted
by HubSpot, RingCentral, Shopify, Cognism, Paddle, Veeam, Rippling and
Miro. <https://wynter.com>

This report is kept for 60 days from 2026-10-08, then deleted along with the personas and their answers.

