Clarity
Do they understand what you do?
15 could name what kind of product this is, unprompted.
https://www.elastic.co/elasticsearch/context-engine15 AI-simulated buyers
Your message needs work: they know what it is, who it's for, and why it's worth their time, but not why to pick you.
Do they understand what you do?
15 could name what kind of product this is, unprompted.
Can they tell what it solves, and who it's for?
12 could quickly tell what problem it solves and who it is for.
Do they actually want it?
10 would take a meeting to learn more.
Is there a reason to pick you over the alternatives?
3 could name a reason to pick you over a similar option.
Your page describes: context engine. They said:
6 couldn't name one; 9 got it right.
Four separate measures, not stages: all 15 personas answered all four questions. Each square is one persona.
Four respondents said the page presumes prior Elasticsearch, Kibana, and ES|QL knowledge and positions the product as a feature update rather than a neutral vendor comparison. One read it as an established vendor extending into agents with unclear maturity. Not one of the four layers, and it does not affect the scores above or the order to fix them in.
These are 15 simulated buyers. Want 15 real ones?
Test with humansThe first is on your weakest layer, the second on the next, the third on the layer the most buyers had a problem with. Each says what to change on the page and why, with one simulated answer behind it.
Why: "Elasticsearch indices, external data sources via Kibana connectors, or ES|QL queries" reads as a requirement to move everything into the Elastic stack. Say plainly which sources are read in place through connectors versus indexed, so buyers can judge…
3 of 15 raised this
“The hard disqualifier risk is the ES|QL/Elasticsearch-index dependency buried in the "How it works" section — if a competing option is store-agnostic, that alone could knock this one out regardless of the numbers”
Why: The 0.625 to 0.917 accuracy and 75% token figures are the reason buyers take the meeting, and they arrive with no source, so they read as a vendor test on vendor data. Name the dataset, the baseline RAG setup and the model beside the numbers.
5 of 15 raised this
“The "Experimental" badge next to "Get access" would actually rule it out for near-term shortlisting over a competitor that's GA — if another tool in this space has production SLAs and this one admits it's not there, that's a real differentiator against it, not for it”
Why: Readers have to reverse-engineer the audience from Elasticsearch, Kibana and ES|QL mentions. Say outright that it is for platform and AI engineering teams running agents on enterprise data.
7 of 15 raised this
“the intended reader is never explicitly named - no "for data platform teams" or "for AI engineering leads"”
These landed. Keep the wording when you edit around it.
The 75% token reduction and accuracy numbers are the one thing that earns a second look
“that 75% token drop (174.8M→42.6M) would show up directly in our inference costs, and the accuracy jump (0.625→0.917) would mean fewer embarrassing wrong answers”
Working with multiple agent harnesses is the one differentiator respondents named
“the "Works with any harness" integration list (LangChain, Claude Code, AWS AgentCore, Gemini) - if that's real and not aspirational, it's a genuine differentiator”
Why: The one thing buyers said competitors can't match, running across Claude Code, Codex, LangChain and AgentCore, is buried mid-page under a generic benefit heading. Put the named harness list in the first screen so the differentiator lands before anyone scrolls.
3 of 15 raised this
“The hard disqualifier risk is the ES|QL/Elasticsearch-index dependency buried in the "How it works" section — if a competing option is store-agnostic, that alone could knock this one out regardless of the numbers”
Why: "Make any agent efficient" is a sentence any context or RAG vendor could run unchanged. Lead the benefits block with the specific combination on offer: precomputed context served to agents you already run, through one query.
3 of 15 raised this
“The hard disqualifier risk is the ES|QL/Elasticsearch-index dependency buried in the "How it works" section — if a competing option is store-agnostic, that alone could knock this one out regardless of the numbers”
Why: The "Experimental" footnote stops budget and pilot conversations because nobody can tell what they are signing up for. State who is eligible, whether it costs anything, what support exists, and expected GA timing.
5 of 15 raised this
“The "Experimental" badge next to "Get access" would actually rule it out for near-term shortlisting over a competitor that's GA — if another tool in this space has production SLAs and this one admits it's not there, that's a real differentiator against it, not for it”
Why: That claim floats with no customer, scale or workload behind it. Say whose support agents, how many tickets or queries, and over what period, or move it next to the benchmark methodology.
5 of 15 raised this
“The "Experimental" badge next to "Get access" would actually rule it out for near-term shortlisting over a competitor that's GA — if another tool in this space has production SLAs and this one admits it's not there, that's a real differentiator against it, not for it”
Why: The term carries four different things, schemas, entities, summaries and memories, and is never defined where it first appears. Give it a one-sentence definition on first mention so readers stop guessing.
5 of 15 raised this
“there's no methodology or benchmark source given, so I'd treat those as marketing until proven otherwise”
Why: The steps assume prior Elastic product knowledge, so the page reads as a feature update for existing customers. Gloss each Elastic term once so a buyer comparing vendors can follow the flow.
4 of 15 raised this
“it assumes I already know what ES|QL and Kibana are, drops "Knowledge Indicators" as a term without much hand-holding at first”
A deliberately adversarial read of the same answers. Each claim was checked back against what the personas said and dropped if nothing supported it.
The only asset carrying the page is also the asset that collapses under scrutiny, leaving nothing standing.
Eight respondents named the token and accuracy figures as the sole reason to take a meeting, while six rejected those same numbers as an undisclosed vendor-run test. The page's one load-bearing claim is the one readers refuse to believe.
The page cannot convert anyone because it never identifies who should read it.
Seven respondents said the audience must be reverse-engineered from Elasticsearch, Kibana, and ES|QL references, and four read the page as a feature update for existing customers. Self-identification is being outsourced to the reader.
Every commercial next step is blocked, so even persuaded readers cannot act.
Five respondents said the Private Preview label, missing pricing, and absent customer references stop budget approval, roadmap commitment, and pilots outright. Interest generated by the metrics has nowhere to go.
The page loses head-to-head comparisons on its own terms.
Four respondents said the Elastic stack requirement reads as lock-in versus store-agnostic rivals, and five said GA alternatives with production SLAs win the comparison. Only one respondent named any capability competitors lack.
Differentiation is effectively absent from the page as written.
A single respondent out of fifteen named multi-harness support as a differentiator, and no other point identified a capability competitors lack. Four respondents instead read the Elastic dependency as a reason to look elsewhere.
Unexplained vocabulary compounds the credibility gap rather than signaling sophistication.
Four respondents could not parse Knowledge Indicators or context engine without definitions, one noting the term conflates schema, entities, summaries, and memories. Jargon stacked on unverifiable numbers gives readers two reasons to disengage.
Elasticsearch dependency reads as lock-in rather than a context layer
3 of 15
“The hard disqualifier risk is the ES|QL/Elasticsearch-index dependency buried in the "How it works" section — if a competing option is store-agnostic, that alone could knock this one out regardless of the numbers”
“if we're not already on Elastic, this probably means adopting their whole stack, not just a context layer, and that's a much bigger ask than the page admits to”
“if a rival on my shortlist shows a named customer case study or a reproducible eval instead of an unlinked citation, they win the meeting. The other lock-in factor is the Elastic-native plumbing itself: ES|QL, Elasticsearch indices, Kibana connectors”
Working with multiple agent harnesses is the one differentiator respondents named
1 of 15 · what worked
“the "Works with any harness" integration list (LangChain, Claude Code, AWS AgentCore, Gemini) - if that's real and not aspirational, it's a genuine differentiator”
Experimental Private Preview status with no pricing disqualifies the product from any…
5 of 15
“The "Experimental" badge next to "Get access" would actually rule it out for near-term shortlisting over a competitor that's GA — if another tool in this space has production SLAs and this one admits it's not there, that's a real differentiator against it, not for it”
“But "Private Preview" and "currently not priced" tells me this is early, and there's not one named customer logo anywhere on this page backing those numbers — it's all generic "support agents" and "Operations Analysis Agents" use cases.”
“But what rules it out, or at least stalls it, is "Experimental" and "currently not priced separately while under Private Preview" sitting right next to those big accuracy numbers — if a competitor on my shortlist is GA, priced, and has even one named reference customer, they win by default because I'm not betting a renewal cycle on something Elastic itself is still calling experimental.”
“it's "Experimental"/Private Preview with no pricing — that's three reasons to not bet a roadmap on it yet”
The 75% token reduction and accuracy numbers are the one thing that earns a second look
6 of 15 · what worked
“that 75% token drop (174.8M→42.6M) would show up directly in our inference costs, and the accuracy jump (0.625→0.917) would mean fewer embarrassing wrong answers”
“accuracy 0.625→0.917, input tokens 174.8M→42.6M, F1 0.561→0.827 — that's specific and checkable in a way most vendor pages aren't”
“the subhead "Context that's complete and fresh" plus the first paragraph about connecting enterprise sources to "build accurate agents" and cutting tokens/latency told me what pain it's addressing within a few lines.”
“that's specific enough to be checkable, and most competitors just say "faster" or "cheaper" with nothing behind it”
“the table claiming 0.625 to 0.917 accuracy and 174.8M to 42.6M input tokens is a concrete enough delta that I'd want it verified on our own data. That's worth a meeting”
The page never says who it is for
7 of 15
“the intended reader is never explicitly named - no "for data platform teams" or "for AI engineering leads"”
“But the "who" is never explicit — there's no line saying "built for platform teams" or "for enterprises running agents at scale," I had to infer it from the use cases (Operations Analysis, Knowledge Base, Investigation Agents) and the integrations list (LangChain, Claude, AWS AgentCore).”
“Who it's for is inferred, not spelled out explicitly — there's no "for platform teams" or "for ML engineers" banner”
“the intended reader is never spelled out explicitly; there's no "for data engineers" or "for platform teams" line anywhere. I had to infer it from the mechanics — mentions of Elasticsearch indices, Kibana connectors, ES|QL, role-based access control”
“The intended reader is never explicitly named as a job title, but it's inferable within seconds from the content — "What does a context engine do for agents?", ES|QL queries, Kibana connectors, framework integrations like LangChain and AWS AgentCore”
The benchmark numbers are not believed because no methodology or source is disclosed
5 of 15
“there's no methodology or benchmark source given, so I'd treat those as marketing until proven otherwise”
“it's one undisclosed "Source" benchmark with no methodology — I'd want to know what counted as baseline RAG, what prompts made up those 96, and whether it holds on our own data before I'd put this ahead of other spend.”
“The numbers (55% cost cut, 75% fewer tokens) are the only thing that gave it substance — without a source or case study link though, I'd want to verify those before I believed them.”
Core terms including Knowledge Indicators and context engine are used without definition
4 of 15
“It's the jargon stacked without definition up front — "Knowledge Indicators," "AI index," "composable ES|QL," "Elastic Workflows" — all introduced as if I already know what they do, so I'm reverse-engineering the product from feature names instead of being told plainly what it is in one sentence.”
“terms like "Knowledge Indicators," "AI index," and "harness" being used confidently without a plain-English definition up front — you have to piece together what they mean from context”
“It's the term "Knowledge Indicator" doing too much work — it's used for schema metadata, extracted entities, document summaries, and agent memories all at once, so I can't tell if it's one data structure or a marketing umbrella over four different ones.”
“Terms like "Knowledge Indicators," "precompute," and "context engine" itself are doing a lot of unexplained work — they're coined-sounding terms stacked on top of each other”
The page reads as an upsell to existing Elastic customers, not a pitch to new buyers
4 of 15
“it assumes I already know what ES|QL and Kibana are, drops "Knowledge Indicators" as a term without much hand-holding at first”
“this is an internal product team shipping a feature update to its installed base, not a sales org trying to win a net-new logo like mine.”
“it reads like an upsell to existing Elastic customers rather than a neutral pitch to someone comparing across vendors — which matters since we're not fully on Elastic”
“an established company testing a new SKU on existing customers rather than a startup desperate to prove itself - so the confidence of the copy is slightly ahead of the product's actual maturity”
15 AI-simulated personas matched to your target market. Each answered independently, without seeing your goal, the scoring criteria, or each other’s answers. Attribution is role, industry and company size only.
Every answer on this page was written by an AI model role-playing a buyer profile, scored on Wynter’s B2B Message Layers framework. The personas were sampled in code across role, industry, company size and behavioral traits; the model wrote only the answers. Scores arrive through fixed verdict categories and the counts are computed in our own code, so no number here was written by a model.
The count is how many personas cleared the bar on each question. A yes can be unhesitating or come with reservations; the scorecard counts both as a yes, and this is the only place the difference is shown. Per layer:
These answers are AI-simulated and directional. Validate anything you’re betting on with real buyers, your ICPs.
A detailed, section-by-section message test report from verified B2B professionals who are actually in-market for what you sell.







