Clarity
Fix firstDo they understand what you do?
8 could name what kind of product this is, unprompted.
https://arcate.io/15 AI-simulated buyers
Your message needs work: they know who it's for, why it's worth their time, and why to pick you, but not what it is.
Do they understand what you do?
8 could name what kind of product this is, unprompted.
Can they tell what it solves, and who it's for?
15 could quickly tell what problem it solves and who it is for.
Do they actually want it?
14 would take a meeting to learn more.
Is there a reason to pick you over the alternatives?
11 could name a reason to pick you over a similar option.
Four separate measures, not stages: all 15 personas answered all four questions. Each square is one persona.
Two respondents said the unsourced precision metrics undercut the rigorous positioning and that no evidence demonstrates successful sales to companies of their size. Not one of the four layers, and it does not affect the scores above or the order to fix them in.
These are 15 simulated buyers. Want 15 real ones?
Test with humansThe first is on your weakest layer, the second on the next, the third on the layer the most buyers had a problem with. Each says what to change on the page and why, with one simulated answer behind it.
Why: "Kendall's τ = 0.924 ... across 60 simulation runs" reads as unsourced noise. State what was simulated, against whose judgment, over how many roadmap items, and link the write-up beside the number.
3 of 15 raised this
“the Kendall's τ = 0.924 and Jaccard = 1.000 numbers have no methodology behind them so I'd discount those until I saw the actual simulation setup”
Why: A €3.7B manufacturer is the only evidence on the page, and a 51-200 person software team reads it as the wrong company. Name a SaaS customer with headcount and outcome.
3 of 15 raised this
“that's a €3.7B industrial company, not a 51-200 person software shop, so I'd need a similarly-sized SaaS customer story to actually trust it applies to me”
These landed. Keep the wording when you edit around it.
The scoring mechanism is understood and repeated back accurately
“It's a tool that ingests customer feedback from Slack, Intercom, Gong, Salesforce and HubSpot, scores it by ARR at risk, and spits out a ranked product roadmap”
The roadmap-defense problem lands as the reader's own problem
“the subhead "Sales holds the signals. Product holds the roadmap. Revenue connects neither" tells you the problem in one line”
Why: 'Agentic' arrives before anything explains it, and 'product intelligence' is a category label. Lead with the job: ranking the roadmap by the revenue at risk behind each request.
3 of 15 raised this
“the Kendall's τ = 0.924 and Jaccard = 1.000 numbers have no methodology behind them so I'd discount those until I saw the actual simulation setup”
Why: "Calibrated weights stop single accounts from hijacking your roadmap" hides the owner of the decision. Write who configures severity multipliers and whether the PM can change them.
3 of 15 raised this
“the Kendall's τ = 0.924 and Jaccard = 1.000 numbers have no methodology behind them so I'd discount those until I saw the actual simulation setup”
Why: Competitors showing three references win on evidence alone. State how many teams run Arcate today and what results they see, rather than resting on a single deployment.
3 of 15 raised this
“that's a €3.7B industrial company, not a 51-200 person software shop, so I'd need a similarly-sized SaaS customer story to actually trust it applies to me”
Why: The page targets PMs only by implication. A line the reader can point at — the PM defending a roadmap to a board — makes the mismatch with the manufacturing proof less jarring.
3 of 15 raised this
“that's a €3.7B industrial company, not a 51-200 person software shop, so I'd need a similarly-sized SaaS customer story to actually trust it applies to me”
No specific edits needed here — this layer held up.
No specific edits needed here — this layer held up.
Why: The heading promises scale and delivers one enterprise manufacturer plus unlabelled statistics, which undercuts the rigorous tone. Show a company-size range you serve.
2 of 15 raised this
“the only proof point they lean on is a €3.7B company, which doesn't tell me they've ever sold to or succeeded with a 51-200 person shop like mine”
A deliberately adversarial read of the same answers. Each claim was checked back against what the personas said and dropped if nothing supported it.
The page has no admissible evidence — every number and reference on it was rejected.
Three respondents dismissed the precision metrics and simulation-run claims as unsourced noise, four called the single Endress+Hauser case insufficient, and three rejected it as the wrong size and industry. The entire proof layer collapses.
Clarity on the mechanic is worthless because it converts into disbelief rather than credibility.
Four respondents restated the revenue-weighted scoring accurately, yet four still said one case study plus unsourced stats cannot prove viability. Readers understand exactly what is claimed and refuse to believe it.
The page loses the deal at the comparison stage, not the comprehension stage.
Three respondents said competitors offering three references would win on evidence, and two saw nothing demonstrating sales to companies of their size. Differentiation rests entirely on a proof point readers disqualify.
Mid-market SaaS readers are given no path to self-identify anywhere on the page.
Three respondents said the title is never stated and no matching SaaS proof point exists, three rejected the enterprise manufacturing reference as mismatched to a 51-200 person shop, and two found no evidence of mid-market sales.
Strong problem framing is squandered by making the reader do the qualifying work.
Four respondents said the roadmap-defense pain lands immediately, but three noted the buyer's title is never named and only implied, with no matching SaaS proof point. The page earns attention and then fails to confirm it.
Unexplained jargon compounds the credibility gap by hiding accountability.
Two respondents flagged 'Agentic' used before definition and vague calibration wording that obscures who owns weighting decisions. Ambiguity about control sits directly on top of metrics three respondents already called unsourced.
Statistics are dismissed because no methodology is shown
3 of 15
“the Kendall's τ = 0.924 and Jaccard = 1.000 numbers have no methodology behind them so I'd discount those until I saw the actual simulation setup”
“"60 simulation runs" is vague enough to rule it back out if I dug in and found it was a synthetic/internal test rather than validated on real customer data — I'd want to know whose judgment, on what dataset, before I let that number carry weight against another vendor's live customer references.”
'Agentic' and the calibration language leave readers guessing
2 of 15
“Phrases like "calibrated weights" and "attribution is visible and editable" sound precise but don't say who calibrates them or edits them, which is exactly the kind of soft language that gets papered over in a demo and falls apart on real data.”
The scoring mechanism is understood and repeated back accurately
4 of 15 · what worked
“It's a tool that ingests customer feedback from Slack, Intercom, Gong, Salesforce and HubSpot, scores it by ARR at risk, and spits out a ranked product roadmap”
“It's a tool that pulls customer feedback out of Slack, Intercom, Gong, Salesforce and HubSpot, weights each signal against the account's ARR, and spits out a revenue-ranked product roadmap with a traceable line back to the original customer quote.”
“the step-by-step in "How it works" (connect channels, score by ARR with the 30x/3x/1x multipliers, rank and decay over time) made the mechanism pretty concrete”
“the subhead "Sales holds the signals. Product holds the roadmap. Revenue connects neither" and "You get a ranked roadmap backed by customer ARR" tell you the problem inside the first screen”
“ranks product roadmap items by revenue-at-risk instead of gut feel — basically an ARR-weighted prioritization layer”
“The Endress+Hauser case (CES drop 3.53 to 1.47, 30% sales capacity freed) is the one concrete proof point that makes me think it does something real”
The only proof point is the wrong company size and industry
3 of 15
“that's a €3.7B industrial company, not a 51-200 person software shop, so I'd need a similarly-sized SaaS customer story to actually trust it applies to me”
“one manufacturing case study propping up the whole page isn't enough to differentiate them from a rival with three relevant references”
“the only proof point they lean on is a €3.7B company, which doesn't tell me they've ever sold to or succeeded with a 51-200 person shop like mine”
“But one case study at a €3.7B industrial company isn't the same as evidence this works for a 51-200 person B2B shop like mine”
One manufacturing case study is not enough to justify changing workflows
4 of 15
“one customer case study and two unsourced stats (Kendall's τ, Jaccard) aren't enough to bet a workflow change on — I'd want two or three more named customers my size”
“But one case study at a €3.7B industrial company isn't the same as evidence this works for a 51-200 person B2B shop like mine”
“The Endress+Hauser number (CES 3.53 to 1.47, 30% technical sales capacity freed) is the kind of proof that makes me want to test it rather than dismiss it, but that's a manufacturing account, not a SaaS org like mine”
“one manufacturing case study propping up the whole page isn't enough to differentiate them from a rival with three relevant references”
The buyer's job title is never named on the page
3 of 15
“A line naming the buyer directly — something like "Built for VPs of Product who answer to revenue and the board" — plus a proof point from a company my size and sector, not just Endress+Hauser”
“I'd want my actual title or a line like "built for VPs of Product reporting to the board" instead of me inferring it from "defend unjustifiable roadmaps alone" — right now I'm doing the work of mapping myself onto the copy rather than the copy doing it for me.”
“The reader is never explicitly named as "VP Product" or "Head of Product," but the language — PMs, roadmaps, board slides, "PMs left to defend unjustifiable roadmaps alone" — makes it obvious within seconds”
The roadmap-defense problem lands as the reader's own problem
4 of 15 · what worked
“the subhead "Sales holds the signals. Product holds the roadmap. Revenue connects neither" tells you the problem in one line”
“"Sales holds the signals. Product holds the roadmap. Revenue connects neither" tells me the problem in one line, and the table comparing Arcate to "Traditional PM Tools (Self-Serve)" nails who this is for without me having to guess too hard”
“the subhead "Sales holds the signals. Product holds the roadmap. Revenue connects neither" and "You get a ranked roadmap backed by customer ARR" tell you the problem inside the first screen”
“The reader is clearly a Head of Product or PM who has to defend a roadmap to a board — "Show the board exactly why you built it"”
“If it worked as promised, my next roadmap review with leadership stops being "trust me, sales was loud about this" and becomes a chain from customer quote to ARR to score to bet — that directly kills the "unjustifiable roadmap" problem I actually live with.”
Nothing on the page shows the company has sold to mid-market
2 of 15
“the only proof point they lean on is a €3.7B company, which doesn't tell me they've ever sold to or succeeded with a 51-200 person shop like mine”
15 AI-simulated personas matched to your target market. Each answered independently, without seeing your goal, the scoring criteria, or each other’s answers. Attribution is role, industry and company size only.
Every answer on this page was written by an AI model role-playing a buyer profile, scored on Wynter’s B2B Message Layers framework. The personas were sampled in code across role, industry, company size and behavioral traits; the model wrote only the answers. Scores arrive through fixed verdict categories and the counts are computed in our own code, so no number here was written by a model.
The count is how many personas cleared the bar on each question. A yes can be unhesitating or come with reservations; the scorecard counts both as a yes, and this is the only place the difference is shown. Per layer:
These answers are AI-simulated and directional. Validate anything you’re betting on with real buyers, your ICPs.
A detailed, section-by-section message test report from verified B2B professionals who are actually in-market for what you sell.







