# Message test — https://arcate.io/

After reading your page, only 9 of 15 personas could name what kind of product this is, unprompted.

- **Page tested:** https://arcate.io/
- **Audience tested against:** For Heads of Product and VPs of Product at B2B SaaS companies, Series A to C, 50 to 500 employees. They deal with customer feedback scattered across Slack, CRM, and support tools, and lack a systematic way to connect product bets to revenue at risk.
- **Personas:** 15 simulated
- **Report:** https://grader.wynter.com/r/arcate-agentic-product-intelligence-for-b2b-te-T15jDp0

> These answers are generated by AI, scored on Wynter's B2B Message
> Layers framework using behaviorally-diverse simulated personas. The
> methodology is real and the critique is directional. What a simulated
> persona cannot have is a live budget, a renewal coming up, or a boss
> asking about this quarter.

---

## 01 · The scores

Every persona answered all four questions. These are four independent
proportions of the same panel, not stages of a funnel.

| Layer | Question | Cleared the bar | Strength | Of those who passed |
| --- | --- | --- | --- | --- |
| 1. Clarity | Do they understand what you do? | 9/15 | 84% | 4 without hesitation, 11 with reservations |
| 2. Relevance | Can they tell what it solves, and who it's for? | 15/15 | 100% | all without hesitation |
| 3. Value | Do they actually want it? | 15/15 | 78% | all with reservations |
| 4. Differentiation | Is there a reason to pick you over the alternatives? | 10/15 | 61% | all with reservations |

**Brand alignment** (a side metric, not one of the four layers) — 15/15, 78% strength (all with reservations). Does the page read like the company you actually are?

**Fix first: Clarity.** Earliest failing layer, walking the sequence in order — not simply the lowest score.

### What they thought you sell

1 of the personas who named a category got it wrong:

- 1× “Revenue-weighted roadmap prioritization tool”

---

## 02 · What to change, layer by layer

Ordered worst-first. Specific edits, not a restatement of the score.

### Clarity

**Replace the τ and Jaccard hero stats with a plain-English line naming what was measured.**

A reader cannot tell what "τ = 0.924" and "Jaccard 1.000" mean or where the numbers came from, so they read as decoration. Say instead that Arcate's ranking matched a senior PM's ranking on the same 700 signals, and put the Greek symbols in a footnote.

*effort medium · impact high · tested against Plain language*

**Add sample, date and method under the statistics block instead of "available on request".**

"Method and data available on request" asks the buyer to trust a number and go hunting for its source. Put the run count, what a run was, who the senior PM benchmark was, and when it ran directly beneath the claim.

*effort medium · impact high · tested against Proof next to the claim*

**Move the CRM data-quality sentences out of the scoring paragraph into their own headed block.**

The answer to "what if our CRM data is a mess" is buried at the end of the Revenue scoring paragraph, so readers finish the page still doubting it. Give it a heading like "If your CRM data is incomplete" and say what a tier default actually is.

*effort low · impact high · tested against Answer the live objection*

**State the run count consistently; the page says 40+ runs in one place and 60 in another.**

The hero says "40+ independent runs" and the Kendall's τ block says "across 60 simulation runs", which reads as sloppy proof. Pick one number and use it everywhere the statistic appears.

*effort low · impact medium · tested against Proof next to the claim*

### Differentiation

**Add a second proof point at SaaS scale beside the Endress+Hauser case study.**

A €3.7B manufacturer is the only customer evidence, so a smaller SaaS buyer cannot tell the method transfers to them. Add a short result from a company closer to their size, or state the account range Arcate is built for.

*effort high · impact high · tested against Proof next to the claim*

**Rewrite the H1 to name the job, not "Agentic product intelligence for B2B teams".**

"Agentic product intelligence" is a category label any AI vendor could use, and it makes the company read as early-stage and vague. Lead with the job the buyer says out loud: ranking the roadmap by the revenue at risk behind each request.

*effort medium · impact medium · tested against Lead with the use case*

**Replace "Traditional PM Tools (Self-Serve)" in the comparison table with the named tools buyers compare.**

An unnamed category lets readers assume their current tool is the exception. Name Productboard, Jira or Canny so the contrast lands against the thing they actually use.

*effort low · impact medium · tested against Give a reason to choose you*

---

## 03 · What is working

### The revenue-weighted scoring mechanism is understood and repeated back

Three respondents described the mechanism accurately: feedback scored against account ARR with decay, ranking roadmap by ARR at risk. They called it transparent and concrete.

> severity multipliers (deal-loss 30x, friction 3x, feature mention 1x) times account ARR, decaying over time, with an audit trail back to the original quote
> 
> — Director of Product, B2B software, 51-200

> They pull customer feedback out of Slack, Intercom, Gong, Salesforce and HubSpot, weight it by the ARR attached to the account, and spit out a ranked product roadmap so you can show the board "this feature maps to €X at risk"
> 
> — Head of Product, Software as a Service, 201-500

### The headline and subhead make the problem and buyer obvious within seconds

Four respondents said the problem statement and target audience were clear immediately from the headline and subhead, with one naming a five-second read.

> the subhead literally says "For product leaders who defend product decisions to the board" right under the headline, so I knew in the first five seconds who this is for and why
> 
> — Director of Product, B2B software, 51-200

> It's spelled out immediately, not something I had to dig for — the subhead says "For product leaders who defend product decisions to the board" and the body line "Sales holds the signals. Product holds the roadmap. Revenue connects neither"
> 
> — Head of Product, Software as a Service, 201-500

> the subhead literally says "For product leaders who defend product decisions to the board," and the whole "information gap" section spells out the problem
> 
> — Senior Product Manager, B2B software, 201-500

### Auditable revenue linkage is the value respondents articulated for board conversations

Two respondents said tying feedback to revenue removes gut-feel justification and that the before/after metrics demonstrate value in board meetings.

> I'd stop walking into board meetings with "sales asked for it" as my justification and instead have an auditable line from customer quote to ARR to roadmap item
> 
> — Senior Product Manager, SaaS, 201-500

> the Endress+Hauser number (CES 3.53 to 1.47, 30% sales capacity freed) is the kind of concrete before/after I'd want to replicate
> 
> — Senior Product Manager, B2B software, 201-500

---

## 04 · What the personas said

### The statistical claims are read as unsourced and undermine credibility rather than build…

Five respondents flagged missing methodology, sample size and attribution behind the statistics. One called them statistics dressed as proof; another said they undermine the credibility claim they support.

> The stats (τ=0.924, Jaccard=1.000) are thrown around without methodology detail, so I'd want the actual validation write-up before I believed the category claim over the marketing.
> 
> — Senior Product Manager, SaaS, 201-500

> The τ=0.924 and Jaccard=1.000 stats against "Senior PM judgment" are interesting but thin on their own—I'd want to know whose judgment, how they picked the 60 simulation runs, and whether it holds outside two industries and 100 accounts
> 
> — Director of Product, B2B software, 51-200

> the Kendall's τ = 0.924 and Jaccard = 1.000 stats are dressed up to look like proof but I don't know the sample or methodology beyond "60 simulation runs"
> 
> — Head of Product, Software as a Service, 201-500

> that's the part I'd need pinned down before I believed the bigger claim
> 
> — VP of Product, SaaS, 51-200

### Respondents doubt the system survives messy or incomplete CRM data

Four respondents raised data quality: vague tier defaults for missing CRM fields, unclear cross-channel weighting, no CRM hygiene messaging, and wanting validation on live rather than curated data.

> "configurable tier defaults" for missing ARR—it's vague enough that I'd want an example of what a tier default actually looks like before I trust it on messy CRM data
> 
> — Director of Product, B2B software, 51-200

> What's harder to pin down is what counts as a 'signal' in practice: does a one-line Slack mention get treated the same as a Gong call transcript, and how do they dedupe the same complaint showing up in three channels? That's the bit the page glosses over.
> 
> — Head of Product, Software as a Service, 201-500

> I'd need a line naming the CRM hygiene problem directly — something like 'works even when 40% of your ARR fields are blank or stale' — because that's the actual barrier at a 200-500 person SaaS company
> 
> — Head of Product, Software as a Service, 201-500

### A €3.7B manufacturer does not generalize to respondents' own company size

Two respondents said the Endress+Hauser example may not transfer to smaller SaaS shops and asked for a concrete example at their own scale.

> I'd want a line naming my situation exactly — something like 'you have a Salesforce CRM and a backlog full of Slack threads and no way to tie the two together' — plus a number showing this works at my company's size, not just at a €3.7B industrial giant
> 
> — Senior Product Manager, B2B software, 201-500

> the Endress+Hauser case is a manufacturing giant, not a 51-200 person SaaS shop — so I don't know if this generalizes to my CRM hygiene or my Slack noise.
> 
> — Director of Product, Software as a Service, 51-200

### One case study is not enough proof of production adoption

Three respondents said a single case study cannot demonstrate real-world use, and a fourth wanted independent reference customers beyond Endress+Hauser.

> The Endress+Hauser CES drop is a real number, so it's not nothing, but one case study and a lot of "book a call" isn't enough to change the roadmap on its own.
> 
> — Head of Product, Software as a Service, 201-500

> one Endress+Hauser CES stat and an unattributed "τ = 0.924" isn't enough to act on
> 
> — VP of Product, SaaS, 51-200

> I'm not signing anything until they show me the τ=0.924/Jaccard=1.000 methodology and a reference customer beyond the Endress+Hauser CES case, which reads more like a services engagement than proof the software generalizes to my stack.
> 
> — Head of Product, SaaS, 201-500

### The page reads as an early-stage vendor with a thin customer base

Six respondents inferred a small, founder-led early-stage company. Three framed this neutrally as positioning; three treated the limited proof points and academic-style claims as a credibility problem.

> I'd guess a small early-stage vendor, maybe under 20 people, probably a couple of years old at most — one flagship reference customer (Endress+Hauser) and "book a diagnostic call" as the only CTA reads like a team still doing high-touch sales
> 
> — Head of Product, Software as a Service, 201-500

> the stats block — τ = 0.924, Jaccard = 1.000 — reads like they're trying to punch above their weight with academic-looking rigor to compensate for having only one real proof point, which is exactly what an early, thin-track-record vendor does
> 
> — Head of Product, B2B software, 201-500

> I picture a small, early-stage B2B SaaS shop—maybe 10-30 people, a year or two old, probably founder-led with an ex-product or ex-consulting background, given how specific the mechanism talk is
> 
> — Director of Product, B2B software, 51-200

> I'd guess a small, early-stage B2B SaaS outfit — maybe 10-30 people, a couple years old, probably founder-led with an ex-consultant or ex-PM background given how precisely they talk about RICE scores and CES methodology
> 
> — Senior Product Manager, B2B software, 201-500

> the stats block undercuts that credibility by trying to sound more rigorous than it can back up
> 
> — Senior Product Manager, Software as a Service, 201-500

> a small, early-stage team — maybe 10-30 people, a couple years old, likely founder-led with an ex-product or ex-consulting background, selling to mid-market/enterprise B2B software companies
> 
> — VP of Product, Software as a Service, 51-200

---

## 05 · The hardest read

An adversarial pass over the findings. Every claim below was checked
against the panel's own answers; unsupported ones were dropped.

- **The proof stack collapses under scrutiny: every credibility asset on the page is contested by more people than endorse it.** *(high)*
  Five respondents call the statistics unsourced, three say one case study cannot show production adoption, and two say the €3.7B manufacturer does not transfer to their scale — versus two who articulated the value.
- **Clarity is being mistaken for persuasion — respondents understand the pitch and still refuse to believe it.** *(high)*
  Four grasp the problem in five seconds and three repeat the scoring mechanism accurately, yet five reject the statistics and six read the page as an unproven early-stage vendor. Comprehension is not the bottleneck; evidence is.
- **The page picked the wrong flagship customer.** *(high)*
  The single Endress+Hauser reference draws fire on two fronts at once: two respondents say a €3.7B manufacturer does not generalize to smaller SaaS buyers, and three plus a fourth want independent references beyond it.
- **The mechanism is explained but never stress-tested, so the explanation invites the objection.** *(high)*
  Three respondents repeat the ARR-decay scoring back accurately, and four immediately attack it on data quality — vague tier defaults for missing CRM fields, unclear cross-channel weighting, no hygiene story, curated rather than live validation.
- **Buyers are pricing in vendor risk the page never addresses.** *(high)*
  Six respondents inferred a small, founder-led early-stage company, and half of them treated the thin proof points and academic-style claims as a credibility problem rather than neutral positioning.
- **The board-meeting value promise is unusable without the sourcing the page withholds.** *(medium)*
  Two respondents value auditable revenue linkage for removing gut-feel from board conversations, but five say the statistics lack methodology, sample size and attribution — the same audit trail a board would demand.

---

## 06 · Who answered

| # | Role | Industry | Company size |
| --- | --- | --- | --- |
| 1 | Senior Product Manager | SaaS | 201-500 |
| 2 | Director of Product | B2B software | 51-200 |
| 3 | Head of Product | Software as a Service | 201-500 |
| 4 | VP of Product | SaaS | 51-200 |
| 5 | Senior Product Manager | B2B software | 201-500 |
| 6 | Director of Product | Software as a Service | 51-200 |
| 7 | Head of Product | SaaS | 201-500 |
| 8 | VP of Product | B2B software | 51-200 |
| 9 | Senior Product Manager | Software as a Service | 201-500 |
| 10 | Director of Product | SaaS | 51-200 |
| 11 | Head of Product | B2B software | 201-500 |
| 12 | VP of Product | Software as a Service | 51-200 |
| 13 | Senior Product Manager | SaaS | 201-500 |
| 14 | Director of Product | B2B software | 51-200 |
| 15 | Head of Product | Software as a Service | 201-500 |

---

## 07 · Before you act on this

The methodology is real, and the critique is directional. What a
simulated persona cannot have is a live budget, a renewal coming up, or
a boss asking about this quarter. **Validate anything you're betting on
with real ICPs who are actually in-market.** Being wrong is more
expensive than you think. Finding out is cheaper than you'd guess.

Wynter runs message testing with verified B2B professionals — trusted
by HubSpot, RingCentral, Shopify, Cognism, Paddle, Veeam, Rippling and
Miro. <https://wynter.com>

This report is kept for 60 days from 2026-09-06, then deleted along with the personas and their answers.

