# Message test — https://arcate.io/

After reading your page, only 8 of 15 personas could name what kind of product this is, unprompted.

- **Page tested:** https://arcate.io/
- **Audience tested against:** For Heads of Product and VPs of Product at B2B SaaS companies, Series A to C, 50 to 500 employees. They own the product roadmap and defend prioritization decisions to leadership. They deal with customer feedback scattered across various channels, and lack a systematic way to connect product bets to revenue at risk.
- **Personas:** 15 simulated
- **Report:** https://grader.wynter.com/r/arcate-agentic-product-intelligence-for-b2b-te-DMy2rHA

> These answers are generated by AI, scored on Wynter's B2B Message
> Layers framework using behaviorally-diverse simulated personas. The
> methodology is real and the critique is directional. What a simulated
> persona cannot have is a live budget, a renewal coming up, or a boss
> asking about this quarter.

---

## 01 · The scores

Every persona answered all four questions. These are four independent
proportions of the same panel, not stages of a funnel.

| Layer | Question | Cleared the bar | Strength | Of those who passed |
| --- | --- | --- | --- | --- |
| 1. Clarity | Do they understand what you do? | 8/15 | 84% | 4 without hesitation, 11 with reservations |
| 2. Relevance | Can they tell what it solves, and who it's for? | 15/15 | 90% | 8 without hesitation, 7 with reservations |
| 3. Value | Do they actually want it? | 14/15 | 74% | all with reservations |
| 4. Differentiation | Is there a reason to pick you over the alternatives? | 11/15 | 63% | all with reservations |

**Brand alignment** (a side metric, not one of the four layers) — 13/15, 70% strength (all with reservations). Does the page read like the company you actually are?

**Fix first: Clarity.** Earliest failing layer, walking the sequence in order — not simply the lowest score.

---

## 02 · What to change, layer by layer

Ordered worst-first. Specific edits, not a restatement of the score.

### Clarity

**Attach methodology to the Kendall's τ and Jaccard figures.**

"Kendall's τ = 0.924 ... across 60 simulation runs" reads as unsourced noise. State what was simulated, against whose judgment, over how many roadmap items, and link the write-up beside the number.

*effort medium · impact high · tested against Proof next to the claim*

**Replace 'Agentic product intelligence for B2B teams' in the H1.**

'Agentic' arrives before anything explains it, and 'product intelligence' is a category label. Lead with the job: ranking the roadmap by the revenue at risk behind each request.

*effort low · impact high · tested against Lead with the use case*

**Say who sets the weights in 'Calibrated weights'.**

"Calibrated weights stop single accounts from hijacking your roadmap" hides the owner of the decision. Write who configures severity multipliers and whether the PM can change them.

*effort low · impact medium · tested against Plain language*

### Differentiation

**Add a mid-market SaaS reference beside the Endress+Hauser case.**

A €3.7B manufacturer is the only evidence on the page, and a 51-200 person software team reads it as the wrong company. Name a SaaS customer with headcount and outcome.

*effort high · impact high · tested against Proof next to the claim*

**Answer the 'one case study' objection under 'Validated at scale'.**

Competitors showing three references win on evidence alone. State how many teams run Arcate today and what results they see, rather than resting on a single deployment.

*effort medium · impact medium · tested against Answer the live objection*

**Name the audience in the hero: B2B SaaS product managers.**

The page targets PMs only by implication. A line the reader can point at — the PM defending a roadmap to a board — makes the mismatch with the manufacturing proof less jarring.

*effort low · impact medium · tested against Name the audience*

### Brand alignment (side metric)

**Ground 'Validated at scale' in customers of the reader's size.**

The heading promises scale and delivers one enterprise manufacturer plus unlabelled statistics, which undercuts the rigorous tone. Show a company-size range you serve.

*effort medium · impact medium · tested against Specifics beat superlatives*

---

## 03 · What is working

### The scoring mechanism is understood and repeated back accurately

Ten respondents restated the mechanic in their own words — revenue-weighted prioritization across five data sources, ARR and decay scoring, severity multipliers, CRM integration, ranking roadmap items by revenue impact.

> It's a tool that ingests customer feedback from Slack, Intercom, Gong, Salesforce and HubSpot, scores it by ARR at risk, and spits out a ranked product roadmap
> 
> — VP of Product, Software, 51-200

> It's a tool that pulls customer feedback out of Slack, Intercom, Gong, Salesforce and HubSpot, weights each signal against the account's ARR, and spits out a revenue-ranked product roadmap with a traceable line back to the original customer quote.
> 
> — Senior Vice President of Product, B2B, 201-500

> the step-by-step in "How it works" (connect channels, score by ARR with the 30x/3x/1x multipliers, rank and decay over time) made the mechanism pretty concrete
> 
> — Head of Product, SaaS, 51-200

> the subhead "Sales holds the signals. Product holds the roadmap. Revenue connects neither" and "You get a ranked roadmap backed by customer ARR" tell you the problem inside the first screen
> 
> — Senior Vice President of Product, B2B, 51-200

> ranks product roadmap items by revenue-at-risk instead of gut feel — basically an ARR-weighted prioritization layer
> 
> — Senior Vice President of Product, B2B, 51-200

> The Endress+Hauser case (CES drop 3.53 to 1.47, 30% sales capacity freed) is the one concrete proof point that makes me think it does something real
> 
> — Senior Vice President of Product, B2B, 51-200

### The roadmap-defense problem lands as the reader's own problem

Five respondents said the opening screen makes the problem and audience immediately obvious, and that defending roadmaps with data instead of gut feel is exactly the product manager's pain.

> the subhead "Sales holds the signals. Product holds the roadmap. Revenue connects neither" tells you the problem in one line
> 
> — Head of Product, SaaS, 51-200

> "Sales holds the signals. Product holds the roadmap. Revenue connects neither" tells me the problem in one line, and the table comparing Arcate to "Traditional PM Tools (Self-Serve)" nails who this is for without me having to guess too hard
> 
> — VP of Product, Software, 201-500

> the subhead "Sales holds the signals. Product holds the roadmap. Revenue connects neither" and "You get a ranked roadmap backed by customer ARR" tell you the problem inside the first screen
> 
> — Senior Vice President of Product, B2B, 51-200

> The reader is clearly a Head of Product or PM who has to defend a roadmap to a board — "Show the board exactly why you built it"
> 
> — Head of Product, SaaS, 201-500

> If it worked as promised, my next roadmap review with leadership stops being "trust me, sales was loud about this" and becomes a chain from customer quote to ARR to score to bet — that directly kills the "unjustifiable roadmap" problem I actually live with.
> 
> — VP of Product, Software, 201-500

---

## 04 · What the personas said

### Statistics are dismissed because no methodology is shown

Four respondents discounted the precision metrics and simulation-run claims outright, calling them unsourced noise rather than evidence, and said the numbers should be ignored absent methodology detail.

> the Kendall's τ = 0.924 and Jaccard = 1.000 numbers have no methodology behind them so I'd discount those until I saw the actual simulation setup
> 
> — VP of Product, Software, 51-200

> "60 simulation runs" is vague enough to rule it back out if I dug in and found it was a synthetic/internal test rather than validated on real customer data — I'd want to know whose judgment, on what dataset, before I let that number carry weight against another vendor's live customer references.
> 
> — VP of Product, Software, 201-500

### 'Agentic' and the calibration language leave readers guessing

Two respondents flagged unclear copy: 'Agentic' is used before it is explained, and vague calibration wording hides who actually owns the weighting decisions.

> Phrases like "calibrated weights" and "attribution is visible and editable" sound precise but don't say who calibrates them or edits them, which is exactly the kind of soft language that gets papered over in a demo and falls apart on real data.
> 
> — VP of Product, Software, 201-500

### The buyer's job title is never named on the page

Three respondents noted the copy targets product managers only by implication, requiring reader inference rather than stating the title or a matching SaaS proof point.

> A line naming the buyer directly — something like "Built for VPs of Product who answer to revenue and the board" — plus a proof point from a company my size and sector, not just Endress+Hauser
> 
> — Senior Vice President of Product, B2B, 201-500

> I'd want my actual title or a line like "built for VPs of Product reporting to the board" instead of me inferring it from "defend unjustifiable roadmaps alone" — right now I'm doing the work of mapping myself onto the copy rather than the copy doing it for me.
> 
> — VP of Product, Software, 201-500

> The reader is never explicitly named as "VP Product" or "Head of Product," but the language — PMs, roadmaps, board slides, "PMs left to defend unjustifiable roadmaps alone" — makes it obvious within seconds
> 
> — Senior Vice President of Product, B2B, 51-200

### One manufacturing case study is not enough to justify changing workflows

Five respondents said a single Endress+Hauser reference plus unsourced stats cannot prove viability, especially for mid-market SaaS, and would need methodology or additional references first.

> one customer case study and two unsourced stats (Kendall's τ, Jaccard) aren't enough to bet a workflow change on — I'd want two or three more named customers my size
> 
> — VP of Product, Software, 51-200

> But one case study at a €3.7B industrial company isn't the same as evidence this works for a 51-200 person B2B shop like mine
> 
> — Senior Vice President of Product, B2B, 51-200

> The Endress+Hauser number (CES 3.53 to 1.47, 30% technical sales capacity freed) is the kind of proof that makes me want to test it rather than dismiss it, but that's a manufacturing account, not a SaaS org like mine
> 
> — Senior Vice President of Product, B2B, 201-500

> one manufacturing case study propping up the whole page isn't enough to differentiate them from a rival with three relevant references
> 
> — Senior Vice President of Product, B2B, 201-500

### The only proof point is the wrong company size and industry

Four respondents rejected the enterprise manufacturing reference as mismatched to a 51-200 person software shop and said competitors offering three references would win on evidence.

> that's a €3.7B industrial company, not a 51-200 person software shop, so I'd need a similarly-sized SaaS customer story to actually trust it applies to me
> 
> — VP of Product, Software, 51-200

> one manufacturing case study propping up the whole page isn't enough to differentiate them from a rival with three relevant references
> 
> — Senior Vice President of Product, B2B, 201-500

> the only proof point they lean on is a €3.7B company, which doesn't tell me they've ever sold to or succeeded with a 51-200 person shop like mine
> 
> — Senior Vice President of Product, B2B, 51-200

> But one case study at a €3.7B industrial company isn't the same as evidence this works for a 51-200 person B2B shop like mine
> 
> — Senior Vice President of Product, B2B, 51-200

### Nothing on the page shows the company has sold to mid-market

Two respondents said the unsourced precision metrics undercut the rigorous positioning and that no evidence demonstrates successful sales to companies of their size.

> the only proof point they lean on is a €3.7B company, which doesn't tell me they've ever sold to or succeeded with a 51-200 person shop like mine
> 
> — Senior Vice President of Product, B2B, 51-200

---

## 05 · The hardest read

An adversarial pass over the findings. Every claim below was checked
against the panel's own answers; unsupported ones were dropped.

- **The page has no admissible evidence — every number and reference on it was rejected.** *(high)*
  Three respondents dismissed the precision metrics and simulation-run claims as unsourced noise, four called the single Endress+Hauser case insufficient, and three rejected it as the wrong size and industry. The entire proof layer collapses.
- **Clarity on the mechanic is worthless because it converts into disbelief rather than credibility.** *(high)*
  Four respondents restated the revenue-weighted scoring accurately, yet four still said one case study plus unsourced stats cannot prove viability. Readers understand exactly what is claimed and refuse to believe it.
- **The page loses the deal at the comparison stage, not the comprehension stage.** *(high)*
  Three respondents said competitors offering three references would win on evidence, and two saw nothing demonstrating sales to companies of their size. Differentiation rests entirely on a proof point readers disqualify.
- **Mid-market SaaS readers are given no path to self-identify anywhere on the page.** *(high)*
  Three respondents said the title is never stated and no matching SaaS proof point exists, three rejected the enterprise manufacturing reference as mismatched to a 51-200 person shop, and two found no evidence of mid-market sales.
- **Strong problem framing is squandered by making the reader do the qualifying work.** *(medium)*
  Four respondents said the roadmap-defense pain lands immediately, but three noted the buyer's title is never named and only implied, with no matching SaaS proof point. The page earns attention and then fails to confirm it.
- **Unexplained jargon compounds the credibility gap by hiding accountability.** *(medium)*
  Two respondents flagged 'Agentic' used before definition and vague calibration wording that obscures who owns weighting decisions. Ambiguity about control sits directly on top of metrics three respondents already called unsourced.

---

## 06 · Who answered

| # | Role | Industry | Company size |
| --- | --- | --- | --- |
| 1 | VP of Product | Software | 51-200 |
| 2 | Senior Vice President of Product | B2B | 201-500 |
| 3 | Head of Product | SaaS | 51-200 |
| 4 | VP of Product | Software | 201-500 |
| 5 | Senior Vice President of Product | B2B | 51-200 |
| 6 | Head of Product | SaaS | 201-500 |
| 7 | VP of Product | Software | 51-200 |
| 8 | Senior Vice President of Product | B2B | 201-500 |
| 9 | Head of Product | SaaS | 51-200 |
| 10 | VP of Product | Software | 201-500 |
| 11 | Senior Vice President of Product | B2B | 51-200 |
| 12 | Head of Product | SaaS | 201-500 |
| 13 | VP of Product | Software | 51-200 |
| 14 | Senior Vice President of Product | B2B | 201-500 |
| 15 | Head of Product | SaaS | 51-200 |

---

## 07 · Before you act on this

The methodology is real, and the critique is directional. What a
simulated persona cannot have is a live budget, a renewal coming up, or
a boss asking about this quarter. **Validate anything you're betting on
with real ICPs who are actually in-market.** Being wrong is more
expensive than you think. Finding out is cheaper than you'd guess.

Wynter runs message testing with verified B2B professionals — trusted
by HubSpot, RingCentral, Shopify, Cognism, Paddle, Veeam, Rippling and
Miro. <https://wynter.com>

This report is kept for 60 days from 2026-09-04, then deleted along with the personas and their answers.

