# Message test — https://arcate.io/

After reading your page, only 9 of 15 personas could name what kind of product this is, unprompted.

- **Page tested:** https://arcate.io/
- **Audience tested against:** For Heads of Product and VPs of Product at B2B SaaS companies, Series A to C, 50 to 500 employees. They own the product roadmap and defend product decisions to leadership. They deal with customer feedback scattered across multiple channels, and lack a systematic way to connect product bets to revenue at risk.
- **Personas:** 15 simulated
- **Report:** https://grader.wynter.com/r/arcate-agentic-product-intelligence-for-b2b-te-90fAYsA

> These answers are generated by AI, scored on Wynter's B2B Message
> Layers framework using behaviorally-diverse simulated personas. The
> methodology is real and the critique is directional. What a simulated
> persona cannot have is a live budget, a renewal coming up, or a boss
> asking about this quarter.

---

## 01 · The scores

Every persona answered all four questions. These are four independent
proportions of the same panel, not stages of a funnel.

| Layer | Question | Cleared the bar | Strength | Of those who passed |
| --- | --- | --- | --- | --- |
| 1. Clarity | Do they understand what you do? | 9/15 | 85% | 5 without hesitation, 10 with reservations |
| 2. Relevance | Can they tell what it solves, and who it's for? | 15/15 | 100% | all without hesitation |
| 3. Value | Do they actually want it? | 14/15 | 74% | all with reservations |
| 4. Differentiation | Is there a reason to pick you over the alternatives? | 11/15 | 63% | all with reservations |

**Brand alignment** (a side metric, not one of the four layers) — 15/15, 78% strength (all with reservations). Does the page read like the company you actually are?

**Fix first: Clarity.** Earliest failing layer, walking the sequence in order — not simply the lowest score.

---

## 02 · What to change, layer by layer

Ordered worst-first. Specific edits, not a restatement of the score.

### Clarity

**Add a source line under the τ and Jaccard stats naming the dataset and runs.**

"τ = 0.924" and "Jaccard 1.000" appear at the top with no dataset, no definition and no way to check them, and "Method and data available on request" pushes the reader away. Say in one line what was scored, by whom, over how many runs, and link the method.

*effort medium · impact high · tested against Proof next to the claim*

**Replace the H1 category label with a plain description of what Arcate does.**

"Agentic product intelligence for B2B teams" tells a reader nothing about the product, and the page later calls it a scoring engine, a roadmap tool and an MCP server. Name one thing in the H1: software that ranks your roadmap by the revenue behind each…

*effort low · impact high · tested against Plain language*

**Cut "Autonomous AI agents" from the comparison table unless the page shows what runs unattended.**

The table claims autonomous agents but every described step is ingest, score and rank on a schedule. State what happens without a human in the loop, or describe the work plainly as continuous ingestion and scoring.

*effort low · impact medium · tested against Concrete over abstract*

### Differentiation

**Add a line under the Endress+Hauser result separating tool impact from process change.**

One customer with one metric reads as a cherry-picked pilot, and nothing says what else changed during those six months. State what stayed constant, the baseline period, and who measured it.

*effort medium · impact high · tested against Answer the live objection*

**Add a pilot offer beside "Book a diagnostic call" that runs on the buyer's own CRM.**

Buyers will not trust ranking quality until it runs against their own messy CRM records, and the only next step is a call. Offer a defined trial on their data and say how long it takes and what they get back.

*effort medium · impact high · tested against Answer the live objection*

**Rewrite the "Traditional PM Tools (Self-Serve)" column header to name the actual alternatives.**

A buyer comparing Arcate to Productboard or a spreadsheet cannot tell which one the column describes, so the contrast reads as a straw man. Name the tools and the workflow being replaced.

*effort low · impact medium · tested against Give a reason to choose you*

### Value

**Show a sample ranked roadmap output near "You receive finished results."**

The page describes the deliverable but never shows it, so a buyer cannot picture what lands in front of the board. Show one ranked list with a quote, an ARR figure and a score.

*effort medium · impact high · tested against Show the product early*

---

## 03 · What is working

### The core mechanism — revenue-weighted prioritization from CRM and feedback signals — is…

Respondents restated the product as weighting sales and feedback signals by ARR to rank the roadmap. Several named the quote-to-ARR audit trail as what makes a board answer defensible.

> a tool that ingests customer feedback signals from places like Slack, Intercom, Gong and your CRM, weights them by the ARR of the account and whether it's a deal-loss versus a feature mention, and spits out a ranked product roadmap
> 
> — Head of Product, B2B Software, 201-500

> It's a tool that hoovers up customer feedback signals from Slack, Intercom, Gong, Salesforce, HubSpot, weights each one by the ARR of the account behind it, and spits out a ranked product roadmap
> 
> — VP of Product, Enterprise Software, 51-200

> the quote-to-ARR audit trail is the specific thing that would save me, not the roadmap itself
> 
> — Head of Product, B2B Software, 201-500

> I'd stop having the same argument every quarter where sales says "the €500K account is churning over this" and I've got no way to weigh that against forty feature-mention Slack messages — the roadmap would already be sorted by revenue at risk
> 
> — Head of Product, Enterprise Software, 201-500

### The headline and subhead land the audience and the pain immediately

Six respondents said the page names its reader — product leaders defending roadmap decisions to the board — and the problem in the opening lines without hunting.

> the subhead literally says "For product leaders who defend product decisions to the board," and the opening line names the problem: "Sales holds the signals. Product holds the roadmap. Revenue connects neither."
> 
> — Head of Product, B2B Software, 201-500

> the subhead literally says "For product leaders who defend product decisions to the board," and the next line, "Get a ranked roadmap backed by customer ARR," tells you the problem and audience in the first five seconds
> 
> — VP of Product, Enterprise Software, 51-200

> the subhead says "For product leaders who defend product decisions to the board" and the line right after, "Get a ranked roadmap backed by customer ARR," tells you the problem (unjustifiable, gut-feel roadmaps) and the reader (product leaders answering to a board) in the first two lines.
> 
> — Senior Product Manager, SaaS, 201-500

> "For product leaders who defend product decisions to the board" is the intended reader stated in plain terms, and the problem is right there too
> 
> — Senior Product Manager, B2B Software, 201-500

> The tone does feel written for someone like me specifically — "For product leaders who defend product decisions to the board" and "Sales holds the signals. Product holds the roadmap. Revenue connects neither" are lines that assume I already live this problem
> 
> — Head of Product, B2B Software, 201-500

### The Endress+Hauser CES 3.53 to 1.47 number is the proof point respondents trusted

Four respondents singled out the named customer and specific metric as more credible than generic logos or simulation-style statistics, and treated it as a real-world reference they could check.

> the CES 3.53 → 1.47 Endress+Hauser number gives me one real-world reference point to press on
> 
> — Head of Product, B2B Software, 201-500

> The thing that would actually tip me toward this one over a competitor is the Endress+Hauser reference with a specific, named metric — "CES 3.53 → 1.47 in six months," independently validated by their CX team
> 
> — Head of Product, B2B Software, 201-500

> Independently validated by Endress+Hauser CX team" is a named company, a named metric, and a named validator, which is more than I usually get
> 
> — VP of Product, Enterprise Software, 51-200

> The one thing that would tip me toward this over a competitor is the Endress+Hauser line — "CES 3.53 → 1.47 in six months... Independently validated by Endress+Hauser CX team" — because it's a named, sizeable industrial company (€3.7B revenue) with a concrete before/after metric, not just a stats claim.
> 
> — Senior Product Manager, SaaS, 201-500

---

## 04 · What the personas said

### Jargon and shifting category naming obscure a simple concept

Three respondents said the copy buries a straightforward idea under jargon, and that the product's category name changes across the page. One noted the agentic label is used without any autonomous behavior shown.

> it's used as a label without ever showing me an agent doing anything autonomous beyond ingest-and-score, so it reads more like a buzzword bolted onto a fairly conventional scoring pipeline
> 
> — Head of Product, B2B Software, 201-500

> the confusion was more that it stacks a lot of jargon ('agentic,' 'MCP server,' 'Kendall's τ,' 'Jaccard') on top of a fairly simple idea, so I had to mentally strip out the stats-jargon to get to the plain-English pitch
> 
> — VP of Product, Enterprise Software, 51-200

> what's slippery is the naming, they call it "Agentic product intelligence" and "customer signal-to-roadmap intelligence" in different places, and that kind of shifting label makes it harder to say in one line what category it belongs to.
> 
> — Senior Product Manager, SaaS, 201-500

### The lead statistics are unusable without dataset, methodology, or attribution

Three respondents said Kendall's τ and Jaccard figures arrive with no public methodology, no dataset definition, and no explanation of PM attribution or validation.

> "τ = 0.924" and "Jaccard = 1.000" are dropped like everyone knows what a 60-run simulation against "Senior PM judgment" actually consisted of
> 
> — Senior Product Manager, B2B Software, 201-500

### The page never shows the artifact the buyer would actually use or reckon with the buying…

Respondents noted there is no specific output, presentation format, or usage moment depicted, and that integration complexity implies multi-stakeholder approval the page does not address. One found the technical founder voice at odds with the enterprise pitch.

> It would need a line that names my actual reporting moment — something like a quarterly business review or board deck — alongside a specific artifact I'd walk in with, not just "defend product decisions," but "here's the slide you'd show."
> 
> — Head of Product, B2B Software, 201-500

> The "MCP server," "CP 1.1/1.2/1.3," and stats dressed up as rigor (τ, Jaccard) read like a technical founder writing for other builders, not a polished go-to-market team selling to VPs.
> 
> — Senior Product Manager, Enterprise Software, 201-500

> But it's Slack/Intercom/Gong/Salesforce/HubSpot integration plus CRM ARR hygiene — that's not a quick swap, and I'd need IT, RevOps, and sales leadership to sign off on data access before I even see a real output.
> 
> — VP of Product, SaaS, 51-200

### One case study is not enough proof and reads as an early-stage vendor leaning on a…

Six respondents flagged the lone reference as possibly a cherry-picked pilot that cannot isolate tool impact from process change, and read the single flagship logo as pre-Series A positioning.

> a small, early-stage B2B SaaS vendor — probably under 20 people, maybe even single-digit, pre-Series A or just past it — selling to mid-market product leaders
> 
> — Head of Product, B2B Software, 201-500

> what would tip me toward Arcate specifically is a second reference customer with a named metric, because right now one Endress+Hauser data point could just as easily be a fluke or a cherry-picked pilot
> 
> — Head of Product, B2B Software, 201-500

> a couple of years old, still leaning on one flagship case study (Endress+Hauser) because they don't have a bench of logos yet
> 
> — Senior Product Manager, B2B Software, 201-500

> What would tip it their way instead is if either of them could hand me a live reference customer I could call this week, or show the same severity-multiplier-style transparency — right now Arcate's specificity edges them out, but only until someone matches the rigor with an actual phone call I can make
> 
> — VP of Product, Enterprise Software, 51-200

> I'd guess a small, early-stage B2B SaaS vendor — maybe 10-30 people, a couple of years old, still in the "one flagship logo" phase given they lean so hard on Endress+Hauser as their proof point rather than a roster of customers
> 
> — Head of Product, SaaS, 201-500

> Small, early-stage B2B SaaS vendor — feels like a seed/Series A shop with maybe one flagship logo (Endress+Hauser) they're leaning on hard because they don't have three more to point to yet.
> 
> — Director of Product, SaaS, 51-200

### Respondents will not believe the result until it runs on their own messy CRM data

Three respondents said a demo or cited case is insufficient and required validation against their own data, plus confirmation that multiplier tier defaults are genuinely tunable.

> I'd go in wanting to see the "configurable tier defaults" and severity-weight editing live, because that's where a generic 30×/3×/1× multiplier either maps to my actual deal sizes or falls apart
> 
> — Head of Product, B2B Software, 201-500

> Show me it correctly ranking a real messy pull from our own CRM and Slack history — not their data — with the deal-loss account actually surfacing above the noise; if that ranking survives contact with our own patchy ARR fields and stale records, that's the one outcome that earns a second meeting
> 
> — Head of Product, Enterprise Software, 201-500

---

## 05 · The hardest read

An adversarial pass over the findings. Every claim below was checked
against the panel's own answers; unsupported ones were dropped.

- **The page's proof rests on a single customer, and that concentration converts its one strong asset into a liability.** *(high)*
  Six respondents flagged the lone reference as a possibly cherry-picked pilot reading as pre-Series A positioning, while the four who trusted the page cited that same Endress+Hauser metric. Remove it and nothing credible remains.
- **The statistics meant to add rigor actively undermine credibility instead of reinforcing it.** *(high)*
  Respondents rejected Kendall's τ and Jaccard figures for lacking methodology, dataset, and attribution, and separately preferred the named customer metric over "simulation-style statistics." The quantitative display reads as decoration.
- **Comprehension does not convert: respondents who correctly restated the mechanism still refused to believe it works.** *(high)*
  Four respondents accurately described revenue-weighted prioritization and named the quote-to-ARR audit trail as defensible, yet three demanded validation on their own messy CRM data and proof that multiplier tiers are tunable. Understanding is not the…
- **The page sells a decision aid but never shows the decision artifact, so the board-defense promise stays abstract.** *(high)*
  Respondents said no specific output, presentation format, or usage moment is depicted, even though the headline explicitly targets product leaders defending roadmap decisions to the board. The claimed moment of value is the one thing missing.
- **Every objection raised points at the same missing thing: evidence the buyer can verify independently.** *(high)*
  Unattributed statistics, one unreplicated case study, no visible output, and demands to test on their own CRM all describe unverifiable claims. The page asks for trust it never earns.
- **Strong audience targeting is squandered by inconsistent naming that makes the product unrecognizable further down the page.** *(medium)*
  Six respondents said the headline names the reader and pain immediately, but three said the category name shifts across the page and the agentic label appears with no autonomous behavior shown. The opening earns attention the body loses.

---

## 06 · Who answered

| # | Role | Industry | Company size |
| --- | --- | --- | --- |
| 1 | Head of Product | B2B Software | 201-500 |
| 2 | VP of Product | Enterprise Software | 51-200 |
| 3 | Senior Product Manager | SaaS | 201-500 |
| 4 | Director of Product | B2B Software | 51-200 |
| 5 | Head of Product | Enterprise Software | 201-500 |
| 6 | VP of Product | SaaS | 51-200 |
| 7 | Senior Product Manager | B2B Software | 201-500 |
| 8 | Director of Product | Enterprise Software | 51-200 |
| 9 | Head of Product | SaaS | 201-500 |
| 10 | VP of Product | B2B Software | 51-200 |
| 11 | Senior Product Manager | Enterprise Software | 201-500 |
| 12 | Director of Product | SaaS | 51-200 |
| 13 | Head of Product | B2B Software | 201-500 |
| 14 | VP of Product | Enterprise Software | 51-200 |
| 15 | Senior Product Manager | SaaS | 201-500 |

---

## 07 · Before you act on this

The methodology is real, and the critique is directional. What a
simulated persona cannot have is a live budget, a renewal coming up, or
a boss asking about this quarter. **Validate anything you're betting on
with real ICPs who are actually in-market.** Being wrong is more
expensive than you think. Finding out is cheaper than you'd guess.

Wynter runs message testing with verified B2B professionals — trusted
by HubSpot, RingCentral, Shopify, Cognism, Paddle, Veeam, Rippling and
Miro. <https://wynter.com>

This report is kept for 60 days from 2026-09-08, then deleted along with the personas and their answers.

