# Message test — https://coworker.ai/

After reading your page, only 8 of 15 personas could name a reason to pick you over a similar option.

- **Page tested:** https://coworker.ai/
- **Audience tested against:** Operations and IT leaders
- **Personas:** 15 simulated
- **Report:** https://grader.wynter.com/r/enterprise-ai-agents-for-every-task-coworker-a-R9tjLXA

> These answers are generated by AI, scored on Wynter's B2B Message
> Layers framework using behaviorally-diverse simulated personas. The
> methodology is real and the critique is directional. What a simulated
> persona cannot have is a live budget, a renewal coming up, or a boss
> asking about this quarter.

---

## 01 · The scores

Every persona answered all four questions. These are four independent
proportions of the same panel, not stages of a funnel.

| Layer | Question | Cleared the bar | Strength | Of those who passed |
| --- | --- | --- | --- | --- |
| 1. Clarity | Do they understand what you do? | 15/15 | 78% | all with reservations |
| 2. Relevance | Can they tell what it solves, and who it's for? | 11/15 | 63% | all with reservations |
| 3. Value | Do they actually want it? | 13/15 | 70% | all with reservations |
| 4. Differentiation | Is there a reason to pick you over the alternatives? | 8/15 | 52% | all with reservations |

**Brand alignment** (a side metric, not one of the four layers) — 8/15, 52% strength (all with reservations). Does the page read like the company you actually are?

**Fix first: Differentiation.** Earliest failing layer, walking the sequence in order — not simply the lowest score.

---

## 02 · What to change, layer by layer

Ordered worst-first. Specific edits, not a restatement of the score.

### Differentiation

**Replace the single case study with two or three customer results naming company size and stack.**

One 300-person customer does not convince an enterprise buyer that this survives their scale. Add deployments with headcount, connector set and measured cost change, or say plainly what size of org is live today.

*effort high · impact high · tested against Proof next to the claim*

**Add benchmark methodology link beside the 9x, 51x and 84.5% numbers in the hero.**

The footnote "Statistically significant benchmarks comparing Coworker MCP vs. Claude Native Tooling" tells a buyer nothing about workload mix, sample size or who ran the test. State the task set, number of runs and date next to the numbers, and link the full…

*effort medium · impact high · tested against Proof next to the claim*

**Pull permission-aware routing, data residency and no-retraining into a named section above connectors.**

The things buyers can only get here are buried in body copy under "Model Routing" while the page leads with price multipliers any vendor could claim. Give permission-aware access, US hosting and no-training-on-your-data their own headed block with one line…

*effort medium · impact high · tested against Give a reason to choose you*

### Relevance

**Name the buyer and team in the hero subhead, above the logo strip.**

Nothing on the page says who this is for, so the reader reverse-engineers it from logos. Write a line naming the role and situation, such as IT and platform leads running Slack, Salesforce and Jira across a few thousand seats.

*effort low · impact high · tested against Name the audience*

**Define OM2 in plain words at first use before the "Introducing OM2" section.**

"OM2 layer deeply understands your enterprise" appears in the first sentence with no explanation, and "knowledge graph" arrives much later. Say in the hero what OM2 is and what a query does end to end.

*effort low · impact high · tested against Plain language*

**Add one worked example under the hero showing a real question, sources used, answer.**

A buyer cannot tell from "right context, right model" what they would actually get back. Show a question, the connectors it touched and the output, in plain English before any claims.

*effort medium · impact medium · tested against Concrete over abstract*

### Value

**Add the baseline and workload behind "9x cheaper" directly under the number.**

Nine times cheaper than what, on which tasks, is never stated, so the number reads as a marketing figure. Name the comparison setup and the task mix beside the claim.

*effort low · impact high · tested against Proof next to the claim*

### Clarity

**Collapse "Book a demo" and "See benchmarks" into one primary action in the hero.**

Two equal buttons split the reader at the moment they are deciding. Make "Book a demo" the button and demote benchmarks to a text link under the numbers.

*effort low · impact low · tested against One clear next action*

### Brand alignment (side metric)

**Rewrite "Your enterprise just started thinking" and the closing "enterprise brain" line in plain terms.**

The brain metaphor and "the platform behind the enterprise brain" read as marketing air next to logos that look mid-market. Say what the layer does, such as answering questions using your company's Slack, Drive, Salesforce and Jira with the user's own…

*effort low · impact medium · tested against Plain language*

**Add a qualifying line under "Trusted by" stating the size range of current customers.**

Logos invite a Fortune 500 reading the customer list does not support, and the gap costs trust. State the headcount band and departments running Coworker today so the claim matches the evidence.

*effort low · impact medium · tested against Specifics beat superlatives*

**Replace "Enterprise ready·SOC 2·50+ connectors" strip with specifics on residency, retention and SSO.**

"Enterprise ready" is a label the reader has to take on faith while the security questions they actually hold go unanswered. Name SOC 2 Type II status, US hosting, data retention policy and SSO/SCIM support.

*effort low · impact medium · tested against Answer the live objection*

---

## 03 · What is working

### The core routing mechanic lands when stated plainly

Respondents repeated back that the platform routes requests to the cheapest or best AI model across existing tools, and that cost savings plus workflow preservation apply to large organizations. The hero conveys problem and buyer, though urgency proof is…

> the problem is clear but the urgency/proof is something I'd have to chase down rather than something the page hands me.
> 
> — Director of IT, Enterprise Software, 5000+

> If the 51x cheaper and 84.5% better quality numbers held up in my own environment, that's real money — we're a 1000+ person shop burning real budget on model calls across teams, and if routing actually cuts that without ripping out Slack/Jira/Salesforce workflows, that's a legitimate line-item win
> 
> — IT Leader, Software Development, 1001-5000

> It's an AI layer that sits across your existing tools (Slack, Salesforce, Jira, etc.) building a company-wide "memory" and then routing each request to whichever model — Claude, GPT, Gemini — is cheapest or best for that task.
> 
> — Director of Operations, Technology, 1001-5000

### Permission-aware routing and data residency are the differentiators respondents could…

Four respondents identified permission-aware routing without SSO sprawl, data residency controls, no-retraining, and model routing flexibility avoiding lock-in as genuine operational differentiators. One noted the differentiation language still lacks…

> The model-routing bit — "as better models ship, your team gets them automatically, no migration, no vendor lock-in" — is the one thing that would actually tip it over a competitor for me, because lock-in to a single model vendor is exactly the kind of disruption risk I'd be weighing against the other options.
> 
> — Director of IT, Software Development, 501-1000

> that's a concrete security/integration promise I can test in a demo, and it's the kind of detail that rules out competitors who hand-wave on access control
> 
> — VP of IT, Enterprise Software, 1001-5000

> The "permission-aware, agents respect the access controls already on your tools" line plus "no new logins, no new SSO maps" would actually pull me toward it, because integration and access-sprawl is normally the thing that kills these rollouts at our scale
> 
> — Operations Leader, Financial Services, 5000+

> The "no training on data... US-hosted. Contractually enforced" line and "permission-aware, agents respect the access controls already on your tools" are the only things here that would actually move me versus a competitor
> 
> — VP of Operations, Technology, 501-1000

---

## 04 · What the personas said

### Undefined technical jargon blocks understanding of what the product actually does

Four respondents said internal and technical terminology is used upfront without definition, obscuring positioning and how a query is processed end to end. One asked for worked examples in plain English instead of marketing language.

> Terms like "knowledge graph" and "model routing" are dropped without a plain-English worked example on first read, so I had to piece the actual function together from the diagram and MCP code snippet further down rather than the headline copy.
> 
> — Director of IT, Software Development, 501-1000

> It's the layering of jargon like "OM2," "Organizational Memory," "model routing," and "MCP" without ever defining them plainly up front
> 
> — Director of Operations, Professional Services, 201-500

> the top-of-page stuff ("Your enterprise just started thinking," "Right context, right model, anywhere you work") is more growth-marketing poetry than IT-buyer language, so it feels like two different people wrote this page: one who's sold into procurement before, one who's still writing pitch-deck taglines.
> 
> — IT Leader, Software Development, 1001-5000

> those are internal jargon that sound precise but don't tell me what actually happens to a query end to end; I had to piece the mechanism together from the code snippets and demo screenshots, not the prose
> 
> — Operations Leader, Professional Services, 501-1000

### Where the product sits relative to incumbent context layers and what it costs is unclear

One respondent said the product occupies an ambiguous position between existing context layers with unclear pricing versus incumbents.

> I'd need to see where OM2 actually beats my current knowledge layer, not just the cost multipliers.
> 
> — Director of IT, Enterprise Software, 5000+

### The target buyer is never stated and has to be reverse-engineered from logos

Six respondents said the intended audience and use case are only inferred from connector logos and the case study rather than named. One asked for specific stack configurations instead of logo implication.

> the audience was inferred from context clues (security badges, connector list, buyer-type logos) rather than spelled out in a sentence like "built for IT directors."
> 
> — Director of IT, Software Development, 501-1000

> The buyer is never explicitly named — no "for CTOs" or "for ops leaders" banner — but the connector logos (Salesforce, Jira, Slack) and "Enterprise ready·SOC 2·50+ connectors" signal it's aimed at mid-to-large enterprise technical/ops buyers, which I inferred rather than read outright
> 
> — VP of Operations, Technology, 501-1000

> the intended reader is never explicitly named — there's no "for IT leaders" or "for ops teams" banner — I had to infer it from the logos (Salesforce, Jira, Gong), the connector list, and lines like "No new logins, no new SSO maps,"
> 
> — IT Leader, Software Development, 1001-5000

> I'd need to see my own stack named explicitly — something like "if you run Salesforce + Slack + Jira across 500+ seats and your AI spend is split across three vendors, this replaces that" — right now it's implied through logos and connector lists, not stated as a direct address to my situation
> 
> — Operations Leader, Professional Services, 501-1000

> Who it's for is never spelled out explicitly — I inferred "enterprise" from the logos (Salesforce, Jira, SOC 2, "enterprise ready") and the customer story about a Service Desk Team Lead, not from any line that says "this is for IT/Ops leaders at mid-size companies."
> 
> — Director of Operations, Technology, 1001-5000

> There's no "built for VPs of Ops" or "built for IT/platform teams" line — I inferred the buyer from the logos (RapidSOS, Harri, Huuuge), the SOC 2/connector badges, and the Huuuge case study quote
> 
> — VP of Operations, Software Development, 5000+

### The headline cost and quality numbers are not believed because no methodology is shown

Eight respondents flagged the cost multipliers and savings claims as unverified — no methodology, sample size, independent audit or before-and-after customer data. Several said they would want a benchmark write-up or head-to-head test before a meeting.

> The 9x/51x/84.5% numbers are asserted against "Coworker MCP vs. Claude Native Tooling" with no methodology shown, so I'd want the actual benchmark write-up before I believed those figures.
> 
> — VP of IT, Enterprise Software, 1001-5000

> If it worked as promised, I'd be collapsing a chunk of model spend and cutting a lot of manual swivel-chair work — pulling Salesforce/Gong/Slack context by hand, drafting follow-ups, updating opportunity stages — into something automated and cheaper
> 
> — VP of IT, Enterprise Software, 1001-5000

> "statistically significant benchmarks comparing Coworker MCP vs. Claude Native Tooling" is their own comparison, not independent, and there's no sample size, no methodology, nothing I can take into a budget meeting and defend
> 
> — Operations Leader, Financial Services, 5000+

> The 9x/51x/84.5% numbers are the headline hook but there's no methodology shown beyond "statistically significant benchmarks vs. Claude Native Tooling"
> 
> — VP of Operations, Technology, 501-1000

> right now it's not worth a meeting on its own merits — it'd be worth a meeting only to get the benchmark methodology and a reference customer with actual before/after cost data, not the Huuuge "4,000+ hours saved" story which is too soft to act on
> 
> — VP of Operations, Technology, 501-1000

> "Statistically significant benchmarks comparing Coworker MCP vs. Claude Native Tooling" is their own benchmark against their own narrow comparison, not against our actual mixed-model stack or our data volume, so I don't trust the number yet.
> 
> — IT Leader, Software Development, 1001-5000

> The 9x/51x cheaper and 84.5% quality numbers are their own benchmark claims with no independent source, so I'd discount those until I saw who verified them.
> 
> — VP of IT, Financial Services, 201-500

> one big flashy stat with no backup loses it — on balance I'd still take the meeting, but I'd walk in planning to ask for the receipts behind the 51x
> 
> — Director of Operations, Professional Services, 201-500

> I'd walk in assuming incumbency wins unless they can show a head-to-head against my current stack, not just against "doing it manually."
> 
> — Director of IT, Enterprise Software, 5000+

### The single 300-person case study is read as too small to prove enterprise scale

Three respondents said one named customer at 300 people is insufficient evidence for enterprise validation, and asked for reference customers of comparable size and stack before taking a meeting.

> I'd want the actual benchmark methodology behind "84.5% better quality" and a reference customer closer to our size/stack than a mobile games studio, since rolling this across 50+ connectors with existing SSO and permissions is a real migration project, not a plug-in.
> 
> — Director of IT, Software Development, 501-1000

> What would rule it out is the Huuuge case study as the only proof point — 300 employees globally isn't remotely our size, and if that's the best customer evidence they've got, I'd assume they haven't proven this at 5000+ yet
> 
> — Operations Leader, Financial Services, 5000+

> "Maurycy Bielawski, Service Desk Team Lead" saying "4,000+ hours saved... up to 94% time reduction." That's a person and a number I could repeat on a call
> 
> — Director of Operations, Professional Services, 201-500

### Enterprise framing outruns a customer base that reads as mid-market early-stage

Six respondents read the logos as mid-market rather than Fortune 500 and the founder pedigree as early-stage or Series B/C still proving credibility. Two said the enterprise framing is ahead of the actual customer base.

> The logo wall (RapidSOS, Harri, Truecaller, Huuuge) skews mid-market to lower-enterprise, not Fortune 500
> 
> — IT Leader, Software Development, 1001-5000

> "Built by operators from Uber, Google, Apple, Glean, BlackRock, Deloitte, Okta, Plaid, Gainsight" is a pedigree flex you only do when you don't have 10 years of market presence to lean on instead.
> 
> — IT Leader, Software Development, 1001-5000

> The logo wall ("/scale, Harri, RapidSOS, Cortex, Huuuge Games, SafetyWing") tells me they sell mid-market to lower-enterprise — recognizable but not Fortune 500 household names
> 
> — Operations Leader, Professional Services, 501-1000

> the single detailed case study is a 300-person gaming company, not a BlackRock-scale deployment, so the SOC 2/GDPR/enterprise framing is slightly ahead of their actual customer base
> 
> — Operations Leader, Professional Services, 501-1000

> "Built by operators from Uber, Google, Apple, Glean, BlackRock, Deloitte, Okta, Plaid, Gainsight" line is doing a lot of work to signal credibility they haven't fully earned yet on their own brand name
> 
> — Operations Leader, Professional Services, 501-1000

> Tone-wise it's written for a technical buyer who's comfortable with MCP, Claude Code, Cursor — more "head of AI" or engineering lead than a traditional IT leader doing vendor risk review.
> 
> — IT Leader, Enterprise Software, 201-500

---

## 05 · The hardest read

An adversarial pass over the findings. Every claim below was checked
against the panel's own answers; unsupported ones were dropped.

- **The page's central number — the cost claim — is dead on arrival as a meeting trigger.** *(high)*
  Seven to eight respondents rejected the cost multipliers and savings as unverified, citing absent methodology, sample size, and audit, and several named a benchmark write-up or head-to-head test as a precondition for a meeting.
- **The page never tells anyone it is for them, so qualification is offloaded onto the reader.** *(high)*
  Six respondents had to reverse-engineer the audience from connector logos and a case study, and one asked for named stack configurations. Audience inference from logos is the weakest possible targeting mechanism.
- **Evidence and framing contradict each other: the page claims enterprise while the proof reads mid-market.** *(high)*
  Six respondents read the logos as mid-market and the founder pedigree as early-stage, while three called a single 300-person case study too small for enterprise validation. The page is actively undercutting its own positioning.
- **The only things that land are the ones stated plainly; everywhere the page reaches for its own vocabulary, comprehension collapses.** *(high)*
  Three respondents repeated back the routing mechanic correctly when stated simply, but four said undefined internal and technical terminology upfront obscured positioning and end-to-end query processing, with one requesting plain-English worked examples.
- **Every negative theme resolves to the same request: show the work, not the claim.** *(high)*
  Respondents asked for worked examples in plain English, specific stack configurations, benchmark methodology, and reference customers of comparable size — four different gaps, one shared root cause of assertion without substantiation.
- **The differentiators that work are buried beneath the claims that don't, so the page spends its credibility before it earns it.** *(medium)*
  Four respondents named permission-aware routing, data residency and no-lock-in as genuine operational differentiators — yet seven rejected the cost claims and six could not identify the buyer, meaning the strongest asset arrives after trust is gone.

---

## 06 · Who answered

| # | Role | Industry | Company size |
| --- | --- | --- | --- |
| 1 | Director of IT | Software Development | 501-1000 |
| 2 | VP of IT | Enterprise Software | 1001-5000 |
| 3 | Operations Leader | Financial Services | 5000+ |
| 4 | Director of Operations | Professional Services | 201-500 |
| 5 | VP of Operations | Technology | 501-1000 |
| 6 | IT Leader | Software Development | 1001-5000 |
| 7 | Director of IT | Enterprise Software | 5000+ |
| 8 | VP of IT | Financial Services | 201-500 |
| 9 | Operations Leader | Professional Services | 501-1000 |
| 10 | Director of Operations | Technology | 1001-5000 |
| 11 | VP of Operations | Software Development | 5000+ |
| 12 | IT Leader | Enterprise Software | 201-500 |
| 13 | Director of IT | Financial Services | 501-1000 |
| 14 | VP of IT | Professional Services | 1001-5000 |
| 15 | Operations Leader | Technology | 5000+ |

---

## 07 · Before you act on this

The methodology is real, and the critique is directional. What a
simulated persona cannot have is a live budget, a renewal coming up, or
a boss asking about this quarter. **Validate anything you're betting on
with real ICPs who are actually in-market.** Being wrong is more
expensive than you think. Finding out is cheaper than you'd guess.

Wynter runs message testing with verified B2B professionals — trusted
by HubSpot, RingCentral, Shopify, Cognism, Paddle, Veeam, Rippling and
Miro. <https://wynter.com>

This report is kept for 60 days from 2026-10-05, then deleted along with the personas and their answers.

