# Message test — https://www.extend.ai/

After reading your page, 12 of 15 personas would take a meeting to learn more.

- **Page tested:** https://www.extend.ai/
- **Audience tested against:** Engineering and operations leaders
- **Personas:** 15 simulated
- **Report:** https://grader.wynter.com/r/extend-document-processing-infrastructure-for-YkdABAk

> These answers are generated by AI, scored on Wynter's B2B Message
> Layers framework using behaviorally-diverse simulated personas. The
> methodology is real and the critique is directional. What a simulated
> persona cannot have is a live budget, a renewal coming up, or a boss
> asking about this quarter.

---

## 01 · The scores

Every persona answered all four questions. These are four independent
proportions of the same panel, not stages of a funnel.

| Layer | Question | Cleared the bar | Strength | Of those who passed |
| --- | --- | --- | --- | --- |
| 1. Clarity | Do they understand what you do? | 15/15 | 82% | 3 without hesitation, 12 with reservations |
| 2. Relevance | Can they tell what it solves, and who it's for? | 15/15 | 81% | 2 without hesitation, 13 with reservations |
| 3. Value | Do they actually want it? | 12/15 | 67% | all with reservations |
| 4. Differentiation | Is there a reason to pick you over the alternatives? | 15/15 | 78% | all with reservations |

**Brand alignment** (a side metric, not one of the four layers) — 13/15, 70% strength (all with reservations). Does the page read like the company you actually are?

**Fix first: Value.** Earliest failing layer, walking the sequence in order — not simply the lowest score.

---

## 02 · What to change, layer by layer

Ordered worst-first. Specific edits, not a restatement of the score.

### Value

**Link the benchmark numbers to a published methodology page with document set and scoring rules.**

The 95.7% and 99.2% figures are Extend's own test results with no clickable methodology, so a buyer cannot check how accuracy was scored. Add a link under the benchmark chart to the full method, document corpus, and prompt set.

*effort medium · impact high · tested against Proof next to the claim*

**Add cost-per-page and latency numbers beside the Fast mode and Light Parse claims.**

"A fraction of the cost" and "low latency" give a buyer nothing to budget or design against. State price per thousand pages and typical response time for each processing mode.

*effort medium · impact high · tested against Specifics beat superlatives*

**Add a trial line under the hero telling readers they can benchmark their own documents.**

Readers want head-to-head results on their documents before believing the benchmark, and the page offers only "Try for free" and "Book demo" with no detail. Say how many pages the free tier covers and that evals run on uploaded samples.

*effort low · impact medium · tested against Answer the live objection*

### Clarity

**Replace "unmatched accuracy" in the hero subhead with the benchmark number and document set.**

"Unmatched accuracy" and "production-ready" carry no units, so the reader must scroll to the benchmark to learn what the product actually does. Put the measured accuracy figure and what was measured in the first sentence.

*effort low · impact high · tested against Specifics beat superlatives*

### Relevance

**Add a line under the hero naming the buyer and their situation explicitly.**

Readers infer the audience from logos and the pip install line instead of reading it. State who this is for, such as engineering teams processing high volumes of PDFs in production.

*effort low · impact medium · tested against Name the audience*

### Brand alignment (side metric)

**Name document types and benchmark coverage under each vertical tab, including real estate.**

"Real estate" and "Logistics" appear as bare tabs with no indication of which documents were tested, so the verticals read as decoration. List the document types scored in each, such as leases, title reports, bills of lading.

*effort medium · impact medium · tested against Concrete over abstract*

**Complete the SOC 2, HIPAA and GDPR block with audit dates, BAA availability, EU region detail.**

The compliance section trails off at "regular third-party penetration" and never says whether a BAA is signed or where EU data physically sits. Spell out certification dates, BAA terms, and the EU hosting region.

*effort low · impact medium · tested against Proof next to the claim*

---

## 03 · What is working

### The engineering tone and self-serve framing successfully signal who the product is for

Five respondents said the technical framing made the problem and developer audience immediately clear, and correctly assumed an engineer buyer at a mid-to-large company with document volume.

> The tone does feel written for someone like me: the pip install snippet, the SDK language list, "View docs," and the benchmark table against named competitors (Gemini, Azure DI, AWS Textract) all assume a technical buyer who wants to self-serve and verify claims, not a procurement generalist who needs hand-holding.
> 
> — Director of Engineering, Real Estate, 201-500

> the hero line "Parse, extract, and split your hardest documents with unmatched accuracy. Ship reliable document agents in minutes, not months" tells me the problem (unreliable/slow document extraction) and the "pip install extend-ai" plus "PythonTypeScriptJavaGoCLI" tabs tell me the reader is an engineer building document pipelines
> 
> — VP of Engineering, Healthcare, 5000+

> I picture a well-funded Series B/C startup, maybe 100-300 people, a few years old—old enough to have real enterprise logos (Brex, Square, Checkr, Amgen, First American, Flatiron Health) but still moving fast enough to ship things like "Composer Agent" and a brand-new "Light Parse" product announced in a banner.
> 
> — Director of Engineering, Real Estate, 201-500

---

## 04 · What the personas said

### 'Unmatched accuracy' and 'production-ready' are asserted without units or definitions

Five respondents said the marketing superlatives carry no units, no definition of accuracy at point of use, and require translation before the real capability is visible.

> the fuzziness is in words like "unmatched accuracy" and "production-ready" — those are marketing adjectives with no unit attached
> 
> — Head of Operations, Financial Services, 1001-5000

> none of those are defined anywhere near where they're used, so I had to go hunting in the benchmark section later to figure out what 'unmatched' was actually measured against
> 
> — Engineering Lead, Healthcare, 5000+

> things like "unmatched accuracy" and "production-ready" that got in the way, because they're asserted as adjectives on the hero rather than defined anywhere, so I had to go hunting in the benchmark section to find out what "accuracy" even means here
> 
> — VP of Engineering, Healthcare, 5000+

> it wasn't the product description itself that was hard, it was that nothing was actually confusing — it just took reading past the slogans ('unmatched accuracy,' 'minutes, not months') to find the real nouns
> 
> — Engineering Lead, Software and Technology, 501-1000

> "Enterprise-grade security," "flag potential errors," "catch regressions" - none of that is specific
> 
> — Senior Engineer, Real Estate, 201-500

### The page never addresses teams already running Textract or Azure

Two respondents said the boundary against OCR incumbents is unclear and there is no explicit callout for switchers, despite the benchmark being built on that comparison.

> Not the category itself—that part was easy—but the exact boundary of what's included was fuzzy: "split" is never defined (splitting by document type? page ranges? both?), and "document workflows" and "Composer Agent" get described with marketing verbs like "orchestration" and "optimization agent" rather than a concrete spec, so I couldn't tell where the parsing API ends and a separate workflow product begins.
> 
> — Director of Engineering, Real Estate, 201-500

### The benchmark is not backed by linked methodology or outside validation, so respondents…

Five respondents flagged that the methodology is not clickable, the numbers are internal test results, and the Vendr bakeoff is referenced without showing outcomes. Three said they would need head-to-head testing on their own documents first.

> Running our own batch of messy lease and title PDFs through it myself and watching it correctly extract fields my current Textract/Azure pipeline chokes on—one real test on our own files beats any number of benchmark tables.
> 
> — Director of Engineering, Real Estate, 201-500

> right now this page gives me numbers but no link to replicate them myself—the "View" and "Read the benchmark" buttons are promising but I haven't actually seen the underlying methodology, just the claim.
> 
> — Director of Engineering, Real Estate, 201-500

> "RealDoc-Bench" and the benchmark charts are their own numbers against their own test set, not against our documents, so I'm not taking that at face value.
> 
> — Head of Operations, Real Estate, 201-500

### Everything past the benchmark drops the numbers respondents came for

Three respondents noted the feature section and cost claims lack latency, cost-per-page, and error rates, despite cost being a primary concern. Security and error-flagging language was called vague.

> "Enterprise-grade security," "flag potential errors," "catch regressions" - none of that is specific
> 
> — Senior Engineer, Real Estate, 201-500

### Vertical and compliance claims are listed, not substantiated

Four respondents said real estate appears without document types or benchmark coverage, healthcare compliance is tacked on at the end, and regulated-industry and EU hosting specifics are missing. The brand reads as generic infrastructure rather than…

> A line naming real estate explicitly alongside the finance/logistics/healthcare verticals already in the benchmark tabs, plus a sample schema or output for something like a lease or title report—right now I have to infer applicability from a generic vertical tab, not from any content written with my documents in mind.
> 
> — Director of Engineering, Real Estate, 201-500

> SOC 2/HIPAA/GDPR, deployment options) is tacked on near the end almost as a checkbox rather than built into the core pitch, which is where I'd want more
> 
> — Engineering Lead, Healthcare, 5000+

> I'd need a line that names my exact context — regulated industry, EU-hosted documents, finance/real-estate document types — not just a pip install and generic logos
> 
> — Director of Engineering, Financial Services, 1001-5000

> The tone is written for someone like me in terms of technical fluency — the pip install line, the SDK language list, the benchmark table — but not in terms of vertical: there's no "built for real estate" anywhere, I have to squint at a "Real estate" tab in a table to see myself in it at all.
> 
> — Director of Engineering, Real Estate, 201-500

> Real estate was even listed as one of their verticals, so relevant enough, but I'd want to see it against actual lease or title docs before caring.
> 
> — Head of Operations, Real Estate, 201-500

### The testimonials read as emotional rather than methodological to an engineering audience

One respondent said the quotes lean on emotional appeal instead of proof; another said the page speaks to developers and not the procurement decision-makers who sign off.

> Where it's less for me specifically is the customer-quote section—"outperformed every solution we tested" and "replicate 6 months of work in 2 weeks" are testimonial-style claims aimed at building trust emotionally rather than giving me the methodology I'd actually need, so that part reads more like a page built for a VP who skims than an engineer who checks sources.
> 
> — Director of Engineering, Real Estate, 201-500

> Tone's written for developers, not for me — "pip install," "ship document agents in minutes" — fine for my eng team, but if I'm signing off, I need the security/compliance page, not the code snippet
> 
> — Head of Operations, Financial Services, 1001-5000

### The audience is signalled but never stated, so it has to be inferred

Three respondents said they worked out the target buyer from logos, code samples, and technical signals rather than any explicit statement on the page.

> The reader is inferred, not stated outright: no line says "for engineering teams migrating off Textract" or similar, but the logos (Brex, Flatiron Health, Square), "ship reliable document agents," and code snippets make it obvious this is aimed at engineers/teams building document pipelines, not business buyers.
> 
> — Senior Engineer, Financial Services, 1001-5000

> it wasn't the product description itself that was hard, it was that nothing was actually confusing — it just took reading past the slogans ('unmatched accuracy,' 'minutes, not months') to find the real nouns
> 
> — Engineering Lead, Software and Technology, 501-1000

---

## 05 · The hardest read

An adversarial pass over the findings. Every claim below was checked
against the panel's own answers; unsupported ones were dropped.

- **The page's central proof point — the benchmark — collapses under the first click, because there is nothing to click.** *(high)*
  Four respondents flagged unlinked methodology and internal-only numbers, and three said they would need head-to-head testing on their own documents before believing anything; the Vendr bakeoff is cited with no outcome.
- **Accuracy superlatives force the reader to do the vendor's work, and that translation step is where belief is lost.** *(high)*
  Five respondents said 'unmatched accuracy' and 'production-ready' carry no units or definition at point of use and require translation before real capability is visible — a claim that needs decoding is a claim that gets discounted.
- **The page front-loads its only quantitative content and then goes silent, so the second half of the scroll carries no persuasive weight.** *(high)*
  Three respondents noted that everything past the benchmark drops latency, cost-per-page, and error rates despite cost being a primary concern, with security and error-flagging language called vague.
- **Vertical and compliance mentions actively damage credibility rather than extend reach.** *(high)*
  Five respondents said real estate appears with no document types or benchmark coverage, healthcare compliance is tacked on at the end, and regulated-industry and EU hosting specifics are absent — the brand reads as generic infrastructure.
- **The benchmark is built on a comparison the page then refuses to make, leaving the highest-intent reader with no reason to switch.** *(medium)*
  Two respondents said the boundary against Textract and Azure is unclear with no explicit switcher callout, despite the comparison being the benchmark's foundation.
- **Audience clarity is the page's only real asset, and it is clarity about who the page is for, not why they should buy.** *(medium)*
  Five respondents credited the engineering tone for signalling the developer buyer, but negative themes on accuracy claims, benchmark backing and missing metrics outnumber that single positive across every other dimension.

---

## 06 · Who answered

| # | Role | Industry | Company size |
| --- | --- | --- | --- |
| 1 | Director of Engineering | Real Estate | 201-500 |
| 2 | VP of Engineering | Software and Technology | 501-1000 |
| 3 | Head of Operations | Financial Services | 1001-5000 |
| 4 | Engineering Lead | Healthcare | 5000+ |
| 5 | Senior Engineer | Real Estate | 201-500 |
| 6 | Engineering Manager | Software and Technology | 501-1000 |
| 7 | Director of Engineering | Financial Services | 1001-5000 |
| 8 | VP of Engineering | Healthcare | 5000+ |
| 9 | Head of Operations | Real Estate | 201-500 |
| 10 | Engineering Lead | Software and Technology | 501-1000 |
| 11 | Senior Engineer | Financial Services | 1001-5000 |
| 12 | Engineering Manager | Healthcare | 5000+ |
| 13 | Director of Engineering | Real Estate | 201-500 |
| 14 | VP of Engineering | Software and Technology | 501-1000 |
| 15 | Head of Operations | Financial Services | 1001-5000 |

---

## 07 · Before you act on this

The methodology is real, and the critique is directional. What a
simulated persona cannot have is a live budget, a renewal coming up, or
a boss asking about this quarter. **Validate anything you're betting on
with real ICPs who are actually in-market.** Being wrong is more
expensive than you think. Finding out is cheaper than you'd guess.

Wynter runs message testing with verified B2B professionals — trusted
by HubSpot, RingCentral, Shopify, Cognism, Paddle, Veeam, Rippling and
Miro. <https://wynter.com>

This report is kept for 60 days from 2026-10-05, then deleted along with the personas and their answers.

