# Message test — https://letsteer.ai/

After reading your page, only 10 of 15 personas would take a meeting to learn more.

- **Page tested:** https://letsteer.ai/
- **Audience tested against:** Developers, Engineering Managers, CTOs at startups trying to optimize efficiency and reduce AI mistakers
- **Personas:** 15 simulated
- **Report:** https://grader.wynter.com/r/steer-ask-every-ai-model-the-same-question-sid-dLIjRsY

> These answers are generated by AI, scored on Wynter's B2B Message
> Layers framework using behaviorally-diverse simulated personas. The
> methodology is real and the critique is directional. What a simulated
> persona cannot have is a live budget, a renewal coming up, or a boss
> asking about this quarter.

---

## 01 · The scores

Every persona answered all four questions. These are four independent
proportions of the same panel, not stages of a funnel.

| Layer | Question | Cleared the bar | Strength | Of those who passed |
| --- | --- | --- | --- | --- |
| 1. Clarity | Do they understand what you do? | 14/15 | 91% | 9 without hesitation, 6 with reservations |
| 2. Relevance | Can they tell what it solves, and who it's for? | 15/15 | 79% | 1 without hesitation, 14 with reservations |
| 3. Value | Do they actually want it? | 10/15 | 59% | all with reservations |
| 4. Differentiation | Is there a reason to pick you over the alternatives? | 10/15 | 59% | all with reservations |

**Brand alignment** (a side metric, not one of the four layers) — 14/15, 74% strength (all with reservations). Does the page read like the company you actually are?

**Fix first: Value.** Earliest failing layer, walking the sequence in order — not simply the lowest score.

---

## 02 · What to change, layer by layer

Ordered worst-first. Specific edits, not a restatement of the score.

### Value

**Add a worked example under the Verdict step showing a real assumption gap being named.**

"Assumption gap" is the page's core promise but a reader cannot tell how Steer decides something is one. Show a short transcript: two model answers, the moderator's challenge, and the missing constraint it surfaced.

*effort medium · impact high · tested against Concrete over abstract*

**Add sample output under "Pick which model to route a task to" with real per-model numbers.**

The routing claim promises evidence from your real prompts but shows no scorecard, no metric, no example. Show a small table of models with win rates or latency on a sample prompt set.

*effort medium · impact high · tested against Proof next to the claim*

**Replace "measurably reduces hallucinations" with the error-reduction figure from the cited paper.**

The citation carries no number, so the accuracy claim rests on the word "measurably." Quote the actual reduction the paper reports, next to the claim.

*effort low · impact medium · tested against Specifics beat superlatives*

### Differentiation

**Add a line under "Representative scenarios" naming teams or usage counts behind the examples.**

Nothing on the page shows anyone has used Steer, and the disclaimer that these are not customer quotes underlines it. State how many decisions have been run through it, or name a team that has.

*effort medium · impact high · tested against Proof next to the claim*

**Add a line under "No token markup" comparing the cost to running the same prompt manually.**

Bring-your-own-keys is a pricing fact, not yet a reason to pick Steer over pasting prompts into three tabs. Say what four models on one question typically cost and how long it takes.

*effort low · impact medium · tested against Give a reason to choose you*

**Name the product category in the H1 area, above "Why not just trust one AI model?"**

A reader has to assemble what Steer is from scattered phrases before they can compare it to anything. Add a one-line descriptor: a browser-based multi-model consensus tool for engineering decisions.

*effort low · impact medium · tested against Lead with the use case*

### Clarity

**Define "assumption gap" in one sentence the first time it appears, in the four-steps paragraph.**

The term carries the verdict but is never defined, so readers guess at what separates it from real disagreement. Write that an assumption gap is disagreement caused by a missing constraint, not by differing judgement.

*effort low · impact high · tested against Plain language*

### Brand alignment (side metric)

**Replace the "in development" tag on paid features with a dated availability commitment.**

An open-ended "in development" label on tamper-evident export reads as a side project rather than a product with a roadmap. Give a quarter or remove the unbuilt feature from the pricing tier.

*effort low · impact medium · tested against Answer the live objection*

---

## 03 · What is working

### The rotating moderator is what respondents could repeat back

Across respondents the core description landed as 'multi-model prompts run in parallel with a rotating moderator that separates real disagreement from assumption gaps.' Several stated it back nearly verbatim without hunting through the page.

> a rotating "moderator" model reviews the independent answers and flags whether disagreements are real factual splits or just an unstated assumption in your prompt
> 
> — Developer, Software Development, 1-10

> a rotating "moderator" model flags where they actually disagree versus where the disagreement is just a missing assumption in the prompt
> 
> — VP of Engineering, Technology Startups, 1-10

> Steer sends the same prompt to multiple LLMs — Claude, GPT, Gemini, DeepSeek, Mistral, Grok, Llama — in parallel, then has a rotating "moderator" model challenge the weak answers and issue a verdict of consensus, real disagreement, or assumption gap.
> 
> — Developer, AI/Machine Learning, 11-50

> browser-based tool that fans your prompt out to multiple AI models (Claude, GPT, Gemini, etc.) at once, then uses a rotating "moderator" model to challenge the weak answers and issue a verdict — consensus, real disagreement, or an assumption gap
> 
> — Engineering Manager, Technology Startups, 201-500

### The confidence-risk framing of single-model deployment landed

Respondents accepted the problem statement: relying on one model to make a technical call is a confidence risk. It was named early enough that the problem was clear before the product was.

> "why not just trust one AI model" and states plainly that "a single AI model is a single point of failure: it can be confidently wrong, and you have no way to know until it matters."
> 
> — Developer, AI/Machine Learning, 11-50

---

## 04 · What the personas said

### 'Assumption gap' is asserted but never defined or scored

Respondents flagged the term as vague with no crisp scoring method, and asked for a worked example showing the moderator correctly identifying one. One wanted it tested on messy prompts rather than the clean TTL example.

> "assumption gap" is doing a lot of work without a crisp definition, and "challenges only the weak answers" begs the question of how weakness is actually scored.
> 
> — Engineering Manager, AI/Machine Learning, 51-200

> I'd want to see the moderator rotation and challenge logic on an actual messy prompt of ours, not the clean TTL example
> 
> — VP of Engineering, Technology Startups, 1-10

> One real example, on an ambiguous prompt, where the moderator correctly flagged 'assumption gap' instead of 'real disagreement' and I could check the reasoning myself — that single demonstrated catch would be worth more than the whole rest of the pitch.
> 
> — Developer, AI/Machine Learning, 11-50

### The category is never stated as a label

One respondent said the category became clear only by piecing together scattered phrases rather than reading a crisp name for what the product is.

> they never give it a crisp one-line name, so I had to synthesize "multi-model verification/consensus tooling" myself from scattered phrases like "moderator review," "targeted rebuttal," and "verdict" rather than reading it in one place
> 
> — CTO, Software Development, 201-500

### No benchmark numbers or example output back the accuracy claims

Respondents said routing and accuracy claims arrive with no data, no live benchmark, and no sample output. One would need a live demo with real metrics before granting a meeting.

> What would rule it out, or at least stall it, is that the "live benchmark against your own prompts" claim in model-selection has no numbers or example output anywhere on the page — if a competitor showed an actual sample verdict or dashboard screenshot, that alone would probably win the comparison for me.
> 
> — Senior Developer, Technology Startups, 11-50

> I'd take the meeting only if they can show me the benchmark-against-real-prompts feature live and put a number on hallucination reduction for something like a code-review or architecture call, not just cite the Du et al. paper
> 
> — CTO, Software Development, 201-500

### There are no named users or customer evidence anywhere

Three respondents noted the absence of named customers, quotes, or teams using the product, and read that absence as undermining the credibility of the claims.

> the total absence of any named team or usage number — every use case is flagged "Representative scenarios — not customer quotes," which tells me nobody real is on record using this yet.
> 
> — Engineering Manager, AI/Machine Learning, 51-200

> No named customers, no "trusted by" logos, "Representative scenarios — not customer quotes" is basically an admission they don't have real users yet
> 
> — VP of Engineering, AI/Machine Learning, 51-200

### The page reads as an early-stage indie project, not a company

Respondents picked up pre-revenue, early-stage signals: the 'in development' tag, rough two-tier pricing, unbuilt paid features like tamper-evident export, and no team or company credibility anchors. One read the tone as indie builder rather than enterprise.

> What's missing for me to trust the company itself, though, is any sense of who's behind it — no team page, no "built by ex-X engineers," nothing to anchor credibility beyond the mechanism they describe.
> 
> — Senior Developer, Software Development, 51-200

> the pricing page shows "proin development" for the paid tier, meaning the tamper-evident export and unlimited history — the parts that'd actually make this usable as a documented team record instead of a toy — don't exist yet
> 
> — VP of Engineering, Technology Startups, 1-10

> It reads like an indie dev-tool shipped by someone who hit this exact pain themselves and built the fix
> 
> — CTO, Technology Startups, 11-50

### The audience is inferred from use cases, not stated

Several respondents worked out that developers making technical calls are the target from code-review and architecture examples rather than being told. Others said the problem and reader were explicit and required no inference, so the page reads…

> "code review," "architecture calls," "prompt regressions," phrases like "endpoint to call in prod" and "null-check actually necessary" — this is written for developers/engineering teams
> 
> — Developer, Software Development, 1-10

> I'd need a line naming the role or team directly — something like "built for engineering managers arbitrating model-routing and code-review disputes" — instead of leaving me to infer it from four use-case cards
> 
> — Engineering Manager, AI/Machine Learning, 51-200

> the header "Why not just trust one AI model?" plus "A single model is a single point of failure" told me the problem in the first two lines, no hunting needed
> 
> — CTO, Software Development, 201-500

> the before/after example (TTL disagreement resolving once you add the missing constraint) made it concrete within seconds
> 
> — VP of Engineering, Technology Startups, 1-10

> the "Why not just trust one AI model?" header plus "A single model is a single point of failure" tells you the problem in one line, and the before/after example (Claude and GPT agreeing on caching but disagreeing on TTL) makes it concrete fast
> 
> — Engineering Manager, Technology Startups, 201-500

---

## 05 · The hardest read

An adversarial pass over the findings. Every claim below was checked
against the panel's own answers; unsupported ones were dropped.

- **The page's central differentiator is a term the page never defines, so the one thing people remember is the one thing they can't evaluate.** *(high)*
  Five respondents repeated back the moderator line including 'assumption gaps' verbatim, while three flagged the term as vague with no scoring method and asked for a worked example on messy prompts. Recall without comprehension is not persuasion.
- **Every accuracy and routing claim on the page is unfalsifiable, which converts the strongest pitch into marketing noise.** *(high)*
  Three respondents found no data, no live benchmark and no sample output behind accuracy claims; three more found no named customers, quotes or teams. Two independent evidence channels are empty, and one respondent required a live demo before granting a…
- **The page disqualifies itself from enterprise consideration before the product is even judged.** *(high)*
  Four respondents read pre-revenue signals — 'in development' tags, rough two-tier pricing, unbuilt paid features like tamper-evident export, no team or company anchors — as indie builder rather than company. Missing customer evidence compounds this.
- **Advertising unbuilt paid features actively damages the credibility of the features that do exist.** *(medium)*
  Respondents cited tamper-evident export as an unbuilt paid feature alongside 'in development' tags, and separately found no benchmark or sample output for shipped capability. A reader cannot tell which claims describe a product and which describe a roadmap.
- **The page makes readers do the work of naming both the product and the audience, and that labor is where prospects leak.** *(medium)*
  Seven respondents inferred the target from code-review and architecture examples rather than being told, and one reached the category only by piecing together scattered phrases. No crisp label for what this is or who it's for appears on the page.
- **The problem framing is the page's only fully earned asset, and it is carrying weight the rest of the page cannot support.** *(medium)*
  The confidence-risk framing of single-model deployment landed cleanly and early, and the moderator description was repeatable. But those two wins sit above undefined terms, zero benchmarks and zero customers — the page sets up a problem it never proves it…

---

## 06 · Who answered

| # | Role | Industry | Company size |
| --- | --- | --- | --- |
| 1 | Developer | Software Development | 1-10 |
| 2 | Senior Developer | Technology Startups | 11-50 |
| 3 | Engineering Manager | AI/Machine Learning | 51-200 |
| 4 | CTO | Software Development | 201-500 |
| 5 | VP of Engineering | Technology Startups | 1-10 |
| 6 | Developer | AI/Machine Learning | 11-50 |
| 7 | Senior Developer | Software Development | 51-200 |
| 8 | Engineering Manager | Technology Startups | 201-500 |
| 9 | CTO | AI/Machine Learning | 1-10 |
| 10 | VP of Engineering | Software Development | 11-50 |
| 11 | Developer | Technology Startups | 51-200 |
| 12 | Senior Developer | AI/Machine Learning | 201-500 |
| 13 | Engineering Manager | Software Development | 1-10 |
| 14 | CTO | Technology Startups | 11-50 |
| 15 | VP of Engineering | AI/Machine Learning | 51-200 |

---

## 07 · Before you act on this

The methodology is real, and the critique is directional. What a
simulated persona cannot have is a live budget, a renewal coming up, or
a boss asking about this quarter. **Validate anything you're betting on
with real ICPs who are actually in-market.** Being wrong is more
expensive than you think. Finding out is cheaper than you'd guess.

Wynter runs message testing with verified B2B professionals — trusted
by HubSpot, RingCentral, Shopify, Cognism, Paddle, Veeam, Rippling and
Miro. <https://wynter.com>

This report is kept for 60 days from 2026-09-23, then deleted along with the personas and their answers.

