# Message test — https://letsteer.ai/

After reading your page, only 10 of 15 personas could name a reason to pick you over a similar option.

- **Page tested:** https://letsteer.ai/
- **Audience tested against:** Engineering Manager, CTO, Developer at a startup company juggling between different AI agents to make better decisions
- **Personas:** 15 simulated
- **Report:** https://grader.wynter.com/r/steer-ask-every-ai-model-the-same-question-sid-9zSNNRQ

> These answers are generated by AI, scored on Wynter's B2B Message
> Layers framework using behaviorally-diverse simulated personas. The
> methodology is real and the critique is directional. What a simulated
> persona cannot have is a live budget, a renewal coming up, or a boss
> asking about this quarter.

---

## 01 · The scores

Every persona answered all four questions. These are four independent
proportions of the same panel, not stages of a funnel.

| Layer | Question | Cleared the bar | Strength | Of those who passed |
| --- | --- | --- | --- | --- |
| 1. Clarity | Do they understand what you do? | 15/15 | 93% | 10 without hesitation, 5 with reservations |
| 2. Relevance | Can they tell what it solves, and who it's for? | 15/15 | 81% | 2 without hesitation, 13 with reservations |
| 3. Value | Do they actually want it? | 11/15 | 63% | all with reservations |
| 4. Differentiation | Is there a reason to pick you over the alternatives? | 10/15 | 61% | all with reservations |

**Brand alignment** (a side metric, not one of the four layers) — 13/15, 70% strength (all with reservations). Does the page read like the company you actually are?

**Fix first: Differentiation.** Earliest failing layer, walking the sequence in order — not simply the lowest score.

---

## 02 · What to change, layer by layer

Ordered worst-first. Specific edits, not a restatement of the score.

### Differentiation

**Rewrite the three-tabs paragraph to name what Steer does that manual comparison cannot.**

Calling multi-tab checking "a manual tax" asserts the advantage without showing it; a reader sees no reason tabs plus a template fall short. Name the specific things tabs cannot do: cross-model challenge, rebuttal, and a recorded verdict per decision.

*effort medium · impact high · tested against Give a reason to choose you*

**Replace vague Pro tier feature descriptions with the specific capabilities and limits each tier includes.**

Pro tier descriptions do not say what a paying user gets that the free path lacks, so the upgrade is unarguable. List concrete differences: number of active models, saved verdict history, team sharing.

*effort medium · impact medium · tested against Tie the feature to the outcome*

**Add a line under "Nothing leaves your browser" on data residency and losing history.**

Local storage reads to some buyers as a compliance and data-loss risk rather than a benefit. Say plainly that no data crosses a border because nothing leaves the device, and what happens to saved verdicts.

*effort low · impact medium · tested against Answer the live objection*

### Value

**Replace the model-selection claim with a stated benchmark output: what metrics Steer reports per model.**

"A live benchmark against your own prompts" never says what the benchmark measures or produces. State what a reader sees after a run, such as agreement rate, catch rate, or cost per prompt per model.

*effort medium · impact high · tested against Specifics beat superlatives*

**Add a line under the Du et al. citation stating what Steer itself measures on your prompts.**

The academic paper is proof about multiagent debate in general, not about this tool's results. Put a sentence next to it saying what error-catch numbers users can see from their own runs.

*effort medium · impact high · tested against Proof next to the claim*

**Move the TTL before/after example above the research citation.**

The concrete Claude-GPT TTL disagreement is the most convincing thing on the page but sits below an academic reference. Lead the problem section with it, then use the paper as backup.

*effort low · impact medium · tested against Concrete over abstract*

### Clarity

**Define "verdict" at first use in the how-it-works paragraph.**

"Verdict" covers consensus, real disagreement, and assumption gaps, so the word reads as a single judgment when it is three different outcomes. Say it is a labelled result, then name the three labels.

*effort low · impact medium · tested against Plain language*

### Relevance

**Name the reader in the opening section: role, team size, and stack.**

The audience is only inferrable from vocabulary and examples, so the right buyer has to work out that it is for them. Add a line naming engineering teams running more than one model provider in production.

*effort low · impact medium · tested against Name the audience*

### Brand alignment (side metric)

**Add a short team or maker line near the footer with names and background.**

Nothing on the page says who builds Steer, which leaves buyers guessing whether it will still exist next quarter. One or two sentences naming the people and their relevant experience answers it.

*effort low · impact medium · tested against Answer the live objection*

---

## 03 · What is working

### The core mechanic — one prompt to multiple models with a rotating moderator — is…

Ten respondents played back the product function unprompted: parallel multi-model comparison via user API keys, with a rotating moderator that challenges weak answers and returns a verdict. Several named it an arbitration layer, not a new model.

> It's a multi-model AI "cross-check" tool — you bring your own API keys, it fires the same prompt at Claude, GPT, Gemini, etc., has one model act as rotating moderator to challenge weak answers, and spits out a verdict: consensus, real disagreement, or "assumption gap."
> 
> — Senior Developer, AI/Machine Learning, 11-50

> a rotating "moderator" model challenges the weak answers, and it spits out a verdict of consensus, real disagreement, or an unstated-assumption gap
> 
> — CTO, Technology, 11-50

> you send one prompt to several models (Claude, GPT, Gemini, etc.) in parallel via your own API keys
> 
> — CTO, Technology, 11-50

> It's a multi-model AI arbitration tool — sends the same prompt to Claude, GPT, Gemini, etc., has one model moderate/challenge the others, and spits out a verdict of consensus, disagreement, or assumption gap. Basically a cross-checking layer on top of AI APIs you already have keys for, not a new model itself.
> 
> — Lead Developer, AI/Machine Learning, 51-200

> Basically automates the "paste into three tabs and compare" thing I already do manually.
> 
> — Senior Developer, Startups, 201-500

> you send one prompt to several models at once (Claude, GPT, Gemini, etc.), a rotating moderator model challenges the weak answers, and it spits out a verdict — consensus, real disagreement, or an unstated-assumption gap.
> 
> — VP of Engineering, SaaS, 1-10

### The cached TTL disagreement example is the proof that lands

Three respondents cited the Claude-GPT TTL disagreement traced to a missing prompt constraint as concrete evidence, valuing the signal that separates genuine model disagreement from underspecified requirements.

> The cached example — Claude and GPT agreeing on caching but disagreeing on TTL until you add the missing staleness-tolerance constraint — is the one bit of proof that made me sit up
> 
> — Director of Engineering, SaaS, 201-500

> I get a verdict that's either consensus, a real factual split, or "you forgot to specify staleness tolerance," which is a genuinely different and useful signal
> 
> — CTO, Technology, 11-50

### Bring-your-own-keys and local-only storage read as low adoption cost

Respondents cited no account requirement, user API keys, and local storage as removing data-residency and security friction, and as a reason the tool is easy to try.

> I'd stop having engineers manually paste the same prompt into three tabs and eyeball the differences
> 
> — VP of Engineering, Startups, 51-200

> the BYO-keys, no-backend, local-storage architecture means adoption cost is low and there's no data-residency fight to have with security
> 
> — CTO, Technology, 11-50

---

## 04 · What the personas said

### "Verdict" is ambiguous when it covers three different outcomes

One respondent flagged that the term "verdict" is applied to consensus, disagreement, and other outcomes without distinction, leaving the word undefined.

> the only fuzzy word is "verdict" itself, since it's used for three quite different outcomes (consensus, disagreement, assumption gap) and the page never shows what the verdict actually looks like on screen.
> 
> — Senior Developer, AI/Machine Learning, 11-50

### Respondents will not believe the model-selection claims without a benchmark on their own…

Four respondents said the academic paper citation is not enough and asked for performance metrics, error-catch rates, or a routing benchmark run on their actual prompts before committing.

> What I'd need before committing budget or workflow time is the model-routing benchmark actually surfaced on real prompts, not just the arXiv citation
> 
> — CTO, Technology, 11-50

> I need to see it running on our actual prompts before I believe the "before/after" TTL example generalizes. Worth a 15-minute look since it's free to try with our own keys, not worth a real commitment yet.
> 
> — Lead Developer, AI/Machine Learning, 51-200

### The page does not show why this beats multi-tab or templated workflows, and the Pro tier…

One respondent said the advantage over manual multi-tab or template approaches is asserted rather than demonstrated; another found Pro tier feature descriptions too vague to justify paying. One read local storage as an EU compliance risk rather than a benefit.

> to beat those, Steer needs to show its moderator catching something a plain side-by-side would've missed, not just a cheaper way to fire the same prompt at multiple models.
> 
> — Senior Developer, AI/Machine Learning, 11-50

> for an EU SaaS company that's a compliance question I'd need answered before signing — is browser local storage actually acceptable for whatever's in these prompts under our data handling policy, or does that just move the risk instead of removing it?
> 
> — Director of Engineering, SaaS, 201-500

> I'd rule it out if the Pro tier stays vague on what "Unlimited Working Summary refreshes" actually does — that phrase means nothing to me and I wouldn't pay $9/mo for a feature I can't picture
> 
> — Engineering Manager, Software, 1-10

### No social proof or team information leaves vendor durability an open question

Two respondents said the absence of social proof and team transparency raises doubts about company maturity and whether the tool is reliable enough for daily use.

> no team page, no company name I recognize, no indication of how long this survives if the one person building it moves on — for a $79 one-time tool that's fine to pilot, but I'd want that answered before it touches anything my team relies on daily
> 
> — Director of Engineering, Technology, 201-500

> the lack of any social proof or "who's using this" signal is the one thing that makes we wonder if this is one person's project rather than a company I'd trust with a production decision yet
> 
> — Senior Developer, SaaS, 11-50

### The page states the problem clearly but never names the reader — audience is inferred…

Six respondents said engineering teams with multi-model setups are the evident target, but four of them noted the audience is inferred from examples and vocabulary rather than stated. One asked for explicit team size and stack.

> this is clearly aimed at engineering teams, probably ones already juggling multiple model APIs, not a generic business buyer.
> 
> — Senior Developer, AI/Machine Learning, 11-50

> I'd want a line that names my actual situation — something like "for teams already running Claude and GPT side-by-side and arguing about which answer to ship" — rather than making me infer it from a caching/TTL example
> 
> — Senior Developer, AI/Machine Learning, 11-50

> I'd need a line naming the team size and stack directly — something like "for 2-15 person engineering teams already paying for multiple model APIs"
> 
> — Engineering Manager, Software, 1-10

> Who it's for isn't spelled out with a title or persona, but it's obvious from the use cases: model selection, code review, architecture calls, prompt regressions — that's an engineering team, specifically the ones arguing about which model to trust in prod. So I inferred the reader from the examples rather than a stated "this is for X," but it took seconds, not hunting.
> 
> — Lead Developer, AI/Machine Learning, 51-200

> Yeah, this was clear fast — the "actual problem" header straight up says "a single AI model is a single point of failure," and the use-cases section spells out who it's for: engineering teams arguing over model routing, code review flags, architecture calls, prompt regressions.
> 
> — Senior Developer, Startups, 201-500

> The use-case section then does the work of naming the reader without ever saying "for engineering teams" outright — model selection, code review, architecture calls, prompt regressions, all framed as things "engineering teams argue about but rarely document."
> 
> — VP of Engineering, SaaS, 1-10

> The intended reader isn't explicitly named as a persona, but it's inferable within seconds from the use cases section — "model selection," "code review," "architecture calls," "prompt regressions" — this is clearly engineering teams who already use multiple LLM APIs
> 
> — Engineering Manager, Technology, 51-200

---

## 05 · The hardest read

An adversarial pass over the findings. Every claim below was checked
against the panel's own answers; unsupported ones were dropped.

- **Comprehension of the mechanic does not convert into belief in the claim, so the page teaches without persuading** *(high)*
  Ten respondents played back the parallel-model-plus-moderator mechanic unprompted, yet four demanded metrics, error-catch rates, or a benchmark on their own prompts before committing. Understanding what it does is not accepting that it works.
- **The academic citation is doing evidence work it cannot carry and should be replaced with operational numbers** *(high)*
  Four respondents rejected the paper citation as insufficient and asked for routing benchmarks on their actual prompts; the only proof that landed was a single anecdotal TTL disagreement example cited by three.
- **The page never argues against the incumbent alternative, which is doing nothing new** *(high)*
  Respondents said the advantage over manual multi-tab and template workflows is asserted rather than demonstrated, and Pro tier descriptions were too vague to justify payment. Without that comparison the product reads as optional.
- **Leaving the audience unstated forces every reader to self-qualify, and some will opt out** *(medium)*
  Six respondents inferred engineering teams with multi-model setups only from examples and vocabulary, with four explicitly noting the reader is never named and one asking for team size and stack.
- **The same architecture choice reads as both a benefit and a liability, so the page has lost control of its own strongest asset** *(medium)*
  Bring-your-own-keys and local storage were cited as removing data-residency and security friction, while one respondent read local storage as an EU compliance risk. Unframed, the feature argues both directions.
- **Absent proof of the vendor, low adoption cost cannot overcome the daily-use decision** *(medium)*
  Two respondents said missing social proof and team information raised doubts about company maturity and reliability for daily use — a durability question that no-account, easy-to-try framing does not answer.

---

## 06 · Who answered

| # | Role | Industry | Company size |
| --- | --- | --- | --- |
| 1 | Senior Developer | AI/Machine Learning | 11-50 |
| 2 | VP of Engineering | Startups | 51-200 |
| 3 | Director of Engineering | SaaS | 201-500 |
| 4 | Engineering Manager | Software | 1-10 |
| 5 | CTO | Technology | 11-50 |
| 6 | Lead Developer | AI/Machine Learning | 51-200 |
| 7 | Senior Developer | Startups | 201-500 |
| 8 | VP of Engineering | SaaS | 1-10 |
| 9 | Director of Engineering | Software | 11-50 |
| 10 | Engineering Manager | Technology | 51-200 |
| 11 | CTO | AI/Machine Learning | 201-500 |
| 12 | Lead Developer | Startups | 1-10 |
| 13 | Senior Developer | SaaS | 11-50 |
| 14 | VP of Engineering | Software | 51-200 |
| 15 | Director of Engineering | Technology | 201-500 |

---

## 07 · Before you act on this

The methodology is real, and the critique is directional. What a
simulated persona cannot have is a live budget, a renewal coming up, or
a boss asking about this quarter. **Validate anything you're betting on
with real ICPs who are actually in-market.** Being wrong is more
expensive than you think. Finding out is cheaper than you'd guess.

Wynter runs message testing with verified B2B professionals — trusted
by HubSpot, RingCentral, Shopify, Cognism, Paddle, Veeam, Rippling and
Miro. <https://wynter.com>

This report is kept for 60 days from 2026-09-23, then deleted along with the personas and their answers.

