← coldeye
Your homepage gets one first impression

Proof coldeye catches what loses buyers.

We put coldeye head to head with the tool it is modeled on, on the same page, judged blind. It won both rounds, 98 to 81. Here is the result, and the raw judge notes, so you can check it yourself.

Audit my site free →

How it ran: both tools audited the same page (https://integpt.com/). The original firsteyes.ai report is its own public sample; ours came from the same class of captured evidence. A judge model scored both blind, labeled only A and B, across six dimensions: coverage, specificity, actionability, rewrite quality, prioritization, insight depth. To cancel position bias, judging ran twice with labels swapped, and a tool only wins by taking both orderings.

Prompt pack v2 · 2026-07-20. The raw, unedited judge output is below.

Round 1 · ordering: ours=A · winner: A

Scores (sum of 6 dimensions, 1-10 each): A 49 · B 40

Both reports find the same core spine (Start Free login-wall, zero social proof, personal Gmail, jargon, buried value) and each contributes a genuine unique insight — A's 'sells an image product but shows no render' and free-tier watermark-contradiction findings, B's INR-only-pricing and buried 'Stop sending renders without deliverables' pain line. A wins decisively on craft integrity: its rewrites refuse to fabricate proof (explicit bracketed slots, 'never a fabricated count or star rating'), whereas B repeatedly hands the owner invented social proof to post verbatim — 'Join 500+ studios', 'Used by 200+ interior studios', and a fabricated named testimonial ('Priya S., Studio Forma') — which is worse than the original and dishonest to ship. B is also self-contradictory and corrupted: its first-impression section states 'no destination mismatch detected' on the very CTA its own conversion section and brutal_truth rank as the #1 killer, and its cta_flow_assessment field is a corrupted placeholder ('$13'). These defects sink B's specificity and rewrite_quality below A despite B's marginally broader element coverage and its concrete Clerk '?intent=signup' fix.

Round 2 · ordering: ours=B · winner: B

Scores (sum of 6 dimensions, 1-10 each): A 41 · B 49

Both reports independently nail the same top-priority defect (the 'Start Free' CTA dumping new visitors onto a 'Welcome back' sign-in wall) and both correctly flag zero social proof and the personal Gmail, so the verdict turns on craft integrity, not overlap. B is decisively cleaner on three counts: it uniquely catches that an image-generation product never shows a single render or before/after (a genuine, high-impact conversion insight A entirely misses), plus the free-tier watermark contradiction that guts the trial's ability to prove value; and critically, B refuses to fabricate social proof, prescribing an honest pilot-plus-bracketed-template approach, whereas A hands the owner ready-to-paste invented numbers ('Trusted by 200+ studios', 'Join 500+ studios') that are CRO malpractice and would be shipped verbatim. A is also marred by a corrupted 'cta_flow_assessment': '$13' placeholder field and a top-3 entry whose 'fix' is actually rewrite copy rather than an instruction, both of which undercut its reliability. A's missing_content list is marginally richer, but that does not offset the integrity and insight gap.

Honest limitations

The judge is an AI model, chosen and run by us. The protocol (blind labels, swapped orderings, win-both requirement) limits that bias, it does not eliminate it. The original's report predates ours, so the live page may have differed between captures. Both full reports are public: ours, and firsteyes.ai's own sample at their site.

Audit my site free →