Blog
What separates expert AI output from the generic kind. Performance data, integration guides, and industry perspectives.
researchJul 12, 20264 min read
We tested our own product for three weeks. Here's why we're winding it down.
We spent three weeks running blind, cross-model benchmarks on our own catalog. The bet didn't hold. This is the honest account of what we found, and why we're winding the product down.
transparencybenchmarksbuilding-in-publicwind-down
researchMar 24, 20268 min read
We Ran 819 API Calls to Find Claude's Signature Catchphrases
We built a simulator that fed 40 developer scenarios into Claude Sonnet across 7 languages. Then we asked Claude to analyze its own output. 332 catchphrases later, we know exactly which phrases Claude reaches for - and why it matters.
experimentclaudemotivationsycophancy