Quesma Blog BenchBenchmarks About Contact
GitHub

404

Page Not Found

The page you’re looking for doesn’t exist or has been moved.

Go to Homepage

While you’re here, check out our latest insights

HN
I burned all my tokens researching how to save tokens

I burned all my tokens researching how to save tokens

How I burned a full Claude limit in 30 minutes and built my own deep research pipeline instead: 3 subscriptions, shared memory, a clear role for each model. You can build the same from what you already pay for.

Bartosz Kotrys17 Jul 2026
Read more
Baba Is Solved by Fable 5 and GPT-5.6 Sol, but at what cost?

Baba Is Solved by Fable 5 and GPT-5.6 Sol, but at what cost?

We ported the puzzle game Baba Is You to the Harbor framework, and benchmarked current models, including Claude, GPT, Gemini, GLM, and DeepSeek. A human Twitch streamer is 4x faster than Claude Fable 5.

Piotr Migdał & Piotr Grabowski16 Jul 2026
Read more
Featured
Tokenflation: When “Hi” triggers 33 tool calls

Tokenflation: When “Hi” triggers 33 tool calls

Tokenflation: simple tasks consuming ever more context, reasoning, and tool calls without more useful output. We benchmarked 14 models; one answered “Hi” with 33 tool calls and an unsolicited commit.

Rafał Strzaliński9 Jul 2026
Read more
View All Blog Posts
©QuesmaInc. 2026·Privacy Policy
BlogAboutContact Media Kit
GitHub