DeepSeek V4.1 Flash beat Haiku 5.5 as my research sub-agent and title model (small eval)

Wait 5 sec.

I ran Haiku 5.5 to see if it could replace DeepSeek V4.1 Flash in two small jobs in my app. It didn't. Small eval, but maybe useful for someone. GPU poor here (M2 Max, 32GB), so DeepSeek runs through OpenRouter with a minimum of 8-bit precision, sorted by throughput. Most runs landed on Baidu, one on Parasail. Job 1: research sub-agent. It gets a brief, searches the web, reads pages and writes a report with sources. Both models had the same tools: Exa and Brave for search, a page fetcher, and image search plus image analysis. I used 4 real briefs (brand visuals, consumer quotes from forums, company background, ad examples in a category). Haiku ran at low, medium and high effort. DeepSeek (current prod incumbent) ran at low. Haiku effort W/T/L vs DeepSeek time cost low 1/0/3 480s $0.09 medium 1/1/2 555s $0.12 high 0/1/3 992s $0.31 DeepSeek low - 721s $0.35 Haiku low is faster and 4x cheaper. But it lost the same two briefs every time. On one, DeepSeek found the brand's own guideline page with exact HEX and Pantone colors. Haiku never found that page, not even on high with 60 tool calls. It guessed the colors from screenshots. Higher effort made it search more, but still not find more. Haiku won one brief: finding real quotes on forums. It never opened a page there and used only the excerpts Exa returns. Cheap win, but a snippet doesn't show who wrote it or the context. Job 2: chat titles. 52 messages, reasoning off for both. DeepSeek won 28, Haiku won 8, 9 split, 7 identical. Same speed (1.1s), and Haiku was a bit cheaper. The problem: Haiku often answers the message instead of titling it. You ask about some topic and the title is the first line of an answer. And one prompt-injection test message became the title. How I judged: Opus judges, blind, both A/B orders. A tie means the two orders disagreed. Yes, Claude judged Claude (same lab bias), and it still picked DeepSeek. One run per setup, so don't treat it as big lab benchmarks, just the way I test agents for my tasks. $1 well spent (or not?)   submitted by   /u/dergachoff [link]   [comments]