DeepSeek 4.1 Flash (as a System One model) vs Jev

Wait 5 sec.

Source: DeepSeek vs Jev. This isn't a normal LLM benchmark. Jev is a "System One" model: no text, it just gives a probability for each option. The benchmark scores that, the right answer plus probabilities. LLMs can do it if you ask for probabilities and turn reasoning off, but it's not what they're built for. Kind of like judging a fish by its ability to climb a tree. Still, DeepSeek holds up pretty well on the pick-one task. A bit behind Jev on accuracy, slower and pricier, but the best calibrated model on the board.   submitted by   /u/frappuccinoCoin [link]   [comments]