I m not sure I would describe this as blanket win over GPT-5.6 Sol. In DeepSeek’s own table, V4.1 Flash is ahead on Terminal-Bench 2.1, DeepSWE, NL2Repo, and AutomationBench, but it is behind on GPQA Diamond, Terminal-Bench 3.0 and 4.0, and SEC-Bench Pro.
The architecture is probably part of the explanation for the lower cost and faster inference. DeepSeek says V4.1 Flash uses a new Causal Encoder–Decoder design, with 8B active parameters for input processing and 16B for decoding, along with much smaller KV caches.
But I hope it is just not benchmaxxed and genuinely good model
I m not sure I would describe this as blanket win over GPT-5.6 Sol. In DeepSeek’s own table, V4.1 Flash is ahead on Terminal-Bench 2.1, DeepSWE, NL2Repo, and AutomationBench, but it is behind on GPQA Diamond, Terminal-Bench 3.0 and 4.0, and SEC-Bench Pro.
The architecture is probably part of the explanation for the lower cost and faster inference. DeepSeek says V4.1 Flash uses a new Causal Encoder–Decoder design, with 8B active parameters for input processing and 16B for decoding, along with much smaller KV caches.
But I hope it is just not benchmaxxed and genuinely good model
benchmarks: https://media2url.com/m/52a77a33347c48