Sept 10 — DeepSeek launched its V4.1 Flash model, the smallest in its new
architecture series, with native multimodal visual understanding. The company
said the model sharply reduces KV cache size: HBM requirements drop to
one-quarter and SSD requirements to one-eighth versus the prior generation.
DeepSeek added that KV cache compression materially lowers operating costs for
agent-style workloads, where cache-hit costs are a significant share of
expenses.