DeepSeek releases V4.1 Flash, an MIT-licensed multimodal model with an unusual architecture
DeepSeek has published V4.1 Flash under the MIT licence: a 552-billion-parameter multimodal mixture-of-experts that reads images and text and activates only 8 to 16 billion parameters per token. It uses an encoder-decoder design rather than the decoder-only shape almost every other current model uses, and it carries a 1-million-token context. It is server-class, but genuinely open.
Featured