Article
AI News Chipmakers

DeepSeek launches V4.1-Flash, a smaller, faster multimodal model

DeepSeek released DeepSeek-V4.1-Flash, a lower-latency model the company says cuts inference costs while preserving multimodal capability.

by TechDefused Newsroom
A close-up image of a smartphone displaying a chat interface for an application called DeepSeek. The screen shows a greeting message and a prompt for user interaction. — Credit: Photo by Solen Feyissa / Unsplash cPhoto by Solen Feyissa / Unsplash
Photo by Solen Feyissa / Unsplash

DeepSeek, the Hangzhou-based AI startup known for large reasoning models, released a new model called DeepSeek-V4.1-Flash.

Caixin Global reported the model is a 552-billion-parameter mixture-of-experts model with native multimodal visual understanding.

And that it uses a new architecture designed to deliver higher performance, faster inference, higher throughput and easier scaling to larger models.

The company framed the Flash variant as a speed- and cost‑focused node in its V4 family, aimed at cutting inference costs and improving latency for production workloads.

That product direction follows DeepSeek’s recent pattern of shipping smaller, efficiency-optimized variants while pursuing vertical integration.

Specifically, it has been adapting models for domestic accelerators, recruiting chip‑design engineers and developing an in‑house inference chip to reduce reliance on third‑party accelerators and lower per‑inference cost.

DeepSeek’s immediate milestones are operational, progress on its in‑house inference chip program and the potential resumption or closing of a paused external funding round, which will shape how quickly the V4 family can scale for wider deployment.

by TechDefused Newsroom