DeepSeek, the Hangzhou-based AI startup known for large reasoning models, released a new model called DeepSeek-V4.1-Flash.
Caixin Global reported the model is a 552-billion-parameter mixture-of-experts model with native multimodal visual understanding.
And that it uses a new architecture designed to deliver higher performance, faster inference, higher throughput and easier scaling to larger models.
The company framed the Flash variant as a speed- and cost‑focused node in its V4 family, aimed at cutting inference costs and improving latency for production workloads.
That product direction follows DeepSeek’s recent pattern of shipping smaller, efficiency-optimized variants while pursuing vertical integration.
Specifically, it has been adapting models for domestic accelerators, recruiting chip‑design engineers and developing an in‑house inference chip to reduce reliance on third‑party accelerators and lower per‑inference cost.
DeepSeek’s immediate milestones are operational, progress on its in‑house inference chip program and the potential resumption or closing of a paused external funding round, which will shape how quickly the V4 family can scale for wider deployment.