
Tech • AI • Robotics
DeepSeek has launched DeepSeek V4.1 Flash, a two-day test build that appears to introduce a new multimodal architecture with much higher speed, improved coding and 3D generation, and lower V4 Flash pricing starting September 10.
DeepSeek V4.1 Flash is available through the existing DeepSeek API under the model ID DeepSeek V4.1 Flash and is scheduled to expire on September 10. Testing is capped at 20 concurrent requests per account. The company released the model without official benchmarks, parameter counts, or weights, effectively inviting developers to stress-test it before a fuller launch.
DeepSeek describes the release as being built on a new architecture with native multimodal support. The temporary rollout has raised expectations because testers are being asked whether this architecture could eventually replace V4 Pro, suggesting the company may be evaluating a broader shift in its model lineup.
Early testing indicates generation speeds of roughly 350 to 400 tokens per second, with a reported peak of 427 tokens per second in one run. That puts the model in unusually fast territory for a reasoning-capable system and makes it competitive on throughput with much smaller or more specialized models, while staying in a lower cost bracket.
DeepSeek is also reducing V4 Flash pricing from September 10 at 12:00 Beijing time. Cached input falls from $0.000075 to $0.00003 per 1 million tokens, cache-miss input drops from about $0.22 to $0.15 per 1 million tokens, and output declines from $0.67 to $0.60 per 1 million tokens. Peak-hour pricing remains at 2x, but the overall direction is toward cheaper high-speed inference.
The model appears especially strong at generating code for interactive browser experiences. In one Three.js test, it built a classical Chinese garden complete with corridors, a reflective pond using shaders, and an explorable environment rather than a static scene. The jump over prior V4 Flash results was described as substantial, especially in visual richness and scene composition.
One repeated weakness is excessive reasoning and unnecessary testing. In a benchmark involving the Seven Wonders in a Three.js environment, V4.1 Flash reportedly took about 2 hours and cost $2.60 to complete two tests, while a competing unreleased model finished in roughly 30 minutes. The trade-off is that the final output quality remained notably strong despite the slower end-to-end completion.
Several game and simulation demos suggest unusually low cost for advanced interactive outputs. A Minecraft-style clone was reportedly generated in about 8 minutes for only pennies, while a Mario Kart-style game with sound effects and multiple environments cost about $1.10. A voxel-art pagoda benchmark used around 112,000 tokens across 20 turns, finished in 7 minutes 41 seconds, and reportedly cost about $0.30.
Beyond raw speed, the model appears to have improved spatial and agentic behavior. It produced a 3D exploded camera view after a few iterations in about 6 minutes, and an autonomous dungeon game was completed in a single attempt, where other models often require retries due to pathfinding failures. A rocket launch simulation showed stronger visuals than its predecessor, though the physics still lagged behind top-tier competitors.
DeepSeek V4.1 Flash looks like a significant interim release: faster, cheaper, and materially better at code-driven multimodal tasks than V4 Flash. If the test build reflects the direction of DeepSeek’s next architecture, the company may be positioning itself aggressively in the market for low-cost, high-speed generative coding models.
Explain this