
DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and high-throughput workloads, while maintaining strong reasoning and coding performance.
The model includes hybrid attention for efficient long-context processing. Reasoning efforts high and xhigh are supported; xhigh maps to max reasoning. It is well suited for applications such as coding assistants, chat systems, and agent workflows where responsiveness and cost efficiency are important.
| $0.002871 | $0.155509 | $0.002871 | 1.50s | 19 tps | ||
| $0.0029 | $1.28 | $0.0029 | 0.71s | 74 tps | ||
44% off | $0.14$0.07854 | $0.28$0.1571 | $0.028$0.01571 | 1.19s | 36 tps | |
43% off | $0.14$0.07924 | $0.28$0.1585 | $0.028$0.01585 | 0.70s | 73 tps | |
| $0.09 | $0.18 | $0.018 | 1.08s | 32 tps | ||
35% off | $0.14$0.091 | $0.28$0.182 | $0.028$0.0182 | 1.59s | 54 tps | |
30% off | $0.138$0.0966 | $0.275$0.1925 | $0.028$0.0196 | 1.40s | 23 tps | |
| $0.098 | $0.196 | $0.0196 | 0.72s | 18 tps | ||
| $0.13 | $0.28 | $0.028 | 1.07s | 37 tps | ||
| $0.134 | $0.268 | $0.0268 | 0.88s | 94 tps | ||
| $0.14 | $0.28 | $0.028 | 1.45s | 62 tps | ||
| $0.14 | $0.28 | $0.028 | 1.38s | 61 tps | ||
| $0.14 | $0.28 | $0.07 | 0.76s | 62 tps | ||
| $0.19 | $0.50 | -- | 1.78s | 12 tps | ||
| $0.21 | $0.56 | $0.031 | 1.53s | 75 tps |
P50, best across providers
P50, best provider
When an error occurs in an upstream provider, we can recover by routing to another healthy provider, if your request filters allow it. You can access per-provider uptime data programmatically through the Endpoints API. Learn more about our load balancing and customization options.