DeepSeek V4 Flash (W&B Inference)

DeepSeek V4-Flash served on W&B Inference (CoreWeave). 284B MoE / 13B active, 1M context, 384K max output. At $0.14/$0.28 per 1M tokens this is among the cheapest broadly-capable models in our catalog — used as the primary backend for `tcs-cheap` and `tcs-balanced`.

wandb-deepseek-v4-flash
STABLEGet Started
1,000,000 context
Starting at $0.14/M input tokens
Starting at $0.28/M output tokens
Streaming
Tools
Reasoning
JSON Output

All Providers for DeepSeek V4 Flash (W&B Inference)

LLM Gateway routes requests to the best providers that are able to handle your prompt size and parameters.

W&B Inference (CoreWeave)
Context: 1M
Input
$0.14
/M tokens
Cached
$0.07
/M tokens
Output
$0.28
/M tokens
Get Started