DeepSeek V4-Flash served on W&B Inference (CoreWeave). 284B MoE / 13B active, 1M context, 384K max output. At $0.14/$0.28 per 1M tokens this is among the cheapest broadly-capable models in our catalog — used as the primary backend for `tcs-cheap` and `tcs-balanced`.
LLM Gateway routes requests to the best providers that are able to handle your prompt size and parameters.