W&B Inference (CoreWeave) Provider

Weights & Biases Inference (powered by CoreWeave) provides OpenAI-compatible inference for open-weight models — including DeepSeek V4 Flash, DeepSeek R1, Llama, and Qwen — at very competitive prices.

Available Models

Z.AI GLM 5.2 (W&B Inference)

glm
wandb-glm-5.2
Streaming
Vision
Tools
Reasoning
JSON Output
W&B Inference (CoreWeave)
Context: 204.8k
Input
$1.39
/M tokens
Cached
$0.26
/M tokens
Output
$4.4
/M tokens

DeepSeek V4 Pro (W&B Inference)

deepseek
wandb-deepseek-v4-pro
Streaming
Tools
Reasoning
JSON Output
JSON Schema
W&B Inference (CoreWeave)
Context: 1.0M
Input
$1.74
/M tokens
Cached
$0.14
/M tokens
Output
$3.46
/M tokens

DeepSeek V4 Flash (W&B Inference)

deepseek
wandb-deepseek-v4-flash
Streaming
Tools
Reasoning
JSON Output
JSON Schema
W&B Inference (CoreWeave)
Context: 1M
Input
$0.14
/M tokens
Cached
$0.07
/M tokens
Output
$0.28
/M tokens

TCS Cheap

tcs
tcs-cheap
Streaming
Tools
Reasoning
JSON Output
JSON Schema
W&B Inference (CoreWeave)
Context: 1M
Input
$0.14
/M tokens
Cached
$0.07
/M tokens
Output
$0.28
/M tokens

TCS Balanced

tcs
tcs-balanced
Streaming
Tools
Reasoning
JSON Output
JSON Schema
W&B Inference (CoreWeave)
Context: 1M
Input
$0.14
/M tokens
Cached
$0.07
/M tokens
Output
$0.28
/M tokens

TCS Reasoning

tcs
tcs-reasoning
Streaming
Vision
Tools
Reasoning
JSON Output
W&B Inference (CoreWeave)
Context: 262.1k
Input
$0.95
/M tokens
Cached
$0.16
/M tokens
Output
$4
/M tokens

TCS Premium

tcs
tcs-premium
Streaming
Vision
Tools
Reasoning
JSON Output
W&B Inference (CoreWeave)
Context: 204.8k
Input
$1.39
/M tokens
Cached
$0.26
/M tokens
Output
$4.4
/M tokens

TCS Vision

tcs
tcs-vision
Streaming
Vision
Tools
Reasoning
JSON Output
W&B Inference (CoreWeave)
Context: 262.1k
Input
$0.95
/M tokens
Cached
$0.16
/M tokens
Output
$4
/M tokens

Kimi K2.6 (W&B Inference)

moonshot
wandb-kimi-k2.6
Streaming
Vision
Tools
Reasoning
JSON Output
W&B Inference (CoreWeave)
Context: 262.1k
Input
$0.95
/M tokens
Cached
$0.16
/M tokens
Output
$4
/M tokens

Z.AI GLM 5.1 (W&B Inference)

glmModel Deactivated
wandb-glm-5.1
Streaming
Vision
Tools
Reasoning
JSON Output
W&B Inference (CoreWeave)
Context: 200k
Deactivated since Aug 25, 2026
Input
$1.4
/M tokens
Cached
$0.26
/M tokens
Output
$4.4
/M tokens

Qwen3 235B A22B Thinking-2507 (W&B Inference)

qwenModel Deactivated
wandb-qwen3-235b-thinking
Streaming
Tools
Reasoning
JSON Output
W&B Inference (CoreWeave)
Context: 262.1k
Deactivated since Aug 4, 2026
Input
$0.1
/M tokens
Cached
/M tokens
Output
$0.1
/M tokens