DeepSeek V4 Flash 0731 版: 高效轻量化 MoE, 百万上下文, 高并发低延迟, 推理能力强
Model ID: deepseek-v4-flash-0731 · Type: chat · Provider: DeepSeek
Endpoints: /v1/chat/completions · /v1/messages
| Input (per 1M tokens) | $0.14 USD |
| Output (per 1M tokens) | $0.28 USD |
| Cache read (per 1M tokens) | $0.028 USD |
from openai import OpenAI
client = OpenAI(api_key="sk-...", base_url="https://api.router.ai/v1")
resp = client.chat.completions.create(
model="deepseek-v4-flash-0731",
messages=[{"role": "user", "content": "Hello"}],
)
print(resp.choices[0].message.content)DeepSeek-V4-Flash-0731 (`deepseek-v4-flash-0731`) is billed per usage at $0.14/1M in · $0.28/1M out, in USD. Current pricing is always listed at https://j8api.com/models/deepseek-v4-flash-0731.
Send a request to https://api.router.ai/v1/v1/chat/completions with the header `Authorization: Bearer <your API key>` and `"model": "deepseek-v4-flash-0731"`. The API is OpenAI-compatible, so any OpenAI SDK works by changing base_url to https://api.router.ai/v1 — no other code change.
DeepSeek-V4-Flash-0731 can be called on: /v1/chat/completions; /v1/messages.
DeepSeek-V4-Flash-0731 accepts up to 1,000,000 input tokens and can return up to 393,216 output tokens. Requests exceeding the input limit are rejected before reaching the model.
DeepSeek-V4-Flash-0731 supports: prompt_caching, thinking, reasoning.
DeepSeek-V4-Flash-0731 is a chat model from DeepSeek, available through the J8API gateway with the same API key as every other model.
Call it through the J8API OpenAI-compatible endpoint (API base: https://api.router.ai/v1). AI agents can discover and call every model on this gateway through MCP (https://mcp.router.ai/mcp) with no manual integration.