Notice
Wan3.0 Officially Launched: Wan3.0 is a native large-scale model for generating 30-second videos, seamlessly orchestrating all audiovisual elements and supporting multi-modal references, delivering viOriginal Manufacturer Authorized · Genuine and TraceableEnterprise‑grade large‑model intelligent gateway platform—provides enterprises with compliant, controllable, stable, and efficient AI model API aggregation and governance services; one‑stop access to⚡️This site has been integrated with the official versions of GLM 5.3 and DeepSeek V4❗️

Text generation / DeepSeek

deepseek-v4-flash

Deep thinking

Overview

A flagship MoE large model with 1.6 trillion parameters and 49 billion activated parameters, natively supporting context lengths of up to one million tokens. Trained on a vast corpus of high-quality data, it excels in advanced mathematical reasoning, complex logical inference, specialized coding, and deep analysis of long-form text, making it well-suited for demanding applications such as cutting-edge research, sophisticated office workflows, and advanced AI agents.

  • Model type: Text generation
  • Input: Text
  • Output: Text
  • Endpoints: chat_completion / chat
Tiered billing details
TierInputOutputCache hit
Fallback tier$0.0001/1K$0.0003/1K$0.0000/1K

Billed bylen (input context token count). Coefficient is the $ / 1M tokens price. Default group price (other groups converted by discount).

Capabilities

  • Near-flagship inference level: On simple Agent tasks, it performs on par with V4-Pro, covering the vast majority of daily intelligent agent and productivity scenarios.
  • Extremely fast response: The activation parameters are only 13B, which is about 1/4 of that of V4 Pro. Its response speed is much faster than that of V4 Pro, and it supports millions of long contexts by default, maintaining stable low latency in long text scenarios.
  • Ultimate cost-effectiveness: Approaching the upper limit of V4 Pro's capabilities with approximately 1/6 of the total parameter count, priced at less than one-tenth of the flagship model.

Code example

Call with deepseek-v4-flash :

curl https://www.starunion.net/v1/chat/completions \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek-v4-flash",
    "messages": [{ "role": "user", "content": "Hello" }]
  }'