Vision / Qwen
qwen3.7-flash
Deep thinkingVisual Understandingfunction callingStructured outputWeb Search
Overview
Qwen3.7 Flash is a vision-language reasoning model from Alibaba. It is suited for multimodal agents, visual coding, search, and computer interaction, with strengths in object recognition, spatial understanding, and real-world visual perception.
- Model type: Vision
- Input: Text / Image / Video
- Output: Text
- Endpoints: chat
Tiered billing details
| Tier | Input | Output | Cache hit | Cache write |
|---|---|---|---|---|
| ≤32K input tokens | $0.0000/1K | $0.0001/1K | $0.0000/1K | $0.0000/1K |
| ≤256K input tokens | $0.0001/1K | $0.0003/1K | $0.0000/1K | $0.0001/1K |
| >256K input tokens | $0.0002/1K | $0.0007/1K | $0.0000/1K | $0.0002/1K |
Billed bylen (input context token count). Coefficient is the $ / 1M tokens price. Default group price (other groups converted by discount).
Code example
Call with qwen3.7-flash :
curl https://www.starunion.net/v1/chat/completions \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3.7-flash",
"messages": [{
"role": "user",
"content": [
{ "type": "text", "text": "Describe this image" },
{ "type": "image_url", "image_url": { "url": "https://example.com/photo.jpg" } }
]
}]
}'