Vision / Google
gemini-3.7-flash
AgenticCodingMultimodal
Overview
Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require responsive performance and reliable multi-step problem solving.
- Model type: Vision
- Input: Text / Image / Video / Audio
- Output: Text
- Endpoints: chat_completion
Tiered billing details
| Tier | Input | Output | Cache hit |
|---|---|---|---|
| Fallback tier | $0.0008/1K | $0.0037/1K | $0.0001/1K |
Billed bylen (input context token count). Coefficient is the $ / 1M tokens price. Default group price (other groups converted by discount).
Code example
Call with gemini-3.7-flash :
curl https://www.starunion.net/v1/chat/completions \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gemini-3.7-flash",
"messages": [{
"role": "user",
"content": [
{ "type": "text", "text": "Describe this image" },
{ "type": "image_url", "image_url": { "url": "https://example.com/photo.jpg" } }
]
}]
}'