≡ All docs
Model API / Batch
Batch
Updated: 2026-09-19
Overview
Pack thousands of requests into one JSONL file, submit it once, and get the results back asynchronously within 24 hours. Good for offline evaluation, bulk labelling and data cleaning. Batch pricing is listed per model: a model whose detail page shows a Batch price is settled at that price, and one without it is settled at the real-time price. Compatible with the OpenAI Batch API, so client.files.* and client.batches.* from the official SDKs work as-is.
End-to-end Flow
- 1. Upload the input file: POST /files (purpose=batch) and keep the returned file id
- 2. Create the batch: POST /batches with that input_file_id
- 3. Poll: GET /batches/{batch_id} until status becomes completed
- 4. Download results: GET /files/{output_file_id}/content
# 1. uploadcurl -X POST "https://www.starunion.net/v1/files" \-H "Authorization: Bearer YOUR_API_KEY" \-F "purpose=batch" \-F "file=@requests.jsonl"# 2. createcurl -X POST "https://www.starunion.net/v1/batches" \-H "Authorization: Bearer YOUR_API_KEY" \-H "Content-Type: application/json" \-d '{"input_file_id": "file-xxxxxxxx","endpoint": "/v1/chat/completions","completion_window": "24h"}'# 3. pollcurl "https://www.starunion.net/v1/batches/batch_xxxxxxxx" \-H "Authorization: Bearer YOUR_API_KEY"# 4. downloadcurl "https://www.starunion.net/v1/files/file-yyyyyyyy/content" \-H "Authorization: Bearer YOUR_API_KEY"
Input File Format
JSONL: one JSON object per line, one request per line. You choose custom_id yourself and use it to map output lines back to input lines, because the output order is not guaranteed to match the input.
{"custom_id":"req-1","method":"POST","url":"/v1/chat/completions","body":{"model":"gpt-4o-mini","messages":[{"role":"user","content":"Explain quantum entanglement"}],"max_tokens":500}}
{"custom_id":"req-2","method":"POST","url":"/v1/chat/completions","body":{"model":"gpt-4o-mini","messages":[{"role":"user","content":"Explain relativity"}],"max_tokens":500}}Files API
/files/files/files/{file_id}/files/{file_id}/files/{file_id}/content| Parameter | Type | Required | Description |
|---|---|---|---|
| file | file | Required | Multipart file field holding the JSONL content |
| purpose | string | Optional | Only batch is supported; defaults to batch |
Batches API
/batches/batches/batches/{batch_id}/batches/{batch_id}/cancel| Parameter | Type | Required | Description |
|---|---|---|---|
| input_file_id | string | Required | The file id returned by the upload step |
| endpoint | string | Optional | The path every line targets, e.g. /v1/chat/completions. Defaults to the url on the first line |
| completion_window | string | Optional | Completion window, defaults to 24h |
| metadata | object | Optional | Custom key-value pairs, echoed back unchanged |
status goes validating → in_progress → finalizing → completed, plus the three failure terminals failed / expired / cancelled. output_file_id and error_file_id only appear once the batch is terminal; successful and failed lines land in those two files respectively.
Billing
How the batch price is set
Batch price = the model's real-time price × its batch discount, with the same coefficient applied to input and output. The discount is listed per model, not site-wide: upstream providers set different batch discounts, and some models get no discount at all.
Where to look: open the model in the model catalogue and check its pricing — "Batch input price" and "Batch output price" are the batch rates. What is listed is what you are charged: if those two lines are there, the batch is settled at them; if they are not, the model has no batch discount and /batches costs the same as a real-time call (it can still be submitted).
Tiered models (those priced by context length) have no single input / output rate, so instead of two price lines they carry a note under the tier table: every tier is multiplied by the discount. Settlement works the same way — each request in the batch is matched to its own tier by its own context length, not by the batch's total token count.
How you are charged
- Quota is estimated from the input file and pre-charged at submit time (input estimated from the text, output from each line's max_tokens); the pre-charge already uses the batch price
- Once the batch finishes, the real per-line usage in the output file is summed and the difference is settled either way
- Your group discount still applies and is multiplied with the batch discount (the detail page shows the default group; another group multiplies once more)
- If the batch fails, is cancelled, or produces no output at all, the pre-charge is refunded in full
- An expired batch is charged for the part that did complete; the rest is refunded
Limits and Differences
- Up to 32MB per file and 50000 lines per batch
- Up to 50 batches in flight per account
- purpose only supports batch (not fine-tune / assistants / vision)
- Every line in a batch must use the same model — split mixed-model workloads into separate batches
- Files (input, output and error) are kept for 30 days by default and then removed, so download your results in time
Didn't find what you were looking for?Contact us →