≡ All docs

Model API / Batch

Batch

Updated: 2026-09-19

Overview

Pack thousands of requests into one JSONL file, submit it once, and get the results back asynchronously within 24 hours. Good for offline evaluation, bulk labelling and data cleaning. Batch pricing is listed per model: a model whose detail page shows a Batch price is settled at that price, and one without it is settled at the real-time price. Compatible with the OpenAI Batch API, so client.files.* and client.batches.* from the official SDKs work as-is.

End-to-end Flow

  • 1. Upload the input file: POST /files (purpose=batch) and keep the returned file id
  • 2. Create the batch: POST /batches with that input_file_id
  • 3. Poll: GET /batches/{batch_id} until status becomes completed
  • 4. Download results: GET /files/{output_file_id}/content
# 1. upload
curl -X POST "https://www.starunion.net/v1/files" \
-H "Authorization: Bearer YOUR_API_KEY" \
-F "purpose=batch" \
-F "file=@requests.jsonl"
# 2. create
curl -X POST "https://www.starunion.net/v1/batches" \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"input_file_id": "file-xxxxxxxx",
"endpoint": "/v1/chat/completions",
"completion_window": "24h"
}'
# 3. poll
curl "https://www.starunion.net/v1/batches/batch_xxxxxxxx" \
-H "Authorization: Bearer YOUR_API_KEY"
# 4. download
curl "https://www.starunion.net/v1/files/file-yyyyyyyy/content" \
-H "Authorization: Bearer YOUR_API_KEY"

Input File Format

JSONL: one JSON object per line, one request per line. You choose custom_id yourself and use it to map output lines back to input lines, because the output order is not guaranteed to match the input.

cURL
{"custom_id":"req-1","method":"POST","url":"/v1/chat/completions","body":{"model":"gpt-4o-mini","messages":[{"role":"user","content":"Explain quantum entanglement"}],"max_tokens":500}}
{"custom_id":"req-2","method":"POST","url":"/v1/chat/completions","body":{"model":"gpt-4o-mini","messages":[{"role":"user","content":"Explain relativity"}],"max_tokens":500}}
Every line must use the same model and the same url, and custom_id must be unique. Break any of these and batch creation returns 400 naming the offending line number.

Files API

POST/files
GET/files
GET/files/{file_id}
DELETE/files/{file_id}
GET/files/{file_id}/content
ParameterTypeRequiredDescription
filefileRequiredMultipart file field holding the JSONL content
purposestringOptionalOnly batch is supported; defaults to batch

Batches API

POST/batches
GET/batches
GET/batches/{batch_id}
POST/batches/{batch_id}/cancel
ParameterTypeRequiredDescription
input_file_idstringRequiredThe file id returned by the upload step
endpointstringOptionalThe path every line targets, e.g. /v1/chat/completions. Defaults to the url on the first line
completion_windowstringOptionalCompletion window, defaults to 24h
metadataobjectOptionalCustom key-value pairs, echoed back unchanged

status goes validating → in_progress → finalizing → completed, plus the three failure terminals failed / expired / cancelled. output_file_id and error_file_id only appear once the batch is terminal; successful and failed lines land in those two files respectively.

Billing

How the batch price is set

Batch price = the model's real-time price × its batch discount, with the same coefficient applied to input and output. The discount is listed per model, not site-wide: upstream providers set different batch discounts, and some models get no discount at all.

Where to look: open the model in the model catalogue and check its pricing — "Batch input price" and "Batch output price" are the batch rates. What is listed is what you are charged: if those two lines are there, the batch is settled at them; if they are not, the model has no batch discount and /batches costs the same as a real-time call (it can still be submitted).

Tiered models (those priced by context length) have no single input / output rate, so instead of two price lines they carry a note under the tier table: every tier is multiplied by the discount. Settlement works the same way — each request in the batch is matched to its own tier by its own context length, not by the batch's total token count.

How you are charged

  • Quota is estimated from the input file and pre-charged at submit time (input estimated from the text, output from each line's max_tokens); the pre-charge already uses the batch price
  • Once the batch finishes, the real per-line usage in the output file is summed and the difference is settled either way
  • Your group discount still applies and is multiplied with the batch discount (the detail page shows the default group; another group multiplies once more)
  • If the batch fails, is cancelled, or produces no output at all, the pre-charge is refunded in full
  • An expired batch is charged for the part that did complete; the rest is refunded
The amount frozen at submit time is an estimate and is usually higher than the final charge; the difference comes back when the batch ends. Keep enough balance to cover the estimate, otherwise batch creation fails with insufficient quota.
Models priced per call (the price list shows a price per request rather than per 1K tokens) cannot be submitted as batches: batches are always settled by token, there is no way to price a per-call model that way, and creation returns 400.

Limits and Differences

  • Up to 32MB per file and 50000 lines per batch
  • Up to 50 batches in flight per account
  • purpose only supports batch (not fine-tune / assistants / vision)
  • Every line in a batch must use the same model — split mixed-model workloads into separate batches
  • Files (input, output and error) are kept for 30 days by default and then removed, so download your results in time

Didn't find what you were looking for?Contact us →