POST /v1/chat/completions behaves exactly like OpenAI's endpoint of the same name. The request
body is forwarded to the upstream model as-is, so tools, vision input, temperature,
max_tokens and friends work whenever the model itself supports them.
curl https://bothub-api.bookab.info/v1/chat/completions \
-H "Authorization: Bearer $BOTHUB_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-5.4",
"messages": [
{"role": "system", "content": "You are a concise assistant."},
{"role": "user", "content": "Explain vector databases in one sentence"}
]
}'
Pass "stream": true for standard SSE:
stream = client.chat.completions.create(
model="gpt-5.4",
messages=[{"role": "user", "content": "Write a short poem"}],
stream=True,
)
for chunk in stream:
print(chunk.choices[0].delta.content or "", end="")
stream_options
yourself — the server adds it.When the model supports it, pass images in OpenAI's format:
{
"model": "gpt-5.4",
"messages": [{
"role": "user",
"content": [
{"type": "text", "text": "What is in this image?"},
{"type": "image_url", "image_url": {"url": "https://example.com/cat.jpg"}}
]
}]
}
The image reference must be something the upstream model can fetch itself: a public URL or a
data: URI. Local file paths do not work.
A single request carries at most 500 messages. A long session that keeps calling tools reaches
that ceiling faster than you'd expect — compact the history instead of appending forever. Going
over returns 400, and the message names the offending field.
POST /v1/responses is OpenAI's newer protocol. It takes input instead of messages:
curl https://bothub-api.bookab.info/v1/responses \
-H "Authorization: Bearer $BOTHUB_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "gpt-5.4", "input": "Hello"}'
store: false on every forwarded request — upstream providers never persist your
prompts.POST /v1/embeddings takes a single string or an array of up to 256 strings.
curl https://bothub-api.bookab.info/v1/embeddings \
-H "Authorization: Bearer $BOTHUB_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "text-embedding-4",
"input": ["first passage", "second passage"]
}'
dimensions and encoding_format (float or base64) are supported when the model supports them.