Quickstart for Sora 2 pipelines
Generate uncensored text for your Sora 2 video pipelines using our OpenAI-compatible API. This quickstart covers authentication, basic requests, streaming, and tool calling to get your text engine running immediately.
Authentication and Base URL
Our API uses standard OpenAI-compatible authentication. You will need an API key from your dashboard and the base URL https://api.sora2apis.com/v1. Pass your key in the Authorization header as a Bearer token. This setup works with any SDK that supports OpenAI compatibility.
Sign up for an account to get your key. New accounts receive $0.50 in trial credit for 7 days, no credit card required. For paid use, top-ups start at $10. The key is generated immediately after signup and can be regenerated at any time.
curl https://api.sora2apis.com/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "uncensored",
"messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
}'
Chat Completions Endpoint
The core endpoint is POST /v1/chat/completions. It accepts a model field set to uncensored, which is an open-weight model tuned to answer without content refusals for lawful adult use. It is not GPT, Claude, or any other vendor's model. The model has a 100,000 token context window for both prompt and completion.
Send your text input via the messages array. The API returns text output ready for your video generation pipeline. There are no embeddings, audio, or video generation capabilities in this endpoint. It is strictly text-in, text-out.
from openai import OpenAI
client = OpenAI(base_url="https://api.sora2apis.com/v1", api_key="YOUR_KEY")
resp = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)
Streaming Responses (SSE)
For real-time text generation, set stream: true in your request. The API returns Server-Sent Events (SSE) allowing you to process tokens as they are generated. This is useful for long scripts or captions where you want to display progress or feed data incrementally.
Each SSE event contains a partial delta of the response. Aggregate these deltas to reconstruct the full text. The streaming response follows the standard OpenAI SSE format, ensuring compatibility with most modern SDKs. This is ideal for dynamic content generation in your Sora 2 workflow.
stream = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Tell the story in second person."}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
Tool and Function Calling
The API supports tool and function calling, allowing you to define custom functions for your model to use. This is valuable for structured data extraction or automated decision-making within your text pipeline. Define your tools in the tools array and specify how the model should respond.
The model will return tool calls when appropriate, which you can then execute and feed back into the conversation. This enables complex workflows where the text engine can interact with external systems or format data precisely for your video generation needs. It is a powerful feature for advanced Sora 2 pipelines.
Rate Limits and Usage
You are limited to 300 requests per minute per key. The maximum request body size is 8 MB. If you exceed the rate limit, you will receive a 429 error. Ensure your application handles retries appropriately. Each account gets one API key, which can be regenerated at any time, revoking the old key.
Pricing is $0.25 per 1M input tokens and $1.00 per 1M output tokens. Credits are prepaid and never expire. If you run out of credit, you will receive a 402 error. Invalid keys result in a 401 error. Top-ups start at $10, with bonuses for larger amounts.
Model Configuration
The model ID is always uncensored. It is an open-weight model running on our own GPU servers, tuned for lawful adult content without refusals. It is not GPT, Claude, Gemini, Grok, or DeepSeek. The model does not generate audio, images, or video. It is strictly for text generation.
You can adjust standard parameters like temperature and max_tokens to control the output. The context window is 100,000 tokens. This flexibility allows you to fine-tune the text output for your specific Sora 2 video generation needs. The model is designed to be a reliable text engine for your pipeline.
Node.js
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://api.sora2apis.com/v1", apiKey: process.env.API_KEY });
const resp = await client.chat.completions.create({
model: "uncensored",
messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);Specs at a glance
Before you integrate, here is exactly what you get with a key.
| Item | Value |
|---|---|
| Compatibility | OpenAI Chat Completions schema; official openai SDKs work unchanged |
| Model | uncensored |
| Endpoints | POST /v1/chat/completions · GET /v1/models |
| Authentication | Authorization: Bearer YOUR_KEY |
| Base URL | https://api.sora2apis.com/v1 |
| Streaming | Supported (stream: true), usage included at the end |
| Context window | 100,000 tokens (prompt + completion together) |
| Completion length | 16,000 tokens max; 2,048 if max_tokens is not set |
| Other parameters | temperature, top_p, stop, seed, presence_penalty, frequency_penalty |
| JSON mode | JSON object mode via response_format json_object |
| Function calling | Supported: tools + tool_choice, tool_calls in the reply (streamed too), tool results as role: tool messages |
| Max body | up to 8 MB per request |
| Headers | X-Request-Id, X-Balance-USD, X-RateLimit-Limit-Requests, X-RateLimit-Limit-Concurrency |
| Parallel requests | up to 8 in parallel per key |
| Rate limit | 300 requests per minute per key |
| Billing | prepaid credit, charged by real token usage; errors and refusals are free |
| Volume bonus | +5% on $50+, +10% on $100+ |
| Credit expiry | paid credit never expires, no subscription |
| Top-up | crypto: USDT on TRON or USDC on Base, $10–$500, any whole sum |
| Price | $0.25 per 1M input tokens · $1.00 per 1M output tokens |
| Free trial | $0.50 of credit valid 7 days, no card needed |
| Key management | one key per account, regenerate any time (the old one stops working) |
| Content | adult content allowed; sexual content involving minors is refused |
| Sign-in | sign in with Google or with e-mail + password |
When a request fails
The type field is stable, the message is for humans. Errors cost nothing.
| HTTP | Type | What to do |
|---|---|---|
400 | bad_request | invalid JSON, empty messages, bad parameter, or prompt + max_tokens over the window — fix and resend |
401 | missing_key · invalid_key · key_revoked | no key, wrong key, or a key replaced by a newer one |
402 | no_credit | balance is empty — top up, requests resume at once |
403 | content_blocked | refused by the content policy |
404 | not_found | only /v1/chat/completions and /v1/models exist |
413 | request_too_large | request body larger than 8 MB |
429 | rate_limited · concurrency | over 300/min or 8 parallel — back off and retry |
503 | upstream_busy | temporary overload, retry shortly |
Questions and answers
What is the context window size?
The context window is 100,000 tokens, which includes both the prompt and the completion. This allows for long-form content generation suitable for scripts and detailed captions.
How do I handle rate limits?
You are limited to 300 requests per minute per key. If you exceed this, you will receive a 429 error. Implement exponential backoff in your application to handle retries gracefully.
Is the content truly uncensored?
The model does not refuse lawful adult, fictional, security-research, or controversial topics. However, sexual content involving minors is always blocked. It is designed for creative freedom within legal bounds.
Your key is one form away
Create an account, copy the key, change the base URL. That is the whole setup.
Get API key