Get API key

Quickstart for Sora 2 pipelines

Generate uncensored text for your Sora 2 video pipelines using our OpenAI-compatible API. This quickstart covers authentication, basic requests, streaming, and tool calling to get your text engine running immediately.

Authentication and Base URL

Our API uses standard OpenAI-compatible authentication. You will need an API key from your dashboard and the base URL https://api.sora2apis.com/v1. Pass your key in the Authorization header as a Bearer token. This setup works with any SDK that supports OpenAI compatibility.

Sign up for an account to get your key. New accounts receive $0.50 in trial credit for 7 days, no credit card required. For paid use, top-ups start at $10. The key is generated immediately after signup and can be regenerated at any time.

curl https://api.sora2apis.com/v1/chat/completions \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "uncensored",
    "messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
  }'

Chat Completions Endpoint

The core endpoint is POST /v1/chat/completions. It accepts a model field set to uncensored, which is an open-weight model tuned to answer without content refusals for lawful adult use. It is not GPT, Claude, or any other vendor's model. The model has a 100,000 token context window for both prompt and completion.

Send your text input via the messages array. The API returns text output ready for your video generation pipeline. There are no embeddings, audio, or video generation capabilities in this endpoint. It is strictly text-in, text-out.

from openai import OpenAI

client = OpenAI(base_url="https://api.sora2apis.com/v1", api_key="YOUR_KEY")

resp = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)

Streaming Responses (SSE)

For real-time text generation, set stream: true in your request. The API returns Server-Sent Events (SSE) allowing you to process tokens as they are generated. This is useful for long scripts or captions where you want to display progress or feed data incrementally.

Each SSE event contains a partial delta of the response. Aggregate these deltas to reconstruct the full text. The streaming response follows the standard OpenAI SSE format, ensuring compatibility with most modern SDKs. This is ideal for dynamic content generation in your Sora 2 workflow.

stream = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Tell the story in second person."}],
    stream=True,
)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)

Tool and Function Calling

The API supports tool and function calling, allowing you to define custom functions for your model to use. This is valuable for structured data extraction or automated decision-making within your text pipeline. Define your tools in the tools array and specify how the model should respond.

The model will return tool calls when appropriate, which you can then execute and feed back into the conversation. This enables complex workflows where the text engine can interact with external systems or format data precisely for your video generation needs. It is a powerful feature for advanced Sora 2 pipelines.

Rate Limits and Usage

You are limited to 300 requests per minute per key. The maximum request body size is 8 MB. If you exceed the rate limit, you will receive a 429 error. Ensure your application handles retries appropriately. Each account gets one API key, which can be regenerated at any time, revoking the old key.

Pricing is $0.25 per 1M input tokens and $1.00 per 1M output tokens. Credits are prepaid and never expire. If you run out of credit, you will receive a 402 error. Invalid keys result in a 401 error. Top-ups start at $10, with bonuses for larger amounts.

Model Configuration

The model ID is always uncensored. It is an open-weight model running on our own GPU servers, tuned for lawful adult content without refusals. It is not GPT, Claude, Gemini, Grok, or DeepSeek. The model does not generate audio, images, or video. It is strictly for text generation.

You can adjust standard parameters like temperature and max_tokens to control the output. The context window is 100,000 tokens. This flexibility allows you to fine-tune the text output for your specific Sora 2 video generation needs. The model is designed to be a reliable text engine for your pipeline.

Node.js

import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://api.sora2apis.com/v1", apiKey: process.env.API_KEY });

const resp = await client.chat.completions.create({
  model: "uncensored",
  messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);

Specs at a glance

Before you integrate, here is exactly what you get with a key.

ItemValue
CompatibilityOpenAI Chat Completions schema; official openai SDKs work unchanged
Modeluncensored
EndpointsPOST /v1/chat/completions · GET /v1/models
AuthenticationAuthorization: Bearer YOUR_KEY
Base URLhttps://api.sora2apis.com/v1
StreamingSupported (stream: true), usage included at the end
Context window100,000 tokens (prompt + completion together)
Completion length16,000 tokens max; 2,048 if max_tokens is not set
Other parameterstemperature, top_p, stop, seed, presence_penalty, frequency_penalty
JSON modeJSON object mode via response_format json_object
Function callingSupported: tools + tool_choice, tool_calls in the reply (streamed too), tool results as role: tool messages
Max bodyup to 8 MB per request
HeadersX-Request-Id, X-Balance-USD, X-RateLimit-Limit-Requests, X-RateLimit-Limit-Concurrency
Parallel requestsup to 8 in parallel per key
Rate limit300 requests per minute per key
Billingprepaid credit, charged by real token usage; errors and refusals are free
Volume bonus+5% on $50+, +10% on $100+
Credit expirypaid credit never expires, no subscription
Top-upcrypto: USDT on TRON or USDC on Base, $10–$500, any whole sum
Price$0.25 per 1M input tokens · $1.00 per 1M output tokens
Free trial$0.50 of credit valid 7 days, no card needed
Key managementone key per account, regenerate any time (the old one stops working)
Contentadult content allowed; sexual content involving minors is refused
Sign-insign in with Google or with e-mail + password

When a request fails

The type field is stable, the message is for humans. Errors cost nothing.

HTTPTypeWhat to do
400bad_requestinvalid JSON, empty messages, bad parameter, or prompt + max_tokens over the window — fix and resend
401missing_key · invalid_key · key_revokedno key, wrong key, or a key replaced by a newer one
402no_creditbalance is empty — top up, requests resume at once
403content_blockedrefused by the content policy
404not_foundonly /v1/chat/completions and /v1/models exist
413request_too_largerequest body larger than 8 MB
429rate_limited · concurrencyover 300/min or 8 parallel — back off and retry
503upstream_busytemporary overload, retry shortly

Questions and answers

What is the context window size?

The context window is 100,000 tokens, which includes both the prompt and the completion. This allows for long-form content generation suitable for scripts and detailed captions.

How do I handle rate limits?

You are limited to 300 requests per minute per key. If you exceed this, you will receive a 429 error. Implement exponential backoff in your application to handle retries gracefully.

Is the content truly uncensored?

The model does not refuse lawful adult, fictional, security-research, or controversial topics. However, sexual content involving minors is always blocked. It is designed for creative freedom within legal bounds.

Your key is one form away

Create an account, copy the key, change the base URL. That is the whole setup.

Get API key