Quickstart for Whisper pipelines
Integrate our uncensored text model into your speech or image pipelines with a standard OpenAI-compatible chat completions API. This guide covers the essential steps to send requests, handle streaming, and manage context windows for post-processing tasks.
Prerequisites
Before starting, ensure you have an active account and API key. Sign up on the Get API key page using only your email and a password. The key appears immediately. You receive $0.50 in trial credit valid for 7 days, no credit card required. For production, top up from $10 via crypto (USDT or USDC). Our pricing is straightforward: $0.25 per 1M input tokens and $1.00 per 1M output tokens. Credits never expire.
Base URL: https://api.whisperapis.com/v1. This endpoint works with any OpenAI-compatible client. You only need to configure the base URL and provide your API key in the authorization header. No special SDKs are required for basic HTTP requests.
Initialize Client
Most developers will use the official OpenAI SDK or a compatible library. Initialize the client by pointing it to our base URL. This ensures all requests route to our uncensored model rather than OpenAI's servers. The model ID is always uncensored.
Ensure your environment variable OPENAI_API_KEY is set to your key from the dashboard. If you are using a different language or library, verify that it supports custom base URLs. The SDK handles the JSON serialization and header injection for you, reducing boilerplate code in your pipeline scripts.
Send a Chat Completion
Send a standard chat completion request to refine your pipeline output. For example, extract structured JSON from speech-to-text transcripts or generate captions for images without content filters.
The endpoint is POST /v1/chat/completions. Include the model ID, your messages array, and optionally your system prompt to define the uncensored behavior.
curl https://api.whisperapis.com/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "uncensored",
"messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
}'Our model handles high-context data well, supporting up to 100,000 tokens total. This is ideal for processing long transcripts or detailed image descriptions in a single pass.
Streaming Responses
For real-time post-processing or interactive pipelines, enable streaming. Set the stream parameter to true in your request. The API returns Server-Sent Events (SSE) instead of a single JSON object.
Parse the delta fields as they arrive. This reduces perceived latency for large outputs, such as generating detailed scripts or extracting long text blocks from noisy OCR results. The streaming format is compatible with standard OpenAI SDK streaming handlers.
stream = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Tell the story in second person."}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)Ensure your client handles partial chunks correctly. Aggregate the deltas to reconstruct the full text response once the stream ends with a finish_reason of stop or length.
Tool Calling Support
Our uncensored model supports function calling, allowing you to define structured outputs for your pipelines. Define your tools in the tools array with JSON schema definitions.
The model returns tool calls in the response, which you can execute to update the conversation context. This is useful for extracting specific data points from unstructured text, such as dates, names, or coordinates from speech transcripts.
from openai import OpenAI
client = OpenAI(base_url="https://api.whisperapis.com/v1", api_key="YOUR_KEY")
resp = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)Pass the tool calls back to the API in subsequent messages. The model will refine its output based on the results. This enables complex multi-step reasoning without leaving the chat interface. Ensure your schema is strict to avoid malformed JSON responses.
Limits, Errors, and Context
Monitor your usage to avoid interruptions. The API enforces a limit of 300 requests per minute per key. If you exceed this, you receive a 429 rate limit error. Implement exponential backoff in your client.
Common Errors:
401: Invalid or expired API key. Regenerate your key from the dashboard.402: Insufficient credit. Top up your account to continue processing.429: Rate limit exceeded. Reduce request frequency or batch smaller tasks.
Context Window: The model supports 100,000 tokens for prompt + completion. If your input exceeds this, truncate older messages or split the data. Request bodies are limited to 8 MB. Ensure your payloads fit within these constraints for reliable performance.
Node.js
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://api.whisperapis.com/v1", apiKey: process.env.API_KEY });
const resp = await client.chat.completions.create({
model: "uncensored",
messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);API facts in one table
Everything the endpoint can and cannot do, in one place — check it before you top up.
| Spec | Value |
|---|---|
| API format | OpenAI Chat Completions schema; official openai SDKs work unchanged |
| Model | uncensored |
| Endpoints | POST /v1/chat/completions · GET /v1/models |
| API key | Bearer token in the Authorization header |
| Base URL | https://api.whisperapis.com/v1 |
| Max output | up to 16,000 tokens per request (default 2,048) |
| SSE streaming | Supported (stream: true), usage included at the end |
| JSON mode | JSON object mode via response_format json_object |
| Other parameters | temperature, top_p, stop, seed, presence_penalty, frequency_penalty |
| Function calling | Yes — tools, tool_choice; replies carry tool_calls, also when streaming; send results back as role: tool |
| Context window | 100,000 tokens, input and output combined |
| Request size | 8 MB request body |
| Rate limit | 300 requests per minute per key |
| Parallel requests | up to 8 in parallel per key |
| Response headers | X-Request-Id, X-Balance-USD, X-RateLimit-Limit-Requests, X-RateLimit-Limit-Concurrency |
| Top-up | crypto: USDT on TRON or USDC on Base, $10–$500, any whole sum |
| Bonus credit | +5% from $50, +10% from $100 |
| How you pay | prepaid credit, charged by real token usage; errors and refusals are free |
| Free trial | $0.50 of credit valid 7 days, no card needed |
| Subscription | no monthly fee; paid credit does not expire |
| Token prices | input $0.25 / 1M tokens, output $1.00 / 1M tokens |
| Keys | one key per account, regenerate any time (the old one stops working) |
| Account | sign in with Google or with e-mail + password |
| Content policy | adult content allowed; sexual content involving minors is refused |
HTTP errors
Every error is JSON with a type you can switch on. You are never charged for an error.
| HTTP | Type | What to do |
|---|---|---|
400 | bad_request | invalid JSON, empty messages, bad parameter, or prompt + max_tokens over the window — fix and resend |
401 | missing_key · invalid_key · key_revoked | no key, wrong key, or a key replaced by a newer one |
402 | no_credit | out of credit; add credit and retry |
403 | content_blocked | sexual content involving minors — refused, not billed |
404 | not_found | only /v1/chat/completions and /v1/models exist |
413 | request_too_large | request body larger than 8 MB |
429 | rate_limited · concurrency | over 300/min or 8 parallel — back off and retry |
503 | upstream_busy | temporary overload, retry shortly |
Questions and answers
What is the context window size?
The model supports a context window of 100,000 tokens, which includes both the prompt and the completion. This allows you to process large transcripts or detailed descriptions in a single request.
Does this API generate images or audio?
No. This is a text-only chat completions API. It is designed for post-processing text outputs from other modalities like speech-to-text or image generation. It does not generate audio, images, or video.
How do I handle streaming errors?
Streaming responses return SSE events. If an error occurs, the stream will typically end with an error message in the error field of the final chunk. Ensure your client handles connection drops gracefully and retries if necessary.
Your key is one form away
Create an account, copy the key, change the base URL. That is the whole setup.