Get API key

Quickstart for Whisper pipelines

Integrate our uncensored text model into your speech or image pipelines with a standard OpenAI-compatible chat completions API. This guide covers the essential steps to send requests, handle streaming, and manage context windows for post-processing tasks.

Prerequisites

Before starting, ensure you have an active account and API key. Sign up on the Get API key page using only your email and a password. The key appears immediately. You receive $0.50 in trial credit valid for 7 days, no credit card required. For production, top up from $10 via crypto (USDT or USDC). Our pricing is straightforward: $0.25 per 1M input tokens and $1.00 per 1M output tokens. Credits never expire.

Base URL: https://api.whisperapis.com/v1. This endpoint works with any OpenAI-compatible client. You only need to configure the base URL and provide your API key in the authorization header. No special SDKs are required for basic HTTP requests.

Initialize Client

Most developers will use the official OpenAI SDK or a compatible library. Initialize the client by pointing it to our base URL. This ensures all requests route to our uncensored model rather than OpenAI's servers. The model ID is always uncensored.

Ensure your environment variable OPENAI_API_KEY is set to your key from the dashboard. If you are using a different language or library, verify that it supports custom base URLs. The SDK handles the JSON serialization and header injection for you, reducing boilerplate code in your pipeline scripts.

Send a Chat Completion

Send a standard chat completion request to refine your pipeline output. For example, extract structured JSON from speech-to-text transcripts or generate captions for images without content filters.

The endpoint is POST /v1/chat/completions. Include the model ID, your messages array, and optionally your system prompt to define the uncensored behavior.

curl https://api.whisperapis.com/v1/chat/completions \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "uncensored",
    "messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
  }'

Our model handles high-context data well, supporting up to 100,000 tokens total. This is ideal for processing long transcripts or detailed image descriptions in a single pass.

Streaming Responses

For real-time post-processing or interactive pipelines, enable streaming. Set the stream parameter to true in your request. The API returns Server-Sent Events (SSE) instead of a single JSON object.

Parse the delta fields as they arrive. This reduces perceived latency for large outputs, such as generating detailed scripts or extracting long text blocks from noisy OCR results. The streaming format is compatible with standard OpenAI SDK streaming handlers.

stream = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Tell the story in second person."}],
    stream=True,
)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)

Ensure your client handles partial chunks correctly. Aggregate the deltas to reconstruct the full text response once the stream ends with a finish_reason of stop or length.

Tool Calling Support

Our uncensored model supports function calling, allowing you to define structured outputs for your pipelines. Define your tools in the tools array with JSON schema definitions.

The model returns tool calls in the response, which you can execute to update the conversation context. This is useful for extracting specific data points from unstructured text, such as dates, names, or coordinates from speech transcripts.

from openai import OpenAI

client = OpenAI(base_url="https://api.whisperapis.com/v1", api_key="YOUR_KEY")

resp = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)

Pass the tool calls back to the API in subsequent messages. The model will refine its output based on the results. This enables complex multi-step reasoning without leaving the chat interface. Ensure your schema is strict to avoid malformed JSON responses.

Limits, Errors, and Context

Monitor your usage to avoid interruptions. The API enforces a limit of 300 requests per minute per key. If you exceed this, you receive a 429 rate limit error. Implement exponential backoff in your client.

Common Errors:

  • 401: Invalid or expired API key. Regenerate your key from the dashboard.
  • 402: Insufficient credit. Top up your account to continue processing.
  • 429: Rate limit exceeded. Reduce request frequency or batch smaller tasks.

Context Window: The model supports 100,000 tokens for prompt + completion. If your input exceeds this, truncate older messages or split the data. Request bodies are limited to 8 MB. Ensure your payloads fit within these constraints for reliable performance.

Node.js

import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://api.whisperapis.com/v1", apiKey: process.env.API_KEY });

const resp = await client.chat.completions.create({
  model: "uncensored",
  messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);

API facts in one table

Everything the endpoint can and cannot do, in one place — check it before you top up.

SpecValue
API formatOpenAI Chat Completions schema; official openai SDKs work unchanged
Modeluncensored
EndpointsPOST /v1/chat/completions · GET /v1/models
API keyBearer token in the Authorization header
Base URLhttps://api.whisperapis.com/v1
Max outputup to 16,000 tokens per request (default 2,048)
SSE streamingSupported (stream: true), usage included at the end
JSON modeJSON object mode via response_format json_object
Other parameterstemperature, top_p, stop, seed, presence_penalty, frequency_penalty
Function callingYes — tools, tool_choice; replies carry tool_calls, also when streaming; send results back as role: tool
Context window100,000 tokens, input and output combined
Request size8 MB request body
Rate limit300 requests per minute per key
Parallel requestsup to 8 in parallel per key
Response headersX-Request-Id, X-Balance-USD, X-RateLimit-Limit-Requests, X-RateLimit-Limit-Concurrency
Top-upcrypto: USDT on TRON or USDC on Base, $10–$500, any whole sum
Bonus credit+5% from $50, +10% from $100
How you payprepaid credit, charged by real token usage; errors and refusals are free
Free trial$0.50 of credit valid 7 days, no card needed
Subscriptionno monthly fee; paid credit does not expire
Token pricesinput $0.25 / 1M tokens, output $1.00 / 1M tokens
Keysone key per account, regenerate any time (the old one stops working)
Accountsign in with Google or with e-mail + password
Content policyadult content allowed; sexual content involving minors is refused

HTTP errors

Every error is JSON with a type you can switch on. You are never charged for an error.

HTTPTypeWhat to do
400bad_requestinvalid JSON, empty messages, bad parameter, or prompt + max_tokens over the window — fix and resend
401missing_key · invalid_key · key_revokedno key, wrong key, or a key replaced by a newer one
402no_creditout of credit; add credit and retry
403content_blockedsexual content involving minors — refused, not billed
404not_foundonly /v1/chat/completions and /v1/models exist
413request_too_largerequest body larger than 8 MB
429rate_limited · concurrencyover 300/min or 8 parallel — back off and retry
503upstream_busytemporary overload, retry shortly

Questions and answers

What is the context window size?

The model supports a context window of 100,000 tokens, which includes both the prompt and the completion. This allows you to process large transcripts or detailed descriptions in a single request.

Does this API generate images or audio?

No. This is a text-only chat completions API. It is designed for post-processing text outputs from other modalities like speech-to-text or image generation. It does not generate audio, images, or video.

How do I handle streaming errors?

Streaming responses return SSE events. If an error occurs, the stream will typically end with an error message in the error field of the final chunk. Ensure your client handles connection drops gracefully and retries if necessary.

Your key is one form away

Create an account, copy the key, change the base URL. That is the whole setup.

Get API key