Uncensored LLM Online API Quickstart
Integrate the uncensored LLM online API into your application with this quickstart guide. Use standard OpenAI-compatible endpoints to send raw, unfiltered text generation requests without managing GPU infrastructure.
- $0.25
- per 1M input tokens
- $1.00
- Output tokens / 1M
- 100,000
- token context
- $0.50
- trial credit
- 300
- requests per minute
Base URL and Authentication
The uncensored llm online API uses standard OpenAI-compatible endpoints. Configure your client to point to the base URL https://api.uncensoredllmonline.com/v1. Authentication is handled via an API key passed in the Authorization header. You can generate your key on the Get API key page using only an email and password. The key is displayed immediately upon signup. No phone number or credit card is required to start. Each account holds one active key, which can be regenerated at any time to revoke access for the old key. Keep your key secure, as it provides direct access to your prepaid credits.
Chat Completions Endpoint
Send text to the uncensored llm by posting to POST /v1/chat/completions. This endpoint accepts a list of messages and returns the model's response. The model ID to use is uncensored. This is an open-weight model tuned to answer without content refusals for lawful adult use. It is not GPT, Claude, or another vendor's model. The API returns raw text generation capabilities suitable for custom integrations. Use the following cURL command to test a basic request:
curl https://api.uncensoredllmonline.com/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "uncensored",
"messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
}'
The request body should include the model ID and your message history. Ensure your client sends the Authorization: Bearer <your-key> header. The response will contain the generated text, which you can parse and use in your application logic.
Python SDK Integration
Use the official OpenAI Python SDK to interact with the uncensored ai model. Set the base_url and api_key in your client initialization. This allows you to use familiar methods like chat.completions.create. The following example demonstrates how to initialize the client and send a simple query:
from openai import OpenAI
client = OpenAI(base_url="https://api.uncensoredllmonline.com/v1", api_key="YOUR_KEY")
resp = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)
This approach lets you leverage existing code structures while switching to an uncensored LLM. The SDK handles JSON serialization and response parsing automatically. Remember that the model ID must be set to uncensored to ensure you are hitting the correct endpoint. This method is ideal for developers who want to minimize boilerplate code.
Node SDK Integration
For JavaScript or TypeScript projects, the OpenAI Node SDK works seamlessly with our API. Initialize the client with your API key and the correct base URL. The endpoint supports standard chat completions, making it easy to integrate into web apps or serverless functions. Here is how you can configure the client:
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://api.uncensoredllmonline.com/v1", apiKey: process.env.API_KEY });
const resp = await client.chat.completions.create({
model: "uncensored",
messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);
Pass the message array and the model ID uncensored to the completion function. The response will contain the generated text. This integration is suitable for developers who need to embed uncensored text generation into Node.js applications. Ensure you handle errors appropriately, as the API may return rate limit or credit exhaustion errors.
Streaming Responses (SSE)
Enable real-time text generation by setting stream: true in your request. The API returns Server-Sent Events (SSE) that deliver the response token by token. This is useful for chat interfaces or applications where latency matters. Use the following example to implement streaming:
stream = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Tell the story in second person."}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
Read from the stream as tokens arrive. The uncensored llm online API supports this natively. Streaming reduces the perceived latency for end-users. Be aware that streaming consumes the same token count as non-streaming requests. Each token is billed according to your prepaid credit balance.
Rate Limits and Context Window
The uncensored llm online API enforces a limit of 300 requests per minute per key. The maximum request body size is 8 MB. The context window is 100,000 tokens, covering both prompt and completion. If you exceed the rate limit, the API returns a 429 error. If your prepaid credit is exhausted, you receive a 402 error. An invalid or expired key results in a 401 error. These limits ensure fair usage across all developers. Pay-as-you-go prepaid credit never expires, so you can use your balance whenever you need it.
Specs at a glance
A quick checklist for developers: format, limits, features, billing.
| Item | Value |
|---|---|
| Protocol | OpenAI-compatible: any OpenAI SDK or client works — change the base URL and the key |
| API key | Bearer token in the Authorization header |
| Model ID | uncensored |
| Endpoints | POST /v1/chat/completions · GET /v1/models |
| Base URL | https://api.uncensoredllmonline.com/v1 |
| SSE streaming | Yes — server-sent events; the last chunk carries token usage |
| Context window | 100,000 tokens (prompt + completion together) |
| Sampling parameters | temperature, top_p, stop, seed and the two penalties are passed through |
| Structured output | JSON object mode via response_format json_object |
| Function calling | Yes — tools, tool_choice; replies carry tool_calls, also when streaming; send results back as role: tool |
| Max output | up to 16,000 tokens per request (default 2,048) |
| Headers | X-Request-Id, X-Balance-USD, X-RateLimit-Limit-Requests, X-RateLimit-Limit-Concurrency |
| Request size | up to 8 MB per request |
| Parallel requests | 8 requests at the same time per key |
| Requests per minute | 300/min per key |
| Price | $0.25 per 1M input tokens · $1.00 per 1M output tokens |
| Free trial | $0.50 for 7 days, no card |
| Subscription | paid credit never expires, no subscription |
| Bonus credit | +5% on $50+, +10% on $100+ |
| Payment | crypto: USDT on TRON or USDC on Base, $10–$500, any whole sum |
| How you pay | prepaid credit, charged by real token usage; errors and refusals are free |
| Sign-in | sign in with Google or with e-mail + password |
| Keys | one key per account, regenerate any time (the old one stops working) |
| Content policy | uncensored for adults; the only hard rule: no sexual content involving minors |
HTTP errors
Every error is JSON with a type you can switch on. You are never charged for an error.
| HTTP | Type | What to do |
|---|---|---|
400 | bad_request | malformed request or too long for the context window |
401 | missing_key · invalid_key · key_revoked | no key, wrong key, or a key replaced by a newer one |
402 | no_credit | out of credit; add credit and retry |
403 | content_blocked | refused by the content policy |
404 | not_found | unknown endpoint |
413 | request_too_large | request body larger than 8 MB |
429 | rate_limited · concurrency | over 300/min or 8 parallel — back off and retry |
503 | upstream_busy | temporary overload, retry shortly |
Questions and answers
What is the context window size?
The context window is 100,000 tokens, which includes both the input prompt and the output completion. This allows for long conversations or large document processing.
How does billing work?
You pay-as-you-go with prepaid credit. The rate is $0.25 per 1M input tokens and $1.00 per 1M output tokens. Credits never expire, and you can top up starting from $10.
Is the model the same as GPT-4?
No, the model ID is <code>uncensored</code>. It is an open-weight model run on our own GPU servers, tuned to answer without content refusals. It is not GPT, Claude, or any other vendor's model.
Your key is one form away
Create an account, copy the key, change the base URL. That is the whole setup.