Get API key

Uncensored LLM Online API Quickstart

Integrate the uncensored LLM online API into your application with this quickstart guide. Use standard OpenAI-compatible endpoints to send raw, unfiltered text generation requests without managing GPU infrastructure.

$0.25
per 1M input tokens
$1.00
Output tokens / 1M
100,000
token context
$0.50
trial credit
300
requests per minute

Base URL and Authentication

The uncensored llm online API uses standard OpenAI-compatible endpoints. Configure your client to point to the base URL https://api.uncensoredllmonline.com/v1. Authentication is handled via an API key passed in the Authorization header. You can generate your key on the Get API key page using only an email and password. The key is displayed immediately upon signup. No phone number or credit card is required to start. Each account holds one active key, which can be regenerated at any time to revoke access for the old key. Keep your key secure, as it provides direct access to your prepaid credits.

Chat Completions Endpoint

Send text to the uncensored llm by posting to POST /v1/chat/completions. This endpoint accepts a list of messages and returns the model's response. The model ID to use is uncensored. This is an open-weight model tuned to answer without content refusals for lawful adult use. It is not GPT, Claude, or another vendor's model. The API returns raw text generation capabilities suitable for custom integrations. Use the following cURL command to test a basic request:

curl https://api.uncensoredllmonline.com/v1/chat/completions \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "uncensored",
    "messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
  }'

The request body should include the model ID and your message history. Ensure your client sends the Authorization: Bearer <your-key> header. The response will contain the generated text, which you can parse and use in your application logic.

Python SDK Integration

Use the official OpenAI Python SDK to interact with the uncensored ai model. Set the base_url and api_key in your client initialization. This allows you to use familiar methods like chat.completions.create. The following example demonstrates how to initialize the client and send a simple query:

from openai import OpenAI

client = OpenAI(base_url="https://api.uncensoredllmonline.com/v1", api_key="YOUR_KEY")

resp = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)

This approach lets you leverage existing code structures while switching to an uncensored LLM. The SDK handles JSON serialization and response parsing automatically. Remember that the model ID must be set to uncensored to ensure you are hitting the correct endpoint. This method is ideal for developers who want to minimize boilerplate code.

Node SDK Integration

For JavaScript or TypeScript projects, the OpenAI Node SDK works seamlessly with our API. Initialize the client with your API key and the correct base URL. The endpoint supports standard chat completions, making it easy to integrate into web apps or serverless functions. Here is how you can configure the client:

import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://api.uncensoredllmonline.com/v1", apiKey: process.env.API_KEY });

const resp = await client.chat.completions.create({
  model: "uncensored",
  messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);

Pass the message array and the model ID uncensored to the completion function. The response will contain the generated text. This integration is suitable for developers who need to embed uncensored text generation into Node.js applications. Ensure you handle errors appropriately, as the API may return rate limit or credit exhaustion errors.

Streaming Responses (SSE)

Enable real-time text generation by setting stream: true in your request. The API returns Server-Sent Events (SSE) that deliver the response token by token. This is useful for chat interfaces or applications where latency matters. Use the following example to implement streaming:

stream = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Tell the story in second person."}],
    stream=True,
)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)

Read from the stream as tokens arrive. The uncensored llm online API supports this natively. Streaming reduces the perceived latency for end-users. Be aware that streaming consumes the same token count as non-streaming requests. Each token is billed according to your prepaid credit balance.

Rate Limits and Context Window

The uncensored llm online API enforces a limit of 300 requests per minute per key. The maximum request body size is 8 MB. The context window is 100,000 tokens, covering both prompt and completion. If you exceed the rate limit, the API returns a 429 error. If your prepaid credit is exhausted, you receive a 402 error. An invalid or expired key results in a 401 error. These limits ensure fair usage across all developers. Pay-as-you-go prepaid credit never expires, so you can use your balance whenever you need it.

Specs at a glance

A quick checklist for developers: format, limits, features, billing.

ItemValue
ProtocolOpenAI-compatible: any OpenAI SDK or client works — change the base URL and the key
API keyBearer token in the Authorization header
Model IDuncensored
EndpointsPOST /v1/chat/completions · GET /v1/models
Base URLhttps://api.uncensoredllmonline.com/v1
SSE streamingYes — server-sent events; the last chunk carries token usage
Context window100,000 tokens (prompt + completion together)
Sampling parameterstemperature, top_p, stop, seed and the two penalties are passed through
Structured outputJSON object mode via response_format json_object
Function callingYes — tools, tool_choice; replies carry tool_calls, also when streaming; send results back as role: tool
Max outputup to 16,000 tokens per request (default 2,048)
HeadersX-Request-Id, X-Balance-USD, X-RateLimit-Limit-Requests, X-RateLimit-Limit-Concurrency
Request sizeup to 8 MB per request
Parallel requests8 requests at the same time per key
Requests per minute300/min per key
Price$0.25 per 1M input tokens · $1.00 per 1M output tokens
Free trial$0.50 for 7 days, no card
Subscriptionpaid credit never expires, no subscription
Bonus credit+5% on $50+, +10% on $100+
Paymentcrypto: USDT on TRON or USDC on Base, $10–$500, any whole sum
How you payprepaid credit, charged by real token usage; errors and refusals are free
Sign-insign in with Google or with e-mail + password
Keysone key per account, regenerate any time (the old one stops working)
Content policyuncensored for adults; the only hard rule: no sexual content involving minors

HTTP errors

Every error is JSON with a type you can switch on. You are never charged for an error.

HTTPTypeWhat to do
400bad_requestmalformed request or too long for the context window
401missing_key · invalid_key · key_revokedno key, wrong key, or a key replaced by a newer one
402no_creditout of credit; add credit and retry
403content_blockedrefused by the content policy
404not_foundunknown endpoint
413request_too_largerequest body larger than 8 MB
429rate_limited · concurrencyover 300/min or 8 parallel — back off and retry
503upstream_busytemporary overload, retry shortly

Questions and answers

What is the context window size?

The context window is 100,000 tokens, which includes both the input prompt and the output completion. This allows for long conversations or large document processing.

How does billing work?

You pay-as-you-go with prepaid credit. The rate is $0.25 per 1M input tokens and $1.00 per 1M output tokens. Credits never expire, and you can top up starting from $10.

Is the model the same as GPT-4?

No, the model ID is <code>uncensored</code>. It is an open-weight model run on our own GPU servers, tuned to answer without content refusals. It is not GPT, Claude, or any other vendor's model.

Your key is one form away

Create an account, copy the key, change the base URL. That is the whole setup.

Get API key