Quickstart: Integrate the Uncensored LLM

Integrate the unrestricted llm into your application with this concise quickstart. Use the OpenAI-compatible API to send text and receive unfiltered responses without managing local infrastructure.

Base URL & Authentication

The Unrestricted API follows the OpenAI chat-completions specification. You interact with it by sending requests to the base URL https://api.unrestricted.cc/v1. Authentication is handled via a Bearer token in the Authorization header. You can generate your API key by signing up with Google or an email address on the Get API key page. The key is displayed immediately and is unique to your account. Each account holds one active key; generating a new one replaces the previous key. Keep your key secure, as it grants access to your prepaid credit.

Chat Completions Endpoint

Send messages to the uncensored llm using the POST /v1/chat/completions endpoint. The model identifier is uncensored. This model is an open-weight model tuned to answer without content refusals for lawful adult use. It is not GPT, Claude, or any other vendor's model. Below is a basic example using curl to request a completion.

curl https://api.unrestricted.cc/v1/chat/completions \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "uncensored",
    "messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
  }'

The response contains the generated text. The context window is 64,000 tokens total. If you do not specify max_tokens, the output is capped at 2,048 tokens. For longer outputs, set max_tokens up to 16,000.

Python SDK Integration

Use the official OpenAI Python SDK to integrate the uncensored ai model into your scripts. Set the base_url and api_key to point to the Unrestricted API. This allows you to use familiar methods like chat.completions.create() while accessing an uncensored llm online that does not apply standard refusal filters. The SDK handles serialization and error parsing automatically.

from openai import OpenAI

client = OpenAI(base_url="https://api.unrestricted.cc/v1", api_key="YOUR_KEY")

resp = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)

This approach is ideal for backend services or data processing pipelines where you need raw text output without additional UI wrappers. The model supports standard parameters like temperature and top_p for controlling randomness.

Node.js SDK Integration

For JavaScript or TypeScript projects, the OpenAI Node SDK works directly with the Unrestricted API. Configure the client with the correct base URL and API key. This enables seamless integration into web servers, APIs, or automated workflows. The uncoded coding llm behavior ensures that code generation and technical explanations are returned without unnecessary conversational filler or safety refusals.

import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://api.unrestricted.cc/v1", apiKey: process.env.API_KEY });

const resp = await client.chat.completions.create({
  model: "uncensored",
  messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);

Remember that the API key is single-use per account. If you rotate keys, update your environment variables accordingly. The SDK supports both synchronous and asynchronous calls, making it suitable for high-throughput applications.

Streaming Responses

Enable streaming by setting stream: true in your request. The API returns a Server-Sent Events (SSE) stream. Each chunk contains partial text. The final chunk includes the token usage statistics. Streaming reduces perceived latency for end-users, especially when generating long outputs. It is particularly useful for chat interfaces or real-time code generation tools.

stream = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Tell the story in second person."}],
    stream=True,
)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)

Handle stream termination and errors carefully. If the connection drops, you may need to retry or reconstruct the response. The API does not support partial refunds for streamed tokens that were successfully generated.

Limits & Quotas

The API enforces strict limits to ensure stability. You are allowed 300 requests per minute per key and 8 concurrent requests. The maximum request body size is 8 MB. If you send an invalid key, you receive a 401 error. A 402 error indicates insufficient prepaid credit. A 429 error means you have exceeded the rate limit. Prepaid credit never expires, and errors are free. Use the GET /v1/models endpoint to verify model availability.

Questions and answers

Does the uncensored model refuse all content?

The model is tuned to answer without refusals for lawful adult, fictional, or controversial topics. However, a hard content limit always applies: requests involving sexual content with minors are refused. This limit is enforced regardless of the model's general uncensored nature.

How does billing work for the unrestricted llm API?

You pay per token usage: $0.25 per 1M input tokens and $1.00 per 1M output tokens. Errors and refusals are free. You can top up with crypto (USDT on TRC20 or USDC on Base) in amounts from $10 to $500. Credit never expires, and no monthly fees apply.

Is this an official OpenAI service?

No. This is an independent service using an open-weight model tuned for unrestricted responses. It is compatible with the OpenAI SDK format but is not GPT-4, GPT-3.5, or any other OpenAI model. Check official documentation for specific OpenAI model details.

Your key is one form away

Create an account, copy the key, change the base URL. That is the whole setup.