Quickstart: Integrate the LLM Proxy
Integrate our uncensored LLM proxy with a single base URL change. Drop-in OpenAI compatibility means your existing code works without modification.
Base URL & Authentication
Replace your standard OpenAI base URL with our endpoint. Authentication is handled via a Bearer token in the Authorization header. This approach works with any OpenAI-compatible client library. You do not need to change your code logic, only the configuration values. Our API mimics the standard OpenAI interface structure, ensuring that your existing integration code remains valid. This reduces integration time to minutes rather than days. You can regenerate your key at any time from your dashboard, which immediately invalidates the old key. This provides security without requiring complex token rotation logic in your application.
First Request
Send a standard chat completion request to test connectivity. The endpoint accepts JSON payloads with a model identifier and message history. Our model ID is "uncensored". This identifier tells the proxy which uncensored model to route your request to. The response structure mirrors the OpenAI format, containing the assistant's message content. You can verify the response by checking the content field in the JSON output. This confirms that your authentication and payload structure are correct. The endpoint supports both streaming and non-streaming modes for standard requests.
curl https://api.llmproxyapi.com/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "uncensored",
"messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
}'
Python SDK Integration
Use the official OpenAI Python library to interact with our API. Set the api_key to your proxy key and the base_url to our endpoint. The rest of the library behavior remains identical to standard OpenAI usage. This includes support for async clients, retries, and standard error handling. You can pass the same parameters you would normally use, such as temperature and max tokens. The proxy handles the routing to the uncensored model behind the scenes. This allows you to maintain your existing Python codebase while gaining access to our model's capabilities.
from openai import OpenAI
client = OpenAI(base_url="https://api.llmproxyapi.com/v1", api_key="YOUR_KEY")
resp = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)
Node SDK Integration
The Node.js SDK works with our API by overriding the base URL. Initialize the client with your proxy API key and the correct endpoint. All other methods, such as chat.completions.create, function as expected. This includes handling standard response objects and error types. You can use the same message structures and parameters as you would with OpenAI. The proxy ensures that the request is processed by the uncensored model. This compatibility allows developers to swap providers quickly without rewriting their client logic. It is ideal for microservices that already use the OpenAI SDK pattern.
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://api.llmproxyapi.com/v1", apiKey: process.env.API_KEY });
const resp = await client.chat.completions.create({
model: "uncensored",
messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);
Streaming Responses
Enable streaming by setting the stream parameter to true in your request. The API returns a Server-Sent Events (SSE) stream. Each chunk contains partial token updates, allowing for real-time display of the response. This is critical for chat interfaces and applications requiring immediate feedback. The stream ends with a final chunk containing the full response metadata. You can parse the chunks using standard SSE libraries in your preferred language. This feature works seamlessly with the OpenAI SDKs, which handle the parsing logic automatically. It provides a smooth user experience without waiting for the entire response to generate.
stream = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Tell the story in second person."}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
Rate Limits & Constraints
Our API enforces a limit of 300 requests per minute per key. If you exceed this, you will receive a 429 status code. Ensure you implement retry logic with exponential backoff for these cases. The request body size is capped at 8 MB. This constraint applies to the total JSON payload, including prompt and context. Pricing is usage-based: $0.25 per million input tokens and $1.00 per million output tokens. Credits are prepaid and do not expire. If your balance is insufficient, requests will return a 402 error. Authentication errors return a 401 status. There is no complex routing or vendor selection; the model is fixed.
Under the hood: specs
A quick checklist for developers: format, limits, features, billing.
| Spec | Value |
|---|---|
| Protocol | OpenAI Chat Completions schema; official openai SDKs work unchanged |
| Methods | POST /v1/chat/completions · GET /v1/models |
| Model | uncensored |
| Authentication | Bearer token in the Authorization header |
| Base URL | https://api.llmproxyapi.com/v1 |
| JSON mode | JSON object mode via response_format json_object |
| Context window | 100,000 tokens (prompt + completion together) |
| Other parameters | temperature, top_p, stop, seed, presence_penalty, frequency_penalty |
| Tools / tool calls | Supported: tools + tool_choice, tool_calls in the reply (streamed too), tool results as role: tool messages |
| Max output | up to 16,000 tokens per request (default 2,048) |
| SSE streaming | Supported (stream: true), usage included at the end |
| Parallel requests | 8 requests at the same time per key |
| Requests per minute | 300/min per key |
| Response headers | X-Request-Id, X-Balance-USD, X-RateLimit-Limit-Requests, X-RateLimit-Limit-Concurrency |
| Request size | 8 MB request body |
| Top-up | USDT (TRC20) or USDC (Base), any whole amount from $10 to $500 |
| Volume bonus | +5% on $50+, +10% on $100+ |
| Credit expiry | no monthly fee; paid credit does not expire |
| Price | $0.25 per 1M input tokens · $1.00 per 1M output tokens |
| How you pay | prepaid credit, charged by real token usage; errors and refusals are free |
| Trial credit | $0.50 of credit valid 7 days, no card needed |
| Content | uncensored for adults; the only hard rule: no sexual content involving minors |
| Account | sign in with Google or with e-mail + password |
| Keys | one key per account, regenerate any time (the old one stops working) |
Errors and what to do
Errors come back as JSON with a stable type; failed and refused requests are not billed.
| Status | Type | Reason |
|---|---|---|
400 | bad_request | malformed request or too long for the context window |
401 | missing_key · invalid_key · key_revoked | check the Authorization header or use your current key |
402 | no_credit | balance is empty — top up, requests resume at once |
403 | content_blocked | sexual content involving minors — refused, not billed |
404 | not_found | only /v1/chat/completions and /v1/models exist |
413 | request_too_large | body over 8 MB |
429 | rate_limited · concurrency | over 300/min or 8 parallel — back off and retry |
503 | upstream_busy | model busy — retry in a few seconds |
Questions and answers
Do you offer embeddings or image generation?
No, we do not offer embeddings, image, audio, or video generation. Our API is strictly for text chat completions. If you need embeddings, you must use a separate service or model. We focus on high-performance text inference with the uncensored model.
Is this an official OpenAI service?
No, we are an independent service. We are compatible with the OpenAI API format, but we host our own uncensored model. It is not GPT-4, GPT-3.5, or any other OpenAI model. You can use the same SDKs, but the underlying model is different.
What happens if I exceed my credit?
Requests will return a 402 Payment Required error. You can top up your account at any time using crypto (USDT or USDC). Credits never expire, so you can add funds when convenient. There is no monthly subscription fee, only pay-as-you-go usage.
Your key is one form away
Create an account, copy the key, change the base URL. That is the whole setup.