Drop-in uncensored LLM proxyhttps://api.llmproxyapi.com/v1

Quickstart: Integrate the LLM Proxy

Integrate our uncensored LLM proxy with a single base URL change. Drop-in OpenAI compatibility means your existing code works without modification.

Base URL & Authentication

Replace your standard OpenAI base URL with our endpoint. Authentication is handled via a Bearer token in the Authorization header. This approach works with any OpenAI-compatible client library. You do not need to change your code logic, only the configuration values. Our API mimics the standard OpenAI interface structure, ensuring that your existing integration code remains valid. This reduces integration time to minutes rather than days. You can regenerate your key at any time from your dashboard, which immediately invalidates the old key. This provides security without requiring complex token rotation logic in your application.

First Request

Send a standard chat completion request to test connectivity. The endpoint accepts JSON payloads with a model identifier and message history. Our model ID is "uncensored". This identifier tells the proxy which uncensored model to route your request to. The response structure mirrors the OpenAI format, containing the assistant's message content. You can verify the response by checking the content field in the JSON output. This confirms that your authentication and payload structure are correct. The endpoint supports both streaming and non-streaming modes for standard requests.

curl https://api.llmproxyapi.com/v1/chat/completions \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "uncensored",
    "messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
  }'

Python SDK Integration

Use the official OpenAI Python library to interact with our API. Set the api_key to your proxy key and the base_url to our endpoint. The rest of the library behavior remains identical to standard OpenAI usage. This includes support for async clients, retries, and standard error handling. You can pass the same parameters you would normally use, such as temperature and max tokens. The proxy handles the routing to the uncensored model behind the scenes. This allows you to maintain your existing Python codebase while gaining access to our model's capabilities.

from openai import OpenAI

client = OpenAI(base_url="https://api.llmproxyapi.com/v1", api_key="YOUR_KEY")

resp = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)

Node SDK Integration

The Node.js SDK works with our API by overriding the base URL. Initialize the client with your proxy API key and the correct endpoint. All other methods, such as chat.completions.create, function as expected. This includes handling standard response objects and error types. You can use the same message structures and parameters as you would with OpenAI. The proxy ensures that the request is processed by the uncensored model. This compatibility allows developers to swap providers quickly without rewriting their client logic. It is ideal for microservices that already use the OpenAI SDK pattern.

import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://api.llmproxyapi.com/v1", apiKey: process.env.API_KEY });

const resp = await client.chat.completions.create({
  model: "uncensored",
  messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);

Streaming Responses

Enable streaming by setting the stream parameter to true in your request. The API returns a Server-Sent Events (SSE) stream. Each chunk contains partial token updates, allowing for real-time display of the response. This is critical for chat interfaces and applications requiring immediate feedback. The stream ends with a final chunk containing the full response metadata. You can parse the chunks using standard SSE libraries in your preferred language. This feature works seamlessly with the OpenAI SDKs, which handle the parsing logic automatically. It provides a smooth user experience without waiting for the entire response to generate.

stream = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Tell the story in second person."}],
    stream=True,
)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)

Rate Limits & Constraints

Our API enforces a limit of 300 requests per minute per key. If you exceed this, you will receive a 429 status code. Ensure you implement retry logic with exponential backoff for these cases. The request body size is capped at 8 MB. This constraint applies to the total JSON payload, including prompt and context. Pricing is usage-based: $0.25 per million input tokens and $1.00 per million output tokens. Credits are prepaid and do not expire. If your balance is insufficient, requests will return a 402 error. Authentication errors return a 401 status. There is no complex routing or vendor selection; the model is fixed.

Under the hood: specs

A quick checklist for developers: format, limits, features, billing.

SpecValue
ProtocolOpenAI Chat Completions schema; official openai SDKs work unchanged
MethodsPOST /v1/chat/completions · GET /v1/models
Modeluncensored
AuthenticationBearer token in the Authorization header
Base URLhttps://api.llmproxyapi.com/v1
JSON modeJSON object mode via response_format json_object
Context window100,000 tokens (prompt + completion together)
Other parameterstemperature, top_p, stop, seed, presence_penalty, frequency_penalty
Tools / tool callsSupported: tools + tool_choice, tool_calls in the reply (streamed too), tool results as role: tool messages
Max outputup to 16,000 tokens per request (default 2,048)
SSE streamingSupported (stream: true), usage included at the end
Parallel requests8 requests at the same time per key
Requests per minute300/min per key
Response headersX-Request-Id, X-Balance-USD, X-RateLimit-Limit-Requests, X-RateLimit-Limit-Concurrency
Request size8 MB request body
Top-upUSDT (TRC20) or USDC (Base), any whole amount from $10 to $500
Volume bonus+5% on $50+, +10% on $100+
Credit expiryno monthly fee; paid credit does not expire
Price$0.25 per 1M input tokens · $1.00 per 1M output tokens
How you payprepaid credit, charged by real token usage; errors and refusals are free
Trial credit$0.50 of credit valid 7 days, no card needed
Contentuncensored for adults; the only hard rule: no sexual content involving minors
Accountsign in with Google or with e-mail + password
Keysone key per account, regenerate any time (the old one stops working)

Errors and what to do

Errors come back as JSON with a stable type; failed and refused requests are not billed.

StatusTypeReason
400bad_requestmalformed request or too long for the context window
401missing_key · invalid_key · key_revokedcheck the Authorization header or use your current key
402no_creditbalance is empty — top up, requests resume at once
403content_blockedsexual content involving minors — refused, not billed
404not_foundonly /v1/chat/completions and /v1/models exist
413request_too_largebody over 8 MB
429rate_limited · concurrencyover 300/min or 8 parallel — back off and retry
503upstream_busymodel busy — retry in a few seconds

Questions and answers

Do you offer embeddings or image generation?

No, we do not offer embeddings, image, audio, or video generation. Our API is strictly for text chat completions. If you need embeddings, you must use a separate service or model. We focus on high-performance text inference with the uncensored model.

Is this an official OpenAI service?

No, we are an independent service. We are compatible with the OpenAI API format, but we host our own uncensored model. It is not GPT-4, GPT-3.5, or any other OpenAI model. You can use the same SDKs, but the underlying model is different.

What happens if I exceed my credit?

Requests will return a 402 Payment Required error. You can top up your account at any time using crypto (USDT or USDC). Credits never expire, so you can add funds when convenient. There is no monthly subscription fee, only pay-as-you-go usage.

Your key is one form away

Create an account, copy the key, change the base URL. That is the whole setup.

Get API key