Compare LLM APIDocs
LLM API Quickstart: OpenAI Compatible Endpoints
Start generating text with our uncensored LLM API in minutes. This quickstart covers the OpenAI-compatible endpoints, SDK setup, and streaming options you need to integrate raw model outputs into your application.
Authentication and Base URL
To begin using the llm api provider services, you need an API key. Sign up on the Get API key page with just an email and password. Your key appears immediately. No phone number or credit card is required for the trial credit. Keep your key secure; you can regenerate it at any time, which invalidates the previous key.
Configure your client to point to our base URL. This ensures all requests are routed to our uncensored model. The endpoint structure mirrors the standard OpenAI chat completions interface.
- Base URL:
https://api.comparellmapi.com/v1 - Header:
Authorization: Bearer YOUR_API_KEY
This setup works with any OpenAI-compatible SDK. You only need to adjust the base URL and inject your key.
Your First Request
Send a simple text completion request to test connectivity. This confirms your API key is valid and credits are available. Use the model ID uncensored to access the raw model. The endpoint accepts POST requests with a JSON body containing messages.
Ensure your request body stays under 8 MB. If you exceed this limit, the server will reject the request. For most text-based applications, a single turn or short conversation fits well within this constraint.
curl https://api.comparellmapi.com/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "uncensored",
"messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
}'If successful, you receive a JSON response with the generated text. This confirms the llm api pricing structure applies to your usage. Each token counts toward your prepaid balance.
Python SDK Integration
Use the official OpenAI Python library for the easiest integration. Install the package via pip, then configure the client with our base URL. This allows you to use familiar methods like chat.completions.create().
Set the model to uncensored. Pass your messages in the standard format. The SDK handles serialization and error parsing automatically.
from openai import OpenAI
client = OpenAI(base_url="https://api.comparellmapi.com/v1", api_key="YOUR_KEY")
resp = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)This approach is ideal for backend services or scripts. The library manages retries and basic error handling. Remember to check the response content for the generated text. You can also access token usage data for billing transparency.
Node.js SDK Setup
For JavaScript or TypeScript environments, the OpenAI Node SDK works seamlessly. Install the package and initialize the client with our base URL. This maintains compatibility with the standard API interface.
Define your system and user messages. The model ID remains uncensored. The SDK returns a promise with the completion result. Handle errors appropriately to manage API limits or insufficient credits.
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://api.comparellmapi.com/v1", apiKey: process.env.API_KEY });
const resp = await client.chat.completions.create({
model: "uncensored",
messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);This method is suitable for web servers or Node.js applications. The code structure mirrors the Python example, ensuring consistency across languages. You can integrate this into Express routes or serverless functions easily.
Streaming Responses (SSE)
For real-time applications, enable streaming by setting the stream parameter to true. The server sends a series of Server-Sent Events (SSE). Each event contains a partial chunk of the response.
This reduces perceived latency for users. The client receives tokens as they are generated. This is useful for chat interfaces or live text generation.
stream = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Tell the story in second person."}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)Parse the SSE stream to extract each token. Accumulate the text to form the complete response. The chatgpt api cost and other providers often charge similarly for streaming, but our uncensored model offers raw output without extra fees for this mode. Monitor the stream end event to finalize the response.
Limits, Errors, and Context
Understand the operational limits to ensure smooth integration. You are allowed 300 requests per minute per API key. If you exceed this, you receive a 429 rate limit error. Wait before retrying.
Common errors include 401 for invalid keys and 402 for insufficient credits. Ensure your prepaid balance is sufficient. The context window supports up to 100,000 tokens for prompt plus completion. Manage your token usage carefully to avoid truncation.
For detailed llm api pricing, refer to the pricing page. Each token is charged according to the input and output rates. No hidden fees apply. This transparency helps you budget accurately for your application.
Under the hood: specs
If your tool speaks the OpenAI API, these are the details that matter.
| Feature | Support |
|---|---|
| API format | OpenAI-compatible: any OpenAI SDK or client works — change the base URL and the key |
| Endpoints | POST /v1/chat/completions · GET /v1/models |
| Base URL | https://api.comparellmapi.com/v1 |
| Model ID | uncensored |
| API key | Authorization: Bearer YOUR_KEY |
| Max context | 100,000 tokens (prompt + completion together) |
| JSON mode | JSON object mode via response_format json_object |
| Other parameters | temperature, top_p, stop, seed, presence_penalty, frequency_penalty |
| Completion length | prompt + completion fit within 100,000 tokens; max_tokens optional, no separate output cap |
| Function calling | Supported: tools + tool_choice, tool_calls in the reply (streamed too), tool results as role: tool messages |
| SSE streaming | Supported (stream: true), usage included at the end |
| Concurrency | up to 8 in parallel per key |
| Response headers | X-Request-Id, X-Balance-USD, X-RateLimit-Limit-Requests, X-RateLimit-Limit-Concurrency |
| Rate limit | 300/min per key |
| Request size | 8 MB request body |
| Volume bonus | +5% on $50+, +10% on $100+ |
| Top-up | crypto: USDT on TRON or USDC on Base, $10–$500, any whole sum |
| Free trial | $0.50 of credit valid 7 days, no card needed · Trial key: 2 parallel requests, 60 req/min; full limits (8 and 300) after first top-up |
| Price | $0.25 per 1M input tokens · $1.00 per 1M output tokens |
| Credit expiry | no monthly fee; paid credit does not expire |
| Billing | pay as you go from prepaid credit; nothing is charged for failed or refused requests |
| Account | sign in with Google or with e-mail + password |
| Content policy | uncensored for adults; the only hard rule: no sexual content involving minors |
| Keys | one active key per account; a new key replaces the old one |
Error reference
Errors come back as JSON with a stable type; failed and refused requests are not billed.
| Code | Type | Meaning |
|---|---|---|
400 | bad_request | invalid JSON, empty messages, bad parameter, or prompt + max_tokens over the window — fix and resend |
401 | missing_key · invalid_key · key_revoked | check the Authorization header or use your current key |
402 | no_credit | balance is empty — top up, requests resume at once |
403 | content_blocked | sexual content involving minors — refused, not billed |
404 | not_found | only /v1/chat/completions and /v1/models exist |
413 | request_too_large | request body larger than 8 MB |
429 | rate_limited · concurrency | over 300/min or 8 parallel — back off and retry |
503 | upstream_busy | temporary overload, retry shortly |
Questions and answers
How do I reset my API key?
You can regenerate your API key at any time from your account dashboard. Generating a new key immediately revokes the old one, so update your applications to use the new key. This is useful if you suspect the key was compromised.
What happens if I run out of credits?
When your prepaid balance reaches zero, you will receive a 402 error on subsequent requests. You can top up from $10 using crypto (USDT or USDC). Bonus credits are added automatically for larger top-ups of $50 or $100.
Is the uncensored model suitable for commercial use?
Yes, the model is designed for lawful adult and creative use cases without content refusals. It is not suitable for content involving minors. The model is open-weight and run on our own servers, providing raw text output for your applications.
Your key is one form away
Create an account, copy the key, change the base URL. That is the whole setup.