NoFilter AI API Documentation
Get started with the NoFilter AI uncensored LLM API in minutes. Use our OpenAI-compatible endpoint for unrestricted text generation with transparent prepaid crypto billing.
Base URL and Authentication
Access the uncensored llm api by pointing your client to https://api.nofilterai.top/v1. Authentication is handled via the Authorization header using a Bearer token. Generate your key instantly on the Get API key page using Google or email; no phone number is required. Each account supports one active key at a time, where a new key replaces the old one. This setup ensures your requests reach the dedicated uncensored model without routing overhead. Keep your key secure, as it grants immediate access to your prepaid credits.
First Request
Send a standard chat completion request to test connectivity. The API accepts text input and returns text output for the uncensored model. Ensure your request includes the model ID and your system prompt. Errors and refusals do not consume credits, so you can test freely. Use the following curl command to verify your setup before building your application.
curl https://api.nofilterai.top/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "uncensored",
"messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
}'
This request targets the chat-completions endpoint directly. If successful, you receive a JSON response with the model's reply. This confirms your account is active and credits are available for further development.
Python SDK Integration
Integrate the uncensored ai api into your Python projects using the official OpenAI SDK. Configure the base URL and API key in your client initialization. This allows you to use familiar methods like chat.completions.create() with the uncensored model. The SDK handles serialization automatically, making it easy to add LLM capabilities to your backend services. Adjust parameters like temperature and top_p to control creativity.
from openai import OpenAI
client = OpenAI(base_url="https://api.nofilterai.top/v1", api_key="YOUR_KEY")
resp = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)
Remember that the context window includes both prompt and completion tokens. Monitor your usage to stay within the 64,000 token limit per request. This approach ensures compatibility with existing codebases while providing unrestricted content generation.
Node SDK Integration
For JavaScript developers, the Node SDK works seamlessly with our endpoint. Initialize the client with your base URL and key, then call the completion method. This enables rapid prototyping for web applications needing uncensored text generation. The SDK supports modern async/await patterns, making error handling straightforward. Use this setup for server-side rendering or API endpoints that require direct model access.
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://api.nofilterai.top/v1", apiKey: process.env.API_KEY });
const resp = await client.chat.completions.create({
model: "uncensored",
messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);
The response includes the generated text and token usage details. You can parse the output directly into your application logic. This integration maintains the simplicity of the OpenAI interface while delivering the specific uncensored model behavior you need.
Streaming Responses
Enable real-time output by setting stream: true in your request. The API returns data via Server-Sent Events (SSE), sending tokens as they are generated. This reduces perceived latency for users in chat applications. The final chunk contains the total token usage, which is critical for billing accuracy. Streaming is ideal for interactive experiences where immediate feedback is required.
stream = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Tell the story in second person."}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
Handle the stream events in your client to display text progressively. This method works with any SSE-compatible library. Ensure your UI can handle partial updates smoothly. Streaming does not affect pricing; you still pay for the total tokens used in the final response.
Limits, Errors, and Context
The API enforces strict limits: 300 requests per minute and 8 concurrent requests per key. The maximum request body size is 8 MB. If you exceed these limits, you receive a 429 rate limit error. Authentication failures return a 401 error, while insufficient credits trigger a 402 error. The model supports a 64,000 token context window, with a max output of 16,000 tokens. These constraints ensure stable performance. Always handle errors gracefully to maintain application reliability.
Questions and answers
How is billing calculated?
Billing is based on real token usage for input and output. Errors and refusals are free. Credits are prepaid via crypto and never expire.
Can I use multiple API keys?
No, each account supports only one active key. Generating a new key replaces the existing one immediately. This simplifies key management for small teams.
What content is blocked?
Sexual content involving minors is always refused. All other lawful adult, fictional, or controversial topics are allowed without refusal.
Your key is one form away
Create an account, copy the key, change the base URL. That is the whole setup.
Get API key