Venice AI API: First request in five minutes
Get your uncensored LLM running in under five minutes with this quickstart guide for the Venice AI API. Use the OpenAI-compatible endpoints to send text, stream responses, and call functions with transparent pay-as-you-go pricing.
https://api.veniceapialternative.com/v1
Base URL and Authentication
Start by creating an account at the Get API key page. You only need an email and a password; no credit card is required to receive the $0.50 trial credit valid for seven days. Your API key is displayed immediately after signup and can be regenerated at any time, which revokes the old key.
Use the base URL https://api.veniceapialternative.com/v1 with the official OpenAI SDKs or any OpenAI-compatible client. Set your API key as the bearer token in the Authorization header. This venice ai api endpoint replaces the default OpenAI base URL in your client configuration.
First Request
Send a simple chat completion request to test the connection. The model ID is always uncensored. This open-weight model runs on our GPU servers and is tuned to answer without content refusals for lawful adult use.
Replace YOUR_API_KEY with your actual key. The response returns text output based on your prompt.
curl https://api.veniceapialternative.com/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "uncensored",
"messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
}'
Python SDK Example
Install the OpenAI Python library and configure it to use our base URL. This allows you to leverage existing codebases that expect OpenAI endpoints.
from openai import OpenAI
client = OpenAI(base_url="https://api.veniceapialternative.com/v1", api_key="YOUR_KEY")
resp = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)
Node SDK Example
For Node.js environments, use the openai package. Initialize the client with the custom base URL and your API key. This ensures all requests route through the venice ai api infrastructure.
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://api.veniceapialternative.com/v1", apiKey: process.env.API_KEY });
const resp = await client.chat.completions.create({
model: "uncensored",
messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);
Streaming Responses
Enable streaming by setting stream: true in your request. The API returns Server-Sent Events (SSE) for real-time token delivery. This is useful for chat interfaces that need to display text as it generates.
stream = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Tell the story in second person."}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
Rate Limits and Constraints
The uncensored model supports a 100,000-token context window for prompt and completion combined. You are limited to 300 requests per minute per API key. Request bodies must not exceed 8 MB.
Common errors include 401 (invalid key), 402 (insufficient credit), and 429 (rate limit exceeded). Top up prepaid credit starting at $10; credits never expire. No monthly subscription locks your usage.
Capabilities and limits
If your tool speaks the OpenAI API, these are the details that matter.
| Item | Value |
|---|---|
| Compatibility | OpenAI Chat Completions schema; official openai SDKs work unchanged |
| Endpoints | POST /v1/chat/completions · GET /v1/models |
| Model | uncensored |
| Authentication | Authorization: Bearer YOUR_KEY |
| Base URL | https://api.veniceapialternative.com/v1 |
| JSON mode | JSON object mode via response_format json_object |
| Other parameters | temperature, top_p, stop, seed, presence_penalty, frequency_penalty |
| Context window | 100,000 tokens, input and output combined |
| Function calling | Yes — tools, tool_choice; replies carry tool_calls, also when streaming; send results back as role: tool |
| Streaming | Supported (stream: true), usage included at the end |
| Completion length | up to 16,000 tokens per request (default 2,048) |
| Request size | 8 MB request body |
| Parallel requests | 8 requests at the same time per key |
| Headers | X-Request-Id, X-Balance-USD, X-RateLimit-Limit-Requests, X-RateLimit-Limit-Concurrency |
| Requests per minute | 300 requests per minute per key |
| Credit expiry | no monthly fee; paid credit does not expire |
| Billing | pay as you go from prepaid credit; nothing is charged for failed or refused requests |
| Bonus credit | +5% from $50, +10% from $100 |
| Token prices | $0.25 per 1M input tokens · $1.00 per 1M output tokens |
| Top-up | crypto: USDT on TRON or USDC on Base, $10–$500, any whole sum |
| Trial credit | $0.50 of credit valid 7 days, no card needed |
| Keys | one active key per account; a new key replaces the old one |
| Account | Google or e-mail and password |
| Content | adult content allowed; sexual content involving minors is refused |
Errors and what to do
The type field is stable, the message is for humans. Errors cost nothing.
| Status | Type | Reason |
|---|---|---|
400 | bad_request | malformed request or too long for the context window |
401 | missing_key · invalid_key · key_revoked | check the Authorization header or use your current key |
402 | no_credit | balance is empty — top up, requests resume at once |
403 | content_blocked | sexual content involving minors — refused, not billed |
404 | not_found | only /v1/chat/completions and /v1/models exist |
413 | request_too_large | request body larger than 8 MB |
429 | rate_limited · concurrency | slow down: rate or parallel limit reached |
503 | upstream_busy | model busy — retry in a few seconds |
Questions and answers
Is this the official Venice.ai API?
No. This is an independent service hosting an uncensored open-weight model. It is not affiliated with or a reseller of Venice.ai's main platform or any other vendor's models.
Does this API support image or audio generation?
No. This is a text-only chat-completion API. It does not support embeddings, image generation, audio, video, or fine-tuning.
What happens if my trial credit expires?
Your trial credit expires after seven days. You can top up with crypto (USDT or USDC) starting at $10. Prepaid credits never expire, and you can regenerate your API key at any time.
Your key is one form away
Create an account, copy the key, change the base URL. That is the whole setup.