Unified AI Compute Gateway

The Unified AI Compute API with 256K Context

Access 256K ultra-long context LLM inference, image synthesis (Z-Image & Qwen-Image-2.1), and video generation (Wan2.2) via a single unified API gateway.

View API Quickstart

Three Engines. One Developer API.

Consolidate your generative stack into a unified endpoint with real-time streaming, structured JSON output, and transparent metering.

💬

LLM Inference (256K Context)

Supports a 256K ultra-long context window capable of processing entire book-length texts in a single pass. Benchmark performance: first-token latency of ~1.4s on short inputs (~63s on full long contexts) and generation speed of ~70 to 86 tokens/second.

  • 256K ultra-long context window (book-scale inputs)
  • Real-time streaming responses with SSE support
  • Structured JSON output mode
🎨

Image Generation

Produce visual assets, concept art, and high-resolution illustrations powered by Z-Image and Qwen-Image-2.1 diffusion models.

  • Powered by Z-Image & Qwen-Image-2.1
  • Text-to-Image synthesis with flexible aspect ratios
  • Direct API response image delivery
🎬

Video Generation

Generate cinematic motion and short clips from text or static visual prompts powered by Wan2.2, executed asynchronously with task status polling.

  • Powered by Wan2.2 video model
  • 5-second to 10-second video segments
  • Async job dispatch with task polling

Developer-Friendly Subscriptions

Choose a monthly tier with included compute points. All tiers share access to single-instance queued API endpoints.

Lite

For indie hackers and developers prototyping with 256K context.

$ 9.99 / month
Includes 4,480 Points / mo
  • 256K Context LLM Inference
  • Z-Image & Qwen-Image-2.1 Synthesis
  • Wan2.2 Video Generation (Async Polling)
  • Single-instance serial request queue
  • Email Support

Max

20× compute points for demanding document processing and regular video generation.

$ 59.99 / month
Includes 89,600 Points / mo
  • 256K Context LLM Inference
  • Z-Image & Qwen-Image-2.1 Synthesis
  • Wan2.2 Video Generation (Async Polling)
  • Single-instance serial request queue
  • Email Support

How Points & Compute are Measured

We use a unified point system so you can balance LLM chats, image generations, and video clips using a single wallet balance.

💬
LLM (256K Context)
Per Token Metered

Input and output tokens are metered separately. Supports up to a 256K context window for ingesting full-book documents in a single request.

🖼️
Image Generation
3 ~ 9 Points / img

Generating a standard image consumes approximately 3 to 9 points depending on the model (fast tier via Z-Image at 3 pts, high quality via Qwen-Image-2.1 at 9 pts).

🎥
Wan2.2 Video Synthesis
~ 224 Points / 5s

Generating a 5-second video clip via Wan2.2 consumes approximately 224 points. 10-second generations scale proportionally. Dispatched asynchronously with polling.

Queue & Quota Policy: All workloads run via single-instance serial execution with request queueing. Unused points carry over for active subscribers.

Quickstart Guide

Drop-in replacement for OpenAI SDKs or raw cURL requests. Zero complex setup.

POST /v1/chat/completions
curl -X POST https://api.tokforge.com/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer tf_live_your_api_key_here" \
  -d '{
    "model": "tf-deepseek-v3",
    "messages": [
      {
        "role": "system",
        "content": "You are a concise technical architect assistant."
      },
      {
        "role": "user",
        "content": "Explain TokForge unified metering in one sentence."
      }
    ],
    "temperature": 0.7,
    "stream": false
  }'