Aion 3.5 suits character chat and long interactive fiction, where its 262K-token context keeps a story's history in view. Skip it for structured extraction, fast replies or regulated data, because reasoning cannot be turned off and, as of 3 October 2026, no independent evaluation exists.
Aion 3.5 is a reasoning-only roleplay and storytelling model that Aion Labs released on 23 September 2026. Aion Labs calls it a multi-model system built on the GLM family. It takes text in, returns text and tool calls, and has no published benchmark scores.
Where it sits
- $3.75/M$ per 1M tokensBlended price (3:1)Lower is better#48 / 76peer median $1.86/Mvendor price, checked by HokAI
Priced around the middle of the 76 GA models with a published price (rank 48), and one of 74 whose vendor states it does not train on customer data.
Ranks are against GA models on HokAI that publish the same figure; ties share a rank.
Provider: Aion Labs · Family: Aion 3
Context window: 262,144 tokens · Max output: 32,768
Input modalities: text · Output: text, tool-calls
About Aion 3.5
Aion 3.5 is a text-only reasoning model from Aion Labs, released on 23 September 2026 as the follow-up to Aion 3.0 (5 May 2026). The vendor describes it as a multi-model system built on the GLM family from Z.ai: several specialised models contribute to each response, which Aion Labs says gives fiction stronger narrative structure, more tension and conflict, and a more nuanced treatment of mature and darker themes. Aion Labs names neither the GLM release it starts from nor a parameter count, and it publishes no weights, so the architecture cannot be checked independently.
The context window is 262,144 tokens, double the 131,072 of Aion 3.0, with up to 32,768 output tokens per request. Reasoning cannot be switched off: reasoning_effort accepts exactly low, high or max (default high), and the model reasons before it answers, so both share the max_tokens budget. OpenRouter measured a median throughput of 58 tokens per second on 3 October 2026, with Aion Labs as the only provider listed there.
Evidence of quality is thin. Aion Labs publishes no benchmark scores or model card for Aion 3.5, and no independent evaluation was found on 3 October 2026, so HokAI lists no scores and does not infer any from the GLM base. The usable signal is adoption: OpenRouter's apps list for the model showed SillyTavern (357 million tokens), Janitor AI (200 million) and Chub AI (66.7 million) as the top three, with the time window unstated, which points to character-chat front ends as the real audience. Hermes Agent, an agent rather than a chat front end, ranked fifth at 49.4 million.
Aion 3.5 is served from the Aion Labs API (OpenAI-compatible chat completions and a Responses-style endpoint, with streaming and tool calls) and through OpenRouter. The privacy policy says prompt and response content is not stored or used for training; the terms say requests are forwarded to upstream AI providers and third-party hosts whose own retention is not described. Accounts are limited to adults, and no security certifications are named.
Best fit is character chat, interactive fiction and long-running roleplay, where the long window keeps a story's history in view. Poor fit is structured extraction (the vendor's API reference documents no response_format parameter), latency-sensitive replies, image or audio input and regulated data. For a longer window and lower rates, compare DeepSeek-V4-Pro-0813 from DeepSeek; for open weights from the Z.ai family it builds on, see GLM-5.2; for self-hosting, Mistral Large 3 from Mistral AI. Aion 3.5 Mini and the older Aion models run on the same API; HokAI's Aion Labs provider page lists the ones it tracks. In the directory, Aion 3.5 sits among proprietary and text-input models, in the 128k-400k context band with GA status.
Pricing
Billed per token from prepaid credits at $3.00 per 1M input, $6.00 per 1M output and $0.75 per 1M cached input tokens, the same rates as Aion 3.0; Aion 3.5 Mini is cheaper per token for the same window. OpenRouter lists identical rates. No batch discount, provisioned throughput or volume price is listed (checked 3 October 2026), and the Business plan with monthly invoicing is custom-quoted. Credits have no expiry date, but the terms let Aion Labs forfeit them after prolonged inactivity, and purchases are generally non-refundable.
What a real job costs
| Job | Input | Output | Total |
|---|---|---|---|
| Summarise a 20-page PDF | $0.090 | $0.0060 | $0.096 |
| Support reply | $0.0060 | $0.0018 | $0.0078 |
| One coding agent run | $0.600 | $0.120 | $0.720 |
Budgets: 20-page PDF = 30k in / 1k out · Support reply = 2k in / 300 out · Coding agent run = 200k in / 20k out. Computed from the vendor's per-token prices at render time; cached-input discounts are not applied.
Key Features
- 262,144-token context window: Double the previous flagship's 131,072 tokens; each request can return at most 32,768 output tokens, per the vendor's model list.
- Always-on reasoning, three levels: The reasoning_effort setting takes low, high or max and defaults to high; reasoning itself cannot be disabled, and any other value returns HTTP 400.
- Multi-model collaborative generation: The vendor says several specialised models contribute to each response, a design aimed at narrative tension; no third-party test of that claim was found.
- Cached input pricing: Cache reads are billed at a quarter of the standard input rate, which helps when a long roleplay history is resent every turn.
- OpenAI-format tool calls: Chat Completions and a Responses-style endpoint accept OpenAI-format tool definitions and stream over server-sent events.
Pros
- Fits long character chats better than its spec sheet suggests: a long window plus cheap cached input is the combination that matters for resent story history.
- Drop-in for existing OpenAI-compatible clients, so trying it in a roleplay front end takes a base URL and a key.
- Narrow, honest scope: the vendor sells it for roleplay and storytelling only, which makes it easy to judge by testing your own scenes.
Cons
- Nothing independent backs the quality claim: no benchmark, model card or third-party evaluation was found on 3 October 2026.
- Reasoning cannot be disabled, so even a short in-character reply waits on reasoning and can come back empty if max_tokens is set too low.
- The data path is partly opaque: prompts go to upstream hosts the privacy policy does not describe, and no security certifications are named.
- Per-token rates sit well above open-weight alternatives such as DeepSeek-V4-Pro-0813 on HokAI's records, with a shorter window.
Frequently Asked Questions
What does Aion 3.5 actually cost?
Aion 3.5 costs $3.00 per 1M input tokens, $6.00 per 1M output tokens and $0.75 per 1M cached input tokens, billed from prepaid credits. On HokAI's records on 3 October 2026 that is about seven times the per-token rates of DeepSeek-V4-Pro-0813 ($0.435 input, $0.87 output). Aion 3.5 Mini, with the same window, is listed at $0.70 and $1.40.
Can you use Aion 3.5 without paying?
Aion Labs offers a free tier with a daily credit allowance and no card, but it does not say which models the allowance covers. The stated limits are 15 requests per minute, 20,000 tokens per minute and 20,000 tokens per day, and that daily figure is lower than the largest output one request may ask for. Any paid top-up moves an account to tier 1 with 50 requests per minute and no daily cap.
What are Aion 3.5's closest competitors?
Choose [DeepSeek-V4-Pro-0813](/hub/models/deepseek-v4-pro-0813) when cost or context length matters more than fiction tuning. Choose [GLM-5.2](/hub/models/glm-5.2) for open weights from the same Z.ai lineage, or [Mistral Large 3](/hub/models/mistral-large-3) if you plan to self-host. Consumer apps such as [Character.ai](/hub/tools/characterai) replace the API for people who only want to chat.
How does Aion 3.5 compare to DeepSeek-V4-Pro-0813 in 2026?
On every published spec DeepSeek-V4-Pro-0813 leads: a 1,000,000-token window, a 384,000-token output cap and open weights. HokAI's record for it lists 80.6% on SWE-bench Verified, whereas Aion 3.5 publishes nothing comparable, so a coding or reasoning comparison is not possible. Aion's only argument is fiction tuning, which neither vendor quantifies, so test both on your own scenes before paying for it.
How do you set up Aion 3.5?
Create an account (adults only), generate an API key in the dashboard and point any OpenAI-compatible client at https://api.aionlabs.ai/v1 with the model id aion-labs/aion-3.5. The same id works on OpenRouter. Set the output limit high enough to cover reasoning as well as the reply, and expect a first reply to arrive after the reasoning pass.
Top Alternatives
- DeepSeek-V4-Pro-0813: Go with DeepSeek-V4-Pro-0813 for a 1,000,000-token window, MIT-licensed weights and lower per-token rates; choose Aion 3.5 only if its fiction tuning wins on your own scenes.
- GLM-5.2: Choose GLM-5.2 for open-source weights from the Z.ai family Aion 3.5 builds on; choose Aion 3.5 for a managed roleplay layer behind one endpoint.
- Mistral Large 3: Mistral Large 3 is the one to self-host under Apache 2.0; Aion 3.5 has no published weights.