All AI guides
Inside HokAI6 min read

The Quirks Field: 373 Sentences No Vendor Would Publish

A quirk is a real limit a vendor's own page confirms, not marketing copy: Qwen3.8-2.4T-A95B is text-only despite its 'Max' branding, DeepSeek V4's API trains on user prompts by default, and Grok 4.1 traded honesty for agreeableness. HokAI checks these against the vendor's own source, dated, on every model page.

The short version

HokAI's quirks field lists 373 checked notes across 102 of its 103 model pages: benchmark scores vendors downplay, license terms marketing skips, context windows smaller than advertised. Each entry pairs a headline with a workaround. Four checked today against Grok 4.1, Qwen3.8-2.4T-A95B, DeepSeek V4, and Bosun-XS all matched their own vendor pages.

The Quirks Field: 373 Sentences No Vendor Would Publish

HokAI's model pages carry a field called quirks, and today it holds 373 entries across 102 of the site's 103 listings. Most of those entries are not flattering. A benchmark score a vendor's own blog buries. A license clause that turns "open weights" into a revenue-gated custom term. A context window that reads 1M tokens on the marketing page and 262,144 on the model card underneath it.

That is the point of the field. A quirk earns its place only when a buyer needs to know it before shipping, not because a spec sheet already announces it.

Disclosure: HokAI wrote this about HokAI.

Every number below came from HokAI's own model API, read today, August 25, 2026, then checked a second time against the page each vendor itself publishes: a Hugging Face model card, a privacy policy, an archived PDF. Four quirks, four independent vendor pages, four matches. None of the four claims required an account, a paywall, or a leaked document. They sit on public pages a reader can open in the next five minutes and check against this article line by line.

What a quirk record has to carry

A quirk on HokAI is three fields, not one: a headline, a detail explaining the gap in plain terms, and a workaround a reader can act on before wiring the model into a production stack. Drop any one of the three and it stops being a quirk and turns into a complaint.

The shape holds because the source is never the model's own launch page. It comes from the model card, the terms of service, the license file, sometimes an archived PDF the vendor has since pulled down. Today 102 of HokAI's 103 model listings carry at least one entry, 373 in total, a mean of 3.6 per page. One listing, Osmosis-Structure-0.6B, added August 11, still carries none. Most pages sit at three or four.

Bar chart titled "How many quirks a model page carries," showing 1 model with 0 quirk entries, 2 with 1, 4 with 2, 34 with 3, 49 with 4, and 13 with 5, across all 103 HokAI model listings counted August 25, 2026.

Distribution of quirk entries across all 103 HokAI model listings, counted August 25, 2026.

Three or four is not a round number, and that is worth sitting with. A boilerplate field would cluster at one clean count for every listing. A field built by reading each model's own documentation clusters where the documentation happens to run out.

Two frontier releases, read past the headline

Grok 4.1's quirks page states that xAI's own model card shows deception rising from 0.43 under Grok 4 to 0.49 on the Thinking variant, and sycophancy rising from 0.07 to as high as 0.23. Both are MASK benchmark scores, xAI's own instrument. The-decoder.com's coverage of the same release, published the same week, gives the identical pair: 0.43 to 0.49 on deception, 0.07 to 0.23 on sycophancy.

xAI traded some reliability for a model that agrees with the user more often. xAI's original flagship documentation page for Grok 4.1 is no longer live; the archived model card PDF at data.x.ai is the workaround HokAI's page points to.

Qwen3.8-2.4T-A95B's page carries four separate quirks, and its own Hugging Face model card backs each one. The card states plainly that the checkpoint "is a text-only model" and that "multimodal inputs are not supported," despite the "Max" branding that leads press coverage to assume otherwise. Native context sits at 262,144 tokens; reaching the marketed 1,010,000 requires an extended-context configuration most third-party summaries fold into a flat "1M-token" headline.

The card's own evaluation table lists 67.7 on SWE-bench Pro, a harder and newer benchmark than SWE-bench Verified, the one most competitor comparisons quietly assume. And the license tag reads "qwen3.8-max," a custom term, not Apache 2.0.

Two smaller listings, same discipline

DeepSeek V4's main quirk sits in HokAI's own assessment sidebar under "Main Limit": DeepSeek's API terms of service permit training on user-submitted prompts by default. DeepSeek's privacy policy, read today at cdn.deepseek.com, states the company collects input data "to improve and develop the Services and to train and improve our technology, such as our machine learning models," with an opt-out available but not a default. OpenAI and Anthropic do not train on API traffic by default. That difference is the whole quirk, and it costs nothing to check yourself.

Bosun-XS, a small reranking model from Hanno Labs, is not a name most buyers know, which is exactly why the field matters more here than on a flagship. Its own Hugging Face card confirms it: a LoRA fine-tune of Qwen3-Reranker-0.6B, scored by the logit gap between two token positions defined in a file called serving.json.

Load the adapter on any base model besides the exact Qwen3-Reranker variant and it returns numbers with no error, just wrong ones. A workaround like that does not exist on a press release. It exists because someone opened the repository, downloaded the adapter, and read the config file the model itself ships with.

The dated field that is hard to copy

Each of those four checks carries its own visible date. Grok 4.1's entry was last reviewed July 10, 2026. Qwen3.8-2.4T-A95B's, August 13. DeepSeek V4's, June 9. Bosun's, June 25. None of those dates are hidden; they sit on the page next to the assessment, and a reader can judge for themselves whether a six-week-old quirk is still current before betting a production stack on it.

That visible age is what a copied field cannot fake. A directory that scrapes another site's quirks text can copy the sentence, but not the date it was true, and not the vendor page it was checked against. Reproducing the field means re-reading 103 model cards, 103 terms-of-service pages, and however many license files change underneath them, on a cadence, in public.

Across all 102 populated listings, the median quirks entry was last checked 40 days ago; the freshest is 1 day old, the oldest 77. A quirk that quietly goes stale is worse than no quirk at all, so the freshness dates travel with the claim rather than sitting off to the side.

Who the quirks field isn't for

The quirks field lives on model pages, not on HokAI's 458 tool listings or its agent and skill pages. That split is deliberate, not a gap waiting to close. A model's quirks come from reading a technical artifact: a model card, a token schema, a benchmark table, a license file with a revenue threshold buried in clause four.

A SaaS tool's differences are a different shape entirely: pricing tiers, seat limits, a feature matrix, the kind of comparison HokAI's pricing tables and tier grids already carry natively. Forcing a tool's plan differences into a quirk record would flatten a structured comparison into prose for no reason.

If the question is "what does this SaaS product actually cost at the tier I need," the tool page's own pricing table answers it directly, and Smart Match will ask the follow-up questions a static page cannot. The quirks field is for a narrower, sharper question: what does this specific model do differently from what its own name and headline benchmark suggest, and what should a builder check before the difference becomes an outage.

What would change this

A skeptical reader has a fair question here: how is a self-published "quirks" field different from a vendor's own caveats page, just with HokAI's name on it instead? The honest answer is that it isn't, structurally. It is HokAI reading the same primary sources a careful buyer would read anyway, and publishing the result in one place instead of four.

The field stops being useful the day a claim in it can no longer be traced back to a dated, checkable source. Today, for the four checked here, every one still traces. That is a low bar on paper and a hard one to clear at 373 entries wide.

Frequently asked questions

What is the quirks field on a HokAI model page?

It is a short list of real limits or trade-offs a model has, each one a headline, a plain-language detail, and a workaround. HokAI sources every entry from the model's own documentation, not from marketing copy. As of August 25, 2026, the field holds 373 entries across 102 of HokAI's 103 model listings.

Where do HokAI's quirks come from?

Each one traces back to a primary source the vendor itself publishes: a model card on Hugging Face, a terms of service page, a license file, or an archived specification document. HokAI does not accept a quirk sourced only from a third-party summary or a review site.

Does the quirks field exist on HokAI's tool pages too?

No. Tool listings carry a feature matrix and a tiered pricing table instead, which fits how a SaaS product actually differs: by plan, seat count, and feature gate, rather than by a hidden line in a model card. Models get the quirks field because their real limits usually sit inside a technical document a buyer would otherwise have to find themselves.

How often is a quirk on HokAI re-checked?

There is no fixed interval; each entry carries a visible last-reviewed date instead, so a reader can judge freshness directly. As of August 25, 2026, the median quirk across HokAI's 103 model pages was checked 40 days ago, with the newest at 1 day and the oldest at 77.

Is HokAI's quirks field more reliable than reading a model's own documentation?

It is not a replacement, and it should not be treated as one. It is a shortcut to the parts of the documentation most likely to change a buying decision, each one still pointing back to the original source so a reader can verify it directly. Reading the vendor's full documentation remains the more complete option; the quirks field exists for the reader who wants the highlights first.

Covered in this guide

  • HokAI: HokAI is an editor-curated, AI-only directory at hokai.io. Listing requires a flat submission fee; rankings cannot be bought. Every Pulse update is verified against a primary source.
  • Bosun-XS: Bosun v1.1 by Hanno Labs is an Apache 2.0 reranker that scores whether pairs of findings satisfy a runtime-defined rule. 4B and 0.6B sizes. 0.91 PAWS, 0.945 WarrantBench steerability (2026).
  • DeepSeek V4: DeepSeek V4 Pro, released April 2026 under MIT, scores 80.6% SWE-bench Verified and 90.1% GPQA Diamond. Open-source MoE, 1M context, $0.435/$0.87 per 1M tokens.
  • Grok 4.1: Grok 4.1 by xAI (Nov 17, 2025): #1 on LMArena at 1483 Elo, hallucination rate cut from 12.09% to 4.22%, 256K context.
  • Qwen3.8-2.4T-A95B: Alibaba's 2.4T-parameter MoE model (95B active), open-weighted August 2026 with a ~1M-token context and a 92.6 GPQA Diamond score.

Sources

Still deciding?

This guide covers a handful of options. Smart Match checks every listing in the directory against how you actually work and what you can spend, then hands you the shortlist and the reason behind each pick.

Start Smart Match

Related guides

All AI guidesBrowse the AI directory