Last updated: 2026-09-25
Touchmark is a forward market for AI inference capacity: buyers lock in a fixed price today for a 1-billion-token allocation on a model like GLM, Kimi or Qwen, delivered in a future month, and draw it down through an OpenAI-compatible API.
About Touchmark
Touchmark brokers forward contracts on AI model inference capacity, built by Y Combinator S26 founders Ilia Bolgov and Roman Yanushevskyi. Bolgov previously worked on Revolut's trading and wealth product team after a mathematics and finance background from Imperial College London; Yanushevskyi is a 2022 International Olympiad in Informatics gold medalist who did quantitative research internships at Citadel Securities and built AI features at Lovable before Y Combinator. The company launched its marketplace to buyers and sellers in August 2026, positioning it against enterprise AI inference spend that roughly doubled to about $1.2 million per organization in 2026, most of it unbudgeted.
The current product is a genuine pivot from Touchmark's earlier concept, a quality-adjusted AI billing platform that scored outputs and adjusted what customers paid per result. Today's Touchmark instead lets a buyer lock in a fixed price now for a block of inference tokens delivered in a future month, on models from families such as GLM, Kimi and Qwen run by third-party providers, on a discount curve that deepens the further out the delivery window sits. Consumption runs through an OpenAI-compatible API, so an existing client only needs a new endpoint, not a rewrite. Buyers who over-provision can relist unused capacity on the order book before drawing on it, and a request-for-quote desk lets a buyer ask for volume, latency or throughput terms the public order book doesn't list. A separate benchmark-resolving contract type settles at delivery into whichever model tops a named public leaderboard, something the vendor's own listings illustrate with a proprietary model like Claude Opus 5 sitting alongside the open-weight supply as one possible resolution target.
Touchmark sits inside a wider shift toward treating AI compute as a purchasable asset rather than a metered bill. Cast.ai vs Together AI covers the older distinction between renting raw GPUs and paying a managed inference bill, and gateway comparisons like AIML API vs OpenRouter cover the pay-as-you-go end of that same market. On the supply side, reserved and on-demand GPU capacity is also available through providers such as Baseten, AtlasCloud, Crusoe Cloud, Pipeshift, Modal, Fireworks AI and Replicate; none of them offer Touchmark's forward-pricing mechanism.
Where this doesn't fit: Touchmark sells nothing to a developer trying a model for a weekend project, or a team unsure how much inference it will burn next month. Purchases can't be reversed once made, capacity left over at the end of the delivery window is gone, and the company discloses no minimum contract size, no security or compliance documentation, and no support channel beyond a FAQ page. A buyer who wants pay-as-you-go flexibility with no forecasting risk is better served by an on-demand gateway like OpenRouter, Together AI or Groq. Touchmark's case is for a buyer who already knows its inference volume well enough in advance to commit capital against it for a real, if modest, discount.
Screenshots


Pricing
Touchmark has no subscription tiers or free tier. Every purchase is a forward contract for a fixed 1-billion-token allocation on one model, delivered in one future month, priced on a curve that caps at 20% below the on-demand list rate and deepens the further out delivery sits. Sales are final at execution and unused tokens expire when the delivery window closes.
Key Features
- Forward Contracts, Not Metered Billing: Each contract locks a price today for a fixed 1-billion-token allocation on one named model, delivered in a chosen future month.
- Discount Curve Up To 20% Off: The published curve discounts up to 20% below the on-demand list rate, deepening the further out the delivery month sits.
- OpenAI-Compatible Router: Consumption runs through an OpenAI-compatible endpoint, so an existing client keeps its code and only swaps in a new base URL.
- Resell Unused Capacity: A contract that hasn't been drawn on yet can be relisted on the order book before its delivery window closes instead of expiring unused.
- Request-for-Quote Desk: Buyers who need non-standard volume, latency or throughput terms outside the public order book can submit an RFQ that providers compete to fill.
- Benchmark-Resolving Contracts: A separate contract type pays out on whichever model leads a named public benchmark at delivery, protecting a buyer if today's top pick falls behind before maturity.
Pros
- A published discount that grows the earlier you commit ahead of the delivery month, for buyers who can forecast their own volume.
- No new SDK to learn: the consumption API already speaks the OpenAI request format most AI integrations use today.
- Resale and RFQ mechanisms give a buyer two ways to recover from over-provisioning or find non-standard terms, rather than eating a bad estimate.
Cons
- A wrong volume forecast is a sunk cost: purchases can't be reversed once made, and any capacity left over disappears when the delivery window closes.
- Two-person team, no published docs site, no visible compliance or security page, and no support channel beyond a FAQ.
- The product itself already pivoted once in 2026, from quality-adjusted billing to this forward market; a multi-month contract is also a bet on the company's direction staying put.
Data Handling
- Training-data policy
- Touchmark does not read, log or store the content of routed prompts or responses, and its privacy policy states it does not use Buyer traffic content to train, fine-tune or improve AI models; providers process routed traffic on a zero-data-retention basis under their own terms.
- Data retention
- Zero retention
Frequently Asked Questions
What are Touchmark's pricing plans in 2026?
Touchmark has no subscription tiers. Every purchase is a single contract: a fixed 1-billion-token block on one named model, delivered in one chosen month, priced by how far ahead you commit, with earlier bookings saving more. Input tokens and cache reads meter at each listing's own conversion ratio, and once a contract executes it can't be cancelled.
Is Touchmark free to use?
No. Touchmark has no free tier or trial; every contract requires funding an account balance or paying by card or wire at purchase. The FAQ and order book are viewable without an account, but consuming inference capacity requires buying a contract first.
What are the best alternatives to Touchmark?
For pay-as-you-go access to the same class of models without locking in a future delivery window, OpenRouter and Together AI both offer an OpenAI-compatible endpoint with on-demand pricing. Groq is the better pick if the priority is low-latency serving of open-weight models rather than advance price discovery.
Touchmark or OpenRouter: which should you pick?
Pick OpenRouter when usage is unpredictable: no commitment, no minimum, pay only for what gets used that month across many providers. Pick Touchmark for the opposite case, a known and sizeable recurring inference bill worth committing to months ahead, in return for a lower fixed price and the risk that leftover capacity goes to waste.
What does it take to start using Touchmark?
Fund an account balance or arrange direct payment, then buy a listed contract at the live best ask or submit a request for quote for non-standard volume or delivery terms. Generate an API key, point an existing OpenAI-compatible client at Touchmark's router, and the purchased allocation draws down automatically as the app makes calls.
Top Alternatives
- OpenRouter: OpenRouter routes pay-as-you-go calls across dozens of providers with no minimum commitment. Touchmark only pays off if you can forecast enough volume to lock a price months ahead and accept that unused tokens expire.
- Together AI: Together AI serves and fine-tunes open-weight models on demand with no forward commitment. Touchmark trades that flexibility for a published discount below list price, in exchange for committing to a delivery month in advance.
- Groq: Groq's edge is inference speed: low-latency serving of open-weight models for workloads where response time matters most. Touchmark doesn't compete on latency; it competes on knowing your price months before you draw on the capacity.
HokAI guides covering Touchmark
- ReasonBlocks vs Touchmark: Which One Actually Fixes Your AI Bill?: ReasonBlocks cuts wasted AI agent tokens by up to 52%. Touchmark, after its August 2026 pivot, sells forward contracts locking in token prices instead.
- Archal vs ReasonBlocks: You're Not Choosing Between Them, You're Choosing When: Archal tests AI agents in a sandbox before they ship. ReasonBlocks fixes them mid-run after they're live. Real pricing, real gaps, and which to buy first.
- The Three Jobs Hiding Inside "AI Agent Observability" in 2026: AI agent observability tools split into three jobs in 2026: tracing, mid-run intervention, and outcome-based billing. Here is how to tell which one you need.