by Jina AI

Jina Search Foundation API review, pricing and limits

Read any URL, search the web, embed and rerank: four HTTP endpoints for retrieval agents.

  • search
  • data
  • browser

Pick this when an agent needs page reading, search, embeddings and reranking from one vendor whose models are published research (23 papers listed). Skip it for interactive browsing where a seven-second page read is too slow, or when you must know the per-token price before you sign up.

What changed

  • · Release Initial launch (1.0) source

All changes across HokAI this week

Jina Search Foundation API is Jina AI's hosted retrieval layer for agents: r.jina.ai converts a URL to Markdown, s.jina.ai returns web search results as readable text, and api.jina.ai serves jina-embeddings and jina-reranker models over REST. One bearer key covers all four endpoints; a free key allows 500 Reader requests a minute.

Maker: Jina AI · Protocol: REST · Auth: api key

Compatible agents: Claude (Claude Code, Desktop) via mcp.jina.ai, Cursor and other MCP clients, LangChain and LlamaIndex (REST), OpenAI function calling (REST), n8n and Make HTTP nodes

Required runtime: Any HTTP client (curl, Python requests, fetch), An MCP client for the mcp.jina.ai route (optional)

About Jina Search Foundation API

Jina Search Foundation API is the hosted retrieval layer from Jina AI, the Berlin research company behind the jina-embeddings and jina-reranker model families. It is a skill rather than a product a person opens: an agent calls it over HTTP to turn a web page into clean text, to search the web, or to embed and rank documents, and gets back JSON it can reason over. The same team publishes the models it serves; the homepage lists 23 peer-reviewed papers, the most recent at EMNLP 2026 and NeurIPS 2026, so the API tends to expose a new model within weeks of the paper.

The four endpoints share one bearer key. r.jina.ai is the Reader: prepend it to any URL and it returns the page as Markdown, with an average latency of 7.9 seconds on the vendor's rate-limit page. s.jina.ai is Search: it runs a web query and returns five results already converted to the same LLM-friendly text, averaging 2.5 seconds. api.jina.ai/v1/embeddings and api.jina.ai/v1/rerank are OpenAI-style POST endpoints documented in a public OpenAPI file (version 2026.10.09 when read). Reader behaviour is steered with request headers rather than a body: X-Engine picks the browser engine, X-Return-Format the output shape, X-Target-Selector and X-Remove-Selector trim the page, X-Token-Budget caps the spend, and X-Respond-With switches on jina-ocr-v1 for scanned documents. The homepage also points agents at mcp.jina.ai as a Model Context Protocol server, which is the route for Claude, Cursor and other MCP clients; the REST endpoints suit LangChain, LlamaIndex and plain function calling.

The embeddings endpoint accepts text, images, video, audio and PDF inputs on the jina-embeddings-v5-omni models and lets the caller choose any output size from 1 to 1024 dimensions, with task flags for retrieval queries, passages, matching, clustering and classification. The reranker runs jina-reranker-v3 or v3.5, the latter scoring documents up to 8192 tokens long. Typical agent jobs: a research agent that reads the top search hits in one call per URL, a RAG pipeline that embeds at index time and reranks the top candidates at query time, and a browsing agent that uses the Buttons and Links summary to decide where to click next. Compare the output style with SearchApi, which returns the raw results page as structured JSON, and Linkup, which returns an answer with citations rather than the page; Exa and Tavily are the nearest neural-search alternatives, and Moss covers the memory side of the same stack.

Billing is by token, not by seat. Reader counts the tokens it returns, embeddings and reranking count the tokens sent in, and the free key tier is enforced by rate limit rather than by a monthly cap: 500 requests a minute on Reader, 100 on Search, 100 requests and 100,000 tokens a minute on embeddings and reranking. A paid key lifts embeddings and reranking to 500 requests and 2,000,000 tokens a minute, and a premium key to 5,000 requests and 50,000,000 tokens, with Reader at 5,000 requests a minute. Token package prices are shown only inside the logged-in dashboard, so they are not listed here. An experimental EU residency switch keeps processing inside the EU per request.

Release velocity is tied to the research calendar: jina-embeddings-v5-text (February 2026), jina-embeddings-v5-omni (May 2026), jina-reranker-v3.5 (July 2026) and jina-ocr-v1 (September 2026) all reached the API in 2026. The batch embeddings endpoints (/v1/batch/embeddings and friends) are the newest addition in the OpenAPI file. If you are choosing between a reader, a search API and a full answer engine for an agent, Smart Match asks the five questions that settle it.

Key Features

  • Header-driven Reader: Thirty-plus request headers steer one endpoint: engine, return format, CSS include and exclude selectors, token budget, cache tolerance, proxy country and locale.
  • OCR for scanned documents: X-Respond-With jina-ocr-v1 turns image-only pages and PDFs into Markdown, walked page by page with X-Page.
  • Multimodal embeddings with adjustable size: jina-embeddings-v5-omni embeds text, images, video, audio and PDFs, and the caller sets any output size from 1 to 1024 dimensions.
  • Long-document reranking: jina-reranker-v3.5 scores documents up to 8192 tokens each, with optional embeddings returned alongside the score.
  • Search that returns readable text: s.jina.ai returns the result pages already converted, so an agent skips the second fetch for the top hits.
  • MCP route: mcp.jina.ai exposes the same APIs as Model Context Protocol tools for Claude, Cursor and other MCP clients.
  • EU residency switch: An experimental per-request option keeps infrastructure and processing inside the EU.

Use Cases

  • Research agent reading search hits: Call s.jina.ai once for the query, then r.jina.ai on the two or three URLs worth reading, and hand the Markdown to the model with X-Token-Budget keeping each page under the context limit.
  • RAG index and rerank: Embed chunks with jina-embeddings-v5-text-small using task retrieval.passage at index time, embed the question with retrieval.query, then send the top 50 candidates to jina-reranker-v3.5 and keep the top 5.
  • Scanned PDF ingestion: POST a PDF to r.jina.ai with X-Respond-With set to jina-ocr-v1 and X-Page to walk the document page by page, producing Markdown tables an agent can quote.
  • Browsing agent navigation: Read a page with X-With-Links-Summary set to all so the agent gets a Buttons and Links list at the end and can pick its next URL without screenshots.

Install

curl "https://r.jina.ai/https://jina.ai/news" -H "Authorization: Bearer YOUR_JINA_KEY"

Requirements

  • A Jina AI account and API key from jina.ai (the free key raises Reader from 20 to 500 requests a minute)
  • Send the key as Authorization: Bearer on every request; Search, embeddings and rerank refuse requests without one
  • Token balance on the key; Reader bills on output tokens, embeddings and rerank on input tokens

Actions

Read URL

Fetches a web page (or an uploaded PDF or HTML file) and returns it as Markdown, text or JSON tuned for LLM input.

curl "https://r.jina.ai/https://www.example.com" \
  -H "Authorization: Bearer $JINA_API_KEY" \
  -H "Accept: application/json" \
  -H "X-Return-Format: markdown" \
  -H "X-With-Links-Summary: all" \
  -H "X-Token-Budget: 20000"
  • url (string) — required: The page to read, appended to https://r.jina.ai/ (or sent as JSON body url on POST).
  • Accept (string): application/json for a JSON object (url, title, content, timestamp); text/event-stream for stream mode on slow pages.
  • X-Engine (string): Browser engine used to fetch the page; trades quality against speed.
  • X-Return-Format (string): Level of detail in the response (markdown, html, text, screenshot and others); default pipeline is tuned for LLM input.
  • X-Timeout (number): Seconds to wait for page load.
  • X-Token-Budget (number): Maximum tokens for this request; the request fails if exceeded.
  • X-Respond-With (string): Set to jina-ocr-v1 to convert images and scanned pages to Markdown (vendor states 40 times the token cost).
  • X-Page (number): Page of a multi-page document (PDF) to process.
  • X-Target-Selector (string): CSS selectors to extract only (e.g. article, .main-content).
  • X-Wait-For-Selector (string): Wait until these elements appear before extracting.
  • X-Remove-Selector (string): CSS selectors to strip before extraction (nav, footer, #ads).
  • X-Retain-Images (string): Set to none to strip all images and save tokens.
  • X-With-Links-Summary (string): Append a Buttons and Links section (all or true) for downstream navigation.
  • X-With-Images-Summary (string): Append an Images section listing the page's visuals.
  • X-With-Generated-Alt (boolean): Caption images that lack alt text.
  • X-Retain-Links (string): Format links as OpenAI citation markers for GPT browsing tools.
  • X-No-Cache (boolean): Bypass the vendor's URL cache and fetch live.
  • X-Cache-Tolerance (number): Accept cached content younger than N seconds.
  • X-Respond-Timing (string): When to treat the page as loaded; later timings capture more dynamic content.
  • X-Set-Cookie (string): Cookie forwarded to the target for authenticated pages; disables caching.
  • X-Proxy-Url (string): Your own proxy for fetching the URL.
  • X-Proxy (string): Country code for a location-based proxy; auto or none.
  • X-Locale (string): Browser locale used to render the page.
  • X-User-Agent (string): Override the browser User-Agent string.
  • X-Referer (string): Set the HTTP Referer header.
  • X-Robots-Txt (string): Bot name to check against the site's robots.txt before fetching.
  • X-With-Iframe (boolean): Include content from embedded iframes.
  • X-With-Shadow-Dom (boolean): Include Shadow DOM content.
  • X-Base (string): Set to final to resolve relative URLs against the post-redirect URL.
  • X-Keep-Img-Data-Url (boolean): Keep inline base64 images instead of converting them to external URLs.
  • X-No-Gfm (string): Opt out of GitHub Flavored Markdown features.
  • X-Md-Heading-Style (string): Markdown heading style passed to Turndown; also X-Md-Hr, X-Md-Bullet-List-Marker, X-Md-Em-Delimiter, X-Md-Strong-Delimiter and X-Md-Link-Style.
  • DNT (boolean): Do not cache or log this request.
  • viewport (object): POST body only: browser window width and height for responsive layouts.
  • injectPageScript (string): POST body only: inline JavaScript or a script URL executed before extraction.

Search Web

Runs a web search and returns five results, each already converted to LLM-friendly text with the same JSON shape as Reader.

import os, requests

r = requests.get(
    "https://s.jina.ai/",
    params={"q": "jina reranker v3.5 max document length"},
    headers={"Authorization": f"Bearer {os.environ['JINA_API_KEY']}", "Accept": "application/json"},
    timeout=60,
)
for hit in r.json()["data"]:
    print(hit["title"], hit["url"], len(hit["content"]))
  • q (string) — required: The search query (query string on GET, JSON body on POST).
  • Accept (string): application/json returns a list of five result objects (url, title, content, timestamp).
  • X-Return-Format (string): Output format for each result's content, as for Reader.
  • X-Token-Budget (number): Maximum tokens for the whole response.
  • X-Locale (string): Locale for the search and the rendered result pages.

Create Embeddings

Converts text, images, video, audio or a PDF into fixed-length vectors with a jina-embeddings model (OpenAI-compatible POST).

curl https://api.jina.ai/v1/embeddings \
  -H "Content-Type: application/json" \
  --header "Authorization: Bearer $JINA_API_KEY" \
  -d '{
    "model": "jina-embeddings-v5-text-small",
    "task": "retrieval.passage",
    "dimensions": 512,
    "input": ["Late chunking keeps context across chunk boundaries.", "Rerankers score query and document together."]
  }'
  • model (string) — required: jina-embeddings-v5-text-nano, jina-embeddings-v5-text-small, jina-embeddings-v5-omni-small, jina-embeddings-v5-omni-nano, plus the v4, v3, v2, CLIP, ColBERT and code models listed in the OpenAPI file.
  • input (array) — required: A string, a TextDoc, ImageDoc, VideoDoc, AudioDoc or PDFDoc, or a list of them; a {content: [...]} item fuses several into one embedding. A PDF must be sent alone.
  • task (string): retrieval.query, retrieval.passage, text-matching, clustering or classification.
  • dimensions (number): Output size, 1 to 1024 on the v5 models.
  • embedding_type (string): float, base64, binary, ubinary, or a list of these.
  • normalized (boolean): L2-normalise the vectors.
  • truncate (boolean): Truncate inputs over the model's token limit instead of erroring.

Rerank Documents

Scores a list of documents against a query with jina-reranker-v3 or v3.5 and returns them sorted by relevance.

const res = await fetch("https://api.jina.ai/v1/rerank", {
  method: "POST",
  headers: { "Content-Type": "application/json", Authorization: `Bearer ${process.env.JINA_API_KEY}` },
  body: JSON.stringify({
    model: "jina-reranker-v3.5",
    query: "how are Reader requests billed",
    top_n: 3,
    documents: ["Reader counts output tokens.", "Embeddings count input tokens.", "The office is in Berlin."],
  }),
});
const { results } = await res.json();
  • model (string) — required: jina-reranker-v3 or jina-reranker-v3.5 (older v2 and v1 and ColBERT rerankers remain available).
  • query (string) — required: The search query to rank against.
  • documents (array) — required: Strings or TextDoc objects to rank.
  • top_n (number): How many results to return; all when omitted.
  • return_documents (boolean): Include the document text in each result.
  • max_doc_length (number): Tokens per document, 1 to 8192; model default is 2048 for v3 and 8192 for v3.5.
  • return_embeddings (boolean): Also return each document's embedding alongside its score.

How to Invoke

Four HTTPS endpoints under one bearer key: GET or POST r.jina.ai/<url> (Reader), GET or POST s.jina.ai/?q= (Search), POST api.jina.ai/v1/embeddings and POST api.jina.ai/v1/rerank. The vendor also lists mcp.jina.ai as an MCP server for clients that prefer tool calls over HTTP.

Pricing

Metered in tokens on a prepaid API key, no subscription (jina.ai and jina.ai/api-dashboard/rate-limit, read 11 October 2026). A free key is issued at signup and is limited by rate rather than by a monthly quota. Reader bills the tokens it returns; each Search request costs a fixed amount starting at 10,000 tokens; embeddings and reranking bill the tokens sent in. The jina-ocr-v1 option on Reader costs 40 times the normal token rate. Token package prices are shown only in the logged-in dashboard and are not reproduced here.

Strengths

  • Search averages 2.5 seconds and returns converted pages, which removes one hop from most research loops.
  • The OpenAPI file is public and versioned by date, so SDK generation and parameter checks need no login.
  • Model updates land as the papers do: four new models reached the API between February and September 2026.

Weaknesses

  • Token package prices are not published outside the logged-in dashboard, so cost per page cannot be estimated before signup.
  • Reader averages 7.9 seconds per page, which is slow for interactive agents without the cache or stream mode.
  • Search is blocked entirely without a key and capped at 100 requests a minute until the premium tier.

Frequently Asked Questions

What does the Jina Search Foundation API cost to run?

There is no subscription. You buy tokens on a prepaid key and each endpoint draws them down: Reader on the tokens it returns, Search as a fixed charge from 10,000 tokens per query, embeddings and reranking on input tokens. The free key is rate-limited rather than capped, and the package prices appear only once you are logged in, so budget after a test run rather than before.

How do I get a first Reader call working?

Create a key at jina.ai, then send a GET to r.jina.ai followed by the full target URL with the key in an Authorization header. Without a key the same call still works at 20 requests a minute, which is enough to test the output format. Add Accept: application/json once you want title, URL and content as separate fields.

Can Claude or Cursor call it without writing HTTP code?

Yes. The vendor lists mcp.jina.ai as a Model Context Protocol server, so any MCP client registers it and the Reader and Search capabilities appear as tools. Frameworks such as LangChain and LlamaIndex use the REST endpoints directly, and the embeddings endpoint follows the OpenAI request shape.

When is a different retrieval skill the better choice?

If the agent needs ranked search results with positions and knowledge panels rather than page text, SearchApi returns that as JSON. If you want a finished answer with citations in one hop, Linkup does that. Exa and Tavily are the closest neural-search alternatives when semantic recall over the open web matters more than reading individual pages.

Jina or SearchApi for an agent that researches the web?

They solve adjacent problems. SearchApi stops at the results page and bills per successful search; Jina's s.jina.ai returns five results already converted to text and r.jina.ai reads any further URL, billed in tokens. Teams that also need embeddings and reranking save a vendor with Jina; SEO and rank-tracking jobs belong with SearchApi.

Top Alternatives

  • SearchApi: Pick SearchApi when you need the results page itself (positions, knowledge graph, AI Mode); pick Jina when you want the pages behind the results already read into text.
  • Linkup: Pick Linkup for a sourced answer in one call; pick Jina when the agent should do its own reading, embedding and reranking.

More Agent Skills on HokAI

View the official Jina Search Foundation API skill page