Choose Olostep when one agent needs scraping, crawling and search behind a single key and you want the know-how installed as skill files, not just tools. Look elsewhere if you need a driveable browser session or published per-minute rate limits; the async batch and crawl jobs also need a webhook or polling.
Olostep is a web data API for AI agents from Olostep Technologies: REST endpoints for scraping, batch scraping, crawling, URL mapping, web search and researched answers, plus an npm CLI and a hosted Model Context Protocol server. Scrapes return html, markdown, text, json, raw_pdf or screenshot, rendered with JavaScript over residential IPs.
Maker: Olostep Technologies, Inc. · Protocol: MCP · Auth: api key
Compatible agents: Claude Code (MCP + skill files), Cursor, Windsurf, VS Code agents, Kilo, Any MCP client (hosted endpoint or stdio), LangChain and LlamaIndex via REST or SDK
Required runtime: Node.js for olostep-cli and the npm SDK (npm install olostep), Python for the pip SDK (pip install olostep), Any HTTP client for the REST API
About Olostep
Olostep is a web data layer that agents call when they need the content of a page, the pages of a whole site, or an answer drawn from the open web. It is made by Olostep Technologies, Inc., a team based in Italy and the United States, and it ships in three forms that share one account: a REST API at api.olostep.com, an npm CLI (olostep-cli) that installs 13 skill files into Claude Code, Cursor, Windsurf, VS Code and Kilo, and a hosted Model Context Protocol server wired in by the same CLI. The vendor publishes an agent-onboarding SKILL.md that tells a coding agent how to authorise itself through a browser hand-off and start calling the API without a human copying keys.
The API is JSON in, JSON out, with a bearer key. Four endpoints are synchronous: /v1/scrapes fetches one URL in any of six formats (html, markdown, text, json, raw_pdf, screenshot) through JavaScript rendering and residential IPs, with a country selector, CSS removal, a Postlight transformer, an LLM extraction block and a max_age cache window; /v1/searches runs a natural-language web search and returns deduplicated links, optionally scraping each hit in the same call; /v1/answers takes a task and returns a researched answer with sources, or NOT_FOUND when the web does not say; /v1/maps lists the URLs of a domain with glob include and exclude patterns. Two are asynchronous and follow create, poll, list, retrieve: /v1/batches processes up to about 10,000 URLs in one job and /v1/crawls follows links from a start URL with depth, page and robots.txt controls, both able to call a webhook when done. A retrieve_id on every scraped item bridges to the final content, and hosted URLs take over when a result is too large to inline.
In practice an agent reaches for Scrape when it already has one URL, Batch when it has many, Map when it only needs the link list, Crawl when it needs the links and their content, Search for results to a query and Answer for a cited conclusion rather than pages. That decision table is written into the vendor's skill files, which is what makes it a skill rather than a scraping product: the know-how for when to use which tool travels with the tools. Compare it with Kernel, which gives an agent a live browser to drive, with Jina Search Foundation API, which reads and searches but has no crawl or batch job, and with SearchApi, which returns Google results pages rather than scraped content. For a sourced answer in one hop, Linkup is the nearer comparison to the Answer endpoint; Exa and Tavily are the neural-search alternatives.
Plans are monthly subscriptions metered in successful requests, with prepaid credit packs for spiky usage and an enterprise tier for hundreds of millions of credits. The free trial needs no card. Every plan renders JavaScript and uses residential IPs; concurrency and the AI-powered browser automations arrive on the paid tiers. Billing runs through Stripe from the dashboard, plans are pro-rated on switch, and the pricing page states that unused plans can be refunded on request.
The product is moving quickly: the CLI, the skill files, the hosted MCP server, the Answer endpoint with structured output, Schedules, Files and a Monitor API all appear in the current docs index. If you are weighing a scraping API against a browser agent or a search API for your stack, Smart Match asks the questions that separate them.
Key Features
- Seven endpoints, one key: Scrape, Search, Answer, Batch, Crawl, Map and Retrieve under api.olostep.com, with Files, Schedules and a Monitor API alongside.
- Skill files for coding agents: olostep init installs 13 skill files into every detected agent so the agent knows which endpoint fits which job.
- Hosted MCP server: olostep mcp install wires the hosted endpoint (or a local stdio server) into Cursor, Claude Code, Windsurf, VS Code and Kilo.
- Six output formats per scrape: html, markdown, text, json, raw_pdf and screenshot, with a Postlight transformer and CSS or class-name removal.
- Async jobs with webhooks: Batches of several thousand URLs and site crawls run as jobs that POST to your webhook when complete.
- Agent self-onboarding: A published SKILL.md walks a coding agent through a PKCE-style browser hand-off that returns an API key without copy-paste.
- Scrape caching: max_age reuses a stored scrape with identical parameters, defaulting to always fresh in the API.
Use Cases
- Competitor catalogue sync: Map the competitor's domain with include_urls /products/**, then send the list to /v1/batches with a parser so each product page comes back as JSON, with a webhook firing when the job completes.
- Deep research without a browser: Call /v1/searches with scrape_options so the top hits arrive as Markdown in one response, then pass the retrieve_ids of the two best pages to the model.
- Site-wide documentation ingest: Crawl a docs site from its start URL with max_pages and exclude_urls for changelogs, let follow_robots_txt stay true, and feed the pages to an embedding pipeline.
- Cited answers inside an agent: Post a task to /v1/answers with a json_format schema and treat NOT_FOUND as a legitimate result, so the agent reports an unknown instead of guessing.
Install
npm install -g olostep-cli && olostep init
Requirements
- An Olostep account and API key (os-...) from olostep.com, or let olostep init sign you in and install the skills and MCP server
- OLOSTEP_API_KEY in .env for app code; the agent-onboarding flow hands the key back after the user clicks Authorize in a browser
- For async jobs (batches, crawls): a publicly reachable HTTPS webhook URL, or a polling loop on the job id
Actions
Create Scrape
Fetches one URL with JavaScript rendering and returns it in the requested formats, optionally parsed to JSON or captured as a screenshot.
curl -X POST https://api.olostep.com/v1/scrapes \
-H "Authorization: Bearer $OLOSTEP_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"url_to_scrape": "https://www.olostep.com/pricing",
"formats": ["markdown", "json"],
"country": "US",
"remove_images": true,
"max_age": 3600
}'url_to_scrape(string) — required: The URL to scrape.formats(array): Any of html, markdown, text, json, raw_pdf, screenshot.wait_before_scraping(number): Milliseconds to wait before scraping starts.country(string): Residential country to load the request from (ISO 3166-1 alpha-2).remove_css_selectors(string): default, none or a JSON-stringified array of selectors to remove.remove_class_names(array): Class names to strip from the content.remove_images(boolean): Strip images from the scraped content.transformer(string): postlight (Mercury Parser, removes ads and clutter) or none.parser(object): With the json format: which built-in or custom parser extracts structured content.llm_extract(object): Schema-driven extraction by a language model.links_on_page(object): Return the absolute links found on the page.actions(array): Browser actions to run before capture.screen_size(object): screen_type desktop (1920x1080), mobile (414x896) or default (768x1024).screenshot(object): Screenshot options when screenshot is among the formats.max_age(number): Reuse a stored scrape with the same parameters if newer than this many seconds; 0 means always fresh.
Create Search
Searches the web for a query and returns a deduplicated link list, with optional scraping of every hit in the same call.
import os, requests
res = requests.post(
"https://api.olostep.com/v1/searches",
headers={"Authorization": f"Bearer {os.environ['OLOSTEP_API_KEY']}"},
json={"query": "olostep batch api webhook payload", "limit": 8, "include_domains": ["docs.olostep.com"]},
timeout=60,
)
for link in res.json()["result"]["links"]:
print(link["url"])query(string) — required: The search query in natural language.limit(number): Maximum links after deduplication.include_domains(array): Restrict results to these bare hosts.exclude_domains(array): Exclude these bare hosts.fast_mode(boolean): Direct low-latency search with the query verbatim instead of the broader default pass.scrape_options(object): Scrape each result page in the same request (formats and other scrape fields).
Create Answer
Performs web research for a task and returns an answer with sources, optionally shaped to a JSON schema.
const res = await fetch("https://api.olostep.com/v1/answers", {
method: "POST",
headers: { Authorization: `Bearer ${process.env.OLOSTEP_API_KEY}`, "Content-Type": "application/json" },
body: JSON.stringify({
task: "What concurrency does the Olostep Standard plan allow?",
json_format: { plan: "", concurrent_requests: "", source_url: "" },
}),
});
const answer = await res.json();task(string) — required: The research task to perform.json_format(object): Desired output as a JSON object with empty values, or a plain description of the fields wanted.
Create Batch
Scrapes a list of URLs (up to about 10,000) asynchronously as one job; poll the batch, list its items, then retrieve each result.
curl -X POST https://api.olostep.com/v1/batches \
--header "Authorization: Bearer $OLOSTEP_API_KEY" \
--header "Content-Type: application/json" \
--data '{
"items": [
{"custom_id": "product-123", "url": "https://example.com/product/123"},
{"custom_id": "sku-456", "url": "https://shop.example.org/items/456"}
],
"country": "DE",
"webhook": "https://hooks.example.com/olostep"
}'
# then: GET /v1/batches/{batch_id}, GET /v1/batches/{batch_id}/items, GET /v1/retrieve?retrieve_id=...items(array) — required: Objects with custom_id (your identifier), url and optional metadata.country(string): ISO 3166-1 alpha-2 country for the batch.parser(object): Parser to turn each page into structured content.links_on_page(object): Return absolute links for every page in the batch.webhook(string): Public HTTPS URL that receives a POST when the batch completes.metadata(object): Tags for the job, echoed back for internal joins.
Create Crawl
Follows links from a start URL and scrapes the pages it finds, with depth, page-count, pattern and robots.txt controls; asynchronous.
import os
from olostep import Olostep # pip install olostep
client = Olostep(api_key=os.environ["OLOSTEP_API_KEY"])
crawl = client.crawls.create(
start_url="https://docs.olostep.com",
max_pages=200,
max_depth=3,
exclude_urls=["/changelog/**"],
follow_robots_txt=True,
)
# poll client.crawls.get(crawl.id), then client.crawls.pages(crawl.id)start_url(string) — required: Where the crawl starts.max_pages(number) — required: Maximum pages to crawl.max_depth(number): Maximum link depth from the start URL.include_urls(array): Glob path patterns to include.exclude_urls(array): Glob path patterns to exclude; excludes win over includes.include_external(boolean): Also crawl first-degree external links.include_subdomain(boolean): Include subdomains.search_query(string): Find and rank links by relevance to a query.top_n(number): Crawl only the top N most relevant links per page for the search_query.timeout(number): End the crawl after this many seconds with the pages completed so far.follow_robots_txt(boolean): Respect robots.txt disallow rules.webhook(string): Public HTTPS URL POSTed when the crawl completes.scrape_options(object): What each page scrape requests (formats and other scrape fields).
Create Map
Returns the URLs of a website, filtered by glob patterns and optionally ranked by a search query, without scraping their content.
curl -X POST https://api.olostep.com/v1/maps \
-H 'Authorization: Bearer os-your-key' \
-H 'Content-Type: application/json' \
-d '{"url": "https://www.olostep.com", "include_urls": ["/blog/**"], "search_query": "mcp server", "top_n": 20}'url(string) — required: The website to map.search_query(string): Order the URL list by how well each matches a query.top_n(number): Cap the ordered list at N URLs.include_subdomain(boolean): Include subdomains of the URL.include_urls(array): Only return paths matching these globs, e.g. /blog/**.exclude_urls(array): Drop paths matching these globs; an exclude beats an include.cursor(string): Pagination cursor from a previous response when the result hit the size limit.
Retrieve Content
Pulls the final html, markdown or json for any scraped item by its retrieve_id, the bridge from scrape, batch and crawl results to content.
curl "https://api.olostep.com/v1/retrieve?retrieve_id=RETRIEVE_ID&formats=markdown" \
-H 'Authorization: Bearer os-your-key'retrieve_id(string) — required: The retrieve_id returned on a scrape, batch item or crawl page.formats(array): Which stored formats to return (html, markdown, json).
How to Invoke
Three routes to one account: REST at https://api.olostep.com (Authorization: Bearer <OLOSTEP_API_KEY>, JSON in and out); the olostep CLI (olostep scrape, map, crawl, answer, batch-scrape, scrape-get); and the hosted MCP server installed with olostep mcp install, whose tools include search_web. Scrapes, searches, answers and maps return inline; batches and crawls are create, poll, list, retrieve.
Pricing
Monthly plans metered in successful requests, no card for the trial (olostep.com/pricing, read 11 October 2026). Trial: $0, 500 requests, JavaScript rendering and residential IPs included, low rate limits. Starter: $9 a month, 5,000 requests, 150 concurrent. Standard: $99 a month, 200,000 requests, 500 concurrent. Scale: $399 a month, 1,000,000 requests, adds AI-powered browser automations. Credit packs for spiky use, valid 6 months: 10,000 credits $20, 250,000 credits $200, 2,000,000 credits $1,000. Enterprise is custom. Plans are pro-rated on switch and the page states refunds on request.
Strengths
- Residential IPs and JavaScript rendering are included on every plan, trial included, rather than sold as add-ons.
- The decision table for Scrape vs Batch vs Crawl vs Map vs Search vs Answer ships inside the skill files, so agents pick the cheap tool first.
- Credit packs valid for 6 months cover spiky workloads without a subscription.
Weaknesses
- Per-minute rate limits are not published; the trial only says 'low rate limits' and paid tiers quote concurrency.
- Batches and crawls need a polling loop or a public webhook, which adds plumbing compared with the synchronous endpoints.
- There is no olostep search CLI command; raw search on the CLI goes through answer, and only the MCP tool search_web returns plain results.
Frequently Asked Questions
What does Olostep cost once the trial runs out?
The trial gives 500 successful requests with no card. After that the Starter plan is $9 a month for 5,000 requests, Standard is $99 for 200,000, and Scale is $399 for a million with browser automations included; the vendor quotes the Standard rate as about $0.40 per thousand requests. Credit packs from $20 cover usage without a subscription.
How does a coding agent get Olostep working on its own?
The vendor publishes a SKILL.md for agents. It generates a session id and code verifier, asks the human to open one authorisation URL, polls /v1/cli-auth-status every few seconds, and receives the API key when the user clicks Authorize. For live use inside an editor, one npm install plus olostep init signs in, drops the skill files into each detected agent and configures the MCP server.
Does Olostep work with Claude, Cursor and other MCP clients?
Yes. The CLI's mcp install command targets Cursor, Claude Code, Windsurf, VS Code and Kilo with the hosted MCP endpoint, or a local stdio server with a flag. Agents without MCP use the REST API or the Node and Python SDKs, and LangChain-style frameworks call the same endpoints over HTTP.
When would another web-data skill suit an agent better?
Kernel is the better fit when the agent has to click through a session rather than fetch pages. Jina's Reader and Search are simpler for read-one-URL jobs and add embeddings and reranking. SearchApi returns Google results pages with positions and knowledge panels, which Olostep's search does not try to replicate.
Olostep or Kernel for a browsing agent?
They sit on different layers. Kernel hands the agent a cloud browser it controls action by action, which suits logins, forms and multi-step flows. Olostep turns URLs and queries into content and lets Crawl, Map and Batch do the fan-out server-side, which suits research and data collection where nobody needs to watch a screen. Many stacks use Olostep for gathering and a browser tool only for the interactive steps.
Top Alternatives
- Kernel: Pick Kernel when the agent must drive a live browser session step by step; pick Olostep when it needs pages, crawls and search results as data.
- Jina Search Foundation API: Pick Jina for read-and-search plus embeddings and reranking; pick Olostep for whole-site crawls, batch jobs and schema-shaped answers.