PacketStream documentation

Read any public web page as Markdown

Send a URL to the PacketStream Fetch API and get the page's main content as Markdown, with its final URL, HTTP status, and title, billed per page that returns content.

The Fetch API reads one public web page and returns it as Markdown. Send a URL to POST /v1/fetch. The API downloads the page live through the PacketStream residential network, follows up to five redirects, and returns the page’s main content with its final URL, HTTP status, title, and description.

Fetch reads static HTML and common text formats. It does not run JavaScript.

Before you begin

You need:

Load the key into your shell without writing it into the command history:

read -r -s -p 'PacketStream API key: ' PACKETSTREAM_API_KEY
printf '\n'
export PACKETSTREAM_API_KEY

Make your first request

curl -s https://api.packetstream.io/v1/fetch \
  -H "Authorization: Bearer $PACKETSTREAM_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://go.dev/doc/effective_go", "max_length": 20000}'

The response, with the page text shortened:

{
  "request_id": "0b9e6c1d-2f4a-4b8c-9d3e-5f6a7b8c9d0e",
  "fetch": {
    "url": "https://go.dev/doc/effective_go",
    "content": "main",
    "include_links": true,
    "country": "us"
  },
  "page": {
    "final_url": "https://go.dev/doc/effective_go",
    "status_code": 200,
    "title": "Effective Go - The Go Programming Language",
    "description": "",
    "content_type": "text/markdown",
    "source_content_type": "text/html",
    "fetched_at": "2026-10-07T12:00:00Z",
    "content": "# Effective Go\n\n## Introduction\n\nGo is a new language. …",
    "start_index": 0,
    "returned_length": 20000,
    "total_length": 104112,
    "truncated": true,
    "next_start_index": 20000,
    "source_truncated": false
  },
  "charged_units": 1,
  "balance_remaining_usd": 12.3475,
  "processing_time_ms": 910
}

page.content holds the Markdown. When page.truncated is true, send the same request with start_index set to page.next_start_index to read the next part. The reference describes every field.

Python

import os
import requests

response = requests.post(
    "https://api.packetstream.io/v1/fetch",
    headers={"Authorization": f"Bearer {os.environ['PACKETSTREAM_API_KEY']}"},
    json={"url": "https://go.dev/doc/effective_go", "max_length": 20000},
    timeout=35,
)
response.raise_for_status()
page = response.json()["page"]
print(page["title"], page["final_url"])
print(page["content"])
if page["truncated"]:
    print("Continue with start_index", page["next_start_index"])

JavaScript

Run this on a server, never in a browser, so the key stays private.

const response = await fetch("https://api.packetstream.io/v1/fetch", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.PACKETSTREAM_API_KEY}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({ url: "https://go.dev/doc/effective_go" }),
  signal: AbortSignal.timeout(35_000),
});
const payload = await response.json();
if (!response.ok) throw new Error(`${payload.error.code}: ${payload.error.message}`);
console.log(payload.page.title, payload.page.content);

Pricing

  • A fetch that returns page content costs $0.0025, the same as one search. charged_units is always 1.
  • Every error is free, including the site’s own errors, timeouts, and rate limits. One rare exception is described under billing.
  • A request needs at least $0.0025 in your balance before it starts.
  • Each part of a long page that you read with start_index is a separate fetch and is billed separately.

Good to know

  • Static pages only. Fetch does not run JavaScript, so many single-page apps return javascript_required. PDFs and images return unsupported_content_type.
  • You choose what to fetch. Fetch does not read robots.txt. You need any permission that the law, a contract, or the site operator requires. See robots.txt and acceptable use.
  • Public URLs only. Only public http and https URLs on ports 80 and 443 are allowed. IP addresses, private or internal hosts, URLs with a user name or password, and URLs that contain an API key are refused.
  • Long pages. A response holds up to 50,000 characters by default, or up to 500,000 with max_length. Each further part is downloaded and billed again.
  • Exit and User-Agent. Requests leave from the United States unless you set country. Sites see a current desktop Chrome User-Agent unless you set user_agent.
  • Time. The server gives up after 25 seconds. Set your client timeout to at least 35 seconds.
  • Privacy. The API records each request’s ID, outcome, latency, and charge. It does not store the URL, your user_agent, or the page.
Questions? We’re here to help. Talk with our support team about your integration.
Contact support