PacketStream documentation
Fetch API reference
PacketStream Fetch API request fields, response fields, long pages, URL rules, redirects, supported content types, and what the target site sees.
POST /v1/fetch takes one JSON object of at most 4,096 bytes. Fetch accepts only POST, so URLs stay in the request body and out of query strings. Other methods return 405 with Allow: POST.
Unknown, repeated, and mistyped fields return 400 invalid_request, and the message names the field. A JSON null counts as an omitted field.
Request fields
| Field | Type | Default | Description |
|---|---|---|---|
url | string | required | The page to read. Absolute http or https URL, up to 2,048 bytes, that follows the URL rules. Percent-encode spaces and control characters. |
content | string | main | main returns the page’s main content, found by a readability algorithm. If no main block with enough text is found, or the page is too large or complex, it returns the whole page. full converts the whole page. main or full, case-insensitive. |
include_links | boolean | true | true keeps Markdown links and images. false keeps only the link text and drops images, which saves tokens. |
max_length | integer | 50000 | Most characters to return. Characters are Unicode code points. 1 to 500,000. |
start_index | integer | 0 | First character to return. 0 or more. A value at or past the page’s total_length returns start_index_out_of_range. See long pages. |
country | string | us | Exit country for the request. Two-letter ISO 3166-1 code of a country the PacketStream network routes, so use gb, not uk. |
user_agent | string | Current desktop Chrome | The User-Agent sent to the site on every request, including redirects. Errors never repeat the value. 1 to 512 printable ASCII characters with no leading or trailing space. An empty string uses the default. |
Response
A successful fetch returns HTTP 200 with a JSON body. See the quick start for a full example. Every response, including errors, has an X-Request-Id header that matches request_id.
| Field | Meaning |
|---|---|
fetch | The accepted request after normalization. The URL has a lowercase scheme and host (international domain names in ASCII form), no default port, and no #fragment. country is lowercase. user_agent appears only if you set it. |
page.final_url | The URL that returned the content, after any redirects, normalized the same way. |
page.status_code | The site’s HTTP status. Always 2xx: any other status is an error. |
page.title, page.description | The page title, and its meta description or, if there is none, its Open Graph description. Whitespace is collapsed and each is capped at 1,000 characters. Empty when the page has none, and for types other than HTML. |
page.content_type | text/markdown for HTML pages. For text types, the same value as the source type. |
page.source_content_type | The media type the site sent, without parameters. If the site sent no usable type, the detected type. |
page.fetched_at | When the site’s response arrived, as an RFC 3339 time in UTC. |
page.content | The requested part of the Markdown or text. |
page.start_index, page.returned_length, page.total_length | Where the part starts, how many characters it holds, and how long the whole converted page is. Characters are Unicode code points. |
page.truncated, page.next_start_index | truncated is true when the page continues past this part. Only then is next_start_index present. |
page.source_truncated | true when only the beginning of the page was converted because it hit a size limit: 3 MiB of body, 50,000 HTML elements, or 8 MiB of Markdown. |
charged_units | Always 1. |
balance_remaining_usd | Your balance after this fetch, rounded down to four decimals. May be omitted. |
processing_time_ms | Server processing time in milliseconds. |
Long pages
To read past the first part, send the same request with start_index set to next_start_index.
- Every call downloads the page again and is billed again, so the page can change between calls.
- Keep
contentandinclude_linksthe same on every call. Both change the converted text and its length. - A
start_indexat or pasttotal_lengthreturns400 start_index_out_of_range. The message gives the currenttotal_length, and the error is not billed.
URL rules
A URL that breaks any of these rules returns 400 url_not_allowed. The error is free, and its message never repeats the URL.
- The scheme is
httporhttps, and the port is 80, 443, or omitted. - The URL contains no user name or password. A
#fragmentis removed. - The host is a public domain name. IP addresses in any form are refused. So are single-label names; local, internal, reserved, and test names such as
localhost,.local,.internal, and.test; PacketStream’s own domains; and some other hosts. - Every address the host resolves to must be public. A host that resolves to a private, loopback, link-local, or reserved address is refused.
- The URL must not contain a PacketStream API key, even percent-encoded.
Redirects
Fetch follows 301, 302, 303, 307, and 308 redirects, up to 5 per request.
- Each new URL is checked against the URL rules and is sent your User-Agent.
- A redirect to a URL that breaks the rules returns
400 url_not_allowed. - A sixth redirect returns
422 too_many_redirects. - A redirect to a different site also counts against that site’s rate limit, so it can return
429 rate_limited.
What fetch reads
- HTML and XHTML (
text/htmlandapplication/xhtml+xml) become Markdown, with links resolved against the final URL. The character encoding comes from a byte-order mark, the Content-Type charset, a meta charset, or sniffing, in that order. - Text (
text/plain,text/markdown,text/csv,text/xml,application/xml,application/json, and any+jsonor+xmltype) is returned as decoded text. A byte-order mark sets the encoding, then the Content-Type charset. Without either, XML follows its own encoding declaration and everything else is read as UTF-8. - Everything else returns
422 unsupported_content_type. That includes PDFs, images, other binary files, and bodies compressed with anything other than gzip. - No JavaScript. A page that shows little text and asks for JavaScript returns
422 javascript_required. A page with no readable text returns422 empty_content, and so does Markdown made only of links or images. - Formatting limits. Lists, quotes, tables, links, and emphasis nested more than 8 levels deep keep their text but lose their Markdown formatting. So does any table over 20,000 cells, and every table after the first 200,000 cells on a page.
- Links. Only
http,https, andmailtolinks andhttporhttpsimages are kept, without link titles, up to 20,000 in all.
What the site sees
- One plain
GETover HTTP/1.1. HTTPS uses TLS 1.2 or later with full certificate checks. - Your
user_agentor the default Chrome User-Agent, a browser-styleAccept,Accept-Encoding: gzip, andAccept-Language: en-US. No cookies and noReferer. - A connection from the PacketStream residential network in the chosen country, which defaults to the United States.
Fetch does not solve CAPTCHAs or switch exits after a block. It retries once, and only after a network or connection failure.