Skip to main content
The Parse API accepts a document, an optional model, and an optional options object. The same core fields work whether you call it synchronously or create a Parse Job, which takes a few additional parameters of its own.

Sample Request

Create a Parse Job on the standard service tier, then collect the parse results from the finished job. Replace YOUR_API_KEY with your API key and document.pdf with the path to your file.

Parameters

Model Version

The model form field selects the parsing model family and snapshot:
  • dpt-3-pro: the highest-quality model. Default.
  • dpt-3-verity: the lower-latency model for digitally created text documents. In preview.
Requests that omit model use the latest DPT-3 Pro snapshot, and the response reports the resolved version in metadata.model_version. For the full comparison, all accepted values, and guidance on pinning snapshots, see Parsing Models. For example, this request selects the latest DPT-3 Pro snapshot:
cURL

Request Options

The optional options form field is a JSON object that customizes the response. All fields are optional; omitted fields take the defaults shown.

Example Options

The following examples all parse the same sample calibration report, so you can see exactly what each option changes. For the complete, unmodified output, download the full default response (parsed with no options). Each example shows the full request with the relevant options line highlighted, followed by a trimmed view of the parse results the finished job returns in result.

Process Specific Pages

Pass pages to parse only the pages you need. Page numbers are 1-indexed: 1 is the document’s first page. metadata.page_count still reports the full document length, while structure.children returns only the requested pages.
When you get the parse results, page_count counts all pages but structure.children holds only pages 1 and 3:

Render Tables as Markdown

Tables are returned as HTML by default. HTML preserves merged cells and nested layouts that pipe syntax cannot represent, such as the Measured Value header spanning the As Found and As Left columns in this calibration table:
Set blocks.table.format to markdown to receive tables as pipe-syntax Markdown instead. Merged cells expand into empty adjacent cells:
When you get the parse results, the markdown field carries the same table in pipe syntax, with the merged header flattened into empty cells:

Suppress a Block Type

Set a block type’s markdown option to false to drop its content from the markdown. Suppressing a block type also speeds up parsing, because suppressed blocks skip model captioning. Here, figures are suppressed, so each figure collapses to a > [!FIGURE] header line. The figure still appears in structure, with a zero-length range and an empty atomic_grounding array.
When you get the parse results, the markdown field shows each figure as a header line instead of its content:

Omit Atomic Grounding

Set atomic_grounding to false to drop the fine-grained atomic_grounding array from every block, leaving only the block-level grounding. Use this to reduce response size when you don’t need the fine-grained coordinates.
When you get the parse results, each block carries only its block-level grounding, and the atomic_grounding field is omitted entirely:

Include Markdown on Each Node

Set inline_markdown to true to add each node’s slice of the document Markdown as a markdown field on the node itself, so you can read a block’s text without slicing the top-level string by range.
When you get the parse results, every structure node (document, page, block, and table cell) carries its markdown slice. Entries in atomic_grounding do not. To get an entry’s text, slice by its range or read the parent block’s markdown: