---
title: "Process documents"
description: "Send a PDF or image to an OCR processor and read each page's text and layout."
---

> Documentation Index
> Fetch the complete documentation index at: https://docs.ezghcloud.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Process documents

Processing sends one document to a [processor](/ocr/processors) and returns the result in the
response. Nothing is stored: OCR keeps neither the document nor its text. You need
`ocr.documents.process` on the processor, and the processor must be active.

## Documents

| | Limit |
| --- | --- |
| Types | PDF, PNG, JPEG, TIFF and WebP, detected from the file's content |
| Size | 10 MiB (10,485,760 bytes) |
| Pages | 1 to 5 pages in the whole document, even when you select fewer |
| Image size | Up to 100 million pixels. A PDF page can be up to 20 million pixels at 200 DPI. |
| Time | 60 seconds per request |

Each PDF page and each TIFF frame is a page. A PNG, JPEG or WebP image is one page.

## Process a document

### Console

1. In the [console](https://console.ezghcloud.com), open **OCR** > **Processors** and select the
processor.
2. Select the **Try it** tab.
3. Choose a file and select **Process**.
4. The result shows each page's text in its own tab. Copy a page, or select **Download text**
(or **Download JSON** for a `json-layout` processor) to save every page.
### CLI

```sh
ezgh ocr process receipt.jpg --processor receipts --project web
ezgh ocr process scan.pdf --processor forms --pages 1-2 -o json --out scan.json --project web
curl -s https://example.com/form.pdf | ezgh ocr process - --processor forms --project web
```

- `--processor` (required) takes the processor's ID, slug or name. `-` reads the document from
standard input.
- By default `ezgh` prints the text of each page that succeeded, with a `--- page N ---` line
before each when there is more than one. `-o json` prints the API's response; `-o yaml`
prints it as YAML. `--out` writes to a file instead.
- A failed page's error goes to standard error and the command still exits `0`. `--strict`
exits `1` if any page failed.
- `ezgh` waits up to 2 minutes for the response unless you set `--timeout`. It retries only after
`503 capacity_unavailable`, never after a response or a request that got none.
### API

Call [ProcessDocument](/ocr-api/ProcessDocument/) with a multipart upload:

```sh
curl -X POST \
  -H "Authorization: Bearer $EZGH_API_KEY" \
  -F "document=@receipt.pdf" \
  -F 'options={"pages": "1-2"}' \
  "https://ocr.ezghcloud.com/v1/organizations/$ORG_ID/projects/$PROJECT_ID/ocr/processors/$PROCESSOR_ID/process"
```

Or with JSON, the document's bytes in standard base64:

```sh
curl -X POST \
  -H "Authorization: Bearer $EZGH_API_KEY" \
  -H "Content-Type: application/json" \
  -d "{\"document\": {\"content\": \"$(base64 < receipt.png | tr -d '\n')\", \"mimeType\": \"image/png\"}}" \
  "https://ocr.ezghcloud.com/v1/organizations/$ORG_ID/projects/$PROJECT_ID/ocr/processors/$PROCESSOR_ID/process"
```

## Request

The request is `multipart/form-data` or `application/json`. Any other `Content-Type` is refused
with `415 unsupported_media_type`.

**Multipart** has these parts, each at most once:

| Part | Required | Content |
| --- | --- | --- |
| `document` | Yes | The document's bytes |
| `options` | No | JSON: `{ "pages": "1-3,5" }` |

**JSON** is this object:

| Field | Required | Content |
| --- | --- | --- |
| `document.content` | Yes | The document's bytes, standard base64 |
| `document.mimeType` | Yes | `application/pdf`, `image/png`, `image/jpeg`, `image/tiff` or `image/webp`. It must match the document's content. |
| `pages` | No | A page selection |

Unknown fields and parts, repeated parts, and `null` values are refused with
`400 invalid_request`.

### Page selection

`pages` lists one-based page numbers and inclusive ranges, separated by commas: `"1-3,5"`. Ranges
go up (`2-4`, not `4-2`), no page can appear twice, and every page must exist in the document. The
selection is at most 64 characters. Results come back in the order you list the pages. Without
`pages`, every page is processed.

`Idempotency-Key` has no effect on processing. Every request processes the document again and
counts again. Don't retry a request that got a response.

## Response

A request with at least one successful page returns `200`:

```json
{
  "processor": "ezgh:us-west-1:org_k3f9a0x2m7qp:prj_g7h8i9j0k1l2:prc_4k2m9x0a7qpe",
  "model": "unlimited-ocr",
  "modelVersion": "1.0.0",
  "outputFormat": "text",
  "status": "partial",
  "pages": [
    { "page": 1, "status": "succeeded", "content": "ACME Supplies\nInvoice 10432\nTotal\t$84.20", "layout": null },
    { "page": 2, "status": "failed", "error": { "code": "page_timeout", "message": "Page processing timed out" } }
  ],
  "pageCount": 2,
  "pagesSucceeded": 1,
  "pagesFailed": 1,
  "usage": { "pages": 1 }
}
```

| Field | Description |
| --- | --- |
| `processor` | The processor's resource name |
| `model`, `modelVersion` | The model and the exact version that processed the document |
| `outputFormat` | `text` or `json-layout` |
| `status` | `succeeded` when every page succeeded, `partial` when some failed |
| `pages` | One entry per processed page, in the order requested |
| `pageCount` | The number of pages processed (the selected pages) |
| `pagesSucceeded`, `pagesFailed` | How many pages succeeded and failed |
| `usage.pages` | The pages you're charged for: the pages that succeeded |

A page that succeeded has `page`, `status: "succeeded"`, `content` and `layout`. A page that
failed has `page`, `status: "failed"` and `error`:

| Page error code | Meaning |
| --- | --- |
| `page_timeout` | The page didn't finish within the request's 60 seconds |
| `page_truncated` | The page's text was longer than the model can return in one response |
| `page_render_failed` | The page couldn't be rendered |
| `page_inference_failed` | The model couldn't read the page |
| `capacity_unavailable` | There was no capacity to process the page |

### Output formats

- **`text`**: `content` is the page's text, and `layout` is `null`.
- **`json-layout`**: `content` is the page's text, and `layout` also gives each text block's
  position:

  ```json
  {
    "page": 1,
    "status": "succeeded",
    "content": "ACME Supplies\nInvoice 10432",
    "layout": {
      "width": 1700,
      "height": 2200,
      "blocks": [
        { "type": "text", "text": "ACME Supplies", "boundingBox": [142, 118, 866, 190] },
        { "type": "text", "text": "Invoice 10432", "boundingBox": [142, 236, 640, 290] }
      ]
    }
  }
  ```

  `width` and `height` are the page's size in pixels as it was read. `boundingBox` is
  `[left, top, right, bottom]` in those pixels. `type` is the block's kind as the model labels
  it, in lowercase letters and underscores. PDF pages are read at 200 DPI. Images are read at
  their own size after EXIF orientation, scaled down to at most 20 million pixels.

`content` joins the page's blocks with line breaks. Tables come back as text: cells separated by
tabs, rows by line breaks. Treat `content` as plain text, never as HTML.

## Errors

| Status | Code | Cause |
| --- | --- | --- |
| `400` | `invalid_request` | A malformed body, invalid base64, an invalid page selection, or a document that can't be read. `issues` lists each problem |
| `400` | `too_many_pages` | The document has more than 5 pages, or none |
| `401` | `unauthenticated` | Missing or invalid credentials |
| `403` | `access_denied` | The caller lacks `ocr.documents.process` on the processor |
| `404` | `not_found` | The organization, project or processor doesn't exist or you can't see it |
| `409` | `processor_disabled` | The processor is disabled |
| `409` | `model_unavailable` | The processor's model version is retired or no longer in the catalog |
| `409` | `detail_unavailable` | The processor's `options.detail` is no longer offered; update it to `high` |
| `409` | `organization_deletion_pending` | The organization is being deleted |
| `413` | `payload_too_large` | The document is over 10 MiB |
| `415` | `unsupported_media_type` | The `Content-Type` isn't JSON or multipart, the document isn't a supported type, or `mimeType` doesn't match its content |
| `429` | `quota_exceeded` | A quota refused the request: pages per billing period, requests per second, or concurrent requests. No `Retry-After`. See [Pricing and limits](/ocr/pricing-and-limits#quotas) |
| `429` | `too_many_requests` | Too many requests. Retry after the `Retry-After` header's seconds |
| `502` | `inference_failed` | No page could be processed |
| `503` | `capacity_unavailable` | No capacity to start processing. Retry after the `Retry-After` header's seconds |
| `503` | `unavailable` | A dependency is unavailable. Retry after the `Retry-After` header's seconds |
| `504` | `processing_timeout` | No page finished within 60 seconds |

Only a `200` response is charged, for its `usage.pages`. Retry a `503` or a
`429 too_many_requests` after the `Retry-After` header's seconds.

Source: https://docs.ezghcloud.com/ocr/process-documents/index.mdx
