Skip to content

Process documents

Send a PDF or image to an OCR processor and read each page's text and layout.

Updated View as Markdown

Processing sends one document to a processor and returns the result in the response. Nothing is stored: OCR keeps neither the document nor its text. You need ocr.documents.process on the processor, and the processor must be active.

Documents

Limit
Types PDF, PNG, JPEG, TIFF and WebP, detected from the file’s content
Size 10 MiB (10,485,760 bytes)
Pages 1 to 5 pages in the whole document, even when you select fewer
Image size Up to 100 million pixels. A PDF page can be up to 20 million pixels at 200 DPI.
Time 60 seconds per request

Each PDF page and each TIFF frame is a page. A PNG, JPEG or WebP image is one page.

Process a document

  1. In the console, open OCR > Processors and select the processor.
  2. Select the Try it tab.
  3. Choose a file and select Process.
  4. The result shows each page’s text in its own tab. Copy a page, or select Download text (or Download JSON for a json-layout processor) to save every page.
ezgh ocr process receipt.jpg --processor receipts --project web
ezgh ocr process scan.pdf --processor forms --pages 1-2 -o json --out scan.json --project web
curl -s https://example.com/form.pdf | ezgh ocr process - --processor forms --project web
  • --processor (required) takes the processor’s ID, slug or name. - reads the document from standard input.
  • By default ezgh prints the text of each page that succeeded, with a --- page N --- line before each when there is more than one. -o json prints the API’s response; -o yaml prints it as YAML. --out writes to a file instead.
  • A failed page’s error goes to standard error and the command still exits 0. --strict exits 1 if any page failed.
  • ezgh waits up to 2 minutes for the response unless you set --timeout. It retries only after 503 capacity_unavailable, never after a response or a request that got none.

Call ProcessDocument with a multipart upload:

curl -X POST \
  -H "Authorization: Bearer $EZGH_API_KEY" \
  -F "document=@receipt.pdf" \
  -F 'options={"pages": "1-2"}' \
  "https://ocr.ezghcloud.com/v1/organizations/$ORG_ID/projects/$PROJECT_ID/ocr/processors/$PROCESSOR_ID/process"

Or with JSON, the document’s bytes in standard base64:

curl -X POST \
  -H "Authorization: Bearer $EZGH_API_KEY" \
  -H "Content-Type: application/json" \
  -d "{\"document\": {\"content\": \"$(base64 < receipt.png | tr -d '\n')\", \"mimeType\": \"image/png\"}}" \
  "https://ocr.ezghcloud.com/v1/organizations/$ORG_ID/projects/$PROJECT_ID/ocr/processors/$PROCESSOR_ID/process"

Request

The request is multipart/form-data or application/json. Any other Content-Type is refused with 415 unsupported_media_type.

Multipart has these parts, each at most once:

Part Required Content
document Yes The document’s bytes
options No JSON: { "pages": "1-3,5" }

JSON is this object:

Field Required Content
document.content Yes The document’s bytes, standard base64
document.mimeType Yes application/pdf, image/png, image/jpeg, image/tiff or image/webp. It must match the document’s content.
pages No A page selection

Unknown fields and parts, repeated parts, and null values are refused with 400 invalid_request.

Page selection

pages lists one-based page numbers and inclusive ranges, separated by commas: "1-3,5". Ranges go up (2-4, not 4-2), no page can appear twice, and every page must exist in the document. The selection is at most 64 characters. Results come back in the order you list the pages. Without pages, every page is processed.

Idempotency-Key has no effect on processing. Every request processes the document again and counts again. Don’t retry a request that got a response.

Response

A request with at least one successful page returns 200:

{
  "processor": "ezgh:us-west-1:org_k3f9a0x2m7qp:prj_g7h8i9j0k1l2:prc_4k2m9x0a7qpe",
  "model": "unlimited-ocr",
  "modelVersion": "1.0.0",
  "outputFormat": "text",
  "status": "partial",
  "pages": [
    { "page": 1, "status": "succeeded", "content": "ACME Supplies\nInvoice 10432\nTotal\t$84.20", "layout": null },
    { "page": 2, "status": "failed", "error": { "code": "page_timeout", "message": "Page processing timed out" } }
  ],
  "pageCount": 2,
  "pagesSucceeded": 1,
  "pagesFailed": 1,
  "usage": { "pages": 1 }
}
Field Description
processor The processor’s resource name
model, modelVersion The model and the exact version that processed the document
outputFormat text or json-layout
status succeeded when every page succeeded, partial when some failed
pages One entry per processed page, in the order requested
pageCount The number of pages processed (the selected pages)
pagesSucceeded, pagesFailed How many pages succeeded and failed
usage.pages The pages you’re charged for: the pages that succeeded

A page that succeeded has page, status: "succeeded", content and layout. A page that failed has page, status: "failed" and error:

Page error code Meaning
page_timeout The page didn’t finish within the request’s 60 seconds
page_truncated The page’s text was longer than the model can return in one response
page_render_failed The page couldn’t be rendered
page_inference_failed The model couldn’t read the page
capacity_unavailable There was no capacity to process the page

Output formats

  • text: content is the page’s text, and layout is null.

  • json-layout: content is the page’s text, and layout also gives each text block’s position:

    {
      "page": 1,
      "status": "succeeded",
      "content": "ACME Supplies\nInvoice 10432",
      "layout": {
        "width": 1700,
        "height": 2200,
        "blocks": [
          { "type": "text", "text": "ACME Supplies", "boundingBox": [142, 118, 866, 190] },
          { "type": "text", "text": "Invoice 10432", "boundingBox": [142, 236, 640, 290] }
        ]
      }
    }

    width and height are the page’s size in pixels as it was read. boundingBox is [left, top, right, bottom] in those pixels. type is the block’s kind as the model labels it, in lowercase letters and underscores. PDF pages are read at 200 DPI. Images are read at their own size after EXIF orientation, scaled down to at most 20 million pixels.

content joins the page’s blocks with line breaks. Tables come back as text: cells separated by tabs, rows by line breaks. Treat content as plain text, never as HTML.

Errors

Status Code Cause
400 invalid_request A malformed body, invalid base64, an invalid page selection, or a document that can’t be read. issues lists each problem
400 too_many_pages The document has more than 5 pages, or none
401 unauthenticated Missing or invalid credentials
403 access_denied The caller lacks ocr.documents.process on the processor
404 not_found The organization, project or processor doesn’t exist or you can’t see it
409 processor_disabled The processor is disabled
409 model_unavailable The processor’s model version is retired or no longer in the catalog
409 detail_unavailable The processor’s options.detail is no longer offered; update it to high
409 organization_deletion_pending The organization is being deleted
413 payload_too_large The document is over 10 MiB
415 unsupported_media_type The Content-Type isn’t JSON or multipart, the document isn’t a supported type, or mimeType doesn’t match its content
429 quota_exceeded A quota refused the request: pages per billing period, requests per second, or concurrent requests. No Retry-After. See Pricing and limits
429 too_many_requests Too many requests. Retry after the Retry-After header’s seconds
502 inference_failed No page could be processed
503 capacity_unavailable No capacity to start processing. Retry after the Retry-After header’s seconds
503 unavailable A dependency is unavailable. Retry after the Retry-After header’s seconds
504 processing_timeout No page finished within 60 seconds

Only a 200 response is charged, for its usage.pages. Retry a 503 or a 429 too_many_requests after the Retry-After header’s seconds.

Navigation

Type to search…

↑↓ navigate↵ selectEsc close