Skip to content

OCR

Extract text and layout from PDFs and images with OCR processors.

Updated View as Markdown

OCR reads the text in PDFs and images. You choose a model from the catalog, save its settings as a processor in one of your projects, and send documents to the processor. Each request returns the text of every page, and optionally each text block’s position, within 60 seconds.

Concepts

Term Meaning
Model An OCR model in the catalog, such as unlimited-ocr, with its versions, capabilities and limits. The catalog is the same for every organization.
Processor A saved configuration in a project: a model and version, the document languages, the output format and the detail level. Processors are regional resources (us-west-1) with an ID, a prc_ slug and a resource name.
Processing Sending one document (up to 10 MiB and 5 pages) to a processor and receiving each page’s result in the response.

The OCR API is at https://ocr.ezghcloud.com. In the console, open OCR; it has Processors (for the current project) and Models. The ezgh ocr commands cover the same tasks.

Permissions

Action Allows OcrFullAccess OcrReadOnlyAccess ReadOnlyAccess
ocr.models.list List models Yes Yes Yes
ocr.models.get Read a model Yes Yes Yes
ocr.processors.list List a project’s processors Yes Yes Yes
ocr.processors.get Read a processor Yes Yes Yes
ocr.processors.create Create processors in a project Yes No No
ocr.processors.update Change a processor’s name, description, tags and configuration Yes No No
ocr.processors.disable Disable a processor Yes No No
ocr.processors.enable Enable a processor Yes No No
ocr.processors.delete Delete a processor Yes No No
ocr.documents.process Process documents with a processor Yes No No

AdministratorAccess and the root user can do all of these. API keys act with their owner’s access. Processor actions can be limited to specific projects or processors by resource name in a custom policy. See IAM.

Trails

OCR calls are recorded in Trails with the event source ocr.ezghcloud.com, under the same rules as every other API. A ProcessDocument event records the number of pages, the document’s size and its type, never its content, file name or extracted text. OCR doesn’t store your documents or their results.

Navigation

Type to search…

↑↓ navigate↵ selectEsc close