Extract Fields
Creates and starts an extract job from files or pre-uploaded handles. Extract pulls structured fields out of documents according to a schema.
Input: Exactly one of file and upload_ids must be provided.
Schema: exactly one of schema (inline JSON schema) and config_id (saved extraction configuration) is required.
Processing: the job runs asynchronously. Poll GET /doc-ai/v1/job/{job_id}/status until a terminal status (completed, partially_completed, failed, rejected), then fetch output via the results or download-url endpoints.
Extract results include result, plus annotations mirroring the result shape where every leaf has confidence and sources.
Authentication
Request
Document file(s) to extract from. Repeatable multipart part. Exactly one of file and upload_ids must be provided.
Comma-separated upload IDs returned by POST /doc-ai/v1/job/upload. Exactly one of file and upload_ids must be provided.
Inline extraction schema as a JSON string. The root must be type: "object" with non-empty properties; every field needs a type and a non-empty description. Supported types: string, number, integer, boolean, object, array (objects need properties, arrays need items); optional enum; maximum nesting depth 4. Exactly one of schema and config_id is required.
Saved extraction configuration ID. Exactly one of schema and config_id is required. To get your config_id, follow the steps listed here.
Language code of the document (BCP-47), e.g. en-IN, hi-IN.
Enable document classification. Boolean sent as text: true or false.