{
"status": "ready",
"text_url": "/v1/documents/550e8400-e29b-41d4-a716-446655440000/download?format=text&strategy=ocr"
}{
"workflow_run_id": "3c90c3cc-0d44-4b50-8888-8dd25736052a",
"workflow_run_item_id": "3c90c3cc-0d44-4b50-8888-8dd25736052a"
}{
"error": {
"code": "validation_error",
"message": "The request is invalid. details lists up to 20 field errors. Request bodies are limited to 2 MiB.",
"request_id": "req_example"
}
}{
"error": {
"code": "invalid_token",
"message": "The bearer token is missing or invalid.",
"request_id": "req_example"
}
}{
"error": {
"code": "forbidden",
"message": "Access denied. The error code says why: forbidden, token_disabled, organization_required, insufficient_role or mfa_required.",
"request_id": "req_example"
}
}{
"error": {
"code": "not_found",
"message": "The resource does not exist in the current organization.",
"request_id": "req_example"
}
}{
"error": {
"code": "extraction_in_progress",
"message": "A different extraction of this document is running (see workflow_run_id), or the uploaded file is missing or has changed.",
"request_id": "req_example"
}
}{
"error": {
"code": "unsupported_media_type",
"message": "This document type cannot be extracted. Extraction supports PDF, DOCX, XLSX, EML, MSG, PNG and JPEG, judged by the declared MIME type.",
"request_id": "req_example"
}
}{
"error": {
"code": "rate_limit_exceeded",
"message": "Too many requests. Wait for the Retry-After delay before retrying.",
"request_id": "req_example"
}
}{
"error": {
"code": "internal_error",
"message": "Unexpected server failure. Include the request ID when contacting support.",
"request_id": "req_example"
}
}{
"error": {
"code": "service_unavailable",
"message": "The service is temporarily unavailable. If workflow_run_id is present, check that run before retrying; otherwise retry after Retry-After.",
"request_id": "req_example"
}
}Extract document text
Converts a document to Markdown text. PDF and DOCX support ocr and deterministic, PNG and JPEG support ocr, and XLSX, EML and MSG support deterministic. Returns 200 when text from this strategy already exists; otherwise follow the returned workflow run, then read the text with GET /v1/documents//text or download it with format=text.
{
"status": "ready",
"text_url": "/v1/documents/550e8400-e29b-41d4-a716-446655440000/download?format=text&strategy=ocr"
}{
"workflow_run_id": "3c90c3cc-0d44-4b50-8888-8dd25736052a",
"workflow_run_item_id": "3c90c3cc-0d44-4b50-8888-8dd25736052a"
}{
"error": {
"code": "validation_error",
"message": "The request is invalid. details lists up to 20 field errors. Request bodies are limited to 2 MiB.",
"request_id": "req_example"
}
}{
"error": {
"code": "invalid_token",
"message": "The bearer token is missing or invalid.",
"request_id": "req_example"
}
}{
"error": {
"code": "forbidden",
"message": "Access denied. The error code says why: forbidden, token_disabled, organization_required, insufficient_role or mfa_required.",
"request_id": "req_example"
}
}{
"error": {
"code": "not_found",
"message": "The resource does not exist in the current organization.",
"request_id": "req_example"
}
}{
"error": {
"code": "extraction_in_progress",
"message": "A different extraction of this document is running (see workflow_run_id), or the uploaded file is missing or has changed.",
"request_id": "req_example"
}
}{
"error": {
"code": "unsupported_media_type",
"message": "This document type cannot be extracted. Extraction supports PDF, DOCX, XLSX, EML, MSG, PNG and JPEG, judged by the declared MIME type.",
"request_id": "req_example"
}
}{
"error": {
"code": "rate_limit_exceeded",
"message": "Too many requests. Wait for the Retry-After delay before retrying.",
"request_id": "req_example"
}
}{
"error": {
"code": "internal_error",
"message": "Unexpected server failure. Include the request ID when contacting support.",
"request_id": "req_example"
}
}{
"error": {
"code": "service_unavailable",
"message": "The service is temporarily unavailable. If workflow_run_id is present, check that run before retrying; otherwise retry after Retry-After.",
"request_id": "req_example"
}
}Authorizations
Personal API key sent as a bearer token, together with X-Organization-Id. The key must have at least the access level the endpoint requires (read, write or admin).
Headers
The organization to act in. Required for personal API keys; list the organizations you can access with GET /v1/organizations.
Path Parameters
Body
ocr recognizes text and tables from the rendered pages; deterministic converts the text embedded in the file. Defaults to ocr for PDF, PNG and JPEG, and deterministic otherwise.
ocr, deterministic A different strategy to try once if the first finds no text. Omit for no fallback.
ocr, deterministic